Evolution solved selective attention.

Reproducing it in machines is the open problem.

The cocktail party problem

Our first problem is the one Cherry named in 1953: the cocktail party effect, how a listener pulls a single conversation from a room full of them. For voice AI, this is the core of addressee detection: knowing which speech is directed at the device versus another person in the room.

We started by cracking the howling in hybrid meetings, then went deep into how conversations converge and diverge. Selective Auditory Attention (SAA) reliably tells whether speech is addressed to the device, whether fusing audio and video or using audio alone, on held-out multi-party sessions (arXiv:2604.08412).

From neuroscience to addressee detection

We work backwards from how these systems filter, bind, and route attention, then rebuild those principles as something that can run continuously, on real devices, in real conversations. The result is a pre-ASR layer for device-directed speech that requires no wake word and no GPU.

Publications

arXiv:2604.08412 [cs.SD]

SAA (Selective Auditory Attention): Device-Addressed Speech Detection for Real-Time Voice AI

David Joohun Kim, Daniyal Anjum, Bonny Banerjee, Omar Abbasi · Submitted April 9, 2026

View on arXiv