Your Voice Agent is Just a Walkie Talkie — Neil Zeghidour, Gradium
On a phone call with someone close to you, as much as 20 percent of the time you are both talking at once. Every real time voice model shipping today is half duplex: it is either listening or speaking, never both. That gap is the whole argument of this talk. Neil Zeghidour is co founder and CEO of Gradium, a Paris company that trains audio foundation models, and he traces the arc from a 2011 phone assistant that mapped transcripts to app actions, through the first open ended voice chat that could converse but take no action, to cascaded voice agents that call tools competently while inheriting the latency and flattened emotion of routing through text. Speech to speech collapsed that pipeline and fixed latency, but he demonstrates the remaining problem live by back channeling at a model, saying mm hmm and yeah the way people do, and watching it stop dead every time. The technical core explains why audio was hard to put in a language model at all. Eight words take about three seconds to say, which at 24 kHz is 72,000 timesteps, and since attention cost grows with the square of sequence length, a sequence thousands of times longer is astronomically more expensive. Neural codecs solve it by compressing audio into token like representations a model can learn on. Full duplex then needs a further step: modeling two token streams at once so both sides can be active, silent, or overlapping. Zeghidour is candid that every gain in naturalness so far has cost intelligence, because a fixed weight budget spent on hearing and speaking is taken from reasoning, and he lays out the two ways forward. Speaker info: - https://x.com/neilzegh - https://gradium.ai Timestamps: 0:00 - Gradium and a research lab built around voice 1:55 - A 2011 assistant, and the pipeline behind it 3:39 - Open ended conversation, zero agency 5:18 - A drive through agent that takes real actions 7:48 - Where speech to speech still falls short 8:37 - Back channeling breaks turn taking 10:20 - Why raw audio will not fit in a language model 11:57 - Two streams instead of one 13:37 - The tension between naturalness and intelligence 15:18 - Path one: scale the speech to speech model 17:00 - Path two: split interface from intelligence




Join the discussion
Sign in to join the discussion
Sign in