Skip to main content
bash TV

Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

AI Engineer

1.9K views15 Sept 2026

YouTube

Asked in Spanish about a mid century sofa, the model answers in Spanish but leaves the phrase mid century in English, because that is how the term is actually used by Spanish speakers. Nobody wrote a rule for that. It falls out of a model trained on audio, video, and text together rather than stitched from separate parts. Valeria Wu Fon leads product for speech to speech in Gemini and Tom Ouyang engineers it, and they open with how far the field moved to get here. Before roughly 2018, turning speech into text meant a chain of hand built components: feature extraction, acoustic modeling, pronunciation modeling, language modeling, and a rescoring pass. End to end models collapsed that chain but still only did the one job. Ask for the speaker's tone, or emotion, or pace, and you were back to building it yourself. Their framing is a three way tension rather than a roadmap. A speech to speech model should be conversational, meaning genuinely low latency; intelligent, meaning it completes tasks and follows instructions; and multimodal in both directions, taking video, screen shares, and documents as readily as speech. Moving any one of those tends to wreck the others, and they give a clean example: turn the thinking budget up and intelligence scores rise while time to first audio falls apart. The demos are chosen to show the corners. Live translation across a multi speaker meeting. A roadside assistance agent reading back a registration plate and a postcode. And a feature they call proactive audio, which is the model knowing that background noise is not its turn. Speaker info: - https://x.com/valeriawu_ - https://www.linkedin.com/in/valeriawu/ - https://www.linkedin.com/in/tom-ouyang-8b5a5142/ Timestamps: 0:00 - Why voice, and why now 1:51 - Speech recognition before 2018 3:27 - Natively multimodal pre training 4:19 - Streaming translation at offline quality 6:01 - Three vectors: conversational, intelligent, multimodal 7:44 - Turning one knob breaks another 8:37 - Live translation in a multi speaker meeting 12:14 - A roadside assistance agent under pressure 14:48 - Adding visual presence

Join the discussion

Sign in to join the discussion

Sign in