SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind
The team took captions from real videos, regenerated the same scenes with their own model, and ran a human eval. People largely preferred the generated version. Dumitru Erhan is quick to deflate that result. The output is not more realistic, it is sharper and more saturated with nicer skin tone, and human preference turns out to be an unreliable thing to optimize against. Nicole Brichtova has a name for the effect, the Instagram filter. The same blind spot surfaces in smaller ways. Their image model quietly began putting wedding rings on hands, and nobody internally caught it until an outside tester asked why, which is reward hacking arriving through the back door. Much of the hour circles one question, whether language is a good enough intermediate representation for any of this. Shane Gu argues it thins out exactly where humans are most sensitive. We carry poor vocabulary for audio, taste, smell and skin tone because those sit close to survival, and a professional wine taster he asked had resorted to borrowing the language people use to describe a date. The tell of AI video, swyx notes, is that everything sounds studio recorded, because studio recordings are the training data, and a model with no representation of standing further away cannot get that sound right either. Evaluation stays stubbornly manual underneath all of it. Ten people in a room, two videos side by side, pick one. Speaker info: Dumitru Erhan: - https://x.com/doomie - https://www.linkedin.com/in/dumitruerhan/ Shane Gu: - https://x.com/shaneguML - https://www.linkedin.com/in/shixiang-shane-gu/ Nicole Brichtova: - https://x.com/nbrichtova - https://www.linkedin.com/in/nicolebrichtova/ Timestamps: 0:00 - Introductions, and why this session exists 2:00 - Nano Banana 2 Lite and the Omni Flash APIs 4:30 - Beyond demos: storyboards, editing, education 7:51 - Video agents, or one model doing it all 10:26 - Video models as zero shot learners 13:01 - Will everything collapse into one model? 14:20 - Is captioning the right intermediate representation? 18:13 - What people actually mean by world models 19:44 - Vision researchers becoming generation researchers 24:51 - Audio, and why joint generation mattered 26:32 - The things language cannot describe 32:03 - Studio quality as the tell of AI video 34:12 - When humans prefer the generated video 36:07 - Prompt engineering and trained sensitivity 38:27 - Setting a default aesthetic for the world 41:09 - The wedding ring nobody noticed 42:10 - How you actually evaluate video 47:35 - What data they are looking for 53:04 - Forward deployed engineers as an eval channel




Join the discussion
Sign in to join the discussion
Sign in