Skip to main content
bash TV

The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian

AI Engineer

1 view22 Sept 2026

YouTube

Show a frontier model part of a chessboard, ask how many white squares are visible, and it answers 32. Andrew Dai's diagnosis: it saw a chessboard, knew chessboards have 32 white squares, and hallucinated the rest. The pattern matching that makes these models superb at naming flowers hurts them once a question needs counting or spatial grounding: they miscount a Catan player's roads from the pieces left off the board, and they miss a robot arm lifting a lid because they cannot hold state across a long video. His test for understanding versus reasoning: if a person can answer in one second, so can the model; if it takes longer, the model falls apart. The benchmarks hide this: one popular reasoning suite uses 32 by 32 pixel images, and a multimodal science exam can mostly be answered without the image. What is missing, he says, is visual thinking. Video generators produce cartoonish explosions because their training data is Hollywood and game engines; detection models are robust but passive. Elorian's approach has four parts: visual reasoning data that does not exist online, a synthetic data flywheel of evals, agents, SFT, and RL, architectural changes on top of the transformer, and native visual chain of thought, where the model draws boxes around every hotel before narrowing to the red ones. Dai spent twelve years at Google Brain and DeepMind, first authored the paper that introduced pretraining and fine tuning, and co led GLaM, PaLM 2 pretraining, and Gemini data. He closes on robotics, construction sites where safety rules live in policy text, and mechanical design, where a testing platform takes thousands of engineering hours and frontier models fail on blueprints and CAD. Speaker info: - https://x.com/andrewdai - https://www.linkedin.com/in/andrewdai/ Timestamps: 0:00 - Frontier models versus human visual reasoning 0:56 - Chessboard hallucination: pattern matching that hurts 2:32 - Catan roads, robot arms, and context amnesia 4:09 - Understanding versus reasoning: the one second test 5:44 - Why the popular benchmarks do not measure it 8:03 - The missing paradigm: visual thinking, not generation or labels 9:38 - Elorian's approach and visual chain of thought 11:01 - The team, from GLaM to Gemini 12:14 - Use cases: robotics, construction, mechanical design 17:36 - Where to find out more

Join the discussion

Sign in to join the discussion

Sign in