World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI
In 2007, Google already had a language model trained on 2 trillion tokens, the same order of magnitude as today's frontier models. Christopher Manning, Stanford professor, former director of the Stanford AI Lab and now at Moonlake AI, uses that as one stop in a tour of AI history. It runs from Dartmouth in 1956 and Shakey the robot to the rise of LLMs, and it ends at his argument for what comes next: embodied intelligence built on simulation. Manning argues that generative video like Genie 3 simulates observations but has no semantics underneath, so it can't support planning. Moonlake takes a single photo or short video and builds an action-conditioned world in code, with objects you can pick up, open and move. It even researches objects on the web to fill in what the camera can't see, like the tea bags inside a closed box. A loop inspired by Claude Code compares renders against reality to shrink the sim-to-real gap, with the goal of replacing 10,000 hours of teleoperation with simulation. The talk ends with Q&A on gaming, ontologies, physics and discovery in latent space. Speaker info: X/Twitter: @chrmanning (https://x.com/chrmanning) LinkedIn: https://www.linkedin.com/in/christopher-manning-011575 Website: https://nlp.stanford.edu/~manning/ Related links: Moonlake AI: https://moonlake.ai Timestamps: 0:00 Intro 1:02 Story time: a history of AI 2:02 Dartmouth 1956 and cybernetics 3:46 The first machine translation demo (1954) 5:41 The Stanford AI Lab, the Stanford Cart and Shakey 8:11 Why early AI happened at Stanford, not Berkeley 11:01 Language models, from Markov to Shannon to IBM 14:10 How we got to large language models 17:40 Why embodied AGI is the North Star 19:30 Why simulation: the 10,000-hour teleop problem 20:50 Shakey's world model and the model in your head 22:30 Action-conditioned world models 23:50 Genie 3 and the problem with pretty pixels 25:19 From a photo to a world you can act in 29:04 Filling in what the camera can't see 30:54 Building world models in code 32:54 Software eats the physical world 35:53 10,000 hours of simulation for free 38:08 Q&A: gaming, ontologies, physics and sim-to-real
More like this




Join the discussion
Sign in to join the discussion
Sign in