Skip to main content
bash TV

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect

AI Engineer

1.0K views26 Sept 2026

YouTube

Big labs say recursive self-improvement is coming, but there's no independent benchmark to check that claim. Elie Bakouch, Research Engineer at Prime Intellect and creator of Hugging Face's SmolLM, set Claude Code and Codex loose on the community's Optimizer Speedrun, a race to train a GPT-2-level model in the fewest steps. Both agents beat the human record. Along the way they behaved very differently. Claude Code kept stopping every nine or ten hours to say the record couldn't be beaten, and sat idle about a third of the time. Codex never stopped, wrote far more notes, spawned more sub-agents and burned more tokens. In a longer six-day run, Kimi turned out to be the most token-efficient, and a paper only Claude found led to the best record. But Bakouch's key finding is sobering: none of the models invented a new optimizer. They combined existing ideas for small gains. He closes with an AlphaEvolve-style loop Prime Intellect is building for real discovery, and makes the case for doing this research in the open. Speaker info: X/Twitter: @eliebakouch (https://x.com/eliebakouch) LinkedIn: https://www.linkedin.com/in/eliebak/ Related links: Prime Intellect: https://www.primeintellect.ai Timestamps: 0:00 Intro: automated AI research 0:37 Why test recursive self-improvement in the open 1:52 Karpathy's GPT-2 speedrun and modded-nanogpt 3:12 The Optimizer Speedrun 4:32 Why speedruns make good environments 5:32 Claude Code and Codex vs. the community 6:52 The setup: goal.md, Slurm and preemptible jobs 7:46 Claude kept giving up; Codex never stopped 8:36 Scratchpads, sub-agents and token burn 10:21 Results: both beat the human record 11:36 Toward a real benchmark: three tracks 12:45 Six days of Claude, Codex, Kimi and GLM 13:45 Measured in tokens, the story changes 14:10 How each model uses research papers 14:35 No novel optimizers 15:40 An AlphaEvolve-style discovery loop 17:49 What Prime Intellect is building 18:54 Why this should happen in the open

Join the discussion

Sign in to join the discussion

Sign in