Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI
To get into the Einstein Arena you have to solve a puzzle proving you are an AI agent. Locking humans out is the point. James Zou and collaborators at Together AI and Stanford built it as an environment rather than a workflow, on the thesis that telling an agent how to work caps what it can do, while setting where it works and what it is rewarded for does not. Agents who log in find curated open problems, each with a deterministic verifier, a forum where they ask each other what has already failed, and a leaderboard that scores submissions live and exposes every solution. Within weeks of launch they held the best known answers to eleven of those problems. One is the kissing number problem: how many spheres can touch a central sphere without overlapping. Trivial in two dimensions, open for centuries in higher ones. Eleven dimensions had stood at 593, and agents on the arena reached 604 in a few days, each refining another's submission. No single frontier agent gets there alone. Swap kernel compilation and benchmarking in behind the same leaderboard and those agents produced kernels more than twice as fast as the prior state of the art, now in production at Together AI. He closes on DSGym, built after he found that twenty to fifty percent of tasks in popular data science benchmarks could be solved without touching the data. Frontier models still score under fifty percent on it, and its execution verified trajectories fine tune open source models that run on a laptop. Speaker info: - https://x.com/james_y_zou - https://www.linkedin.com/in/james-zou-2123a4133 - https://www.cs.stanford.edu/people/james-zou Timestamps: 0:00 - Design environments, not workflows 1:54 - Einstein Arena: prove you are an agent to enter 3:25 - A forum, a verifier, and a live leaderboard 5:19 - The kissing number problem 7:24 - 604 spheres in eleven dimensions 9:06 - The same arena, pointed at GPU kernels 10:57 - DSGym: a gym for data science agents 12:12 - Benchmarks you can beat without the data 14:30 - Small models trained on verified runs
More like this

The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs

Laundromat POS System & Laundry Software: Wash-Dry-Fold POS | SourceForge Podcast, episode #135

Aiarty Video Enhancer Review (2026) I Tried to Upscale 3 Blurry Videos to 4K — Here Are the Results

Stop Building AI Slop – Build High-End Web Apps with AI
Join the discussion
Sign in to join the discussion
Sign in