An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases
Live on stage, Weights & Biases' new agent ARIA kicks off a batch of autoresearch experiments on Karpathy's autoresearch project, running on real GPUs, and nearly beats its own best result before the talk ends. Tim Sweeney, Principal Engineer at Weights & Biases by CoreWeave, shows what ARIA does inside W&B. It launches and monitors training jobs, summarizes the best runs, finds patterns across 200+ experiments, and builds reports and dashboards. It's now on the W&B iOS app too. Sweeney then shows how the team builds ARIA. It logs 100% of traces to Weave and runs LLM judges on live traffic to catch signals like user frustration. Tasks are written as YAML "unit tests," and a nightly eval suite (73% vs. 72% for the latest candidate) drives go/no-go decisions. His tips: invest in agent-focused observability to catch behavioral bugs, treat evals as your new CI, keep humans reviewing traces, and add value through context and tools before over-engineering the harness. Speaker info: LinkedIn: https://www.linkedin.com/in/tssweeney Related links: Weights & Biases: https://wandb.ai W&B Weave: https://wandb.ai/site/weave Timestamps: 0:00 Intro 1:28 Agenda 2:03 About Weights & Biases 2:33 Demo: meet ARIA 3:13 Running Karpathy's autoresearch 4:33 Kicking off a live experiment batch 5:17 W&B Launch and the GPU cluster 5:57 Summarizing runs and finding patterns 7:02 Reports and workspaces 8:37 Recap: a data science companion 9:17 ARIA on iOS 9:42 Toward end-to-end automated research 10:22 How we built ARIA: the architecture 13:02 Demo: the agent dashboard in Weave 14:57 ARIA analyzing its own conversations 15:12 Signals: LLM judges on live traffic 16:07 Tasks as unit tests 17:17 Nightly evals 18:07 Recap 18:47 Tips for productionizing agents 20:12 Did the live run beat the record?




Join the discussion
Sign in to join the discussion
Sign in