From Zero to Leaderboard: Agent Evaluation — Wolfram Ravenwolf, Weights & Biases
An agent repairs a broken git repository and earns a perfect score, but the run stays out of the default leaderboard because it tested only one task. Wolfram Ravenwolf uses that distinction to build an evaluation pipeline where every score can be traced back to its conditions. WolfBench compares the average, best and worst runs, the tasks solved at least once, and the tasks solved every time. Repeating the same configuration exposes the gap between a lucky success and dependable behavior. The configuration includes six layers: benchmark, runner, sandbox, agent, model service and settings. Versions, provider routing, system prompts, resource limits and timeouts all affect what the result actually measures. The live workflow launches a Terminal Bench task through Harbor, collects the results, filters incomplete runs in a Marimo dashboard, and turns valid runs into an interactive leaderboard. Weave preserves the trajectories behind the bars so reviewers can inspect commands, messages and failures. Ravenwolf examines two traps in interpreting the results: a cheaper model can consume enough tokens to raise total cost, and greater reasoning effort can lower the score when it leaves too little time to execute commands. Infrastructure failures need different treatment from genuine task failures, while zero temperature still does not guarantee identical runs. The resulting comparison helps teams choose a configuration for their own constraints, with costs and consistency visible alongside the headline score. Speaker info: - https://www.linkedin.com/in/wolframravenwolf/ - https://x.com/WolframRvnwlf - https://wolfbench.ai/ - https://github.com/wandb/WolfBench Timestamps: 0:00 - Workshop setup 0:52 - Meet Wolfram Ravenwolf 3:31 - Evaluation mindset and methodology 4:12 - Why one average hides reliability 5:48 - The five WolfBench metrics 6:30 - Repeat runs under uniform conditions 7:24 - Evaluate a configuration, not just a model 9:53 - Six layers of an evaluation pipeline 10:50 - Choose and pin the benchmark 12:37 - Harbor and evaluation runner versions 14:31 - Isolated sandboxes 16:37 - Compare agent harnesses 18:38 - Settings that change the result 20:00 - Prompts, tools and resource limits 21:33 - Timeouts, concurrency and retries 25:35 - Fix inference provider routing 27:39 - Launch the git recovery task 30:41 - Preserve evidence beyond the score 31:49 - The WolfBench collection toolkit 34:19 - Inspect runs in the Marimo dashboard 34:59 - Filter incomplete and invalid runs 36:09 - Export results and generate charts 36:50 - A perfect score on one task 37:18 - Read the interactive leaderboard 39:06 - Consistency, average and ceiling 41:22 - Token usage and total cost 45:24 - When more reasoning lowers the score 47:39 - Inspect the underlying Weave traces 49:13 - Review configurations and failure causes
More like this

AI Security Engineer Foundations + Certificate — Javier Garza, Snyk

Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI

SonarQube + OpenAI: Agentic Development — Killian Carlsen-Phelan, Sonar

Let Your Agent Cook: Using Skills to Evaluate and Improve Your App — Ankur Duggal, Arize AI
Join the discussion
Sign in to join the discussion
Sign in