Skip to main content
bash TV

Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase

AI Engineer

3.2K views5 Oct 2026

YouTube

The official judge said the agent succeeded 74% of the time. A better verifier said 38%. Miguel González Fernández, tech lead for Browserbase's agent platform, and Corby Rosset, researcher at Microsoft Research, present their research on verifiers for computer-use and web agents. Deterministic evals stopped scaling as agents improved, and the LLM judges bundled with popular web benchmarks turned out to be confidently wrong. Train against them and you get a more confident liar, not a better agent. They explain how the Universal Verifier works: task-specific rubrics, ranking the most relevant screenshots as evidence for each criterion, isolating errors, catching hallucinations and separating controllable from uncontrollable failures. It cut false positives from about half to near zero and agreed with humans as often as humans agree with each other. They also test whether an autoresearch loop can rebuild the verifier. In this talk: • Why LLM-as-judge verifiers on WebVoyager-style benchmarks overstate success • Rubric design: grade only what was asked, don't cascade errors, check the screenshots • Validating a verifier against human labels with Cohen's kappa • What autoresearch got right building a verifier in a day, and where humans still mattered SPEAKERS Miguel González Fernández, Tech Lead, Browserbase LinkedIn: https://www.linkedin.com/in/miguelgfz/ Corby Rosset, Researcher, Microsoft Research X: https://x.com/corby_rosset LinkedIn: https://www.linkedin.com/in/corbyrosset/ LINKS Paper: The Art of Building Verifiers for Computer Use Agents: https://arxiv.org/abs/2604.06240 Browserbase blog: https://www.browserbase.com/blog/building-verifiers-for-computer-use-agents Fara + CUAVerifierBench (GitHub): https://github.com/microsoft/fara Stagehand: https://www.browserbase.com/stagehand CHAPTERS 0:00 Intro 0:22 Browserbase and Microsoft Research 1:02 The research 1:36 Why deterministic evals stopped scaling 2:45 LLM judges: confident liars 3:47 74% vs 38%: the verifier gap 4:22 Why existing verifiers fail 4:57 How the Universal Verifier works 6:26 Four guiding principles 7:11 Good vs bad rubrics 7:51 Don't cascade errors 8:41 Catching subtle hallucinations 9:26 Controllable vs uncontrollable failures 10:28 Verifying the verifier with humans 11:32 It started correcting humans 12:23 Better verifier, better training data 13:37 Can autoresearch build the verifier? 16:15 Paper, code and benchmark 17:45 Q&A: beyond the browser 19:00 Q&A: avoiding overfitting Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #AgentEvals #ComputerUse #AIEngineer

Join the discussion

Sign in to join the discussion

Sign in