From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI
An LLM correctness judge rejects all 13 reports from a financial analysis agent. Laurie Voss examines why: the agent used live web research, while the judge graded from its own knowledge without that research context. Supplying the collected sources to a faithfulness evaluator produces a more useful split of six faithful reports and seven unfaithful ones. The notebook builds the agent with the Claude Agent SDK and instruments it with OpenTelemetry and OpenInference to send traces to Arize AX. Reading those traces exposes failed file writes, missing report content, and excessive web searches. A deterministic ticker check adds a cheap first layer of evaluation. Voss builds a custom actionability rubric with explicit criteria, tagged inputs, examples, and binary labels. Each evaluator checks one dimension, so a failure identifies what needs attention. Human annotations provide a comparison set for evaluating the judge, with deliberately random demo labels illustrating how to investigate disagreement. The workshop covers held out tests, precision and recall, judge biases, and the distinction between capability and regression evals. Failing traces become datasets, their explanations guide prompt revisions, and controlled experiments compare the changed agent on the same cases. Online evaluations and production monitoring extend that loop to new traffic, with requirements and recurring failure themes directing subsequent coding agent fixes. The central practice is to read actual traces and define success before automating the grading. Speaker info: - https://x.com/seldo - https://seldo.com - https://github.com/Arize-ai/arize-skills Timestamps: 0:00 - Workshop overview and notebook setup 7:31 - Traces, evals, and the limits of vibes 11:21 - Code, LLM, and human evaluations 14:27 - Agent failures and grading valid solutions 17:18 - Capability and regression evals 21:19 - Instrumenting the Claude Agent SDK 28:47 - Building and tracing the financial analyst 36:26 - Generating diverse test queries 38:54 - Reading traces and defining success 45:31 - Categorizing failures and choosing priorities 49:16 - Writing a deterministic ticker check 58:14 - A correctness judge that rejects every report 62:11 - Faithfulness with the research context 66:25 - Writing a custom actionability rubric 77:35 - Focused evaluators and production criteria 80:00 - Comparing judges with human annotations 86:25 - Precision, recall, and judge biases 90:27 - Turning failures into evaluation datasets 92:39 - Using explanations to improve prompts 95:24 - Comparing changes with controlled experiments 102:29 - Online evaluations and production monitoring 106:01 - Coding agent fixes guided by failure themes 108:27 - Putting the full improvement loop together




Join the discussion
Sign in to join the discussion
Sign in