Skip to main content
bash TV

From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 — Laurie Voss, Arize AI

AI Engineer

2.4K views5 Oct 2026

YouTube

A toy store agent fails a request for toys under $6 because its search tool has no price filter. Laurie Voss turns that missing feature into a live demonstration of an improvement loop: add observability, inspect what happened, find the problem, and ask a coding agent to fix it. The Wonder Toys application starts without traces flowing into Arize AX. Skills installed in the workshop repository help the coding agent add tracing and query the resulting data. The diagnosis reaches beyond exceptions and HTTP errors to search quality, including empty results and overly narrow filters. After adding price filtering, Voss repeats the shopping request and checks the new input and output in the trace view. The broader argument is that reading traces individually stops working when an application produces millions of them. Evals compress behavior into scores and explanations, but people can still become the bottleneck when they must interpret every failure. Voss introduces Signal as another layer that groups recurring problems and suggests fixes. In the demonstration, findings can become GitHub issues, evaluation datasets, new evaluators, or proposed pull requests. Audience questions press on deployment, trust, and preventing a fix from breaking behavior that already works. Regression evals provide that check, while Signal itself has tracing and an evaluation suite. The workshop moves from a manually requested repair toward continuous improvement, with definitions of good behavior directing what the automation should change. Speaker info: - https://x.com/seldo - https://seldo.com - https://github.com/Arize-ai/project-rosetta-stone Timestamps: 0:00 - From 101 to continuous improvement 3:12 - Traces as the source of truth 6:13 - Scaling beyond manual trace review 9:02 - Finding signals and closing the loop 12:03 - Workshop setup 14:13 - Wonder Toys and trust questions 16:31 - Adding observability with a coding agent 20:19 - Testing and inspecting live traces 22:25 - Exposing the missing price filter 24:41 - Asking an agent to analyze traces 28:51 - Diagnosing search quality problems 30:28 - Adding price filtering 33:05 - Continuous monitoring with Signal 36:06 - Deployment and regression evaluation questions 40:22 - Evaluating Signal itself

Join the discussion

Sign in to join the discussion

Sign in