Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust
A batch of 968 sample support traces exposes order lookup failures and broken escalation calls. Doug Guthrie uses those cases to connect Braintrust observability with the work of improving an agent. The workshop starts with a Python support agent built on the OpenAI Agents SDK, adds tracing, establishes an offline evaluation baseline, and publishes custom scorers. Online automations grade incoming traces, while Topics groups patterns in tasks, sentiment, and workflow issues. A custom facet targets failures that the default issue classification misses. Guthrie configures conversation grouping and sampling, imports the sample traces, and inspects the resulting labels and topic maps to find cases worth investigating. The next step is to turn those findings into changes. The embedded Loop assistant queries trace data with SQL and examines representative failures. Selected examples return to evaluation datasets so subsequent changes can be checked against real failure cases. A coding agent with the Braintrust CLI and an improvement skill extends that process into the repository: investigate logs, propose code and scorer changes, add regression examples, and run evaluations. Guthrie shows an automation that packages findings, explanations, and evaluation evidence into a pull request for review. Audience questions cover deployment, sampling rates, project scope, access controls, and how evaluation assets should evolve with the agent. A final demonstration exposes an agent evaluation to a playground, letting collaborators adjust prompts and models while execution stays on the evaluation server. Speaker info: - https://www.linkedin.com/in/doug-guthrie-07994a48 - https://github.com/dpguthrie - https://github.com/dpguthrie/mastering-ai-observability-workshop Timestamps: 0:00 - Workshop overview and companion repository 5:08 - Tracing as the foundation of observability 8:00 - Connecting production feedback to development 9:46 - Scorers, judges, and human calibration 15:06 - Topics and previously unknown failure modes 18:24 - The Topics processing pipeline 22:54 - Questions on Topics and deployment 32:37 - Creating an organization and project 39:13 - Repository and environment setup 42:27 - Braintrust CLI and coding agent skills 47:41 - Tracing the support agent and baseline evals 52:05 - Publishing custom scorers 54:45 - Online automation scope and sampling 57:23 - Custom topic facets and processing settings 63:08 - Importing the sample support traces 67:21 - Reading topic maps and workflow failures 71:48 - Loop and SQL analysis of trace data 73:52 - Bringing failures back to evaluation datasets 75:40 - Automating agent improvement with coding skills 84:15 - Deployment and model questions 87:55 - Topic dashboards, sampling, and score composition 92:47 - Project scope, changing evals, and trace access 101:50 - Reviewing code changes and new regression cases 104:32 - GitHub automation and reviewable pull requests 107:01 - Exposing remote evaluations to a playground




Join the discussion
Sign in to join the discussion
Sign in