Skip to main content
bash TV

Let Your Agent Cook: Using Skills to Evaluate and Improve Your App — Ankur Duggal, Arize AI

AI Engineer

125 views11 Oct 2026

YouTube

A financial agent invents answers, repeats tool calls and produces summaries that its own evaluations flag as wrong. Ankur Duggal connects those failures to a coding agent through Arize's CLI and skills, then asks it to investigate traces, change the application and run experiments. The demo begins with an investment question that looks reasonable at the surface but lacks enough context to support its recommendation. Traces reveal the supervisor, specialist agents, model calls and tools underneath that answer. They turn a vague complaint about quality into evidence about the exact steps that failed. Duggal sets up the CLI profile and skills, then shows the coding agent pulling spans and evaluation results into its local workspace. The investigation uncovers numeric grounding problems, fabricated responses, redundant searches and failures that the agent silently works around. An earlier experiment targets stock queries outside the United States: changing the country context alone is insufficient until the underlying tool is also updated. Dataset, evaluator and experiment skills make those proposed fixes testable rather than a collection of unverified edits. Questions cover local experiment execution, judge models, evaluation costs, sampling and the risk that a coding agent spends its effort on linting instead of the actual defect. The live repair is still running at the close. Duggal recommends human review of proposed changes and evaluations tailored to the application's failures, so later regressions can be detected as traffic and models change. Speaker info: - https://github.com/Arize-ai/arize-skills Timestamps: 0:00 - From vibes to trusted evaluations 1:58 - Traces reveal an agent's actual decisions 3:19 - A refund question and the wrong tool 4:15 - The improvement loop 6:03 - Bring traces, evaluations, and code together 7:27 - CLI and skills as the interface 8:06 - A flawed financial agent 11:05 - Set up a profile and install skills 13:19 - Inspect supervisor and specialist traces 16:03 - Grounding, fabrication, and silent failures 18:24 - Build evaluations and run experiments 20:05 - An earlier fix for international stock queries 22:23 - Where experiments and judges run 23:31 - Choose an evaluation model 24:37 - More context than the code alone 26:13 - What makes a summary good? 28:35 - Avoid focusing only on linting 30:21 - Evaluation costs and sampling 31:28 - Review the proposed changes 33:19 - Tailor evaluations to your application

Join the discussion

Sign in to join the discussion

Sign in