Skip to main content
bash TV

Microsoft Foundry Evaluation: How to Test AI Agents

Udzial (By Gaurav Khurana)

26 Sept 2026

YouTube

How do you test an AI agent when the same question gets a different answer every time? Assertions fail, so testers use evaluation. Here's how in Microsoft Foundry. This clip is from a live session in my Microsoft Foundry series, and it's the part made for testers. We look at why groundedness is the #1 metric enterprises care about (remember the Air Canada chatbot case), then run a real evaluation on an agent in Foundry. What you'll learn: - Why actual == expected doesn't work for AI, and what evaluation means instead - Groundedness, relevance and safety filters: what customers check first - Running an agent evaluation in Foundry with synthetic test data - Choosing evaluators, why each one costs an AI call, and dropping the ones you don't need - Custom thresholds (e.g. strict groundedness) and regex-style checks - How temperature affects answers, and why newer models like GPT-5 remove it - Reading scores, drilling into failures and exporting raw JSON Chapters: 0:00 Why Assertions Fail for AI Testing 0:35 Groundedness: The #1 Metric 0:54 The Air Canada Chatbot Case 1:26 Relevance, Safety & Red Teaming 2:24 Run an Evaluation in Foundry 3:21 Evaluators & What They Cost 3:38 Custom Evaluators, Thresholds & Regex 4:56 Temperature: Creative vs Consistent 6:03 Read Scores & Raw JSON Results Watch next: ▶ Microsoft Foundry - AI Platform (full playlist): https://www.youtube.com/playlist?list=PLeGiFBPpRC04 Connect: 🎓 Courses: https://gauravkhurana.com/courses 🌐 Website: https://gauravkhurana.com 📅 1:1 mentoring: https://topmate.io/gauravkhurana #AITesting #MicrosoftFoundry #LLMEvaluation #SoftwareTesting #SharingIsCaring

Join the discussion

Sign in to join the discussion

Sign in