Skip to main content
bash TV

When Software Starts Thinking: The Now of Quality Engineering

TestMu AI (Formerly LambdaTest)

1 Sept 2026

YouTube

In this session, π‚π‘π’π§π¦πšπ²πš 𝐊𝐨𝐭𝐑𝐚𝐫𝐒, Director - Quality Engineering and AI CoE at QualiZeal, marks the end of the rule that quality engineering was built on: given the same input, software produces the same output. LLM-based applications, RAG systems and AI agents reason, adapt, use memory, invoke tools and collaborate with other agents - producing a class of enterprise applications where correctness is no longer binary. Chinmaya also argues that quality engineering has to evolve into AI Assurance, a discipline focused on building confidence in AI systems rather than verifying functionality, using enterprise examples covering non-deterministic outputs, RAG pipeline validation, hallucination measurement, autonomous agent workflows and prompt injection defense. 𝐇𝐒𝐠𝐑π₯𝐒𝐠𝐑𝐭𝐬: 0:04 Session Opens - When Software Starts Thinking 0:35 Generative AI Broke the Rule: Same Input, Same Output 1:06 Speaker Intro - Chinmaya Kothari, Director of QE and AI COE, QualiZeal 2:08 The Opening Example: "Will This Contract Clause Hold Up?" 2:38 Same Question, Same System, Opposite Answers 3:08 The Deterministic World: Printing the Same Paragraph Every Time 4:12 How We Tested in That World - Any Change Was a Defect 4:43 The Translator Example: Same Paragraph, Three Different Translations 5:44 Interpretation Enters, and the Whole Testing Paradigm Changes 6:15 The Generative AI Example: "Delivery Within 30 Days" 6:46 Words Added, Meaning Changed - That's Hallucination 7:18 Confidence Does Not Mean Correctness 7:48 The Agentic Example: "Ship This Assignment by Friday" 8:19 Message Drift as a Distinct Failure Mode 8:50 Bias, Inaccuracy, Safety, PII and PHI Exposure 9:21 The Tester's Role Changes: From Correct to Trustable 9:51 The Paradigm Shift From Pass/Fail to Overall Trust 10:21 The Trust Attributes You Have to Evaluate 10:51 The NIST AI Risk Management Framework and Its Six Attributes 11:21 If Even One Attribute Fails, the System Is Not Trustable 11:53 The Four-Step Framework: Risk, Metric, Test, Evidence 12:23 Step One - Risk: Where Can Each Trust Attribute Go Wrong? 12:54 Citation Sufficiency as a RAG-Specific Reliability Risk 13:24 Step Two - Metric: How Do You Measure Hallucination? 13:54 170-Plus Metrics Defined So Far, and Still Growing 14:55 AI Systems Hallucinate by Design - You Cannot Stop It 15:26 Which Brings You to Thresholds 16:28 Five Hallucinations in 100 Is Acceptable; Six Is Not 17:30 One Risk, One Metric - Now Multiply It Across Everything 18:00 Step Four - Test: Actually Measuring the Metric 19:00 Packaging Everything as Evidence and an Audit Trail 19:30 The Trap: One Passing Trust Attribute Is Not a Passing System 20:31 Is 100% Trustable a Fair Requirement? 21:32 A Composite Trust Score Above 95% as the Release Gate 22:03 What Changed: We Used to Do Proofreading, Now We Fact-Check 23:37 Why Citation Becomes the Key RAG Metric 24:39 Citation Sufficiency and Citation Accuracy 25:40 Recapping RMTE: Risk, Metric, Test, Evidence 26:43 Back to the Contract Clause - That Wasn't a Bug 27:13 An LLM Hardly Ever Says "I Don't Know" 28:17 Too Much Creativity and Translation Becomes Fiction 30:37 Q&A: Validating an Agent's Tool Selection Before Production 32:08 Q&A: How Much Human Intervention Before an AI Test Can Be Trusted 33:41 Internal Chatbot vs Customer-Facing - Very Different Bars 34:13 Q&A: What Will Have Changed Most by Next Year? 35:15 Q&A: Testing an Agent's Decision Process, Not Just Its Answer 36:47 Only When Every Layer Is Trustable Is the System Trustable 37:17 Q&A: Are Confidence and Correctness Together the Best Combination? 38:18 Why Their Framework Lets You Weight Each Trust Attribute Register for TestMuConf 2027: https://www.testmuai.com/testmuconf-2027/?utm_source=youtube&utm_medium=organic&utm_term=&utm_campaign=when_software_starts_thinking #TestMuConf #TestMuAI #AIAssurance #QualityEngineering #RAG #LLMTesting #AgenticAI #AIQuality

Join the discussion

Sign in to join the discussion

Sign in