Evals in AI: A Deep Dive — Tejas Kumar, IBM
A customer support answer passes a test because it contains the word “cannot,” even though it approves a return outside the store's policy. Tejas Kumar builds that failure live to show where ordinary assertions and fuzzy matching stop being useful. The example grows into an LLM judge, then a dataset of customer scenarios, agent responses, and human verdicts. Along the way, the judge starts favoring an answer generated by its own model family. Adding the actual return policy improves its decisions, but missing details and a loyal customer's appeal still expose gaps. Kumar keeps examining disagreements, supplying context, and eventually changing the model rather than treating a green test as proof that the system works. The workshop connects those experiments to an evaluation process that can survive production. Test cases, expected behavior, a judge, and aggregation form the basic structure. Position bias, sycophancy, self preference, and verbosity can distort the result, so the instrument itself needs calibration. Kumar sets an 80% agreement threshold for a CI gate, then explains why fresh traffic and human review must keep the dataset current. An OpenRAG demonstration retrieves an uploaded refund policy instead of relying on a rule buried in a prompt. He also distinguishes checks performed before deployment from harness protections at runtime, and recommends starting with deterministic checks before paying for model judgments. Audience questions explore synthetic examples, shared policy, and where evaluation belongs in the development loop. Speaker info: - https://x.com/tejaskumar_ - https://linkedin.com/in/tejasq - https://github.com/langflow-ai/openrag Timestamps: 0:00 - Introduction 2:03 - Why evals and harnesses work together 6:09 - Evals and unsafe agent behavior 9:47 - Fuzzy unit tests and evaluation components 15:21 - Evaluation techniques 18:30 - Four ways judges can mislead 21:11 - Calibrating a judge against human verdicts 23:30 - CI gates and keeping datasets current 26:15 - Retrieving living policy with OpenRAG 28:55 - Live coding from a unit test 31:26 - A misleading substring match 32:08 - Building and testing an LLM judge 36:19 - Calibrating against a scenario dataset 43:00 - Refining policy and changing the model 46:08 - Turning agreement into a CI gate 48:12 - A policy retrieval demo 51:53 - Recap and evaluation costs 54:36 - Audience questions




Join the discussion
Sign in to join the discussion
Sign in