AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash
The people annotating DoorDash's eval data are not engineers, and they build their own annotation tools. Because the GenAI platform team went API first, strategy and operations staff can point a coding agent at those endpoints and vibe code whatever interface their use case needs, whether that is grading restaurant menus or reviewing images. The platform team stopped trying to anticipate every UI, and shipped stable APIs instead. Nachiket Paranjape and Swaroop Chitlur Haridas make the broader case that evals stopped being an engineering harness for them and became a cross functional job. That reframing has an org chart attached. Strategy and operations set the quality bar, product managers turn it into rubrics, operations run the annotations, and engineering supplies telemetry, datasets and judges. Which group actually owns a judge prompt varies by team, and they treat that variation as a sign the org is still learning rather than a problem to standardize away. The loop underneath is deliberately plain: trace, sample down to something a human will really look at, annotate, promote a golden set, calibrate the judge against it, then monitor and go again. Judge calibration runs self serve through a UI, showing the original and optimized prompts side by side so a product manager can see what changed and decide whether to trust it. Per annotation cost fell sharply. Speaker info: Nachiket Paranjape: - https://x.com/nmparanjape - https://www.linkedin.com/in/nachiketparanjape/ Timestamps: 0:00 - The GenAI platform team, and its three forces 2:05 - Why eval became the fourth pillar 3:05 - UI first, then API first, then workflow first 4:01 - Evals as a team sport, not an engineering harness 4:57 - Who owns which part of quality 5:53 - The continuous loop: trace, sample, annotate, calibrate 7:42 - Telemetry and workflow as two surfaces 9:32 - Operators vibe coding their own annotation UIs 11:21 - Calibrating judge prompts, self serve 13:10 - Different teams, different prompt owners 14:04 - What it did to annotation cost
More like this

How Warmwind Automates Apps That Don't Have APIs

AI-Native Organisations Run on Skills: How to Structure and Scale Them — Imad Touil, QuantumBlack

Your Code Has Bugs. Lean4 Has Proofs: Formal Verification for Engineers — Varun Pant, AWS

How do you diffuse AI into the real world? — Varun Shenoy, Long Lake
Join the discussion
Sign in to join the discussion
Sign in