Skip to main content
bash TV

Run AI Eval with Jev | LLM-as-a-Judge 5-Min Tutorial with Python

Geosley Andrades

134 views2 Oct 2026

YouTube

How do you know if your AI app is actually giving good answers? In this video, I walk through running a real AI evaluation with Jev, TypeSafe's decision model, using a medical diagnosis app as the example. We go from a simple rubric to automated scoring you can run from VS Code. What you'll learn: - How to turn quality criteria into a rubric (Bad / Average / Great) - How to build an evaluation dataset with common cases, failures, and edge cases - How to set up State and Questions in the Jev Playground - The difference between Jev's Noul, Choice, and Score question types -- How to run the whole eval from Python with the Jev API - How to compare different AI models fairly with blind scoring 🔗 Resources Jev / TypeSafe docs: https://docs.typesafe.ai/introduction TypeSafe console (get your API key): https://console.typesafe.ai/keys #AIEvals #LLMasaJudge #Jev #TypeSafe #AITesting #GenAI #Python

Join the discussion

Sign in to join the discussion

Sign in