Skip to main content
bash TV

Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab

AI Engineer

2.0K views9 Oct 2026

YouTube

Your agent scores great on benchmarks. Then real customers use it and it breaks. Felipe Blanes from the Amazon AGI Lab shares what he learned working directly with customers of Nova Act, Amazon's service for building browser agents, from research preview to general availability on AWS. He calls the core problem the benchmark illusion: static evals look great until customers do things nobody expected, and then all you have left is hope. His fix is an eval flywheel built around the customer: define success the way the customer does, capture signals (including by talking to customers), diagnose gaps in the model, harness or product, and feed them into decisions. He explains the trust cliff (80% reliability feels like more work; around 92% feels trustworthy), why being open about limits builds more trust than higher benchmarks, surprising use cases from Amazon Leo, Hertz and Sola, and four principles for keeping evals fresh. In this talk: • The benchmark illusion, and why static evals leave you hoping • A four-step eval flywheel driven by real customer signals • The trust cliff: why 80% reliability isn't enough • Why being open about limits builds more trust than higher scores SPEAKER Felipe Blanes, Amazon AGI Lab LinkedIn: https://www.linkedin.com/in/felipeblanes/ LINKS Amazon Nova Act: https://nova.amazon.com/act Nova Act docs: https://docs.aws.amazon.com/nova-act/latest/userguide/what-is-nova-act.html CHAPTERS 0:00 Intro 0:47 What Nova Act is 1:57 Working directly with customers 2:22 The benchmark illusion 3:21 Static evals and hope 4:26 Close the customer loop 4:41 The eval flywheel 4:56 Step 1: define success 5:11 Step 2: capture signals 5:31 Talk to your customers 5:51 Step 3: diagnose gaps 6:56 Step 4: feed decisions 7:26 Four stages of the customer journey 9:35 Lesson 1: the trust cliff 11:10 Lesson 2: be open about limits 12:45 Lesson 3: surprising use cases (Amazon Leo, Hertz, Sola) 14:44 Four principles 15:59 Evals should get smarter every week 16:14 Two takeaways Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #LLMEvals #AIAgents #AIEngineer

Join the discussion

Sign in to join the discussion

Sign in