Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now. The question is who can run them fast and reliably. Yunmo Koo, founding engineer at FriendliAI, argues that open-weight models like GLM 5.2 and MiniMax M3 now compete with the best proprietary models. He shows GLM 5.2 and Claude Opus 4.8 building similar tower defense games from one prompt, with GLM 5.2 more than 5.6x cheaper. Then he runs the same model and the same coding agent on two providers: FriendliAI builds a Candy Crush-style game in about two minutes, the other takes more than three. He explains where the gap comes from: prefix caching with cache-aware routing and a GPU, CPU and NVMe cache hierarchy for long, repetitive agent inputs, plus hybrid speculative decoding that adapts to live traffic. He also covers why reliability matters for agents (one failed call can sink a whole run) and shares Kilo Code's results of up to 7x faster responses with fewer errors. In this talk: • Why open-weight models are now good enough for many workloads • Why the same model feels different on different providers • Prefix caching, cache-aware routing and a tiered KV cache • Hybrid speculative decoding and why reliability is part of speed SPEAKER Yunmo Koo, Founding Engineer, FriendliAI LinkedIn: https://www.linkedin.com/in/yunmokoo/ GitHub: https://github.com/kooyunmo LINKS FriendliAI: https://friendli.ai Kilo Code case study: https://friendli.ai/customers/kilo CHAPTERS 0:00 Intro 0:40 Open vs closed models 1:09 Open weights crossed the threshold 2:24 GLM 5.2 vs Opus: a tower defense game 2:49 5.6x cheaper 3:09 Provider quality 3:53 Same model, different provider 4:33 It's the whole stack 5:28 Long, repetitive agent inputs 6:27 Cache-aware routing 6:47 A cache hierarchy 7:37 Speculative decoding 9:06 Hybrid, self-improving speculation 9:41 Reliability is part of speed 11:06 Test your own workload 11:45 Kilo Code: up to 7x faster Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #LLMInference #OpenSourceAI #AIEngineer
More like this

AI Security Engineer Foundations + Certificate — Javier Garza, Snyk

SonarQube + OpenAI: Agentic Development — Killian Carlsen-Phelan, Sonar

Let Your Agent Cook: Using Skills to Evaluate and Improve Your App — Ankur Duggal, Arize AI

AI Apps in a Flash: Ship to GPUs Without Docker — Dean Quiñanola, Runpod
Join the discussion
Sign in to join the discussion
Sign in