Skip to main content
bash TV

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

AI Engineer

3.3K views3 Oct 2026

YouTube

Open models have nearly caught up with proprietary ones. Serving them fast and cheaply is the hard part. Dylan Bristot, who leads product marketing for Nebius Token Factory, and developer advocate Sujee Maniyam explain what it takes to run open LLMs in production. Dylan covers the trade-off between closed APIs and self-hosting, and the inference → data → post-training → deployment loop. Sujee then goes layer by layer through the optimizations behind fast inference: NVFP4 on the latest NVIDIA hardware, engine selection, cache-aware routing, speculative decoding with custom draft models, KV cache offloading, disaggregated prefill and decode, and finding the quantization sweet spot. In this talk: • Why open models are now competitive, and what that means for cost and lock-in • Cache-aware routing versus naive load balancing for LLMs • Speculative decoding and training draft models on your own traffic • KV cache offloading, prefill/decode disaggregation and quantization trade-offs SPEAKERS Sujee Maniyam, AI Developer Advocate, Nebius LinkedIn: https://www.linkedin.com/in/sujeemaniyam/ X: https://x.com/sujee_dev Website: https://sujee.dev/ Dylan Bristot, Senior AI Product Marketing Manager, Nebius LinkedIn: https://www.linkedin.com/in/dylanbristot/ LINKS Nebius Token Factory: https://tokenfactory.nebius.com/ Token Factory docs: https://docs.tokenfactory.nebius.com/quickstart Nebius: https://nebius.com/ CHAPTERS 0:00 Intro 0:12 Engineering open LLMs for production 0:52 Who is Nebius 2:27 Nebius Token Factory 2:42 Closed APIs vs self-hosting 4:06 The full loop: inference, data, post-training, deploy 6:21 What we handle so you don't have to 8:06 From running models to improving them 10:01 Scaling open models to production 10:26 How good are open models? 11:26 Teacher and student models 11:46 Hardware layer: NVFP4 12:26 Choosing the serving engine 12:55 Cache-aware routing 14:24 Speculative decoding 15:29 Training custom draft models 15:59 KV cache and offloading 17:33 Disaggregated prefill and decode 18:13 Finding the quantization sweet spot 18:43 Recap 19:22 Try Token Factory Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #LLMInference #OpenSourceAI #AIEngineer

Join the discussion

Sign in to join the discussion

Sign in