Is Speculative Decoding Worth It? Profiling vLLM on NVIDIA Blackwell — Akamai
A small model guesses the next few tokens. The big model checks them in one pass. When is that worth it? Sheilah Kirui, developer advocate at Akamai, explains speculative decoding and how to tell whether it's worth turning on. She walks through the prefill and decode phases of inference, how a small draft model proposes tokens that the target model verifies in a single forward pass, and the memory cost of hosting a second model and its KV cache. She covers how to choose a draft model, then demos both setups side by side with vLLM on a single NVIDIA Blackwell GPU: structured output ran 1.6x faster with a high acceptance rate, while creative writing saw far lower acceptance. She also explains why long-context and high-concurrency workloads gain less. In this talk: • Prefill vs decode, and what speculative decoding speeds up • Choosing a draft model: size, shared tokenizer, cost and accuracy • Demo: acceptance rate on structured vs creative tasks in vLLM • When it's worth it: spare VRAM, small batches, structured output, short context SPEAKER Sheilah Kirui, Developer Advocate, Akamai LINKS Akamai developers on GitHub: https://github.com/akamai-developers vLLM docs: https://docs.vllm.ai Akamai: https://www.akamai.com Akamai on X: https://x.com/Akamai CHAPTERS 0:00 Intro 0:37 Prefill and decode 1:37 The idea behind speculative decoding 2:42 The cost: two models in memory 3:26 Choosing a draft model 3:56 Demo setup on one Blackwell GPU 5:41 Which workloads benefit 6:21 Demo: structured output 9:00 Demo: 1.6x faster 9:50 Demo: creative writing 10:55 Long context limits the gains 11:50 Checklist 12:39 Q&A: tools and resources Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #LLMInference #vLLM #AIEngineer
More like this

Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex

Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg

The 6 Pillars of an Agentic Harness for Production — Varun Krovvidi, Resolve AI

Move Fast and Don't Break Things: Scaling Databases for the AI Era — PlanetScale
Join the discussion
Sign in to join the discussion
Sign in