Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
A single token of KV cache on Mistral 7B costs 131 KB. Multiply that by 16,000 tokens of context and 80 concurrent users and the cache alone wants 42 GB of GPU memory, which is why requests start failing on a 24 GB card. Harshul Jain, a senior software engineer at Audible, and Tanmay Sah, an independent AI researcher, spend this workshop building that number up from first principles. They open on three symptoms every team hits: memory that climbs with context length, time to first token that degrades as prompts grow, and throughput that collapses because a naive server answers requests one after another. The rest of the session explains what causes each one, working down through the inference pipeline into the attention layer, then into how a GPU actually splits its memory between fixed model weights and the cache that has to grow. From there the workshop splits optimization in two. Sah takes the model side, using two deliberately silly teaching devices, the ostrich algorithm for assumptions you wave through and the world cup algorithm for problems you cut into brackets, to move from quantization through multi head, multi query, grouped query, and latent attention, then flash attention and tiling. Jain takes the serving side: paged attention borrowed from operating system paging, continuous batching, prefix caching, and KV quantization, each benchmarked against a plain baseline. They close on engine selection, where their own testing found no statistical difference between vLLM and SGLang on standard workloads but a three to four times gap once agentic branching enters the picture. Slides and runnable notebooks are linked in the repo. Speaker info: - https://x.com/hj1393 - https://www.linkedin.com/in/hjain1393/ - https://harshuljain.substack.com/ - https://www.linkedin.com/in/tanmay-sah/ Timestamps: 0:00 - Introductions and what the workshop covers 2:43 - What LLM inference is, and why it costs so much 5:21 - The repo, the slides, and the free GPU notebooks 8:23 - Three pain points, memory, first token, throughput 12:18 - Foundations, the pipeline and the attention layer 14:53 - KV cache math and how GPU memory divides 20:41 - Prefill and decode, and why decode is memory bound 28:41 - Throughput, concurrency, and the trade off triangle 34:02 - The capacity calculator and picking a GPU 37:10 - Model optimization, quantization and attention variants 48:04 - Flash attention, the scorecard, and a demo 1:00:14 - Serving optimizations, paging, batching, caching 1:06:12 - Benchmarking vLLM, and speculative decoding 1:17:27 - vLLM against SGLang, choosing an engine, what to study next
More like this

Hedra AI Review - (2026) I Gave Hedra One Business Question — It Created the Entire Video

Cherry Servers (2026) Earn Recurring Affiliate Commissions - Refer Once, Earn Monthly?

High-Performance, Scalable Cloud Hosting: Nexcess Managed Cloud | SourceForge Podcast, episode #139

An Ancient Guide to Cybersecurity: 8 Lessons from The Art of War
Join the discussion
Sign in to join the discussion
Sign in