Skip to main content
bash TV

How AI Models Scale Beyond a Single GPU Across LLM Workloads

IBM Technology

24.6K views6 Oct 2026

YouTube

Learn more about AI Models here → https://ibm.biz/~GnXROtDog The biggest AI models can't fit on one GPU. Grace Ableidinger explains how distributed AI inference scales LLMs across GPUs using data, pipeline, tensor, and expert parallelism. Learn how KV cache, throughput, and prefill/decode bottlenecks shape production AI serving. 00:00 – See Why AI Models Need Distributed Inference 01:50 – Scale LLM Traffic with Data Parallelism 02:45 – Split AI Models with Pipeline Parallelism 03:45 – Divide LLM Layers with Tensor Parallelism 04:32 – Scale Mixture-of-Experts Across GPUs 05:44 – Separate LLM Prefill and Decode Workloads 07:51 – Combine GPU Parallelism for Production AI 08:29 – Orchestrate Distributed LLM Inference AI news moves fast. Sign up for a monthly newsletter for AI updates from IBM → https://ibm.biz/~86jhlfVag AI was used in the creation of the transcript and metadata for this video. #llm #aimodels #gpu #machinelearning --------------------------------------------------------------------------------------------------------- Find us on YouTube: 🔵 IBM Technology: https://www.youtube.com/@IBMTechnology 🔵 IBM: https://www.youtube.com/@IBM 🔵 IBM Developer: https://youtube.com/@IBMDeveloperAdvocates

Join the discussion

Sign in to join the discussion

Sign in