From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model
Your GPUs might be waiting on your data, not your model. This pipeline went from 15% to 90% GPU utilization. Tarun Sunkaraneni from Amazon AGI shows that the hardest part of fast multimodal training often isn't the GPU at all, but keeping it fed with data. Training a Qwen3-VL-style model on images stored in S3, the baseline pipeline spent about 85% of its time waiting on data, mostly loading, decoding and resizing images one at a time. He fixes the bottlenecks one by one: concurrency with asyncio and Ray actors, prefetching so data is ready before the trainer asks, and Ray's object store to stop copying big image tensors between processes. Then, at production scale, a new bottleneck appears on a single machine's network card, solved by spreading workers across nodes and using zero-copy reads, which gave another 50% throughput. In this talk: • Why multimodal training is often CPU- and IO-bound, not GPU-bound • Concurrency, prefetching and zero-copy transfer with Ray • Choosing between asyncio, threads and processes around the GIL • Why some settings only pay off at scale SPEAKER Tarun Sunkaraneni, Amazon AGI GitHub: https://github.com/TSunny007 Website: https://tsunny007.github.io/ LINKS Ray: https://www.ray.io CHAPTERS 0:00 Intro 0:45 Is training really GPU-bound? 1:10 Four bottlenecks 1:50 Why multimodal is CPU-heavy 2:45 The wait-time ratio 3:35 Our stack 4:10 The data 4:35 System layout 5:15 Baseline: 15% utilization 6:20 Fix 1: concurrency 7:05 Async, threads or processes? 8:10 Asyncio plus Ray actors 9:15 Fix 2: prefetching 10:40 Fix 3: zero-copy with Ray's object store 12:20 Recap: 85% to 20% wait time 12:55 Scaling up: a new bottleneck 13:25 The network card bottleneck 14:35 Spread scheduling and zero copy 15:35 Lessons Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #MLInfrastructure #GPUs #AIEngineer
More like this

Java 25 killed the boilerplate? void main() explained #shorts

Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI

Why Your Company Needs a Context Graph (and How to Build It) — Gil Feig, Merge

Generation Is Cheap, Review Is Expensive: How to Stop Shipping AI Slop — Gabriel Martinez, G2i
Join the discussion
Sign in to join the discussion
Sign in