Skip to main content
bash TV

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

AI Engineer

2.7K views10 Oct 2026

YouTube

Your GPUs might be waiting on your data, not your model. This pipeline went from 15% to 90% GPU utilization. Tarun Sunkaraneni from Amazon AGI shows that the hardest part of fast multimodal training often isn't the GPU at all, but keeping it fed with data. Training a Qwen3-VL-style model on images stored in S3, the baseline pipeline spent about 85% of its time waiting on data, mostly loading, decoding and resizing images one at a time. He fixes the bottlenecks one by one: concurrency with asyncio and Ray actors, prefetching so data is ready before the trainer asks, and Ray's object store to stop copying big image tensors between processes. Then, at production scale, a new bottleneck appears on a single machine's network card, solved by spreading workers across nodes and using zero-copy reads, which gave another 50% throughput. In this talk: • Why multimodal training is often CPU- and IO-bound, not GPU-bound • Concurrency, prefetching and zero-copy transfer with Ray • Choosing between asyncio, threads and processes around the GIL • Why some settings only pay off at scale SPEAKER Tarun Sunkaraneni, Amazon AGI GitHub: https://github.com/TSunny007 Website: https://tsunny007.github.io/ LINKS Ray: https://www.ray.io CHAPTERS 0:00 Intro 0:45 Is training really GPU-bound? 1:10 Four bottlenecks 1:50 Why multimodal is CPU-heavy 2:45 The wait-time ratio 3:35 Our stack 4:10 The data 4:35 System layout 5:15 Baseline: 15% utilization 6:20 Fix 1: concurrency 7:05 Async, threads or processes? 8:10 Asyncio plus Ray actors 9:15 Fix 2: prefetching 10:40 Fix 3: zero-copy with Ray's object store 12:20 Recap: 85% to 20% wait time 12:55 Scaling up: a new bottleneck 13:25 The network card bottleneck 14:35 Spread scheduling and zero copy 15:35 Lessons Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #MLInfrastructure #GPUs #AIEngineer

Join the discussion

Sign in to join the discussion

Sign in