GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe
Across thousands of GPUs, hardware failures are guaranteed. Fixing them by hand at 3 a.m. doesn't scale. Connor Guerrero, Young Jeong and Nikhil Gupta from Crusoe explain why they built Managed Slurm on top of Kubernetes. Slurm gives researchers gang scheduling, topology awareness and familiar sbatch workflows, but falls short on dynamic resources, node health and observability. Kubernetes fills those gaps without either team changing how it works. They walk through AutoClusters' fully automatic remediation of an XID 79 GPU failure, then demo killing a GPU mid-training and resuming from checkpoint in under 15 minutes with no human action. In this talk: • Where traditional Slurm excels for training and where it falls short • Why one stack beats separate Slurm and Kubernetes infrastructure • The automatic remediation flow: notify, drain, replace, requeue, resume • Sharing GPUs between training and inference as demand shifts SPEAKERS Connor Guerrero, Sr. Developer Relations Manager, Crusoe LinkedIn: https://www.linkedin.com/in/connor-guerrero/ GitHub: https://github.com/c-jg Young Jeong, Staff Solutions Engineer, Crusoe LinkedIn: https://www.linkedin.com/in/youngsjeong/ Nikhil Gupta, Senior Software Engineer, Managed Orchestration, Crusoe LINKS Crusoe: https://www.crusoe.ai/ Managed Slurm docs: https://docs.crusoecloud.com/orchestration/slurm/overview/index.html AutoClusters docs: https://docs.crusoecloud.com/orchestration/cmk/autoclusters Self-healing PyTorch training (blog): https://www.crusoe.ai/resources/blog/self-healing-distributed-pytorch-training-with-slurm-on-crusoe-managed-kubernetes Slinky (Slurm on Kubernetes): https://github.com/SlinkyProject/slurm-operator CHAPTERS 0:00 Intro 0:12 GPU failures are inevitable 0:56 What Crusoe Cloud is 1:56 Why this architecture 2:13 Where Slurm excels 3:31 Where Slurm falls short 4:21 Node health checks 5:04 Observability gaps 5:51 Why Kubernetes 6:45 Two stacks, double the burden 7:35 Managed Slurm on Kubernetes 8:09 What researchers and platform teams see 8:58 Sharing GPUs between training and inference 9:43 AutoClusters: automatic remediation 10:03 What happens on an XID 79 error 11:27 Demo: a GPU fails mid-training 12:56 Back to training in under 15 minutes 13:16 One-click Slurm 14:01 Recap 15:55 Design for failure Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #MLInfra #Kubernetes #AIEngineer




Join the discussion
Sign in to join the discussion
Sign in