How MiniMax M3 Was Built: Sparse Attention and Native Multimodality — Olive Song
Most models learn text first and get vision bolted on later. MiniMax M3 learned both from the very first training step. Olive, who works on reinforcement learning research at MiniMax, goes inside MiniMax M3, an open-weight model with frontier-level coding and agentic ability, a 1 million token context window and native multimodality. She explains MSA, MiniMax's sparse attention design (about 9x faster prefill and 15x faster decoding than full attention): an index branch that picks which blocks matter and a sparse branch that attends only to those, and why existing designs like DeepSeek's DSA didn't fit grouped-query attention or GPU memory access. Then she covers training with text and vision from step zero, why adding vision later gave unstable results, what the attention maps show, and why 3D attention inside the vision transformer improves visual understanding. In this talk: • What makes MiniMax M3 different: agentic coding, 1M context, native multimodal • How MSA sparse attention works, and why DSA didn't fit • Block-level retrieval and kernels built for real GPUs • Why training on text and vision from step zero beats adding vision later SPEAKER Olive Song, RL Lead, MiniMax X: https://x.com/olive_jy_song LINKS MiniMax: https://www.minimax.io MiniMax M3 (GitHub): https://github.com/MiniMax-AI/MiniMax-M3 MiniMax M3 (Hugging Face): https://huggingface.co/MiniMaxAI/MiniMax-M3 MiniMax M3 tech report: https://arxiv.org/abs/2606.13392 MSA (GitHub): https://github.com/MiniMax-AI/MSA CHAPTERS 0:00 Intro 1:03 Meet MiniMax M3 1:37 1M token context 2:57 Sparse attention: index and sparse branches 4:17 Why not DeepSeek's DSA? 5:07 Problems with existing sparse attention 5:37 GPU memory access 7:01 MSA: what we changed 7:46 Block-level retrieval 8:21 Efficient kernels 9:01 Multimodal from step zero 9:46 Training strategies compared 11:45 What the attention maps show 13:10 3D attention in the ViT Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #MiniMax #MultimodalAI #AIEngineer




Join the discussion
Sign in to join the discussion
Sign in