Skip to main content
bash TV

Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax

AI Engineer

3.2K views4 Sept 2026

YouTube

An intern designed the sparse-attention architecture behind MiniMax M3. That detail comes after Olive Song explains the larger problem the team was trying to solve: short context windows aren’t enough when agents must work across long conversations, tool responses and complex environments. M3 combines a functional one-million-token context window with coding, agentic and multimodal capabilities, allowing it to understand text, images and video within the same model. Thomas Wolf and Olive Song unpack how MiniMax made that context window efficient, why the company trained M3 as multimodal from its very first step and what happens when that training goes wrong. They also discuss MiniMax’s open research culture, products used by hundreds of millions of people, the role of community feedback in improving open-source models and the internal agent harnesses accelerating the company’s own research. By the end, M3 is no longer just the model being discussed—it is already helping the team build M3.1. About Thomas Wolf Thomas Wolf is the co-founder and Chief Science Officer of Hugging Face, the open platform and community for building, sharing and collaborating on machine learning models, datasets and applications. About Olive Song Olive Song is the RL Lead at MiniMax, where she works on the research and training behind frontier open-source models including MiniMax M3. Recorded at the AI Engineer World’s Fair in San Francisco. Stay up to date on our upcoming events and content by joining our newsletter here: https://www.ai.engineer/newsletter Timestamps: 0:00 - Introduction 0:39 - The race between the leading open-source models 2:34 - MiniMax M3: coding, vision and one million tokens 3:46 - Why agents need longer context windows 6:22 - From GPT-2’s 1,024 tokens to million-token models 7:30 - The intern who designed MiniMax’s sparse attention 8:29 - How research projects work inside MiniMax 10:01 - Training a multimodal model from the first step 13:09 - Going beyond one trillion parameters 13:25 - Building AI products for 300 million users 15:20 - Why MiniMax plans to keep open-sourcing its models 17:29 - Agents that understand presentations and long videos 18:27 - Automating AI research with agent harnesses 19:27 - M3 is already helping build M3.1 19:55 - Why multi-agent systems come next

Join the discussion

Sign in to join the discussion

Sign in