Skip to main content
bash TV

Robot Demos Are Easy. Reliability Is Hard — Jason Ma, Dyna Robotics

AI Engineer

1.6K views24 Sept 2026

YouTube

A generalist robot that succeeds 80 to 90% of the time makes a great video but a poor product. Jason Ma, co-founder and CTO of Dyna Robotics, shows how Dyna got its model, Dyna-1, to a 99.4% success rate folding restaurant napkins for 24 hours straight, including recovering after it pulled over the whole stack. The key is a reward model that watches the robot and scores its progress. When the score dips, the robot is making a mistake, so the team can collect targeted recovery data and fine-tune again in a human-in-the-loop active-learning cycle. Ma also covers Dyna's pre-training data pyramid of more than 200,000 hours, and its architecture that pairs a reasoning model with a world action model. He shows deployments in restaurants and a Sacramento laundromat, and a robot folding T-shirts for three days straight at CoRL in Korea with no site-specific data. Speaker info: X/Twitter: @JasonMa2020 (https://x.com/JasonMa2020) Website: https://jasonma2016.github.io/ Related links: Dyna Robotics: https://www.dyna.co Timestamps: 0:00 Intro 0:32 About Dyna Robotics 1:12 The research and deployment flywheel 2:16 How robot foundation models work 3:11 The pre-training data pyramid 4:21 Reasoning model plus world action model 5:01 New tasks with under an hour of data 6:10 Why 80–90% success isn't enough 7:40 Case study: restaurant napkin folding 9:00 Dyna-1: 99.4% success over 24 hours 9:59 What makes napkin folding hard 11:04 Why standard post-training stalls at 80% 11:39 Reward models that score robot progress 13:04 Scalable supervision and active learning 14:08 Error recovery highlights 15:43 Real deployments: restaurants and a laundromat 16:28 Working at new sites with no new data 17:13 Three days of T-shirt folding at CoRL 18:13 Opening Red Bull cans at live events 18:58 Summary 19:52 Q&A

Join the discussion

Sign in to join the discussion

Sign in