From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Feed a model one hour of video and roughly a million visual tokens go in, yet the loss lands on about two percent of them, since the only ground truth is a transcript or a few labeled frames. Armen Aghajanyan calls that a humongous waste, and predicting every pixel treats a background pixel with the same weight as a gripper tip or a contact point. Perceptron's answer is a perceptive objective that learns which percepts will matter, rather than hardcoding the gripper. The second problem is context bloat from cameras that never switch off. Patch averaging buys ten times compression, but his fix is data sparse mixture of experts, a router that decides per layer which tokens to read and which to skip. Left alone, the model zooms into the graph in a figure and spends more tokens on fruit when asked to segment fruit. Put together, that produced the model his team released a few weeks earlier, trained on a petabyte spanning text, images, video, and trajectories from desktop use to video games, which he says beats a frontier lab's embodied reasoning model at a fraction of the cost. Detection turns into an agentic task, with the model tiling the image, raising the contrast, and proposing boxes until it finds the bird. The biggest result is a new scaling law: training jointly on perception, reasoning, and control lets ten times more video pretraining substitute for ten times less teleop data, which costs about a hundred dollars an hour. He closes with one model emitting control tokens to sort books by reading their titles, an open release promised for July, and questions on temporal context, background robustness, and structured extraction. Speaker info: - https://x.com/ArmenAgha - https://www.linkedin.com/in/armenag - https://perceptron.inc Timestamps: 0:00 - Perceptron's north star: one model that perceives, reasons, and acts 1:23 - Early fusion, and the VLM to VLA to world model ladder 3:40 - Challenge one: an hour of video has almost no ground truth 5:41 - Challenge two: context bloat from cameras that never switch off 7:04 - Data sparse mixture of experts lets the model pick its tokens 8:54 - The first embodied foundation model and the petabyte behind it 9:50 - Detection as an agentic task: tile, zoom, raise the contrast 10:45 - Orchestrators, tactile policies, and self verifying annotation 12:35 - New scaling laws: trade teleop hours for video pretraining 14:11 - One model emitting control tokens, and open weights in July 16:17 - Q&A: temporal context, background robustness, structured extraction
More like this

5 Best Email Template Builders in 2026 (For Beginners - No Coding Needed)

AI Engineer Paris 2026 Main Stage: Google DeepMind, ElevenLabs, Hugging Face & Stripe | Day 2

Hugo AI Review (2026) - How to Create an AI Agent for Customer Support

From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam
Join the discussion
Sign in to join the discussion
Sign in