Skip to main content
bash TV

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

AI Engineer

899 views23 Sept 2026

YouTube

Feed a model one hour of video and roughly a million visual tokens go in, yet the loss lands on about two percent of them, since the only ground truth is a transcript or a few labeled frames. Armen Aghajanyan calls that a humongous waste, and predicting every pixel treats a background pixel with the same weight as a gripper tip or a contact point. Perceptron's answer is a perceptive objective that learns which percepts will matter, rather than hardcoding the gripper. The second problem is context bloat from cameras that never switch off. Patch averaging buys ten times compression, but his fix is data sparse mixture of experts, a router that decides per layer which tokens to read and which to skip. Left alone, the model zooms into the graph in a figure and spends more tokens on fruit when asked to segment fruit. Put together, that produced the model his team released a few weeks earlier, trained on a petabyte spanning text, images, video, and trajectories from desktop use to video games, which he says beats a frontier lab's embodied reasoning model at a fraction of the cost. Detection turns into an agentic task, with the model tiling the image, raising the contrast, and proposing boxes until it finds the bird. The biggest result is a new scaling law: training jointly on perception, reasoning, and control lets ten times more video pretraining substitute for ten times less teleop data, which costs about a hundred dollars an hour. He closes with one model emitting control tokens to sort books by reading their titles, an open release promised for July, and questions on temporal context, background robustness, and structured extraction. Speaker info: - https://x.com/ArmenAgha - https://www.linkedin.com/in/armenag - https://perceptron.inc Timestamps: 0:00 - Perceptron's north star: one model that perceives, reasons, and acts 1:23 - Early fusion, and the VLM to VLA to world model ladder 3:40 - Challenge one: an hour of video has almost no ground truth 5:41 - Challenge two: context bloat from cameras that never switch off 7:04 - Data sparse mixture of experts lets the model pick its tokens 8:54 - The first embodied foundation model and the petabyte behind it 9:50 - Detection as an agentic task: tile, zoom, raise the contrast 10:45 - Orchestrators, tactile policies, and self verifying annotation 12:35 - New scaling laws: trade teleop hours for video pretraining 14:11 - One model emitting control tokens, and open weights in July 16:17 - Q&A: temporal context, background robustness, structured extraction

Join the discussion

Sign in to join the discussion

Sign in