From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam
Well under one percent of the Common Crawl corpus that frontier models train on is in an Indian language, even as frontier labs call India a fast growing market. Krishna Prasad Srinivasan says the knowledge exists but was never digitized, and Sarvam's answer is a three billion parameter vision language model, small enough for one GPU, that he says beats document AI models a hundred times larger. The model is unusual twice over. Its backbone is a state space model rather than a transformer, because a page can run to ten thousand visual tokens and quadratic attention gets expensive, while an SSM keeps constant memory. And in late 2025, when most OCR models were monolithic page level VLMs, the team bet on block level OCR wrapped in a layout harness and a reading order harness, which many 2026 releases have since converged on. Training runs as a four stage curriculum: thirteen trillion text tokens across English, 22 Indian languages, math, and code before the model sees a pixel, so a language prior can resolve a smudged word; continual pretraining on three hundred million image text pairs; supervised fine tuning on a hundred million OCR samples; then reinforcement learning. OCR correctness is machine checkable, so rewards are unit tests on character error rate, table structure, or grammar, which makes verifiable reward RL scalable here. The moat is underneath: a data engine for languages with no labeled data, and evals built for real usefulness. Four months after launch the model is digitizing 35 million pages for insurers, banks, and governments, with a public Indic benchmark from the 1800s to today promised soon. Questions cover low resource transfer, synthetic data, and sovereignty. Speaker info: - https://x.com/fewshotlearner - https://www.linkedin.com/in/krishnapsrinivasan/ Timestamps: 0:00 - A 3B model that beats models 100 times larger on document AI 1:14 - Sarvam, and why India is missing from the machine readable world 2:37 - Why Indic document intelligence is hard 3:18 - The contrarian bet: block level OCR with a harness 4:31 - Why a state space model instead of a transformer 6:08 - Four stage curriculum: 13 trillion text tokens first 8:05 - The moat underneath: data engine and evals 9:13 - RL with verifiable rewards for OCR 9:54 - Results on English benchmarks and 22 Indian languages 10:50 - The agentic workbench and 35 million pages in production 12:39 - Q&A: low resource transfer, synthetic data, the Indic benchmark, sovereignty
More like this

5 Best Email Template Builders in 2026 (For Beginners - No Coding Needed)

AI Engineer Paris 2026 Main Stage: Google DeepMind, ElevenLabs, Hugging Face & Stripe | Day 2

Hugo AI Review (2026) - How to Create an AI Agent for Customer Support

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Join the discussion
Sign in to join the discussion
Sign in