Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI
A 4 billion parameter model, trained with RL for under $500, beat its 235 billion parameter sibling on real financial questions. Charles Dickens, research scientist at Snorkel AI, shares how Snorkel and UC Berkeley's Sky Computing Lab trained a small model to beat a much bigger one. Working from SEC 10-K filings, they built FinQA, a financial question-answering dataset with expert-validated answers, and found even frontier models hallucinating table schemas, flooding their own context, and repeating failed strategies. Using the open-source rLLM framework and reinforcement learning with a simple pass/fail reward, they trained a Qwen3 4B model for under $500 that reached about 60% versus 51% for the 235B model. The skills transferred to harder multi-table questions without hurting general tool use. Surprisingly, simpler training data and simpler rewards worked best, because the real bottleneck was disciplined tool use, not reasoning depth. In this talk: • How the FinQA dataset was built and verified from 10-K filings • Failure modes: hallucinated schemas, context flooding, poor recovery • Training a 4B model with RL in rLLM for under $500 • Why simple data and binary rewards beat fancier setups SPEAKER Charles Dickens, Senior Applied Research Scientist, Snorkel AI LinkedIn: https://www.linkedin.com/in/charles-dickens/ GitHub: https://github.com/dickensc LINKS Snorkel blog: how a 4B model outsmarted a 235B giant: https://snorkel.ai/blog/how-tool-discipline-let-a-4b-model-outsmart-a-235b-giant-on-financial-tasks/ rLLM-FinQA-4B on Hugging Face: https://huggingface.co/rLLM/rLLM-FinQA-4B rLLM framework: https://github.com/rllm-org/rllm Snorkel Open Benchmarks Grants: https://benchmarks.snorkel.ai/ Snorkel AI: https://snorkel.ai CHAPTERS 0:00 Intro 0:32 Specialization can beat scale 1:07 About Snorkel AI 2:32 Working with UC Berkeley 3:51 A tale of two models 4:41 Specialists, not polymaths 5:06 Building FinQA from 10-K filings 6:01 Three layers of verification 7:16 Where models fail 8:16 Training with rLLM 9:16 The training environment 9:40 Training for under $500 10:20 Result: 4B beats 235B 10:55 Does it transfer to harder problems? 12:05 General tool use held up 12:35 Simple data won 13:20 Simple rewards won 14:00 A blueprint for enterprise agents 14:45 How to evaluate agents 15:45 Open Benchmarks Grants Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #ReinforcementLearning #FinanceAI #AIEngineer
More like this

Java 25 killed the boilerplate? void main() explained #shorts

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

Why Your Company Needs a Context Graph (and How to Build It) — Gil Feig, Merge

Generation Is Cheap, Review Is Expensive: How to Stop Shipping AI Slop — Gabriel Martinez, G2i
Join the discussion
Sign in to join the discussion
Sign in