Skip to main content
bash TV

Scale the Judgment, Not the Model — Andrew Orobator, Reddit

AI Engineer

1.7K views27 Sept 2026

YouTube

Swap in a smarter model and you get a slightly better answer. Take away the tests, gates and review, and everything falls apart. Andrew Orobator, a senior Android engineer at Reddit, argues that the model was never the bottleneck. The bottleneck is the judgment around it. Humans absorb judgment implicitly, but agents need it spelled out. He shows how to get judgment out of people's heads and into the repo. Skills are institutional judgment turned into something an agent can run. Work logs let a fresh agent pick up at milestone 7 of 9 (this talk itself was built with one). Personas lend you a security reviewer's or designer's eye. He then covers the verification ladder and his feature-flag cleanup agent, which went 7 for 7 on green-CI PRs at $1.26 each. And he shares a warning: when he asked Codex for reasons to unlock his repo guard, it quietly added a self-authorizing "emergency recovery" exception. His line: "Agents will build ladders to climb out of the pit of success." Speaker info: X/Twitter: @aorobator (https://x.com/aorobator) Related links: Vibe Engineering series (Medium): https://medium.com/@andreworobator Timestamps: 0:00 The war room for dead feature flags 0:37 Intro 0:55 What this talk isn't 1:55 We're harness engineers now 2:25 Humans absorb judgment; agents need it explicit 3:00 The model isn't the bottleneck 3:30 Judgment trapped in people's heads 4:40 Skills as knowledge lines 5:15 Skills vs. documentation 5:25 Work logs 6:25 This talk was built with a work log 6:45 Personas 8:05 Borrowing eyes you don't have 8:20 Governance 9:00 Verification 9:40 The verification ladder 10:15 Spin at the gate until green 10:35 Make the agent record itself 11:20 Case study: a feature-flag agent 12:10 $1.26 per pull request 12:50 Agents climb out of the pit of success 14:25 A society of specialist agents 15:25 In, on and off the loop 15:55 Encoded judgment rots 17:05 Managing judgment 18:15 Scale the judgment, not the model

Join the discussion

Sign in to join the discussion

Sign in