Skip to main content
bash TV

The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI

AI Engineer

6.8K views30 Sept 2026

YouTube

Developers using autonomous agents wrote 741% more code but shipped only 30% more software. The bottleneck is human review, and "just review harder" doesn't scale. Reviewer effectiveness collapses past about 400 lines, and agents now open 10,000-line PRs. Laurie Voss (Head of Developer Relations at Arize AI, co-founder of npm) reviews what the industry is actually doing about it. The evidence covers OpenAI's zero-human-code product, METR's finding that about half of SWE-bench-passing PRs wouldn't be merged, and Cognition's FrontierCode (88% on SWE-bench Pro vs. 29% on real mergeability). She explains why a mergeability benchmark would immediately become a training signal for frontier models. She also covers how Cursor and GitHub run review in production, why multi-pass review and default suspicion cut false positives, and what Carlini's agent-built C compiler and Bun's million-line Zig-to-Rust port (13,044 unsafe blocks) reveal about taking humans out of the loop. Automated reviewers can be fooled by prompt injection that humans catch, which leaves production as the last reviewer standing. Speaker info: X: https://x.com/seldo Bluesky: https://bsky.app/profile/seldo.com LinkedIn: https://www.linkedin.com/in/seldo/ GitHub: https://github.com/seldo COMPANY Arize AI: https://arize.com X: https://x.com/arizeai LinkedIn: https://www.linkedin.com/company/arizeai Timestamps: 0:00 The new bottleneck: human review 1:37 741% more code, 30% more software 2:37 Generation is no longer the bottleneck 3:22 Why "review harder" fails: the Cisco study 4:52 Stop reading code? Loops and OpenAI's zero-human-code product 6:42 Do passing tests mean mergeable? METR's SWE-bench study 8:01 FrontierCode: 88% vs. 29% 9:11 Mergeability as the next training signal 11:06 Automated review today: GitHub Copilot and Cursor 13:16 Review is fusing with repair 13:51 CodeRabbit, Greptile, Graphite and the acceptance metric 15:00 Can you skip the human? Carlini's C compiler 16:00 Bun's Zig-to-Rust port and 13,044 unsafe blocks 17:30 How OpenAI rebuilt review as a system 18:45 Dex Horthy: "Please read the code" 19:30 Context the tests can't see 19:55 Where the human checkpoint survives 20:40 Who reviews the reviewers? Prompt injection 22:04 Production: the last reviewer standing 22:44 Code review is being rebuilt, not killed 23:29 What to do today: build a review harness

Join the discussion

Sign in to join the discussion

Sign in