How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI
Laurie Voss reran a year old benchmark and the models walked straight through its ceiling without noticing it was there. IFScale asks a model to write a business report containing a list of exact words, then counts how many actually appear. A year ago frontier models started dropping instructions somewhere around 200 to 300, which is a hard limit on how much you can put in a skills file. Voss, head of developer relations at Arize AI and a cofounder of npm, first replicated that result on the three models from the original paper still reachable by API, then pointed the same test at the current frontier. They scored 100 percent immediately. He had to raise the benchmark from 500 words to 10,000 before anything broke at all. The boundary now sits near 2,000 instructions, and for the best model closer to 5,000, roughly a tenfold gain in twelve months. The stranger finding is that the models no longer fail the same way. One simply forgets, shedding nearly half its rules by 2,000. One refuses at the API level, because a pile of random words trips its safety classifier. One spends its whole thinking budget verifying instructions and leaves no room to answer. The last writes five thousand words of the report, decides the request is stupid, says so politely, and stops, which reads like a finished answer unless you get to the end. Voss argues this moves the work: fitting rules into a small budget was a compression problem and it is gone, while knowing whether the model obeyed is a verification problem that only checking outputs solves. The whole study cost 29 dollars. Speaker info: - https://x.com/seldo - https://www.linkedin.com/in/seldo/ - https://seldo.com Timestamps: 0:00 - The 200 instruction ceiling and where it came from 1:49 - Not knowing whether it followed your rules 2:58 - How the IFScale benchmark works 4:19 - Replicating last year's result 6:12 - Pointing the same test at current models 7:07 - The headline, a tenfold jump in a year 9:01 - Four models, four different ways to fail 14:27 - What this changes in your workflow 16:29 - Caveats, and what the test does not prove 19:32 - Capacity went up, reliability did not 21:06 - Compression solved, verification remains
More like this

Hyper3D Review (2026) - Is Rodin Gen 2.5 Good Enough for Real 3D Work?

The Universal Remote Control for AI — Alex Hancock, Block

MCP Apps: Give the Model Data, Give the User a UI — Dustin Mihalik, Indeed

Your agents lack context: Here's how to fix "You're absolutely right!" — Brandon Waselnuk, Unblocked
Join the discussion
Sign in to join the discussion
Sign in