Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA
Ask your own company what it actually contracted for with a given vendor and nobody can tell you. The number exists, negotiated carefully, sitting in a pricing table inside a signed PDF no system ever read back. Hiral Shah puts a figure on the aggregate: roughly two trillion dollars of negotiated value locked in agreements that organizations never return to, because recovering it means human reading, human review, and manual work across disconnected systems. The scale on Docusign's side is its own engineering problem. Close to two million paying customers, a billion users, and about a million agreements flowing through each day, all of it needing to come out structured, queryable, and useful. Agreements also refuse to be flat. One governs another, which amends a third, so answering a simple question often means traversing twenty years of a business. The specific thing that breaks is tables. Pricing tiers, rate cards, SKUs and service levels are exactly the terms people need, and they are exactly what generic extraction handles worst, because reading a page line by line destroys a merged cell or a nested column. Shah and Sean Sodha describe the model they built together for that one job, an unusually small vision language model at roughly 900 million parameters, designed as an extractor rather than a generator. It replaces a stack of separate layout and table models with a single pass that returns reading order, semantic structure and preserved tables. Their headline lesson is about restraint: a purpose built small model ran table extraction around twenty times faster than the general alternatives they tested, and lower context meant lower latency and lower cost at their volume. Speaker info: - https://www.linkedin.com/in/shahhiral/ - https://www.linkedin.com/in/sean-sodha/ Timestamps: 0:00 - Why agreement data is an engineering problem 1:42 - A million agreements a day 2:19 - Two trillion dollars nobody goes back for 3:59 - Why tables break generic extraction 5:12 - Building retrieval models in the open 7:03 - An extractor, not a generator 8:32 - Demo: an order form becomes structured data 10:30 - Structuring agreements across an organization 11:18 - Purpose built models beat general ones 13:12 - What comes next 13:44 - Q&A: is OCR going away?




Join the discussion
Sign in to join the discussion
Sign in