Skip to main content
bash TV

From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss

AI Engineer

2.1K views5 Oct 2026

YouTube

A chatbot pilot on 40 hand-picked files works great. Point it at 80,000 SharePoint documents and it falls apart. Jeff Koss and Leo Platzer, co-founder and former CTO of Deasy Labs (acquired by Collibra), walk through taking an unstructured-data AI project from stalled POC to production. Using a manufacturer building a legal-ops chatbot over 80,000 SharePoint files, they demo AI-assisted taxonomies and metadata tagging with evidence and confidence scores, sensitive-data detection, and quality checks for duplicates, conflicts and freshness. Data slices refresh on a schedule. Leo then shows why this matters: when 30% of a corpus is stale or duplicated, up to 80% of an agent's retrieved context can be useless, and cleaning it roughly doubled recall in multi-hop RAG evals. He finishes by generating a context file for coding agents with the SDK. In this talk: • Why document data quality means duplicates, conflicts and freshness • AI-assisted taxonomies and metadata tagging with evidence and confidence • Finding and filtering sensitive data before it reaches a chatbot • How stale, duplicated data crowds out good context, and what cleaning it does to recall SPEAKERS Leo Platzer, Co-founder & former CTO, Deasy Labs (Collibra) LinkedIn: https://www.linkedin.com/in/leonardplatzer/ Jeff Koss, Staff Customer Engineer, Deasy Labs (Collibra) LINKS Collibra unstructured data for AI: https://www.collibra.com/use-cases/solution/unstructured-data-for-ai Collibra acquires Deasy Labs: https://www.collibra.com/company/newsroom/press-releases/collibra-acquires-deasy-labs Deasy Labs: https://www.deasylabs.com Collibra: https://www.collibra.com/ CHAPTERS 0:00 Intro 0:18 Deasy Labs and Collibra 1:01 From stalled POC to production 1:11 Scenario: a legal-ops chatbot 2:11 Who needs what 3:25 40 files works 3:55 80,000 files doesn't 4:25 Data quality for documents 5:05 The symptoms 6:35 Demo: tagging with a taxonomy 9:04 Metadata with evidence 10:04 Finding sensitive data 10:54 Duplicates, conflicts, freshness 11:54 Data slices and suggested taxonomies 13:33 Keeping data fresh 14:13 Results: four months to days 15:00 Why duplicates poison retrieval 16:54 Stale data can fill 80% of context 17:39 Results: about 2x recall 18:29 Demo: a context file for coding agents 20:36 Wrap-up Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #RAG #DataQuality #AIEngineer

Join the discussion

Sign in to join the discussion

Sign in