Skip to main content
bash TV

Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI

AI Engineer

5.3K views2 Oct 2026

YouTube

Token budgets are out of control, and limiting tokens cripples your developers. Cast AI took a different route. Laurent Gil, co-founder and president of Cast AI, and Žilvinas Urbonas, who leads engineering for Kimchi, show how Cast AI built its own open-source coding harness after its coding-agent bill took off. The key idea is to measure cost per task instead of cost per token, and to let the harness choose the right proprietary or open model for each task based on the outcome. They share three months of internal results with 300 employees, then demo Ferment for multi-hour autonomous tasks, Teleport for agent sessions that keep running after you close your laptop, and Studio for team collaboration. In this talk: • Why cost per task matters more than cost per token • How an outcome-aware harness switches models automatically as new ones ship • Ferment: milestone-based, self-scoring autonomous coding runs that deploy to staging • Teleport and Studio: remote sandboxes and a shared Kanban board for agent sessions SPEAKERS Laurent Gil, Co-founder & President, Cast AI LinkedIn: https://www.linkedin.com/in/laurentgil/ X: https://x.com/laurentgil Žilvinas Urbonas, Engineering lead, Kimchi (Cast AI) LinkedIn: https://www.linkedin.com/in/zilvinasurbonas/ LINKS Kimchi: https://kimchi.dev/ Kimchi on X: https://x.com/getkimchi Kimchi (GitHub): https://github.com/getkimchi/kimchi Cast AI: https://cast.ai CHAPTERS 0:00 Intro 0:13 Meet Kimchi 0:52 Token costs out of control 1:31 Why rationing tokens is the wrong answer 2:26 Make tokens unlimited and cheap 3:05 Cost per token vs cost per task 4:30 A harness that picks the model 4:49 Results: 2.5x savings in 3 months 5:54 Watching the harness switch models 7:23 How the harness works 8:15 Ferment: multi-hour autonomous tasks 9:05 Build, verify, deploy to staging 11:03 Reading diffs isn't enough anymore 11:48 Kimchi Teleport 14:02 Why 62% of engineers code this way 14:23 Kimchi Studio for teams 17:04 Wrap-up Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #CodingAgents #OpenSourceAI #AIEngineer

Join the discussion

Sign in to join the discussion

Sign in