Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI
The web has roughly 30 trillion tokens of usable text. Frontier pre-training needs more. Here's how DatologyAI generates the rest. Bogdan Gaza, co-founder and CTO of DatologyAI, shares the engineering lessons from running synthetic data jobs at trillion-token scale, including a recent run of about 12 trillion tokens across web, math and code. He covers BeyondWeb, DatologyAI's rephrasing-based synthetic data recipe, and the move from a split Slurm/Kubernetes setup to a single Ray, KubeRay and vLLM pipeline on EKS on HyperPod. He then walks through four bottlenecks: S3 metadata, GPU failures, cross-cluster scheduling and inference tuning. In this talk: • Why seeded rephrasing beats asking a model for synthetic data from scratch • Batching S3 metadata fetches to cut 9–11 days down to about 2 hours • Right-sized partitions plus checkpointing for recoverable GPU failures • Scheduling CPU and GPU together across clusters, and vLLM flag sweeps for about 40% more throughput SPEAKER Bogdan Gaza, Co-founder & CTO, DatologyAI LinkedIn: https://www.linkedin.com/in/bogdangaza/ X: https://x.com/hurrycane LINKS DatologyAI: https://www.datologyai.com BeyondWeb (blog): https://www.datologyai.com/blog/beyondweb BeyondWeb (paper): https://arxiv.org/pdf/2508.10975 CHAPTERS 0:00 Intro 0:12 Synthetic data at trillion-token scale 0:27 Why synthetic data: the data wall 1:47 What DatologyAI does 2:22 BeyondWeb results 4:11 How rephrasing works 5:51 Where we started: Slurm vs Kubernetes 6:31 The stack: Ray, KubeRay, vLLM 7:41 Today's setup on HyperPod 8:40 One workflow: curate, synthesize, train, eval 9:30 Engineering lessons at scale 9:40 Bottleneck 1: S3 metadata (11 days → 2 hours) 12:29 Bottleneck 2: GPU instability and checkpointing 14:59 Bottleneck 3: Cross-infrastructure orchestration 17:53 Bottleneck 4: Inference tuning (+40% throughput) 18:53 Recap 19:13 From 30B to 12T synthetic tokens 20:02 We're hiring Recorded at the AI Engineer World's Fair 2026 in San Francisco. Subscribe for more talks from the engineers building with AI. AI Engineer: https://ai.engineer YouTube: https://www.youtube.com/@aiDotEngineer X: https://x.com/aiDotEngineer LinkedIn: https://www.linkedin.com/company/aidotengineer/ #SyntheticData #MLInfra #AIEngineer




Join the discussion
Sign in to join the discussion
Sign in