Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify
One in four US Spotify Premium subscribers now use a single system every day, which Spotify calls the Large Taste Model. Yves Raimond, Spotify's SVP and GM of AI & Personalization, explains how Spotify turned its recommendation system "on its head" to be LLM-native. It went from curated playlists, to recommendations like Discover Weekly, to what Spotify calls generative personalization: a steerable DJ, prompted playlists, an editable taste profile and personal podcasts, across a catalog of more than 100 million tracks plus podcasts and audiobooks. Staff ML Engineer Jacqueline Wood then shows how the models are trained with NEO, Spotify's four-stage recipe. Semantic IDs are added to an open-weight LLM like Qwen. The new tokens are grounded while the backbone stays frozen, which keeps its language ability intact, where continued pre-training wiped it out. Then the model is instruction-tuned across many Spotify tasks, which even helped cold-start audiobook recommendations. She also covers decoding choices (98% of semantic IDs come out valid even without constrained decoding), and how grounding LLM judges with user profiles and behavior raised their agreement with humans by 91% on ambiguous queries. Speaker info: Jacqueline Wood, LinkedIn: https://www.linkedin.com/in/jacquelinewood Related links: Spotify: https://www.spotify.com Timestamps: 0:00 Intro: making LLMs speak Spotify 0:47 Spotify's scale: 760M users and 100M+ tracks 1:52 From curation to recommendations 2:37 Generative personalization 3:02 From guessing to reasoning, from black box to steerable 4:07 A Spotify DJ you can steer 4:37 Prompted playlists 5:32 Taste profile 6:22 Personal podcasts 6:52 The Large Taste Model 7:27 One in four US Premium subscribers use it daily 8:03 Jacqueline Wood: how the models are trained 8:18 Semantic IDs for Spotify's catalog 9:03 Natural-language podcast recommendations 9:33 NEO: a four-stage training recipe 10:02 Domain grounding with a frozen backbone 10:52 Capability induction: multitask tuning 11:32 Multitask gains and cold-start audiobooks 12:32 Why frozen grounding beats continued pre-training 14:02 Beam search vs. constrained decoding 15:02 One system in production 15:47 Breaking habitual listening in podcast discovery 16:11 Grounded LLM judges 18:16 Scaling evaluation sets with LLM judges 18:56 Summary




Join the discussion
Sign in to join the discussion
Sign in