Spotify Research
Read post

Teaching Large Language Models to Speak Spotify: How Semantic IDs Enable Personalization

Spotify developed a method to adapt large language models for personalized content recommendations by introducing Semantic IDs—compact tokens that encode relationships between catalog items and user behaviors. The approach involves building catalog-native representations from textual and behavioral signals, aligning these with an open-weight LLM's vocabulary, and fine-tuning on personalization tasks. The domain-adapted 1B-parameter model achieved up to 1.96× improvement over baselines in episode recommendations, with multi-task training providing an additional 22% boost. The system enables explainable recommendations while maintaining real-time performance through efficient serving infrastructure using vLLM and Redis-backed key-value stores.

    #machine-learning#llm#spotify#recommendation-systems
Nov 25, 2025•15m read time•From research.atspotify.com
Post cover image
Table of contents
A Spotify catalog-native vocabulary for LLMsDomain Specific Training: Learning personalization tasksEvaluating LLMs that speak SpotifyScalingServingConclusionAcknowledgments
231 Impressions
Spotify Research's image
Spotify Research

Spotify_Research's publication is a hub for academic research and industry insights in the field of ...

18 Followers

•

5 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard