Spotify Research
Read post

Personalizing Agentic AI to Users' Musical Tastes with Scalable Preference Optimization

Spotify Research presents a hybrid approach for personalizing AI-powered music recommendations using LLM-based agentic systems. The method combines reward models with Direct Preference Optimization (DPO) to create a continuous learning flywheel that adapts to user preferences from listening behavior. The system interprets natural language queries, orchestrates music search tools, and learns from user interactions like plays, skips, and saves. Production A/B tests showed 4% increase in listening time, higher playlist saves, and 70% reduction in erroneous tool calls while maintaining quality standards.

    #machine-learning#llm#spotify#reinforcement-learning#recommendation-systems
Sep 23, 2025•9m read time•From research.atspotify.com
Post cover image
Table of contents
Limitations of traditional approachesA hybrid approach: Reward Models + Direct Preference OptimizationThe Preference Tuning FlywheelWhy reward models matterStable, scalable fine-tuningOnline experimentsEngineering practices that made the differenceLooking aheadAcknowledgments
335 Impressions
Spotify Research's image
Spotify Research

Spotify_Research's publication is a hub for academic research and industry insights in the field of ...

18 Followers

•

5 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard