Spotify Research
Read post

Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations

Spotify Research introduces a framework that uses Multimodal Large Language Models (MLLMs) to generate rich text descriptions from video and audio content, significantly improving video recommendation systems. The approach converts raw video frames and audio into semantically dense descriptions that capture intent, humor, and world knowledge - elements traditional encoders miss. Testing on the MicroLens-100K dataset showed performance improvements of up to 60% when integrated with standard recommendation architectures like two-tower models and SASRec, with particularly strong gains for longer videos.

    #machine-learning#llm#spotify#multimodal#recommendation-systems
Sep 22, 2025•6m read time•From research.atspotify.com
Post cover image
Table of contents
The FrameworkEmpirical EvaluationTakeawaysReferences
281 Impressions
Spotify Research's image
Spotify Research

Spotify_Research's publication is a hub for academic research and industry insights in the field of ...

18 Followers

•

5 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard