---
title: "Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations"
url: https://daily.dev/posts/describe-what-you-see-with-multimodal-large-language-models-to-enhance-video-recommendations-wxtipsf8b
source_url: https://research.atspotify.com/2025/9/describe-what-you-see-with-multimodal-large-language-models-to-enhance-video/
type: article
source: "Spotify Research"
published: 2025-09-22T15:14:13.698Z
updated: 2025-09-22T15:14:37.109Z
tags: ["machine-learning", "llm", "spotify", "multimodal", "recommendation-systems"]
reading_time: 6
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations

**[Spotify Research](https://daily.dev/sources/spotify_research)** · 6 min read · 0 upvotes · 0 comments

## Summary

Spotify Research introduces a framework that uses Multimodal Large Language Models (MLLMs) to generate rich text descriptions from video and audio content, significantly improving video recommendation systems. The approach converts raw video frames and audio into semantically dense descriptions that capture intent, humor, and world knowledge - elements traditional encoders miss. Testing on the MicroLens-100K dataset showed performance improvements of up to 60% when integrated with standard recommendation architectures like two-tower models and SASRec, with particularly strong gains for longer videos.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://research.atspotify.com/2025/9/describe-what-you-see-with-multimodal-large-language-models-to-enhance-video/>

## Similar posts on daily.dev

- [Multimodal LLMs Basics: How LLMs Process Text, Images, Audio & Videos](https://daily.dev/posts/multimodal-llms-basics-how-llms-process-text-images-audio-videos-fw0j9el9k) · ByteByteGo · 5 upvotes · 0 comments
- [Teaching Large Language Models to Speak Spotify: How Semantic IDs Enable Personalization](https://daily.dev/posts/teaching-large-language-models-to-speak-spotify-how-semantic-ids-enable-personalization-i4u6pcpnq) · Spotify Research · 1 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#spotify](https://daily.dev/tags/spotify), [#multimodal](https://daily.dev/tags/multimodal), [#recommendation-systems](https://daily.dev/tags/recommendation-systems)

[View this post on daily.dev](https://daily.dev/posts/describe-what-you-see-with-multimodal-large-language-models-to-enhance-video-recommendations-wxtipsf8b)
