<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8" -->

---
title: How Netflix Taught an LLM to Recommend Movies So That...
description: Netflix built GenRec, a large language model-based recommendation system, to replace parts of its complex, feature-engineered ranking stack. GenRec is trained...
canonical: https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How Netflix Taught an LLM to Recommend Movies So That You Keep Watching | daily.dev
og:description: Netflix built GenRec, a large language model-based recommendation system, to replace parts of its complex, feature-engineered ranking stack. GenRec is trained...
og:url: https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8
og:image: https://api.daily.dev/og/posts/7nSw95Tq8.png
og:image:alt: How Netflix Taught an LLM to Recommend Movies So That You Keep Watching
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How Netflix Taught an LLM to Recommend Movies So That You Keep Watching

**[ByteByteGo](https://daily.dev/sources/bytebytego)** · 17 min read · 45 upvotes · 8 comments

## Summary

Netflix built GenRec, a large language model-based recommendation system, to replace parts of its complex, feature-engineered ranking stack. GenRec is trained in two phases: first adapting an open-source foundation LLM with Netflix-specific data, then post-training it to rank catalog items using reward-weighted objectives. User history is converted into text (verbalization) and carefully trimmed via context engineering to control token costs, which were cut to roughly a third of the original size with a proportional serving cost reduction. At inference, GenRec uses prefill-only scoring (no token-by-token generation) with a ranking head that scores catalog item embeddings, served via vLLM with distilled models and prefix caching. A 10%-traffic, 4-week A/B test showed statistically significant but small improvements (0.115% short-term, 0.006% long-term engagement metrics).

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blog.bytebytego.com/p/how-uber-built-a-genie-to-answer>

## Questions this post answers

### How does Netflix use an LLM to generate movie and TV show recommendations with GenRec?

GenRec adapts an open-source foundation LLM in two phases: first learning Netflix content and user behavior, then post-training to rank items using reward signals. User watch history is converted to text (verbalization), processed through the LLM, and a scoring head combines the pooled hidden state with learned item embeddings to produce ranking scores via softmax, restricting recommendations to the Netflix catalog.

_Explore how teams apply LLMs to ranking and recommendation systems through daily.dev._

### What is prefill-only inference and why does Netflix use it for GenRec recommendations?

Prefill-only inference means the model processes the input prompt once and derives scores directly from the ranking head, skipping autoregressive token-by-token generation entirely. Netflix uses this to avoid the cost of writing out recommendation text, serving GenRec via vLLM with distilled models and prefix caching, since requests sharing a prefix can reuse prior computation and cut serving cost.

_Developers optimizing LLM inference costs can track techniques like this on daily.dev._

### How much did Netflix reduce token usage and serving cost through context engineering for GenRec?

Netflix reduced tokens describing user context to roughly one-third of the original amount using cleaning, compression, and wording changes, which produced a similar roughly one-third reduction in serving cost. This context engineering involved weighing strong versus weak behavioral evidence, compressing repetitive actions, and adding richer metadata only for cold-start items.

_Engineers tuning LLM context budgets for cost and quality can follow similar case studies on daily.dev._

## Community discussion

Top comments from developers on daily.dev.

**@digeomel** · 1 upvotes

> As interesting as it may be to read from a technical point of view, I wonder, why does Netflix care if we keep watching? I mean, you pay a standard monthly subscription anyway, it's not like they will charge you more if you watch more.

**@kartiknvj** · 1 upvotes

> The number that jumps out is the A/B result: statistically significant but 0.115% short-term, basically a rounding error for a full ranking-stack rewrite. It makes me wonder how they decided that cleared the bar over the feature-engineered baseline. When you are scoring prefill-only with a ranking head, what did they lean on to catch quality regressions between the offline reward objective and that live test?

**@rappayne** · 0 upvotes

> Very practical use of the technology. Everybody wins (unless it becomes addictive). Clever and innovative, too.

**@ristotoldsep** · 0 upvotes

> 0.115% sounds tiny until Netflix-scale money enters the chat.

## Similar posts on daily.dev

- [GenRec: Towards LLM-Native Recommendation at Netflix](https://daily.dev/posts/genrec-towards-llm-native-recommendation-at-netflix-wygnl3h1u) · Netflix TechBlog · 2 upvotes · 0 comments
- [Generative AI for Recommendations: what YouTube, Netflix and Meta are Moving to, Built From Scratch](https://daily.dev/posts/generative-ai-for-recommendations-what-youtube-netflix-and-meta-are-moving-to-built-from-scratch-jpbxplzky) · Medium · 1 upvotes · 0 comments
- [How LinkedIn Feed Uses LLMs to Serve 1.3 Billion Users](https://daily.dev/posts/how-linkedin-feed-uses-llms-to-serve-1-3-billion-users-wymxiezwf) · ByteByteGo · 20 upvotes · 1 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#netflix](https://daily.dev/tags/netflix), [#vllm](https://daily.dev/tags/vllm), [#recommendation-systems](https://daily.dev/tags/recommendation-systems)

[View this post on daily.dev](https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How Netflix Taught an LLM to Recommend Movies So That You Keep Watching","url":"https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8"},"datePublished":"2026-10-07T15:34:25.595Z","dateModified":"2026-10-07T16:05:15.793Z","description":"Netflix built GenRec, a large language model-based recommendation system, to replace parts of its complex, feature-engineered ranking stack. GenRec is trained...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f986c4fdc4143d5f4bccbb6f6cc31fbc?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f986c4fdc4143d5f4bccbb6f6cc31fbc?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"ByteByteGo","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"ByteByteGo","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/35be29234ee14d01a9cd049c52e12753","url":"https://daily.dev/sources/bytebytego"},"commentCount":8,"discussionUrl":"https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":45},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":8}],"keywords":"machine-learning,llm,netflix,vllm,recommendation-systems","timeRequired":"PT17M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"ByteByteGo","item":"https://daily.dev/sources/bytebytego"},{"@type":"ListItem","position":3,"name":"How Netflix Taught an LLM to Recommend Movies So That You Keep Watching"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8","comment":[{"@type":"Comment","text":"As interesting as it may be to read from a technical point of view, I wonder, why does Netflix care if we keep watching? I mean, you pay a standard monthly subscription anyway, it’s not like they will charge you more if you watch more.","datePublished":"2026-10-08T17:00:11.872Z","url":"https://daily.dev/posts/7nSw95Tq8#c-mc1V9RXnJ","author":{"@type":"Person","name":"Dimitris","url":"https://daily.dev/digeomel","image":"https://avatars3.githubusercontent.com/u/3295889?v=4"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"The number that jumps out is the A/B result: statistically significant but 0.115% short-term, basically a rounding error for a full ranking-stack rewrite. It makes me wonder how they decided that cleared the bar over the feature-engineered baseline. When you are scoring prefill-only with a ranking head, what did they lean on to catch quality regressions between the offline reward objective and that live test?","datePublished":"2026-10-08T18:13:21.642Z","url":"https://daily.dev/posts/7nSw95Tq8#c-zaMSKJG6d","author":{"@type":"Person","name":"kartik-nvjk","url":"https://daily.dev/kartiknvj","image":"https://media.daily.dev/image/upload/s--3gGgsVCw--/f_auto/v1781456774/avatars/avatar_TvTVeiMdkRCqWUDullFmy?_a=BAMAMiWQ0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"Very practical use of the technology. Everybody wins (unless it becomes addictive). Clever and innovative, too.","datePublished":"2026-10-08T15:03:28.716Z","url":"https://daily.dev/posts/7nSw95Tq8#c-LCJeVek47","author":{"@type":"Person","name":"Rap Payne","url":"https://daily.dev/rappayne","image":"https://avatars.githubusercontent.com/u/1122906?v=4"}},{"@type":"Comment","text":"0.115% sounds tiny until Netflix-scale money enters the chat.","datePublished":"2026-10-09T10:51:34.261Z","url":"https://daily.dev/posts/7nSw95Tq8#c-QUEsQcOSF","author":{"@type":"Person","name":"Risto Tõldsep","url":"https://daily.dev/ristotoldsep","image":"https://lh3.googleusercontent.com/a/ACg8ocLDWc6mZn0JwNmXXw6WY0L_HJ6pegzRttooC5VgtXESHSj3MNYx=s96-c"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-netflix-taught-an-llm-to-recommend-movies-so-that-you-keep-watching-7nsw95tq8#faq","mainEntity":[{"@type":"Question","name":"How does Netflix use an LLM to generate movie and TV show recommendations with GenRec?","acceptedAnswer":{"@type":"Answer","text":"GenRec adapts an open-source foundation LLM in two phases: first learning Netflix content and user behavior, then post-training to rank items using reward signals. User watch history is converted to text (verbalization), processed through the LLM, and a scoring head combines the pooled hidden state with learned item embeddings to produce ranking scores via softmax, restricting recommendations to the Netflix catalog. Explore how teams apply LLMs to ranking and recommendation systems through daily.dev."}},{"@type":"Question","name":"What is prefill-only inference and why does Netflix use it for GenRec recommendations?","acceptedAnswer":{"@type":"Answer","text":"Prefill-only inference means the model processes the input prompt once and derives scores directly from the ranking head, skipping autoregressive token-by-token generation entirely. Netflix uses this to avoid the cost of writing out recommendation text, serving GenRec via vLLM with distilled models and prefix caching, since requests sharing a prefix can reuse prior computation and cut serving cost. Developers optimizing LLM inference costs can track techniques like this on daily.dev."}},{"@type":"Question","name":"How much did Netflix reduce token usage and serving cost through context engineering for GenRec?","acceptedAnswer":{"@type":"Answer","text":"Netflix reduced tokens describing user context to roughly one-third of the original amount using cleaning, compression, and wording changes, which produced a similar roughly one-third reduction in serving cost. This context engineering involved weighing strong versus weak behavioral evidence, compressing repetitive actions, and adding richer metadata only for cold-start items. Engineers tuning LLM context budgets for cost and quality can follow similar case studies on daily.dev."}}]}
```

