---
title: "Profile-aware LLM-as-a-Judge for Podcasts: A Better Middle Ground Between Offline Metrics and A/B Tests"
url: https://daily.dev/posts/profile-aware-llm-as-a-judge-for-podcasts-a-better-middle-ground-between-offline-metrics-and-a-b-te-8karfizw2
source_url: https://research.atspotify.com/2025/9/profile-aware-llm-as-a-judge-for-podcasts-a-better-middle-ground-between/
type: article
source: "Spotify Research"
published: 2025-09-19T12:19:52.821Z
updated: 2025-09-19T12:20:13.510Z
tags: ["machine-learning", "llm", "spotify", "recommendation-systems"]
reading_time: 6
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Profile-aware LLM-as-a-Judge for Podcasts: A Better Middle Ground Between Offline Metrics and A/B Tests

**[Spotify Research](https://daily.dev/sources/spotify_research)** · 6 min read · 0 upvotes · 0 comments

## Summary

Spotify Research introduces a profile-aware LLM-as-a-Judge approach for evaluating podcast recommendations that bridges the gap between fast offline metrics and expensive A/B tests. The method creates human-readable user profiles from 90 days of listening history, then uses LLMs to score candidate episodes against these profiles. In a 47-user study, the approach achieved 75% alignment with human judgments and successfully differentiated between production recommendation models, offering a scalable middle ground for recommendation system evaluation.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://research.atspotify.com/2025/9/profile-aware-llm-as-a-judge-for-podcasts-a-better-middle-ground-between/>

## Similar posts on daily.dev

- [From Models to Products: LLMs for Recommendation at Spotify Scale](https://daily.dev/posts/from-models-to-products-llms-for-recommendation-at-spotify-scale-rnpxxwmdh) · Spotify Research · 0 upvotes · 0 comments
- [When Can LLMs Replace Humans in A/B Tests?](https://daily.dev/posts/when-can-llms-replace-humans-in-a-b-tests--joccfpevz) · Spotify Labs · 6 upvotes · 0 comments
- [Who watches the watchers? LLM on LLM evaluations](https://daily.dev/posts/who-watches-the-watchers-llm-on-llm-evaluations-qxvsetxuu) · Stack Overflow Blog · 0 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#spotify](https://daily.dev/tags/spotify), [#recommendation-systems](https://daily.dev/tags/recommendation-systems)

[View this post on daily.dev](https://daily.dev/posts/profile-aware-llm-as-a-judge-for-podcasts-a-better-middle-ground-between-offline-metrics-and-a-b-te-8karfizw2)
