Spotify engineers argue that LLM evals and A/B experiments should form a funnel, not competing alternatives. LLM evals (automated judges assessing relevance, coherence, tone) belong before experiments to filter out weak candidates and raise the hit rate of what gets tested. Experiments then validate whether real users respond as predicted and catch regressions in secondary metrics that evals miss. A key insight is that evals need continuous calibration against online outcomes — without this, eval scores are opinions, not evidence. The post describes a feedback loop where running evals on A/B test data helps calibrate judges over time, making both evals and experiments progressively smarter.
Table of contents
What evals give us, and what they don’tTwo calibration layers, one feedback loopClose the loop407 Impressions