On June 24, 2026, Spotify's video transcoding infrastructure hit maximum capacity, causing multi-hour delays in podcast episode publication. Five converging factors caused the outage: insufficient headroom for traffic spikes, a concurrent batch reprocessing job, recent quality improvements that increased per-item processing cost, a prioritization system malfunction that treated batch jobs equally to new submissions, and a resource scheduling bug underutilizing compute by ~10%. The team stopped the batch job, deployed a scheduling fix, and brought additional capacity online overnight. All queues cleared by 01:02 UTC on June 25. Follow-up actions include a 67% capacity increase, fixed resource scheduling, earlier monitoring alerts, improved prioritization, and better backpressure mechanisms throughout the pipeline.

5m read timeFrom engineering.atspotify.com
Post cover image
Table of contents
What happened?What caused it?Timeline (UTC)Where do we go from here?
1.1K Impressions