Discord experienced a major voice outage affecting 17% of global sessions after engineers attempted to vertically scale the session service. A safety check ran longer than the Kubernetes termination grace period, causing pods to be marked dead and triggering millions of simultaneous reconnection attempts — a classic thundering herd. The surge exhausted memory in the US East region, cascaded globally, and overwhelmed the Erlang-based voice connection service, dropping 14 of 15 voice instances. Targeted and full cluster restarts both failed under the connection flood. Recovery was achieved by applying aggressive rate limits and doubling cluster size, with the outage lasting approximately 3 hours. Key takeaways: rate limiting and back-pressure are critical safeguards in distributed systems.