Zalando's engineering team built an in-process client-side load balancer (CSLB) to handle over a million requests per second of internal fan-out traffic for their Product Read API, replacing shared Skipper ingress hops. The implementation replicates Skipper's xxHash64 consistent-hash ring for cache locality, uses a Kubernetes watch-based informer for pod discovery, and adds N-ring fade-in to prevent cold-cache spikes on scale-up. A key innovation is occupancy-based bounded load using Little's Law (seconds of work per second) rather than in-flight counts or throughput, combined with a latency multiplier borrowed from Finagle. Results include eliminating Skipper's fleet from 50+ pods to 8, reducing their own pod fleet by 25%, and saving over $1,000/day. AZ-aware routing was prototyped but paused due to edge cases around bounded-load threshold miscalculation during dual fade-in. The post also covers pipeline improvements, retry hardening, FIFO buffering, and how detailed logging revealed mysterious node-level network freezes that had previously been invisible.
Table of contents
Skipper and the Fan-Out ProblemBuilding the Same Hash RingKubernetes DiscoveryFixing the Pipeline FirstRolling It Out SafelyEliminating Scale-Up Spikes: N-Ring Fade-InTaming Pod Occupancy with Bounded LoadAZ-Aware RoutingHardening the Fan-Out PathLessonsWhat's NextShould You Build Your Own?Acknowledgements35.2K Impressions1 Comment