Zalando
Read post

Client-Side Load Balancing at a Million Requests Per Second

Zalando's engineering team built an in-process client-side load balancer (CSLB) to handle over a million requests per second of internal fan-out traffic for their Product Read API, replacing shared Skipper ingress hops. The implementation replicates Skipper's xxHash64 consistent-hash ring for cache locality, uses a Kubernetes watch-based informer for pod discovery, and adds N-ring fade-in to prevent cold-cache spikes on scale-up. A key innovation is occupancy-based bounded load using Little's Law (seconds of work per second) rather than in-flight counts or throughput, combined with a latency multiplier borrowed from Finagle. Results include eliminating Skipper's fleet from 50+ pods to 8, reducing their own pod fleet by 25%, and saving over $1,000/day. AZ-aware routing was prototyped but paused due to edge cases around bounded-load threshold miscalculation during dual fade-in. The post also covers pipeline improvements, retry hardening, FIFO buffering, and how detailed logging revealed mysterious node-level network freezes that had previously been invisible.

    #kubernetes#java#distributed-systems
Jun 23•31m read time•From engineering.zalando.com
Post cover image
Table of contents
Skipper and the Fan-Out ProblemBuilding the Same Hash RingKubernetes DiscoveryFixing the Pipeline FirstRolling It Out SafelyEliminating Scale-Up Spikes: N-Ring Fade-InTaming Pod Occupancy with Bounded LoadAZ-Aware RoutingHardening the Fan-Out PathLessonsWhat's NextShould You Build Your Own?Acknowledgements
35.2K Impressions1 Comment
Zalando's image
Zalando

Zalando Tech Blog is a platform where Zalando's engineering team shares insights, experiences, and b...

56 Followers

•

377 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard