A detailed production retrospective on building a Java/Kafka-based cloud contact center platform handling 80k BHCC across 10k agents. The author documents five major failure modes discovered in production: async latency violating real-time contracts, cache mismatch across service instances, Kafka partition limits blocking horizontal scaling, cross-cluster deduplication adding 200ms+ latency, and blocking REST calls inside consumer threads causing 30+ minutes of consumer lag. Each problem is traced through its root cause and resolved solution, including a three-generation state management evolution from Kafka Global State Stores to local in-memory caches to Redis shared cache with a background recovery thread. JVM-specific concerns are addressed including Spring Boot startup overhead, GC pressure under high throughput, and the potential of JDK 21 virtual threads. The article concludes with five architectural decisions the author would make differently from day one.

20m read timeFrom infoq.com
Post cover image
Table of contents
The Fundamental Tension: Async by Default, Real-Time by RequirementThe Cache Mismatch Problem: Distributed State in DisguiseThe Three-Generation State Management EvolutionPartition Limits: The Hidden Ceiling on Horizontal ScalingCross-Cluster Event Propagation: The Deduplication TaxJVM-Specific Considerations: What Java Adds to These TradeoffsKafka Streams and the RocksDB Performance TrapThe Cascading Failure: When Downstream Latency Freezes Upstream ConsumersWhat I Would Design DifferentlyConclusionAbout the Author
22K Impressions1 Comment