Salesforce Engineering
Read post

Reducing Single-Region Risk at 4B Metrics Per Minute

Salesforce's Argus team shares how they re-architected their internal observability platform to eliminate single-region AWS dependency while handling 4 billion metrics per minute. The original centralized model meant a single regional outage could take down global metrics visibility. Over six months, three teams redesigned the system into a geo-local model that processes and stores metrics closer to their origin, reducing cross-region data transfer costs and blast radius. A federation query layer backed by Elasticsearch routes requests only to relevant geographies, handles partial results gracefully, and returns HTTP 206 responses when a region is unavailable. The new architecture is live in 4 of 5 production regions, with OpenTSDB and HBase as core storage layers and metadata caching keeping wildcard query latency low.

    #observability#distributed-systems#elk
Aug 05•7m read time•From engineering.salesforce.com
Post cover image
1.5K Impressions
Salesforce Engineering's image
Salesforce Engineering

The Salesforce Engineering Blog offers a deep dive into Salesforce technologies, providing technical...

189 Followers

•

925 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard