Uber Engineering
Read post

From Archival to Access: Config-Driven Data Pipelines

Uber's Compliance Data Store team developed a config-driven archival and retrieval framework to manage terabytes of regulatory data efficiently. The system automatically moves data from hot storage (HDFS) to cold storage (S3) based on configurable policies, while providing on-demand retrieval capabilities. The solution reduced manual intervention by 90% and successfully handles over 500 regulatory reports, addressing challenges like schema evolution, data consistency during backfills, and resource optimization. The framework uses MySQL for metadata management, Apache Airflow for orchestration, and includes a user-friendly interface for self-service data retrieval.

    #backend#compliance#data-engineering#apache-spark#apache-hadoop
Jun 05, 2025•13m read time•From uber.com
Post cover image
Table of contents
ChallengesArchitectureUI/DesignUse Cases at UberNext Steps
153 Impressions
Uber Engineering's image
Uber Engineering

The Uber Engineering Blog offers insights, technical deep dives, and updates on the engineering chal...

178 Followers

•

268 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard