Grab Tech Blog
Read post

Enabling near real-time data analytics on the data lake

This post discusses the challenges of handling frequent updates in a data lake and introduces the Hudi format as a solution. It explains the configurations optimized for high and low throughput sources and how to connect to Kafka and RDS data sources. It also highlights the importance of indexing for Hudi tables and the impact of the Hudi Data Ingestion solution on business metrics and fraud detection.

    #data-science#data-analysis#kafka#spark#apache-flink#data-lake
Feb 23, 2024•7m read time•From engineering.grab.com
Post cover image
Table of contents
IntroductionHigh throughput sourceLow throughput sourceConnecting to our Kafka (unbounded) data sourceConnecting to our RDS (bounded) data sourceIndexing for Hudi tablesImpactWhat’s next?References
10 Impressions
Grab Tech Blog's image
Grab Tech Blog

Grab is a leading technology company in Southeast Asia, offering a wide range of services, including...

51 Followers

•

225 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard