<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h" -->

---
title: Enabling near real-time data analytics on the data lake
description: This post discusses the challenges of handling frequent updates in a data lake and introduces the Hudi format as a solution. It explains the configurations...
canonical: https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Enabling near real-time data analytics on the data lake | daily.dev
og:description: This post discusses the challenges of handling frequent updates in a data lake and introduces the Hudi format as a solution. It explains the configurations...
og:url: https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h
og:image: https://api.daily.dev/og/posts/ijM6Wwv1H.png
og:image:alt: Enabling near real-time data analytics on the data lake
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Enabling near real-time data analytics on the data lake

**[Grab Tech Blog](https://daily.dev/sources/grab)** · 7 min read · 0 upvotes · 0 comments

## Summary

This post discusses the challenges of handling frequent updates in a data lake and introduces the Hudi format as a solution. It explains the configurations optimized for high and low throughput sources and how to connect to Kafka and RDS data sources. It also highlights the importance of indexing for Hudi tables and the impact of the Hudi Data Ingestion solution on business metrics and fraud detection.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://engineering.grab.com/enabling-near-realtime-data-analytics>

---

Tags: [#data-science](https://daily.dev/tags/data-science), [#data-analysis](https://daily.dev/tags/data-analysis), [#kafka](https://daily.dev/tags/kafka), [#spark](https://daily.dev/tags/spark), [#apache-flink](https://daily.dev/tags/apache-flink), [#data-lake](https://daily.dev/tags/data-lake)

[View this post on daily.dev](https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Enabling near real-time data analytics on the data lake","url":"https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h"},"datePublished":"2024-02-23T07:08:51.939Z","dateModified":"2024-05-09T09:06:08.895Z","description":"This post discusses the challenges of handling frequent updates in a data lake and introduces the Hudi format as a solution. It explains the configurations...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/85f7efb48c041420cf6a23448d5385b4?_a=AQAEufR","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/85f7efb48c041420cf6a23448d5385b4?_a=AQAEufR","isAccessibleForFree":true,"articleSection":"Grab Tech Blog","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Grab Tech Blog","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/c8b8850354ee4494a05777b62e42cc62","url":"https://daily.dev/sources/grab"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/enabling-near-real-time-data-analytics-on-the-data-lake-ijm6wwv1h","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"data-science,data-analysis,kafka,spark,apache-flink,data-lake","timeRequired":"PT7M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Grab Tech Blog","item":"https://daily.dev/sources/grab"},{"@type":"ListItem","position":3,"name":"Enabling near real-time data analytics on the data lake"}]}
```

