Small files are a silent killer in Lakehouse architectures. Streaming ingestion, CDC pipelines, concurrent writers, and poor partitioning strategies all generate thousands of tiny Parquet files. Each file adds metadata overhead, slows query planning, and inflates cloud costs. The solution is compaction — a process that merges small files into larger, more efficient ones. The post introduces the problem and references compaction strategies in Delta Lake, Apache Hudi, and Apache Iceberg, with the full deep-dive available via a paid Substack subscription.

4m read timeFrom luminousmen.com
Post cover image
20 Impressions