DeepL shares how they adopted ClickHouse as their central data warehouse starting in 2020, driven by a need for privacy-conscious analytics. They began with a single-node MVP using Kafka, a custom sink, and Metabase for visualization, then scaled to a 3-shard × 3-replica cluster ingesting ~500 million rows per day. Key investments included schema automation via protobuf definitions, an A/B testing experimentation framework leveraging ClickHouse's statistical computation capabilities, and an ML infrastructure for website personalization using user history stored in ClickHouse.

1m read timeFrom clickhouse.com
Post cover image
1.9K Impressions