Databricks has announced the General Availability of the Variant data type, which allows teams to ingest semi-structured data (JSON, XML, CSV) without sacrificing query performance. Variant Shredding, also GA, uses Predictive Optimization and machine learning to automatically identify frequently queried fields and store them as columns in Parquet files, enabling up to 30x faster reads compared to storing JSON as strings and 4x faster than unshredded Variant. Over 5,000 teams are already using Variant, executing 500M+ queries per month across 160+ TB of data. The feature integrates with Auto Loader and Lakeflow Pipelines, and supports both Delta and Iceberg table formats.
Table of contents
Flexible ingestion at scaleFaster, smarter queries with Predictive OptimizationUsing Variant in DatabricksGet started with Variant today69 Impressions