Databricks has announced the Public Preview of AI Runtime (AIR), a serverless training stack that provides on-demand distributed GPU training on NVIDIA A10s and H100s. It eliminates infrastructure management overhead by pre-installing PyTorch, CUDA, and popular distributed training frameworks like Ray, Hugging Face Transformers, and Composer. AIR integrates natively with the Databricks Lakehouse, supports MLflow for GPU observability, and connects with Lakeflow for job orchestration and CI/CD via Declarative Automation Bundles. Current public preview supports up to 8x H100s in a single node, with multi-node support in private preview. Early adopters include Rivian, FactSet, and YipitData, using it for LLM fine-tuning, computer vision, and recommendation systems.

5m read timeFrom databricks.com
Post cover image
Table of contents
On-demand NVIDIA H100 and A10 GPUs in notebooksLakeflow for production-ready workloadsRuntime optimized for distributed deep learning
667 Impressions