NVIDIA has released Cosmos 3, an open omni-model for physical AI built on a Mixture-of-Transformers (MoT) architecture. Unlike previous Cosmos releases that required separate models for different tasks, Cosmos 3 unifies world generation, physical reasoning, and action generation in a single model. It supports text, image, video, audio, and action modalities and can function as a video generator, VLM, forward/inverse dynamics model, or robot policy. Two sizes are available: Cosmos 3 Nano (8B parameters, workstation-grade) and Cosmos 3 Super (32B parameters, for large-scale SDG and research). Both are available on Hugging Face with Diffusers integration, post-training scripts, and open synthetic datasets for robotics, autonomous vehicles, and warehouse scenarios.

10m read timeFrom huggingface.co
Post cover image
Table of contents
SECTION 1: What's new with Cosmos 3?SECTION 2: Cosmos 3 CapabilitiesSECTION 3: Using Cosmos 3 with DiffusersSECTION 4: Datasets for physical AISection 5: Cosmos FrameworkSECTION 6: ResourcesAcknowledgments
24 Impressions