NVIDIA releases Nemotron 3.5 Content Safety, a 4B-parameter multimodal safety model built on Google Gemma 3 4B IT. Key additions over Nemotron 3 include unified multimodal evaluation (user prompt + image + assistant response in one pass), custom enterprise policy enforcement at inference time, auditable reasoning traces via THINK mode, and coverage across 12 explicitly trained languages plus ~140 via zero-shot transfer. The model supports three output modes ranging from low-latency binary verdicts to full reasoning traces, uses the Aegis 2.0 taxonomy (13 core + 10 subcategories), and achieves ~85% average accuracy across multilingual and multimodal benchmarks with 3x lower latency than comparable multimodal safety models. NVIDIA is also releasing the training dataset, a rarity in the multimodal safety space. The model is available on Hugging Face under NVIDIA Open Model License and as a production NIM microservice.

11m read timeFrom huggingface.co
Post cover image
Table of contents
What's New in Nemotron 3.5 Content SafetyModel ArchitectureReasoningTraining DataBenchmarkingLatencyAddressing the Benchmark GapGetting Started
51 Impressions