Google has released Gemma 4 12B, a new model that fits within 16GB of VRAM or unified memory, making it runnable on standard consumer laptops. Despite its smaller size, it benchmarks nearly on par with the Gemma 4 26B model and even surpasses it on DocVQA. A key differentiator is its unified, encoder-free architecture that supports native audio inputs by projecting raw audio signals directly into the LLM backbone — a first for a mid-sized Gemma model. Developer communities have responded positively, particularly around the native audio support and local inference potential. Some skepticism exists around coding performance compared to alternatives like Qwen. The release highlights a broader trend toward on-device AI, where local inference offers privacy and eliminates per-token cloud costs.