Google has released Gemma 4 12B, a new model that fits within 16GB of VRAM or unified memory, making it runnable on standard consumer laptops. Despite its smaller size, it benchmarks nearly on par with the Gemma 4 26B model and even surpasses it on DocVQA. A key differentiator is its unified, encoder-free architecture that supports native audio inputs by projecting raw audio signals directly into the LLM backbone — a first for a mid-sized Gemma model. Developer communities have responded positively, particularly around the native audio support and local inference potential. Some skepticism exists around coding performance compared to alternatives like Qwen. The release highlights a broader trend toward on-device AI, where local inference offers privacy and eliminates per-token cloud costs.

5m read timeFrom thenewstack.io
Post cover image
Table of contents
Almost as good as Gemma 4 26B, but much smallerThe star attraction: native audio inputsSo far, so goodIs the future local?
4 Impressions