A hands-on comparison of three local vision-capable models — Gemma 4 E4B, Qwen 3.5 9B, and Ministral 3 3B — tested on identical tasks: a Docker error screenshot, a busy UI with a counting task, and a low-quality photo of a medicine bottle. Qwen 3.5 9B came out clearly ahead, correctly reading fine print, counting models, identifying active ingredients, and even suggesting UI interactions. Gemma 4 E4B handled simpler tasks but got lossy on busy screens. Ministral 3 3B performed well on the debugging screenshot but hallucinated on more complex visual tasks. The takeaway: vision capability varies significantly among local models, and only Qwen 3.5 9B consistently understood what it was looking at.

5m read timeFrom xda-developers.com
Post cover image
Table of contents
Gemma 4 E4B and its small brainQwen 3.5 9B and the eyes to matchMinistral 3 3B knows what it's looking at, mostly
70 Impressions