I tested Gemma 4, Qwen 3.5, and Ministral 3 for vision tasks, and only one understood the assignment
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
A hands-on comparison of three local vision-capable models — Gemma 4 E4B, Qwen 3.5 9B, and Ministral 3 3B — tested on identical tasks: a Docker error screenshot, a busy UI with a counting task, and a low-quality photo of a medicine bottle. Qwen 3.5 9B came out clearly ahead, correctly reading fine print, counting models, identifying active ingredients, and even suggesting UI interactions. Gemma 4 E4B handled simpler tasks but got lossy on busy screens. Ministral 3 3B performed well on the debugging screenshot but hallucinated on more complex visual tasks. The takeaway: vision capability varies significantly among local models, and only Qwen 3.5 9B consistently understood what it was looking at.