A hands-on comparison of three sub-2B local LLMs — Google's Gemma 4 E2B, Alibaba's Qwen 3.5 0.8B, and Liquid AI's LFM2.5-1.2B — tested on a structured writing task and a real-time weather query via Brave Search MCP. Qwen 3.5 0.8B hallucinated confidently and badly, Gemma 4 E2B was inconsistent with units, and LFM2.5-1.2B won the real-time tool-calling test with a clean, accurate response despite failing the creative task. The takeaway: tiny models aren't toys, they're specialists, and matching them to their intended use case matters more than raw benchmark scores.

5m read timeFrom xda-developers.com
Post cover image
Table of contents
Gemma 4 E2BQwen 3.5 0.8BLFM2.5-1.2B-Instruct
40 Impressions