A developer conducted a personal evaluation of 11 LLMs using 130 real prompts from their bash history, covering programming, sysadmin, technical explanations, and general knowledge tasks. The evaluation found that open models like DeepSeek and Qwen3 often outperformed expensive closed models like Claude Sonnet and Gemini Pro in terms of accuracy, cost, and speed. Key findings include that reasoning models rarely help for simple tasks, Gemini Flash is exceptionally fast, and closed models are overpriced. The author now uses multiple models simultaneously via tmux scripts for different use cases.

11m read timeFrom darkcoding.net
Post cover image
Table of contents
What I learntPrompts and winnersOverall resultsMy decision: Use several at onceCaveatsBonus - the poem
952 Impressions