Skip to main content

Which model should I actually use?

The labs, the benchmarks and their catches, and what models cost. Ends with how to choose.

Check yourself.

One question per step. Take it cold to find where to start, or after reading to see what stuck. Nobody's grading you.

  1. As of mid-2026, who leads the open-weight frontier?

  2. How should you actually use a frontier model comparison table?

  3. You've picked an open-weight model to build a product on. What should you check first?

  4. A model tops SWE-bench and LMArena. What do those scores tell you?

  5. Two models list similar per-token prices. Why can real costs still differ sharply?

  6. Starting a new AI feature, which model choice does the handbook recommend by default?