Simon Willison
Read post

Kimi K3, and what we can still learn from the pelican benchmark

Moonshot AI released Kimi K3, a 2.8 trillion parameter model claiming to be the first open 3T-class model. It's priced at $3/million input and $15/million output tokens — the most expensive model from a Chinese AI lab to date — and leads Arena.ai's Frontend Code arena. Simon Willison tested it using his long-running 'pelican riding a bicycle' SVG benchmark, revealing that K3 only has one reasoning effort level (max), uses 13,241 reasoning tokens for a simple prompt, and appears to have an ~85-token hidden system prompt. The post also reflects on the benchmark's diminishing correlation to overall model quality after 21 months, while arguing it still provides useful signal: confirming API access, estimating cost, and testing basic SVG/spatial reasoning capabilities.

    #ai#kimi-k3#llm
Jul 16•6m read time•From simonwillison.net
Post cover image
20 Impressions
Simon Willison's image
Simon Willison

Simon Willison's blog offers a mix of technical tutorials, data analysis projects, and reflections o...

215 Followers

•

1.4K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard