---
title: "Kimi K3, and what we can still learn from the pelican benchmark"
url: https://daily.dev/posts/kimi-k3-and-what-we-can-still-learn-from-the-pelican-benchmark-avvfo2ysr
source_url: https://simonwillison.net/2026/Jul/16/kimi-k3
type: article
source: "Simon Willison"
published: 2026-07-16T20:20:42.758Z
updated: 2026-07-23T02:12:21.337Z
tags: ["ai", "kimi-k3", "llm"]
reading_time: 6
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Kimi K3, and what we can still learn from the pelican benchmark

**[Simon Willison](https://daily.dev/sources/simonwillison)** · 6 min read · 0 upvotes · 0 comments

## Summary

Moonshot AI released Kimi K3, a 2.8 trillion parameter model claiming to be the first open 3T-class model. It's priced at $3/million input and $15/million output tokens — the most expensive model from a Chinese AI lab to date — and leads Arena.ai's Frontend Code arena. Simon Willison tested it using his long-running 'pelican riding a bicycle' SVG benchmark, revealing that K3 only has one reasoning effort level (max), uses 13,241 reasoning tokens for a simple prompt, and appears to have an ~85-token hidden system prompt. The post also reflects on the benchmark's diminishing correlation to overall model quality after 21 months, while arguing it still provides useful signal: confirming API access, estimating cost, and testing basic SVG/spatial reasoning capabilities.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://simonwillison.net/2026/Jul/16/kimi-k3>

---

Tags: [#ai](https://daily.dev/tags/ai), [#kimi-k3](https://daily.dev/tags/kimi-k3), [#llm](https://daily.dev/tags/llm)

[View this post on daily.dev](https://daily.dev/posts/kimi-k3-and-what-we-can-still-learn-from-the-pelican-benchmark-avvfo2ysr)
