Per-token pricing breaks down for long-horizon agentic AI workloads because a single user task can trigger dozens of inference calls with growing context windows, turning a predictable cost into a wide distribution. The post proposes replacing per-token budgeting with a task-envelope framework built around four metrics: median task cost, P95 task cost, cost per successful outcome, and headroom ratio. It also introduces Kimi K2.6, Moonshot AI's 1-trillion-parameter MoE model (32B active params, 256K context, 58.6% SWE-Bench Pro), now available on DigitalOcean Serverless Inference via an OpenAI-compatible endpoint — positioned as a cost-efficient frontier model for agentic coding workloads with bursty, unpredictable call patterns.

23m read timeFrom digitalocean.com
Post cover image
Table of contents
The invoice that ended the per-token eraKey TakeawaysMeet Kimi K2.6The old mental model (and why it held up)What changes with long-horizon agentic workflowsWhy per-token budgeting quietly breaksThe unit shift: from price-per-token to price-per-completed-taskA practical budgeting framework (4 numbers every CTO should track)Why Kimi K2.6 is the model worth budgeting forWhy serverless inference fits this spend patternK2.6 on DigitalOcean Serverless Inference: the concrete exampleClosing thoughts: budget the envelope, ship the agentFrequently Asked QuestionsFurther reading
95 Impressions