<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx" -->

---
title: Open-weight models now handle most tokens on Vercel&#x27;s AI...
description: Open-weight models surpassed 56% of all tokens routed through Vercel&#x27;s AI Gateway in August, up from just 7% in December 2024, with some days reaching 78.4%....
canonical: https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Open-weight models now handle most tokens on Vercel&#x27;s AI Gateway, but Anthropic still captures 64% of spending | daily.dev
og:description: Open-weight models surpassed 56% of all tokens routed through Vercel&#x27;s AI Gateway in August, up from just 7% in December 2024, with some days reaching 78.4%....
og:url: https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx
og:image: https://api.daily.dev/og/posts/N0evrINNx.png
og:image:alt: Open-weight models now handle most tokens on Vercel&#x27;s AI Gateway, but Anthropic still captures 64% of spending
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Open-weight models now handle most tokens on Vercel's AI Gateway, but Anthropic still captures 64% of spending

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

Open-weight models surpassed 56% of all tokens routed through Vercel's AI Gateway in August, up from just 7% in December 2024, with some days reaching 78.4%. Yet they captured only 14% of dollar spend, while Anthropic took 64% of spending despite lower token volume, a share that hasn't dropped below 61% since December 2024. Within Anthropic, Opus 5 spend rose to 22.5% while Fable 5 fell to 4.9%. Z.ai's GLM-5.3-Flash tripled its predecessor's volume within five days, Google's gateway share dropped from 30% to 5% after Gemini 3 Flash lost volume, and Moonshot AI, DeepSeek, and Z.ai combined now exceed OpenAI's spend on the gateway.

## Content

Open-weight models crossed 56% of all tokens routed through Vercel's AI Gateway in August, the first time they've held a majority. Back in December 2024 they were at 7%. On some days the split has pushed even further, with open models briefly hitting 78.4% of token volume while closed models fell to 21.6%.

But token share and dollar share are telling very different stories.

Despite handling the majority of tokens, open-weight models captured only 14 cents of every dollar spent on the gateway. Anthropic, by contrast, took 64 cents of every dollar despite far lower token volume, a figure that hasn't dropped below 61% since December 2024. Cheap tokens at high volume still don't add up to expensive tokens at moderate volume.

Within Anthropic's own lineup, there's been some internal reshuffling. Fable 5 spend dropped from 13.2% to 4.9%, while Opus 5 rose to 22.5%. Most of that switching stayed inside Anthropic rather than moving to competitors.

Elsewhere, Z.ai's GLM-5.3-Flash tripled its predecessor's daily volume within five days of launch. Google had a rougher month: Gemini 3 Flash lost most of its volume to competitors, pulling Google's overall gateway token share down from 30% to 5%.

On the open-weight side, Moonshot AI and DeepSeek have climbed to third and fourth place by spend on some days. Combined with Z.ai, their spend now surpasses OpenAI's on the gateway. Worth noting: that's spend on inference across providers (mostly US-based), not revenue flowing directly to the open-weight labs themselves.

The pattern here is pretty clear. Open models are winning on volume, especially for high-throughput, cost-sensitive workloads. Anthropic is winning on revenue, which suggests its models are being used for tasks where output quality justifies the price. Whether that gap closes depends on whether open models can move upmarket, or whether Anthropic's pricing pressure eventually forces a shift.

## Questions this post answers

### What percentage of tokens on Vercel's AI Gateway are now handled by open-weight models versus closed models like Anthropic?

Open-weight models handled 56% of all tokens routed through Vercel's AI Gateway in August, up from just 7% in December 2024, with some days reaching as high as 78.4%. Despite this token majority, open-weight models captured only 14% of dollar spend, while Anthropic took 64% of spending on far lower token volume.

_Track shifting LLM cost and usage patterns on daily.dev before locking in a model provider._

### Why does Anthropic capture most AI Gateway spending despite open-weight models handling more tokens?

Anthropic's models are apparently used for tasks where output quality justifies a higher price, so even with much lower token volume than open-weight models, Anthropic captures 64 cents of every dollar spent on Vercel's AI Gateway. Open-weight models are winning cheap, high-throughput workloads, while Anthropic wins on higher-value revenue per token.

_Developers weighing quality versus cost across LLM providers can follow these tradeoffs on daily.dev._

### Which open-weight AI labs are gaining the most spend share on Vercel's AI Gateway?

Moonshot AI and DeepSeek have climbed to third and fourth place by spend on some days, and combined with Z.ai, their spend now surpasses OpenAI's on the gateway. Z.ai's GLM-5.3-Flash tripled its predecessor's daily volume within five days of launch, while this spend reflects inference cost across mostly US-based providers, not revenue to the labs themselves.

_Compare emerging open-weight model providers as options mature by following updates on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 1 discussion and 65 comments across x (as of 2026-09-19).

**TL;DR:** The community's main takeaway is that token volume and dollar spend are two different stories: open-weight models dominate raw throughput while closed models (mainly Anthropic) still capture the expensive, high-value calls, so the real signal is the divergence between the two metrics rather than either one alone.

**Sentiment:** 30% positive · 60% mixed · 10% skeptical

**The case for**

- Open models are seen as handling high-throughput, low-entropy tasks (parsing, retries, summaries, test gen) cheaply and effectively.
- Some report large real-world cost savings from routing bulk workloads (e.g. via DeepSeek/Kimi) to open weights instead of closed APIs.
- Combined open-lab spend (Moonshot, DeepSeek, Z.ai) overtaking OpenAI is viewed as a meaningful shift worth tracking.

**The pushback**

- Several argue token volume share is a misleading or 'least useful' KPI since open models are cheaper and more verbose per task, inflating their token counts.
- Skepticism that headline volume numbers reflect a self-selecting sample or Vercel-specific usage rather than an industry-wide trend.
- One commenter questions the unit economics behind the open-model narrative, calling it marketing that masks negative margins.

**By community**

- x (mixed): Replies largely agree the volume/spend divergence is the real story, with some cost-savings anecdotes but recurring pushback that token share alone is a weak or misleading metric.

**Hottest debate:** Whether token volume share is a meaningful signal of adoption or just an artifact of open models being cheaper/more verbose per task.

**Open questions**

- Is the 78.4% open-weight volume figure global or skewed toward specific (e.g. US-heavy) usage?
- Does this reflect Vercel's customer base specifically or an industry-wide trend?
- How does the split hold up when measured by completed-work or cost-per-successful-task rather than raw token count?

**Highlights**

> @rauchg wild shift in real time
> — [jeheskielsunloy on x · 1 points](https://x.com/jeheskielsunloy/status/2101194413519389064)

> @rauchg I keep coming back to cost per successful task here. Open models can own cheap, long-context volume while closed models still capture the high-value agent steps, so token share alone may be the least useful KPI.
> — [Sagarvd01 on x](https://x.com/Sagarvd01/status/2101191006536286284)

> @rauchg wild shift in real time
> — [jeheskielsunloy on x · 1 points](https://x.com/jeheskielsunloy/status/2101194413519389064)

> @rauchg Self selecting sample.
> — [PanglossWasHere on x](https://x.com/PanglossWasHere/status/2101272127991070991)

> @rauchg Moon is #3 by volume. moon's margin on inference: -14%. vgpai just bought a billboard for open models to mask negative unit economics with vanity metrics. it’s not adoption; it’s arson marketing and the fire marshal hasn’t clocked in yet
> — [TheAIShrink on x](https://x.com/TheAIShrink/status/2101193794892841218)

**Source threads**

- [x](https://x.com/rauchg/status/2101186741042663579) · 0 points · 65 comments

---

Tags: [#open-source](https://daily.dev/tags/open-source), [#llm](https://daily.dev/tags/llm), [#anthropic](https://daily.dev/tags/anthropic), [#vercel](https://daily.dev/tags/vercel)

[View this post on daily.dev](https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Open-weight models now handle most tokens on Vercel's AI Gateway, but Anthropic still captures 64% of spending","url":"https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx"},"datePublished":"2026-09-19T05:49:27.784Z","dateModified":"2026-09-19T12:35:33.322Z","description":"Open-weight models surpassed 56% of all tokens routed through Vercel's AI Gateway in August, up from just 7% in December 2024, with some days reaching 78.4%....","image":"https://pbs.twimg.com/media/HSjnTgfaQAAHWbk.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HSjnTgfaQAAHWbk.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"open-source,llm,anthropic,vercel","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Open-weight models now handle most tokens on Vercel's AI Gateway, but Anthropic still captures 64% of spending"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/open-weight-models-now-handle-most-tokens-on-vercel-s-ai-gateway-but-anthropic-still-captures-64-o-n0evrinnx#faq","mainEntity":[{"@type":"Question","name":"What percentage of tokens on Vercel's AI Gateway are now handled by open-weight models versus closed models like Anthropic?","acceptedAnswer":{"@type":"Answer","text":"Open-weight models handled 56% of all tokens routed through Vercel's AI Gateway in August, up from just 7% in December 2024, with some days reaching as high as 78.4%. Despite this token majority, open-weight models captured only 14% of dollar spend, while Anthropic took 64% of spending on far lower token volume. Track shifting LLM cost and usage patterns on daily.dev before locking in a model provider."}},{"@type":"Question","name":"Why does Anthropic capture most AI Gateway spending despite open-weight models handling more tokens?","acceptedAnswer":{"@type":"Answer","text":"Anthropic's models are apparently used for tasks where output quality justifies a higher price, so even with much lower token volume than open-weight models, Anthropic captures 64 cents of every dollar spent on Vercel's AI Gateway. Open-weight models are winning cheap, high-throughput workloads, while Anthropic wins on higher-value revenue per token. Developers weighing quality versus cost across LLM providers can follow these tradeoffs on daily.dev."}},{"@type":"Question","name":"Which open-weight AI labs are gaining the most spend share on Vercel's AI Gateway?","acceptedAnswer":{"@type":"Answer","text":"Moonshot AI and DeepSeek have climbed to third and fourth place by spend on some days, and combined with Z.ai, their spend now surpasses OpenAI's on the gateway. Z.ai's GLM-5.3-Flash tripled its predecessor's daily volume within five days of launch, while this spend reflects inference cost across mostly US-based providers, not revenue to the labs themselves. Compare emerging open-weight model providers as options mature by following updates on daily.dev."}}]}
```

