<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb" -->

---
title: Ling 3.0 Flash is now available on Vercel AI Gateway and...
description: Ling 3.0 Flash, an open-weight Sparse Mixture-of-Experts model from inclusionAI (Ant Group), is now available on Vercel AI Gateway and Kilo, both offering it...
canonical: https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Ling 3.0 Flash is now available on Vercel AI Gateway and Kilo | daily.dev
og:description: Ling 3.0 Flash, an open-weight Sparse Mixture-of-Experts model from inclusionAI (Ant Group), is now available on Vercel AI Gateway and Kilo, both offering it...
og:url: https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb
og:image: https://api.daily.dev/og/posts/z2xSo5bBB.png
og:image:alt: Ling 3.0 Flash is now available on Vercel AI Gateway and Kilo
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Ling 3.0 Flash is now available on Vercel AI Gateway and Kilo

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 3 upvotes · 1 comments

## Summary

Ling 3.0 Flash, an open-weight Sparse Mixture-of-Experts model from inclusionAI (Ant Group), is now available on Vercel AI Gateway and Kilo, both offering it free for a limited time. The model has 124B total parameters but activates only ~5.1B per token, giving it large-model capacity at lower inference cost. It features a 256K native context window (extendable to 1M tokens) and a hybrid reasoning mode suited for agentic coding, document work, and long-context conversations. Vercel's AI Gateway offers it free through August 3rd with BYOK support, automatic failover, and Zero Data Retention. Kilo also offers it free with no specified end date.

## Content

Ling 3.0 Flash, the latest model from inclusionAI (Ant Group's AI lab), is now available across several platforms including Vercel's AI Gateway and Kilo.

## What it is

Ling 3.0 Flash uses a sparse Mixture-of-Experts architecture with 124B total parameters but activates only ~5.1B per token during inference. (One correction worth noting: a community observer pointed out the active parameter count is closer to 104B with 93 layers, so take the official 5.1B figure with some skepticism until the weights are public.) It has a native 256K token context window, reportedly extendable to 1M, and supports both standard and "thinking" modes that blend fast responses with step-by-step reasoning.

The model is aimed at agentic coding tasks, document processing, and long-context multi-turn conversations where token efficiency matters more than raw parameter count.

## Availability

Vercel's AI Gateway has it live now, free for three weeks through August 3rd, with BYOK support, no markup on provider pricing, automatic failover, and Zero Data Retention. Kilo is also offering it free for a limited time.

The model weights aren't public yet. The team announced first and will open-source the weights separately - a pattern vLLM explicitly endorsed in their own announcement. Their argument: freezing the checkpoint before release gives inference projects a stable window to test correctness, tune performance, and validate serving configs, rather than scrambling on day zero with a moving target. vLLM says their support for Ling 3.0 Flash will go live when the weights drop.

## Worth watching

The efficiency angle is the real story here. A model that activates 5.1B parameters per token while carrying 124B total is designed to be cheap to run at scale, not just impressive on a spec sheet. Whether it holds up in production agentic workloads is something the community will figure out once the weights are actually available.

## Community discussion

Top comments from developers on daily.dev.

**@trevorsuna** · 1 upvotes

> Only activating 5.1B of 124B parameters is the interesting bit here. Gateway failover and reporting make trials easier, but I’d still compare cold-start latency and real token cost before moving an agent workflow over.

## Similar posts on daily.dev

- [Ling 3.0 Flash Fin: A 124B Parameter Model Punching at the Weight Class of Trillion-Parameter Giants](https://daily.dev/posts/ling-3-0-flash-fin-a-124b-parameter-model-punching-at-the-weight-class-of-trillion-parameter-giants-qmtf8wq40) · Medium · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-coding](https://daily.dev/tags/ai-coding), [#vercel](https://daily.dev/tags/vercel), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts), [#ai-gateway](https://daily.dev/tags/ai-gateway)

[View this post on daily.dev](https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Ling 3.0 Flash is now available on Vercel AI Gateway and Kilo","url":"https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb"},"datePublished":"2026-07-24T00:12:05.843Z","dateModified":"2026-07-27T15:26:34.276Z","description":"Ling 3.0 Flash, an open-weight Sparse Mixture-of-Experts model from inclusionAI (Ant Group), is now available on Vercel AI Gateway and Kilo, both offering it...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/c1ce90f79f41b3851927672539cc00ae?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/c1ce90f79f41b3851927672539cc00ae?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"llm,ai-coding,vercel,mixture-of-experts,ai-gateway","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Ling 3.0 Flash is now available on Vercel AI Gateway and Kilo"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/ling-3-0-flash-is-now-available-on-vercel-ai-gateway-and-kilo-z2xso5bbb","comment":[{"@type":"Comment","text":"Only activating 5.1B of 124B parameters is the interesting bit here. Gateway failover and reporting make trials easier, but I’d still compare cold-start latency and real token cost before moving an agent workflow over.","datePublished":"2026-07-24T02:29:10.776Z","url":"https://daily.dev/posts/z2xSo5bBB#c-ZpZc0k73M","author":{"@type":"Person","name":"Trevor Suna","url":"https://daily.dev/trevorsuna","image":"https://media.daily.dev/image/upload/s--dZ7gXxpp--/f_auto/v1784081551/avatars/avatar_EMoP47rpuw8DNjhp6R1b6?_a=BAMAMicg0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}}]}
```

