<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg" -->

---
title: Mercury 2: Inception&#x27;s Breakthrough in Real-Time...
description: Inception Labs has launched Mercury 2, a reasoning diffusion LLM that uses a diffusion-based architecture instead of autoregressive decoding to generate...
canonical: https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Mercury 2: Inception&#x27;s Breakthrough in Real-Time Language Processing | daily.dev
og:description: Inception Labs has launched Mercury 2, a reasoning diffusion LLM that uses a diffusion-based architecture instead of autoregressive decoding to generate...
og:url: https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg
og:image: https://api.daily.dev/og/posts/ea5kyYvcg.png
og:image:alt: Mercury 2: Inception&#x27;s Breakthrough in Real-Time Language Processing
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Mercury 2: Inception's Breakthrough in Real-Time Language Processing

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Inception Labs has launched Mercury 2, a reasoning diffusion LLM that uses a diffusion-based architecture instead of autoregressive decoding to generate multiple tokens in parallel. It achieves over 1,000 tokens per second on NVIDIA Blackwell GPUs, more than five times faster than sequential models. Designed for latency-sensitive environments, it targets use cases like agentic loops, real-time voice interfaces, coding assistants, and RAG pipelines. Mercury 2 is OpenAI API-compatible, supports native tool use, tunable reasoning, and schema-aligned JSON output, with a 128K token context window priced at $0.25/$0.75 per million input/output tokens. Early access and free public testing are available.

## Content

Inception Labs has recently unveiled Mercury 2, heralded as a pioneering reasoning diffusion large language model (LLM) that significantly outpaces traditional language processing models. Leveraging a diffusion-based architecture instead of the conventional autoregressive decoding, Mercury 2 is capable of generating multiple tokens simultaneously through parallel refinement. This innovative approach allows the model to achieve processing speeds exceeding 1,000 tokens per second on NVIDIA Blackwell GPUs, particularly more than five times faster than sequential decoding models.

With a focus on reducing latency, Mercury 2 has been tailored for latency-sensitive production environments, making it ideal for use cases such as agentic loops, real-time voice interfaces, coding assistants, and retrieval-augmented generation (RAG) pipelines. Its impressive capabilities make it not only a rapid model but also one that does not compromise on quality, delivering reasoning-grade outputs within strict real-time latency parameters.

Mercury 2 is OpenAI API-compatible and supports advanced features such as native tool use, tunable reasoning, and schema-aligned JSON output. The model also boasts a substantial context window of 128,000 tokens and is competitively priced at $0.25 per million input tokens and $0.75 per million output tokens.

To provide users with early exposure, Inception Labs is offering early access and free public testing via their website, encouraging widespread experimentation and feedback collection to refine the model further for varied industrial applications.

## Similar posts on daily.dev

- [Inception Labs: Making LLMs Faster and More Cost-Efficient](https://daily.dev/posts/inception-labs-making-llms-faster-and-more-cost-efficient-jwseakaal) · The New Stack · 1 upvotes · 0 comments
- [Inception Labs says its diffusion LLM is 10x faster than Claude, ChatGPT, Gemini](https://daily.dev/posts/inception-labs-says-its-diffusion-llm-is-10x-faster-than-claude-chatgpt-gemini-x9mtyj6ul) · The New Stack · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#diffusion-models](https://daily.dev/tags/diffusion-models)

[View this post on daily.dev](https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Mercury 2: Inception's Breakthrough in Real-Time Language Processing","url":"https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg"},"datePublished":"2026-02-25T22:35:14.829Z","dateModified":"2026-02-25T22:35:34.580Z","description":"Inception Labs has launched Mercury 2, a reasoning diffusion LLM that uses a diffusion-based architecture instead of autoregressive decoding to generate...","image":"https://pbs.twimg.com/profile_images/1899962463115751424/i-6MBWau_normal.jpg","thumbnailUrl":"https://pbs.twimg.com/profile_images/1899962463115751424/i-6MBWau_normal.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/mercury-2-inception-s-breakthrough-in-real-time-language-processing-ea5kyyvcg","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,diffusion-models","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Mercury 2: Inception's Breakthrough in Real-Time Language Processing"}]}
```

