<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p" -->

---
title: DeepSeek V4.1 Flash released on Hugging Face, replacing...
description: DeepSeek V4.1 Flash has launched on Hugging Face as a 552B parameter Mixture-of-Experts model featuring a new Encoder-Decoder architecture, a new pre-training...
canonical: https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: DeepSeek V4.1 Flash released on Hugging Face, replacing V4 Pro | daily.dev
og:description: DeepSeek V4.1 Flash has launched on Hugging Face as a 552B parameter Mixture-of-Experts model featuring a new Encoder-Decoder architecture, a new pre-training...
og:url: https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p
og:image: https://api.daily.dev/og/posts/zIKP1p10p.png
og:image:alt: DeepSeek V4.1 Flash released on Hugging Face, replacing V4 Pro
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# DeepSeek V4.1 Flash released on Hugging Face, replacing V4 Pro

**[Collections](https://daily.dev/sources/collections)** · 4 min read · 6 upvotes · 0 comments

## Summary

DeepSeek V4.1 Flash has launched on Hugging Face as a 552B parameter Mixture-of-Experts model featuring a new Encoder-Decoder architecture, a new pre-training method, and larger-scale reinforcement learning post-training. It supports native visual understanding and is the smallest model in DeepSeek's new architecture family. DeepSeek says the model outperforms V4 Pro on performance, cost, speed, and compute usage, and as a result V4 Pro is being retired and phased out once V4.1 Flash is fully live.

## Content

## What it is

DeepSeek released V4.1 Flash, a 552B MoE model that only activates 8B parameters during prefill and 16B during decode. Despite the "flash" label and minor version number, this is a substantial architectural overhaul - several people have noted it probably deserved to be called V5.

Weights are open, it's available on the API, and as of launch it's the default model on HuggingChat.

## The architecture changes

The biggest structural change is a **Causal Encoder-Decoder (CED)** design. The decoder's KV cache is built from the encoder's final hidden states rather than running full decoder computation. This makes it particularly efficient for input-heavy workloads like agentic tasks, where you're feeding in a lot of context before generating much output.

A few other things worth noting:

- **Compressed Sparse Attention 2 (CSA2)** with three operating modes for how it handles KV, indexer K, and Top-K indices
- **Engram module** - 196-197B parameters of n-gram memory the model looks up rather than computes, decoupling memorization from computation. vLLM notes this is a quarter of the checkpoint.
- Only four layers write compressed KV; the rest of the model shares it, which is how they got KV cache down roughly 400x versus their first model
- FP4 KV cache with quantization-aware training
- Native multimodal support via DeepSeek-ViT (a standard ViT architecture), trained first with SigLIP loss then autoregressive loss for the combined vision+LLM
- Trained on 45T tokens

On post-training, DeepSeek was unusually candid: they didn't introduce new algorithms. Their view is that at this stage, engineering the data and environment pipeline returns more than algorithmic novelty. They used the model itself to construct training environments from internal usage data, and introduced a scalar effort level `b` as an explicit RL conditioning signal (max=100, high=75, low=50). The final stage used over 40 teacher models for distillation.

## Benchmarks

Results are strong on coding and agent tasks - it matches or beats Claude Opus 5 and GPT-5.6 Sol on DeepSWE v1.1, CyberGym, and a few others. It's the new top open-weight model on the Vals Index, beating Kimi K3. One coding benchmark retest (with thinking enabled at max effort) scored 81.25%, above V4 Flash (72.5%) and V4 Pro (76.25%).

It's not uniformly dominant though. It trails on Humanity's Last Exam and ProgramBench, and on the Artificial Analysis Index it scores 40 - above V4 Pro but below GLM-5.3-Flash. The honest read is: very strong on agentic/coding tasks, more mixed elsewhere.

## Pricing and API routing

Prices dropped as much as 32% from August levels. Off-peak API pricing is $0.15 per million input tokens and $0.60 per million output tokens, doubling at peak. One comparison: building a landing page with V4.1 Flash cost $0.026 versus $1.21 with Claude Fable 5 for similar quality output.

Starting September 14 at 04:00 UTC, all V4 Pro traffic routes to V4.1 Flash at flash pricing. V4 Pro is being retired - the reasoning being that V4.1 Flash outperforms it on essentially every metric while costing less. No date yet for V4.1 Pro.

## Running it locally

This is where it gets interesting for people with high-end consumer hardware. antirez has been testing it in DwarfStar:

- **Single M5 Max 128GB**: SSD streaming gets around 15 t/s generation. The encoder blocks (50% of the model) fit comfortably in memory; the loader takes 8 seconds to cache them from SSD, then prefill hits 800 t/s.
- **Dual M5 Max 128GB via RDMA**: tensor parallel execution gets ~25 t/s
- **Single DGX Spark**: ~9 t/s with SSD streaming; dual-Spark RDMA gets ~22 t/s

The SSD streaming performance is better than expected, partly because recent changes to retain the right experts in cache help, and partly because V4.1 may reuse the same experts more consistently than V4.

One caveat: it's not really a fit for a single 128GB system in the traditional sense - you're relying on SSD streaming. For comfortable local use you'd want 2x128GB or a large Mac Studio. Support is now on DwarfStar's main branch.

vLLM also supports it from day 0, verified on both NVIDIA and AMD GPUs.

## The broader picture

DeepSeek's response to increased demand was to make the model cheaper and faster rather than raise prices - which is the opposite of what happened in August when prices quadrupled. Chinese AI stocks (MiniMax, Z.ai, Alibaba) dropped on the news, and DeepSeek is reportedly preparing for a Shanghai STAR Market listing.

The "flash" naming is doing a lot of work to undersell what this actually is. A 552B model with a new encoder-decoder architecture, 45T token pretraining, and large-scale RL post-training that beats the previous flagship at a fraction of the cost isn't a minor update - it just happens to be cheap to run.

## Questions this post answers

### What is DeepSeek V4.1 Flash and how big is it?

DeepSeek V4.1 Flash is a 552 billion parameter Mixture-of-Experts model released on Hugging Face, using a new Encoder-Decoder structure built with a new pre-training method and larger-scale reinforcement learning post-training. It supports native visual understanding and is the smallest model in DeepSeek's new architecture family.

_Developers picking a model for multimodal or agentic workloads can track releases like this one on daily.dev._

### Is DeepSeek V4 Pro being deprecated?

Yes, DeepSeek V4 Pro is being retired and will be phased out after V4.1 Flash goes fully live. DeepSeek states that V4.1 Flash outperforms V4 Pro on performance, cost, speed, and total compute usage, making it no longer worthwhile to keep offering V4 Pro at a higher price and slower speed.

_Teams relying on a pinned model id should watch for deprecations like this via daily.dev._

## Community take

How the wider developer community reacted, aggregated from 1 discussion and 22 comments across x (as of 2026-09-16).

**TL;DR:** Most of the discussion is hands-on hardware testing of running the model via SSD streaming on setups like Spark and M3 Ultra, with people reporting decode speeds ranging from under 1 to around 25 tokens/sec and debating whether that's actually usable.

**Sentiment:** 15% positive · 25% mixed · 60% skeptical

**The case for**

- Dual-Spark RDMA setups reportedly more than double throughput compared to a single Spark.
- Some see 9-22 t/s via SSD streaming as a workable one-box or scale-up option for certain jobs.

**The pushback**

- Several report the speeds achieved (as low as 0.3-9 t/s decode) as too slow to be worth using.
- IOPS and SSD wear/warranty are raised as likely bottlenecks or risks with heavy SSD-streaming inference.
- Concerns that prefill time on long contexts (e.g. 100K prompts) may be a hidden bottleneck that SSD streaming can't hide.

**By community**

- x (skeptical): Hands-on testers largely report the SSD-streaming setup as slow and marginal, with ongoing debate over whether any configuration is genuinely worth using.

**Hottest debate:** Whether current SSD-streaming inference speeds (roughly 9-25 t/s) are actually 'worth it' for practical use versus just a technical curiosity.

**Open questions**

- What is the actual prefill time at long context (e.g. 100K tokens) on SSD-streaming setups?
- Is memory bandwidth or interconnect the true bottleneck for dual-Spark configurations?
- How much does sustained SSD streaming degrade SSD lifespan/warranty in practice?

**Highlights**

> @antirez half a trillion via SSD streaming means IOPS is going to be the absolute bottleneck long before compute even wakes up
> — [bullbear\_info on x](https://x.com/bullbear_info/status/2099238700957999171)

> @antirez The SSD's warranty is about to become the real bottleneck here, not the token speed.
> — [onefinalprompt on x](https://x.com/onefinalprompt/status/2099254563425440092)

> @antirez 9 t/s from SSD on one Spark is a real number and the one beside it decides the use case: time to first token on a 100K prompt, since the KV build is what SSD streaming cannot hide. Have you measured prefill at long context on that setup?
> — [arcyton on x](https://x.com/arcyton/status/2099348260854837506)

> @TheDavidTai Yep still too slow. I'll try to improve things in the next days, maybe for this specific model it could be worthwhile to release a Q2 model with Q4 projections and shared experts (currently both at Q8) for the dual spark setup.
> — [antirez on x · 3 points, 1 comments](https://x.com/antirez/status/2099238052006686962)

> @antirez with 9 tok sec , it can be used without reasoning. is it worth?
> — [morandalex0\_0 on x](https://x.com/morandalex0_0/status/2099231511866171870)

**Source threads**

- [x](https://x.com/antirez/status/2099226068083204300) · 0 points · 22 comments

## Similar posts on daily.dev

- [DeepSeek-V4: a million-token context that agents can actually use](https://daily.dev/posts/deepseek-v4-a-million-token-context-that-agents-can-actually-use-uczd4kivx) · Hugging Face · 27 upvotes · 2 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#deepseek](https://daily.dev/tags/deepseek), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"DeepSeek V4.1 Flash released on Hugging Face, replacing V4 Pro","url":"https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p"},"datePublished":"2026-09-10T06:13:48.383Z","dateModified":"2026-09-16T14:25:47.632Z","description":"DeepSeek V4.1 Flash has launched on Hugging Face as a 552B parameter Mixture-of-Experts model featuring a new Encoder-Decoder architecture, a new pre-training...","image":"https://pbs.twimg.com/media/HR1ZhA3acAA7OsQ.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HR1ZhA3acAA7OsQ.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":6},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,deepseek,mixture-of-experts","timeRequired":"PT4M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"DeepSeek V4.1 Flash released on Hugging Face, replacing V4 Pro"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/deepseek-v4-1-flash-released-on-hugging-face-replacing-v4-pro-zikp1p10p#faq","mainEntity":[{"@type":"Question","name":"What is DeepSeek V4.1 Flash and how big is it?","acceptedAnswer":{"@type":"Answer","text":"DeepSeek V4.1 Flash is a 552 billion parameter Mixture-of-Experts model released on Hugging Face, using a new Encoder-Decoder structure built with a new pre-training method and larger-scale reinforcement learning post-training. It supports native visual understanding and is the smallest model in DeepSeek's new architecture family. Developers picking a model for multimodal or agentic workloads can track releases like this one on daily.dev."}},{"@type":"Question","name":"Is DeepSeek V4 Pro being deprecated?","acceptedAnswer":{"@type":"Answer","text":"Yes, DeepSeek V4 Pro is being retired and will be phased out after V4.1 Flash goes fully live. DeepSeek states that V4.1 Flash outperforms V4 Pro on performance, cost, speed, and total compute usage, making it no longer worthwhile to keep offering V4 Pro at a higher price and slower speed. Teams relying on a pinned model id should watch for deprecations like this via daily.dev."}}]}
```

