<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr" -->

---
title: DeepSeek launches V4.1-Flash and retires V4-Pro, its...
description: DeepSeek released V4.1-Flash, a 552-billion-parameter model that activates only 8 billion parameters for input and 16 billion for output, aimed at cutting the...
canonical: https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model | daily.dev
og:description: DeepSeek released V4.1-Flash, a 552-billion-parameter model that activates only 8 billion parameters for input and 16 billion for output, aimed at cutting the...
og:url: https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr
og:image: https://api.daily.dev/og/posts/mauD6lKXR.png
og:image:alt: DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

**[The Next Web](https://daily.dev/sources/tnw)** · 5 min read · 1 upvotes · 0 comments

## Summary

DeepSeek released V4.1-Flash, a 552-billion-parameter model that activates only 8 billion parameters for input and 16 billion for output, aimed at cutting the cost of AI agent workloads. The company claims it beats its own flagship V4-Pro on coding and agent benchmarks while needing roughly a quarter of the KV cache memory of V4-Flash. Starting 04:00 UTC on 14 September, all V4-Pro traffic will be routed to V4.1-Flash at cheaper rates until a V4.1-Pro ships, with no date given. Prices dropped as much as 32% according to Bloomberg Intelligence, reversing a quadrupling of prices in August. DeepSeek's benchmark table shows it matching or beating Claude Opus 5 and GPT-5.6 Sol on some coding and cybersecurity tests, but trailing badly on Humanity's Last Exam and ProgramBench. Shares of Chinese rivals MiniMax and Z.ai fell over 8%, and Alibaba slid more than 2%, as DeepSeek reportedly prepares for a Shanghai STAR Market listing.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://thenextweb.com/news/deepseek-v4-1-flash-launch-v4-pro-retired-price-cut>

## Questions this post answers

### What happens to DeepSeek V4-Pro after V4.1-Flash launches?

DeepSeek is retiring V4-Pro by routing all its requests to V4.1-Flash starting 04:00 UTC on September 14, billed at the cheaper Flash rates. This arrangement continues until a V4.1-Pro model arrives, though DeepSeek has not given a release date for it. Every existing V4-Pro customer effectively becomes a V4.1-Flash customer that Monday.

_daily.dev helps teams tracking DeepSeek's pricing and model changes stay ahead of migration deadlines._

### How much memory does DeepSeek V4.1-Flash use for its KV cache compared to previous models?

DeepSeek V4.1-Flash needs 890 bytes per token for its KV cache, about a quarter of what V4-Flash required and roughly 437 times less than DeepSeek's first model released in 2023. The model has 552 billion total parameters but activates only 8 billion for input and 16 billion for output per token, cutting memory costs for long agent sessions.

_engineers comparing LLM inference costs can follow architecture shifts like this on daily.dev._

### How does DeepSeek V4.1-Flash's coding performance compare to Claude Opus 5 and GPT-5.6 Sol?

On the DeepSWE v1.1 software engineering benchmark, DeepSeek V4.1-Flash scores 74.2, narrowly ahead of Claude Opus 5's 74.0 and OpenAI's GPT-5.6 Sol at 73.0, according to DeepSeek's own benchmark table. However, Opus 5 leads decisively on Humanity's Last Exam (56.3 vs 36.8) and ProgramBench (37.0 vs 20.3), and both US models beat V4.1-Flash on Terminal-Bench 3.0.

_developers weighing model choices for coding agents can track benchmark comparisons like this on daily.dev._

---

Tags: [#open-source](https://daily.dev/tags/open-source), [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#claude](https://daily.dev/tags/claude), [#deepseek](https://daily.dev/tags/deepseek)

[View this post on daily.dev](https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model","url":"https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr"},"datePublished":"2026-09-10T15:56:37.312Z","dateModified":"2026-09-10T19:33:56.913Z","description":"DeepSeek released V4.1-Flash, a 552-billion-parameter model that activates only 8 billion parameters for input and 16 billion for output, aimed at cutting the...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/adb348fddc1ad8e5ac59846a187d4582?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/adb348fddc1ad8e5ac59846a187d4582?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"The Next Web","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"The Next Web","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/tnw","url":"https://daily.dev/sources/tnw"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"open-source,llm,ai-agents,claude,deepseek","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"The Next Web","item":"https://daily.dev/sources/tnw"},{"@type":"ListItem","position":3,"name":"DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/deepseek-launches-v4-1-flash-and-retires-v4-pro-its-flagship-model-maud6lkxr#faq","mainEntity":[{"@type":"Question","name":"What happens to DeepSeek V4-Pro after V4.1-Flash launches?","acceptedAnswer":{"@type":"Answer","text":"DeepSeek is retiring V4-Pro by routing all its requests to V4.1-Flash starting 04:00 UTC on September 14, billed at the cheaper Flash rates. This arrangement continues until a V4.1-Pro model arrives, though DeepSeek has not given a release date for it. Every existing V4-Pro customer effectively becomes a V4.1-Flash customer that Monday. daily.dev helps teams tracking DeepSeek's pricing and model changes stay ahead of migration deadlines."}},{"@type":"Question","name":"How much memory does DeepSeek V4.1-Flash use for its KV cache compared to previous models?","acceptedAnswer":{"@type":"Answer","text":"DeepSeek V4.1-Flash needs 890 bytes per token for its KV cache, about a quarter of what V4-Flash required and roughly 437 times less than DeepSeek's first model released in 2023. The model has 552 billion total parameters but activates only 8 billion for input and 16 billion for output per token, cutting memory costs for long agent sessions. engineers comparing LLM inference costs can follow architecture shifts like this on daily.dev."}},{"@type":"Question","name":"How does DeepSeek V4.1-Flash's coding performance compare to Claude Opus 5 and GPT-5.6 Sol?","acceptedAnswer":{"@type":"Answer","text":"On the DeepSWE v1.1 software engineering benchmark, DeepSeek V4.1-Flash scores 74.2, narrowly ahead of Claude Opus 5's 74.0 and OpenAI's GPT-5.6 Sol at 73.0, according to DeepSeek's own benchmark table. However, Opus 5 leads decisively on Humanity's Last Exam (56.3 vs 36.8) and ProgramBench (37.0 vs 20.3), and both US models beat V4.1-Flash on Terminal-Bench 3.0. developers weighing model choices for coding agents can track benchmark comparisons like this on daily.dev."}}]}
```

