<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/meanwhile-in-china-deepseek-cut-its-kv-cache-hbm-need-75-and-ssd-need-87-5--kbco5kipz" -->

---
title: Meanwhile in China: DeepSeek Cut Its KV-Cache HBM Need...
description: DeepSeek&#x27;s September 10 release of V4.1-Flash cuts KV-cache HBM requirements by 75% and SSD requirements by 87.5% versus its prior architecture, a...
canonical: https://daily.dev/posts/meanwhile-in-china-deepseek-cut-its-kv-cache-hbm-need-75-and-ssd-need-87-5--kbco5kipz
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Meanwhile in China: DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%. | daily.dev
og:description: DeepSeek&#x27;s September 10 release of V4.1-Flash cuts KV-cache HBM requirements by 75% and SSD requirements by 87.5% versus its prior architecture, a...
og:url: https://daily.dev/posts/meanwhile-in-china-deepseek-cut-its-kv-cache-hbm-need-75-and-ssd-need-87-5--kbco5kipz
og:image: https://api.daily.dev/og/posts/Kbco5kIPZ.png
og:image:alt: Meanwhile in China: DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%.
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Meanwhile in China: DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%.

**[Jacob B. Bonde](https://daily.dev/sources/5cqvizkr5tdfdotvdjulg)** · [@byteoutlaw](https://daily.dev/byteoutlaw) · 16 upvotes · 11 comments

## Summary

DeepSeek's September 10 release of V4.1-Flash cuts KV-cache HBM requirements by 75% and SSD requirements by 87.5% versus its prior architecture, a memory-efficiency gain during inference rather than across the full model or training pipeline. The piece frames this as a potential risk to the memory-shortage investment thesis underpinning Micron and Sandisk stocks, citing recent Micron Cloud Memory and Sandisk data-center revenue figures, gross margins, and hedge fund holdings data, while noting the bull case that cheaper tokens could stimulate enough additional usage to offset the efficiency gain.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://finance.yahoo.com/technology/ai/articles/deepseek-cut-kv-cache-hbm-011603128.html>

## Community discussion

Top comments from developers on daily.dev.

**@devgenx** · 1 upvotes

> This time next year local LLMs will out perform GPT6.

**@akkitto** · 1 upvotes

> Yeah, DeepSeek is really amazing. I'm using the older models like DeepSeek V4 Flash 0731 a lot, because they are super cheap, but used also DeepSeek V4 Pro 0813 & since recently V4.1 Flash and they are really amazing. Comparatively very cheap, yet they do a very good job. They take off a lot of load.

**@lezli01** · 1 upvotes

> Chinese are not playing around...

**@markcoleman** · 0 upvotes

> Gated delta net is in a few models. Qwen uses it to great effect. It's been a great year for local llms.

**@theacademe** · 0 upvotes

> KV-cache reduction combined with FlashML's freetoken CPU/GPU splitting would be an interesting synthesis to easily allow 35B parameter models to run on consumer 8GB VRAM cards. And 4 node NVidia DGX clusters (supported as of March 2026) having around ~512 GB unified memory would support up to 700B models. Locally. But it's going to really heat up your office!

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments

---

Tags: [#ai-inference](https://daily.dev/tags/ai-inference), [#deepseek](https://daily.dev/tags/deepseek)

[View this post on daily.dev](https://daily.dev/posts/meanwhile-in-china-deepseek-cut-its-kv-cache-hbm-need-75-and-ssd-need-87-5--kbco5kipz)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/meanwhile-in-china-deepseek-cut-its-kv-cache-hbm-need-75-and-ssd-need-87-5--kbco5kipz","headline":"Meanwhile in China: DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%.","text":"Shared: DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%. Micron and Sandisk Investors Should Pay Attention","url":"https://daily.dev/posts/meanwhile-in-china-deepseek-cut-its-kv-cache-hbm-need-75-and-ssd-need-87-5--kbco5kipz","datePublished":"2026-09-13T08:40:02.376Z","dateModified":"2026-09-13T08:43:50.750Z","author":{"@type":"Person","name":"Jacob B. Bonde","url":"https://daily.dev/byteoutlaw","image":"https://media.daily.dev/image/upload/s--veTChHK7--/f_auto/v1733220070/avatars/avatar_5cQvIZKr5tDFDotVDjulg","description":"I am about to change this world with my fellow programmers! 😊❤️","interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":22865}},"image":"https://media.daily.dev/image/upload/s--HRgLpUt6--/f_auto/v1722860399/public/Placeholder%2003","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":16},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":11}],"sharedContent":{"@type":"WebPage","url":"https://api.daily.dev/r/QLObrWXnZ"},"comment":[{"@type":"Comment","text":"This time next year local LLMs will out perform GPT6.","datePublished":"2026-09-13T10:09:19.598Z","url":"https://daily.dev/posts/Kbco5kIPZ#c-DvSusBilu","author":{"@type":"Person","name":"Mike","url":"https://daily.dev/devgenx","image":"https://media.daily.dev/image/upload/s--PrWOWtSq--/f_auto/v1789486214/avatars/avatar_ytCVp6dQicn9P0yKfhQCI?_a=BAMAMicg0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"Yeah, DeepSeek is really amazing. I’m using the older models like DeepSeek V4 Flash 0731 a lot, because they are super cheap, but used also DeepSeek V4 Pro 0813 &amp; since recently V4.1 Flash and they are really amazing. Comparatively very cheap, yet they do a very good job. They take off a lot of load.","datePublished":"2026-09-13T10:53:30.581Z","url":"https://daily.dev/posts/Kbco5kIPZ#c-6KA8v2Zto","author":{"@type":"Person","name":"Daniel","url":"https://daily.dev/akkitto","image":"https://media.daily.dev/image/upload/s--FtwJqX4c--/f_auto/v1754900041/avatars/avatar_29TCpY2hJR72V3BlxPXzX?_a=BAMClqZW0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"Chinese are not playing around…","datePublished":"2026-09-13T14:05:19.742Z","url":"https://daily.dev/posts/Kbco5kIPZ#c-Qxd18SNp5","author":{"@type":"Person","name":"László Szabó","url":"https://daily.dev/lezli01","image":"https://lh3.googleusercontent.com/a/ACg8ocKj3v7EFYJoXUUuro6ALF9fD3RTRATiBpBOVcOzqro4fy6bWqMm=s96-c"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"Gated delta net is in a few models. Qwen uses it to great effect. It’s been a great year for local llms.","datePublished":"2026-09-13T15:50:11.760Z","url":"https://daily.dev/posts/Kbco5kIPZ#c-T3eQh0qeo","author":{"@type":"Person","name":"Mark Coleman","url":"https://daily.dev/markcoleman","image":"https://lh3.googleusercontent.com/a/ACg8ocJ7hP8deNUggdRJg_Ty4VaCSFEskKSp_TewbQCEl7lE0NDA5fbn-A=s96-c"}},{"@type":"Comment","text":"KV-cache reduction combined with FlashML’s freetoken CPU/GPU splitting would be an interesting synthesis to easily allow 35B parameter models to run on consumer 8GB VRAM cards. And 4 node NVidia DGX clusters (supported as of March 2026) having around ~512 GB unified memory would support up to 700B models. Locally. But it’s going to really heat up your office!","datePublished":"2026-09-15T04:58:38.780Z","dateModified":"2026-09-15T05:03:12.748Z","url":"https://daily.dev/posts/Kbco5kIPZ#c-61Lnt7k9x","author":{"@type":"Person","name":"theacademe","url":"https://daily.dev/theacademe","image":"https://media.daily.dev/image/upload/s--C29_q-xW--/f_auto/v1756852849/avatars/avatar_wNAjl1VBaVx2CCjbkkb5h?_a=BAMClqZW0"}}],"isPartOf":{"@type":"WebPage","url":"https://daily.dev/sources/5cqvizkr5tdfdotvdjulg","name":"Jacob B. Bonde"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Jacob B. Bonde","item":"https://daily.dev/sources/5cqvizkr5tdfdotvdjulg"},{"@type":"ListItem","position":3,"name":"Meanwhile in China: DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%."}]}
```

