<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm" -->

---
title: DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since...
description: Within three weeks of DeepSeek-V4.1-Flash's release, Inferact and the vLLM community optimized its serving performance, achieving a 1.9x speedup at low...
canonical: https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0 | daily.dev
og:description: Within three weeks of DeepSeek-V4.1-Flash's release, Inferact and the vLLM community optimized its serving performance, achieving a 1.9x speedup at low...
og:url: https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm
og:image: https://api.daily.dev/og/posts/Js5LNa7tm.png
og:image:alt: DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0

**[vLLM](https://daily.dev/sources/vllm)** · 9 min read · 0 upvotes · 0 comments

## Summary

Within three weeks of DeepSeek-V4.1-Flash's release, Inferact and the vLLM community optimized its serving performance, achieving a 1.9x speedup at low concurrency and a 5.3x throughput improvement under a 150 TPS latency constraint on the SemiAnalysis AgentX agentic benchmark. Gains come from SWA bounded replay (trimming sliding-window attention recomputation with CUDA graphs, cutting TTFT 30-40%), integration of DeepSeek's new kernels (MegaAttention, Mega-mHC, Mega-Gate, DeepSelect, sparse MQA logits), and aggressive kernel fusion including mHC side-stream overlap and fused all-reduce. The model's causal encoder-decoder architecture activates 16B params during decode but only 8B during prefill, and uses FP4 KV cache, CSA2, and inter-layer KV sharing to reach an 890-byte-per-token KV footprint. Accuracy checks on GSM8K and GPQA showed no meaningful degradation from the approximate SWA replay technique.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://vllm.ai/blog/2026-10-07-deepseek-v41-flash>

## Questions this post answers

### What is SWA bounded replay in DeepSeek V4.1 and how does it speed up prefill?

SWA bounded replay is a technique that reruns only the last 128 tokens of a sliding-window attention cache instead of exactly recomputing it, clipping the window at the replay start. It trades bit-exactness for efficiency, letting layers 21-39 skip most prompt tokens during prefill. Combined with CUDA graphs on the trimmed layers, it cuts prefill computation time by 30-40% with negligible accuracy loss on GSM8K and GPQA.

_Engineers tuning agentic LLM serving pipelines can track inference optimization techniques like this via daily.dev._

### How much does NVFP4 compressed KV cache reduce memory versus FP8 in MegaAttention?

NVFP4 compressed KV cache is 45% smaller than the previous FP8 KV cache format used in FlashMLA's MegaAttention kernel. MegaAttention performs query RoPE, sparse attention, inverse RoPE, and FP8 casting in one kernel launch, improving kernel efficiency by 1.45x through fusion that removes HBM writes between operations, which also increases per-GPU concurrency in high-throughput serving.

_Teams comparing KV cache compression strategies for GPU memory efficiency can follow developments like this on daily.dev._

### What throughput gains did vLLM achieve serving DeepSeek-V4.1-Flash on the SemiAnalysis AgentX benchmark?

vLLM achieved a 1.9x speedup in low-latency serving and roughly a 5.3x throughput improvement under a 150 TPS constraint, compared to its day-0 implementation, within three weeks of the model's release. Low-latency serving used TP4 with FlashInfer attention, while high-throughput serving switched to DEP2 with DP attention and MegaAttention's NVFP4 KV compression to boost per-GPU concurrency.

_Developers benchmarking agentic LLM serving stacks can stay on top of results like these through daily.dev._

## Similar posts on daily.dev

- [DeepSeek-V4: a million-token context that agents can actually use](https://daily.dev/posts/deepseek-v4-a-million-token-context-that-agents-can-actually-use-uczd4kivx) · Hugging Face · 27 upvotes · 2 comments
- [Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0 – PyTorch](https://daily.dev/posts/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0-mmxalpv2n) · PyTorch · 1 upvotes · 0 comments

---

Tags: [#ai-inference](https://daily.dev/tags/ai-inference), [#deepseek](https://daily.dev/tags/deepseek), [#cuda](https://daily.dev/tags/cuda), [#vllm](https://daily.dev/tags/vllm)

[View this post on daily.dev](https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0","url":"https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm"},"datePublished":"2026-10-07T19:40:25.399Z","dateModified":"2026-10-07T19:48:13.911Z","description":"Within three weeks of DeepSeek-V4.1-Flash's release, Inferact and the vLLM community optimized its serving performance, achieving a 1.9x speedup at low...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/059e05f23410172aa084e2791c3e7f6f?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/059e05f23410172aa084e2791c3e7f6f?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"vLLM","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"vLLM","logo":"https://media.daily.dev/image/upload/s--hTxEuls9--/f_auto/v1744613054/logos/vllm","url":"https://daily.dev/sources/vllm"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-inference,deepseek,cuda,vllm","timeRequired":"PT9M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"vLLM","item":"https://daily.dev/sources/vllm"},{"@type":"ListItem","position":3,"name":"DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/deepseek-v4-1-flash-on-vllm-5x-agentic-throughput-since-day-0-js5lna7tm#faq","mainEntity":[{"@type":"Question","name":"What is SWA bounded replay in DeepSeek V4.1 and how does it speed up prefill?","acceptedAnswer":{"@type":"Answer","text":"SWA bounded replay is a technique that reruns only the last 128 tokens of a sliding-window attention cache instead of exactly recomputing it, clipping the window at the replay start. It trades bit-exactness for efficiency, letting layers 21-39 skip most prompt tokens during prefill. Combined with CUDA graphs on the trimmed layers, it cuts prefill computation time by 30-40% with negligible accuracy loss on GSM8K and GPQA. Engineers tuning agentic LLM serving pipelines can track inference optimization techniques like this via daily.dev."}},{"@type":"Question","name":"How much does NVFP4 compressed KV cache reduce memory versus FP8 in MegaAttention?","acceptedAnswer":{"@type":"Answer","text":"NVFP4 compressed KV cache is 45% smaller than the previous FP8 KV cache format used in FlashMLA's MegaAttention kernel. MegaAttention performs query RoPE, sparse attention, inverse RoPE, and FP8 casting in one kernel launch, improving kernel efficiency by 1.45x through fusion that removes HBM writes between operations, which also increases per-GPU concurrency in high-throughput serving. Teams comparing KV cache compression strategies for GPU memory efficiency can follow developments like this on daily.dev."}},{"@type":"Question","name":"What throughput gains did vLLM achieve serving DeepSeek-V4.1-Flash on the SemiAnalysis AgentX benchmark?","acceptedAnswer":{"@type":"Answer","text":"vLLM achieved a 1.9x speedup in low-latency serving and roughly a 5.3x throughput improvement under a 150 TPS constraint, compared to its day-0 implementation, within three weeks of the model's release. Low-latency serving used TP4 with FlashInfer attention, while high-throughput serving switched to DEP2 with DP attention and MegaAttention's NVFP4 KV compression to boost per-GPU concurrency. Developers benchmarking agentic LLM serving stacks can stay on top of results like these through daily.dev."}}]}
```

