<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r" -->

---
title: Rethinking KV Caching For Production Inference | daily.dev
description: Stanford research found ~62% of tokens sent to AI agents on every call is repeated content, yet most inference systems recompute KV vectors from scratch each...
canonical: https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Rethinking KV Caching For Production Inference | daily.dev
og:description: Stanford research found ~62% of tokens sent to AI agents on every call is repeated content, yet most inference systems recompute KV vectors from scratch each...
og:url: https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r
og:image: https://api.daily.dev/og/posts/1Px3UV98R.png
og:image:alt: Rethinking KV Caching For Production Inference
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Rethinking KV Caching For Production Inference

**[Daily Dose of Data Science \| Avi Chawla \| Substack](https://daily.dev/sources/dailydoseofds)** · 9 min read · 0 upvotes · 0 comments

## Summary

Stanford research found ~62% of tokens sent to AI agents on every call is repeated content, yet most inference systems recompute KV vectors from scratch each time. Prefix caching helps but has a hard ceiling: any change to the cached prefix causes a full miss, breaking RAG multi-document queries, reordered documents, and growing conversation histories. LMCache is an open-source project (10k+ stars) that disaggregates cache management into a separate process, eliminating resource contention with inference. It uses shared GPU memory, zero-copy cross-GPU sharing, and parallel multi-tier loading (GPU, CPU, SSD, remote storage). On H200 GPUs with Qwen3-235B and 50 concurrent users, it delivers 14x faster time-to-first-token and 4x faster decoding. The companion research CacheBlend (EuroSys 2025 Best Paper) solves the prefix-matching limitation by selectively recomputing only the small fraction of tokens with cross-document attention, giving 2-4x faster multi-document processing with no quality loss. LMCache ships with Prometheus/OpenTelemetry, a Kubernetes operator, and fault-tolerant failover.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blog.dailydoseofds.com/p/rethinking-kv-caching-for-production>

## Questions this post answers

### Why does prefix caching fail for multi-document RAG queries?

Prefix caching requires the cached portion to be an exact byte-for-byte prefix of the new request, so it fails when a query needs multiple documents cached independently. If document A and document B were each cached alone, combining them causes a cache miss because the second document's cached KV state was computed without awareness of the first, and the same problem occurs when document order changes or conversation history grows.

_Anyone architecting RAG pipelines can track caching techniques like these as they evolve on daily.dev._

### How much faster is LMCache compared to in-process KV caching?

On H200 GPUs running the Qwen3-235B model with 50 concurrent users, LMCache delivers 14x faster time-to-first-token and 4x faster decoding compared to in-process caching, with startup time dropping from over 3 minutes to about 30 seconds. This comes from running cache management as a separate process so it never contends with inference for GPU resources.

_Teams evaluating inference infrastructure can follow benchmarks like this on daily.dev before choosing a caching approach._

### What is CacheBlend and how does it speed up multi-document queries?

CacheBlend is a technique from the LMCache team that won the EuroSys 2025 Best Paper Award, addressing the problem of combining independently cached documents. It identifies the small fraction of tokens with strong cross-document attention connections and selectively recomputes only those, reusing everything else from independent caches, giving 2 to 4x faster processing for multi-document RAG queries without quality loss.

_Developers building RAG systems can keep up with techniques like CacheBlend through daily.dev._

## Similar posts on daily.dev

- [GKE Inference Gateway prefix caching accelerates AI inference](https://daily.dev/posts/gke-inference-gateway-prefix-caching-accelerates-ai-inference-qyjrx3igz) · Google Cloud · 1 upvotes · 0 comments
- [The Complete Guide to Inference Caching in LLMs](https://daily.dev/posts/the-complete-guide-to-inference-caching-in-llms-uh6o61eic) · Machine Learning Mastery · 1 upvotes · 0 comments
- [The KV Cache Tax: Why Inference Servers Run Out of Memory Before Compute](https://daily.dev/posts/the-kv-cache-tax-why-inference-servers-run-out-of-memory-before-compute-yzoonznpj) · Towards Data Science · 0 upvotes · 0 comments
- [The Inference Tax: How Prefix-Aware Routing Eliminates the Hidden Cost of LLMs at Scale](https://daily.dev/posts/the-inference-tax-how-prefix-aware-routing-eliminates-the-hidden-cost-of-llms-at-scale-eej7jbkjc) · DigitalOcean · 0 upvotes · 0 comments
- [KV vs Prefix vs Prompt vs Semantic Caching](https://daily.dev/posts/kv-vs-prefix-vs-prompt-vs-semantic-caching-k9hzdns5l) · Daily Dose of Data Science \| Avi Chawla \| Substack · 0 upvotes · 0 comments

---

Tags: [#rag](https://daily.dev/tags/rag), [#ai-inference](https://daily.dev/tags/ai-inference), [#vllm](https://daily.dev/tags/vllm)

[View this post on daily.dev](https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Rethinking KV Caching For Production Inference","url":"https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r"},"datePublished":"2026-07-07T16:56:57.534Z","dateModified":"2026-09-14T08:00:38.553Z","description":"Stanford research found ~62% of tokens sent to AI agents on every call is repeated content, yet most inference systems recompute KV vectors from scratch each...","image":"https://substackcdn.com/image/fetch/$s_!fFcv!,w_1200,h_675,c_fill,f_jpg,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3469bd3a-08a1-4389-89fb-59434b6cb71c_960x671.gif","thumbnailUrl":"https://substackcdn.com/image/fetch/$s_!fFcv!,w_1200,h_675,c_fill,f_jpg,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3469bd3a-08a1-4389-89fb-59434b6cb71c_960x671.gif","isAccessibleForFree":true,"articleSection":"Daily Dose of Data Science | Avi Chawla | Substack","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Daily Dose of Data Science | Avi Chawla | Substack","logo":"https://media.daily.dev/image/upload/s--4IHQgTOw--/f_auto/v1710503712/logos/dailydoseofds","url":"https://daily.dev/sources/dailydoseofds"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"rag,ai-inference,vllm","timeRequired":"PT9M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Daily Dose of Data Science | Avi Chawla | Substack","item":"https://daily.dev/sources/dailydoseofds"},{"@type":"ListItem","position":3,"name":"Rethinking KV Caching For Production Inference"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/rethinking-kv-caching-for-production-inference-1px3uv98r#faq","mainEntity":[{"@type":"Question","name":"Why does prefix caching fail for multi-document RAG queries?","acceptedAnswer":{"@type":"Answer","text":"Prefix caching requires the cached portion to be an exact byte-for-byte prefix of the new request, so it fails when a query needs multiple documents cached independently. If document A and document B were each cached alone, combining them causes a cache miss because the second document's cached KV state was computed without awareness of the first, and the same problem occurs when document order changes or conversation history grows. Anyone architecting RAG pipelines can track caching techniques like these as they evolve on daily.dev."}},{"@type":"Question","name":"How much faster is LMCache compared to in-process KV caching?","acceptedAnswer":{"@type":"Answer","text":"On H200 GPUs running the Qwen3-235B model with 50 concurrent users, LMCache delivers 14x faster time-to-first-token and 4x faster decoding compared to in-process caching, with startup time dropping from over 3 minutes to about 30 seconds. This comes from running cache management as a separate process so it never contends with inference for GPU resources. Teams evaluating inference infrastructure can follow benchmarks like this on daily.dev before choosing a caching approach."}},{"@type":"Question","name":"What is CacheBlend and how does it speed up multi-document queries?","acceptedAnswer":{"@type":"Answer","text":"CacheBlend is a technique from the LMCache team that won the EuroSys 2025 Best Paper Award, addressing the problem of combining independently cached documents. It identifies the small fraction of tokens with strong cross-document attention connections and selectively recomputes only those, reusing everything else from independent caches, giving 2 to 4x faster processing for multi-document RAG queries without quality loss. Developers building RAG systems can keep up with techniques like CacheBlend through daily.dev."}}]}
```

