<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh" -->

---
title: Anthropic Introduces Prompt Caching to Reduce Latency...
description: Anthropic has introduced prompt caching in the Claude AI models, significantly reducing costs and latency for developers. Available in public beta for Claude...
canonical: https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Anthropic Introduces Prompt Caching to Reduce Latency and Costs for Claude Models | daily.dev
og:description: Anthropic has introduced prompt caching in the Claude AI models, significantly reducing costs and latency for developers. Available in public beta for Claude...
og:url: https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh
og:image: https://api.daily.dev/og/posts/EBwqqj0eH.png
og:image:alt: Anthropic Introduces Prompt Caching to Reduce Latency and Costs for Claude Models
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Anthropic Introduces Prompt Caching to Reduce Latency and Costs for Claude Models

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Anthropic has introduced prompt caching in the Claude AI models, significantly reducing costs and latency for developers. Available in public beta for Claude 3.5 Sonnet and Claude 3 Haiku models, prompt caching allows reuse of frequently used prompts, cutting down reprocessing needs. Early reports show 90% cost and 85% latency reductions for long prompts. The cache has a 5-minute lifetime and operates effectively for long instructions and embedded documents, though developers must consider security implications.

## Content

Anthropic has rolled out a new prompt caching feature in the API for its Claude family of generative AI models, aimed at significantly decreasing costs and latency for developers. Available now in public beta for Claude 3.5 Sonnet and Claude 3 Haiku models, prompt caching allows frequently used prompts to be saved and reused, which cuts down on the need to continually reprocess these prompts. Future support for Claude 3 Opus is also planned.

This new capability promises remarkable efficiency improvements, offering cost reductions of up to 90% and latency reductions of up to 85% for long prompts. Cached contexts can be reused in future API calls, which is especially useful for applications such as coding assistants, long-form document processing, and detailed instruction sets. While there is a 25% increase in cost for writing to the cache, using cached content costs 10% less than generating new prompts.

Developers have shared early reports of notable speed and cost improvements with the caching feature. The system works particularly well for long instructions and embedding documents within prompts. A notable feature of the cache is its 5-minute lifetime, which is refreshed upon each use, ensuring that frequently accessed prompts remain available.

However, the new capability isn't without potential concerns. There are some security implications related to the sharing of cached prompts, which developers will need to address. Nevertheless, the dramatic efficiency gains make prompt caching a compelling addition for those using Claude 3.5 Sonnet and Claude 3 Haiku in their applications.

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments
- [CNCF Unveils Schedule for KubeCon \+ CloudNativeCon Europe 2026](https://daily.dev/posts/cncf-unveils-schedule-for-kubecon-cloudnativecon-europe-2026-ikhcoa5cb) · CNCF · 2 upvotes · 0 comments
- [CNCF Debuts KubeCon \+ CloudNativeCon Japan 2026 Schedule](https://daily.dev/posts/cncf-debuts-kubecon-cloudnativecon-japan-2026-schedule-xp5pyudub) · CNCF · 1 upvotes · 0 comments

---

Tags: [#tech-news](https://daily.dev/tags/tech-news), [#ai](https://daily.dev/tags/ai), [#cloud](https://daily.dev/tags/cloud), [#machine-learning](https://daily.dev/tags/machine-learning)

[View this post on daily.dev](https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Anthropic Introduces Prompt Caching to Reduce Latency and Costs for Claude Models","url":"https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh"},"datePublished":"2024-08-15T19:46:05.911Z","dateModified":"2024-08-15T21:19:48.227Z","description":"Anthropic has introduced prompt caching in the Claude AI models, significantly reducing costs and latency for developers. Available in public beta for Claude...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/07acc1ed263f36ff7f50eacf1771497d?_a=AQAEuiZ","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/07acc1ed263f36ff7f50eacf1771497d?_a=AQAEuiZ","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/anthropic-introduces-prompt-caching-to-reduce-latency-and-costs-for-claude-models-ebwqqj0eh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"tech-news,ai,cloud,machine-learning","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Anthropic Introduces Prompt Caching to Reduce Latency and Costs for Claude Models"}]}
```

