<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap" -->

---
title: LLM Cost Optimization Techniques: A Production Playbook
description: A vendor-neutral playbook for reducing LLM production costs organized around three levers: caching (provider prompt caching, semantic response caching, KV...
canonical: https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: LLM Cost Optimization Techniques: A Production Playbook | daily.dev
og:description: A vendor-neutral playbook for reducing LLM production costs organized around three levers: caching (provider prompt caching, semantic response caching, KV...
og:url: https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap
og:image: https://api.daily.dev/og/posts/8Dc4ADyaP.png
og:image:alt: LLM Cost Optimization Techniques: A Production Playbook
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM Cost Optimization Techniques: A Production Playbook

**[BigData Boutique blog](https://daily.dev/sources/bigdataboutique)** · 11 min read · 3 upvotes · 1 comments

## Summary

A vendor-neutral playbook for reducing LLM production costs organized around three levers: caching (provider prompt caching, semantic response caching, KV caching), model routing (static routing, dynamic routing, cascading via RouteLLM and FrugalGPT), and token compression (trimming system prompts, tighter RAG retrieval, capping output tokens). It also covers batch APIs for a flat 50% discount on non-interactive workloads, the fine-tune-vs-prompt and self-host-vs-API architecture decisions, and governance practices like per-team budgets and chargeback. Instrumentation - tracking cost per feature and per resolved task - is framed as the essential first step before optimizing.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://bigdataboutique.com/blog/llm-cost-optimization-techniques>

## Questions this post answers

### How much discount does Anthropic prompt caching give and what does the cache write cost?

A cache read on Anthropic's API costs 0.1x the base input price, a 90% discount, while the initial cache write costs 1.25x for a 5-minute TTL or 2x for a 1-hour TTL. Because of the write premium, caching pays off after just a single cache read on the short TTL, making it worthwhile for any prompt reused more than once.

_Anyone tuning prompt caching strategy can track pricing details like this on daily.dev._

### How much can semantic caching reduce LLM API calls on repetitive workloads?

Semantic response caching, which returns a stored response when a new query is close enough in vector space to a prior one, has been reported to reduce API calls by up to 68.8% on repetitive query workloads, with hit rates above 97%. The open-source GPTCache library implements this with pluggable vector backends like Redis, PostgreSQL, Milvus, and FAISS, and similarity thresholds are typically tuned conservatively around 0.75 to 0.85 cosine similarity to avoid false-positive matches.

_Teams evaluating caching layers for LLM apps compare approaches like this on daily.dev._

### How much can LLM model routing and cascading reduce inference costs without hurting accuracy?

The RouteLLM framework from LMSYS reports reaching 95% of GPT-4 performance while cutting cost by over 85% on MT Bench by routing only hard queries to the strongest model. Separately, the FrugalGPT approach from Stanford, which cascades from a cheap model to a stronger one only when confidence is low, reports matching the best individual LLM's accuracy with up to 98% cost reduction.

_Developers deciding between static routing and cascading weigh trade-offs like these on daily.dev._

## Community discussion

Top comments from developers on daily.dev.

**@agustinbarrientos** · 0 upvotes

> I'd replay near-neighbor queries that require different answers across policy, tenant, or freshness boundaries

## Similar posts on daily.dev

- [LLM Cost Optimization Strategies That Cut 60-90%](https://daily.dev/posts/llm-cost-optimization-strategies-that-cut-60-90--wv1vkltn8) · portkey · 0 upvotes · 0 comments
- [The systems guide to production token optimization](https://daily.dev/posts/the-systems-guide-to-production-token-optimization-y3mhhtlgk) · The New Stack · 0 upvotes · 0 comments
- [How To Reduce LLM Token Costs by 70–90%](https://daily.dev/posts/how-to-reduce-llm-token-costs-by-70-90--lqaosovr2) · C\# Corner · 1 upvotes · 0 comments
- [LLM routing in production: Choosing the right model for every request](https://daily.dev/posts/llm-routing-in-production-choosing-the-right-model-for-every-request-zj72a5hr3) · LogRocket · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#finops](https://daily.dev/tags/finops), [#ai-inference](https://daily.dev/tags/ai-inference), [#ai-gateway](https://daily.dev/tags/ai-gateway)

[View this post on daily.dev](https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"LLM Cost Optimization Techniques: A Production Playbook","url":"https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap"},"datePublished":"2026-09-02T14:06:51.116Z","dateModified":"2026-09-02T14:12:27.774Z","description":"A vendor-neutral playbook for reducing LLM production costs organized around three levers: caching (provider prompt caching, semantic response caching, KV...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1848adafe8e99840944d44349f4cba12?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1848adafe8e99840944d44349f4cba12?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"BigData Boutique blog","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"BigData Boutique blog","logo":"https://media.daily.dev/image/upload/s--3BDyon-q--/f_auto/v1717941801/logos/bigdataboutique","url":"https://daily.dev/sources/bigdataboutique"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"llm,finops,ai-inference,ai-gateway","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"BigData Boutique blog","item":"https://daily.dev/sources/bigdataboutique"},{"@type":"ListItem","position":3,"name":"LLM Cost Optimization Techniques: A Production Playbook"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap","comment":[{"@type":"Comment","text":"I’d replay near-neighbor queries that require different answers across policy, tenant, or freshness boundaries","datePublished":"2026-09-03T17:30:18.884Z","url":"https://daily.dev/posts/8Dc4ADyaP#c-UCUNirrcx","author":{"@type":"Person","name":"Agustin Barrientos","url":"https://daily.dev/agustinbarrientos","image":"https://media.daily.dev/image/upload/s--5ayxQnqn--/f_auto/v1788281802/avatars/avatar_wQYYVe5Tbj0NJ7C7qPoa8?_a=BAMAMicg0"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/llm-cost-optimization-techniques-a-production-playbook-8dc4adyap#faq","mainEntity":[{"@type":"Question","name":"How much discount does Anthropic prompt caching give and what does the cache write cost?","acceptedAnswer":{"@type":"Answer","text":"A cache read on Anthropic's API costs 0.1x the base input price, a 90% discount, while the initial cache write costs 1.25x for a 5-minute TTL or 2x for a 1-hour TTL. Because of the write premium, caching pays off after just a single cache read on the short TTL, making it worthwhile for any prompt reused more than once. Anyone tuning prompt caching strategy can track pricing details like this on daily.dev."}},{"@type":"Question","name":"How much can semantic caching reduce LLM API calls on repetitive workloads?","acceptedAnswer":{"@type":"Answer","text":"Semantic response caching, which returns a stored response when a new query is close enough in vector space to a prior one, has been reported to reduce API calls by up to 68.8% on repetitive query workloads, with hit rates above 97%. The open-source GPTCache library implements this with pluggable vector backends like Redis, PostgreSQL, Milvus, and FAISS, and similarity thresholds are typically tuned conservatively around 0.75 to 0.85 cosine similarity to avoid false-positive matches. Teams evaluating caching layers for LLM apps compare approaches like this on daily.dev."}},{"@type":"Question","name":"How much can LLM model routing and cascading reduce inference costs without hurting accuracy?","acceptedAnswer":{"@type":"Answer","text":"The RouteLLM framework from LMSYS reports reaching 95% of GPT-4 performance while cutting cost by over 85% on MT Bench by routing only hard queries to the strongest model. Separately, the FrugalGPT approach from Stanford, which cascades from a cheap model to a stronger one only when confidence is low, reports matching the best individual LLM's accuracy with up to 98% cost reduction. Developers deciding between static routing and cascading weigh trade-offs like these on daily.dev."}}]}
```

