<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq" -->

---
title: RTK reports huge token savings, but our cost benchmarks...
description: A benchmark comparing token and cost outcomes with and without RTK (Rust Token Killer), a terminal-output-compressing plugin for AI coding agents, tested with...
canonical: https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: RTK reports huge token savings, but our cost benchmarks disagree | daily.dev
og:description: A benchmark comparing token and cost outcomes with and without RTK (Rust Token Killer), a terminal-output-compressing plugin for AI coding agents, tested with...
og:url: https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq
og:image: https://api.daily.dev/og/posts/7wx1vDcuq.png
og:image:alt: RTK reports huge token savings, but our cost benchmarks disagree
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# RTK reports huge token savings, but our cost benchmarks disagree

**[Quesma](https://daily.dev/sources/quesma)** · 7 min read · 0 upvotes · 0 comments

## Summary

A benchmark comparing token and cost outcomes with and without RTK (Rust Token Killer), a terminal-output-compressing plugin for AI coding agents, tested with Claude Code on Fable 5.0 and OpenCode with DeepSeek V4 Pro on Terminal-Bench 2.1. Results show mixed to negative effects: Fable's apparent savings depended almost entirely on a single task, while DeepSeek got 17% more expensive on average due to extra agent turns triggered by RTK's rewrites. The analysis argues RTK's self-reported

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://quesma.com/blog/does-rtk-make-ai-coding-cheaper>

## Questions this post answers

### Does RTK actually reduce the cost of running Claude Code or OpenCode agents?

Not reliably. Testing on Terminal-Bench 2.1 found Fable 5.0 with Claude Code was only 1% more expensive on a task-weighted basis, with almost all apparent savings coming from a single task (winning-avg-corewars), while OpenCode with DeepSeek V4 Pro cost 17% more on average with RTK enabled, mainly because RTK's output rewriting triggered more agent turns.

_Anyone deciding whether to adopt RTK for agent cost control can track benchmark writeups like this on daily.dev._

### Why does RTK's reported token savings metric (rtk gain) not match actual cost savings?

rtk gain measures raw minus filtered command output in bytes divided by four, not billed tokens, so it can massively overstate savings. In one test, two head -1 calls accounted for 69% of a comparison's reported savings by comparing limited reads against a whole file that was never going to be returned, making an actually more expensive attempt look optimized.

_Developers weighing AI-agent cost metrics can follow methodology critiques like this via daily.dev._

### Can a bug in an AI coding agent tool like RTK cause runaway costs?

Yes. A DeepSeek git-multibranch attempt using RTK 0.45.0 hit an unsupported find flag that got rewritten incorrectly on every retry, producing 339 consecutive errors before timing out and costing about 9 times as much as the matching baseline attempt that also passed. The bug was fixed in RTK 0.46.0, after the benchmark runs.

_Teams relying on agent tooling plugins can watch for gotchas like this by following coverage on daily.dev._

## Similar posts on daily.dev

- [Don’t Break the Agent: Lessons in Token Optimization](https://daily.dev/posts/don-t-break-the-agent-lessons-in-token-optimization-fgdkm8bwd) · JFrog · 1 upvotes · 1 comments
- [toktrack: Track AI CLI spending across Claude, Codex & Gemini in 40ms](https://daily.dev/posts/toktrack-track-ai-cli-spending-across-claude-codex-gemini-in-40ms-qajwlumdx) · Product Hunt · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#claude-code](https://daily.dev/tags/claude-code), [#opencode](https://daily.dev/tags/opencode)

[View this post on daily.dev](https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"RTK reports huge token savings, but our cost benchmarks disagree","url":"https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq"},"datePublished":"2026-09-11T08:50:32.283Z","dateModified":"2026-09-11T08:50:57.021Z","description":"A benchmark comparing token and cost outcomes with and without RTK (Rust Token Killer), a terminal-output-compressing plugin for AI coding agents, tested with...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/bbf05f9662be2ba71acb8f719a0f7eed?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/bbf05f9662be2ba71acb8f719a0f7eed?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Quesma","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Quesma","logo":"https://media.daily.dev/image/upload/s--I-Be0YJY--/f_auto,q_auto/v1774964372/logos/quesma?_a=BAMAMiWQ0","url":"https://daily.dev/sources/quesma"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,claude-code,opencode","timeRequired":"PT7M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Quesma","item":"https://daily.dev/sources/quesma"},{"@type":"ListItem","position":3,"name":"RTK reports huge token savings, but our cost benchmarks disagree"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/rtk-reports-huge-token-savings-but-our-cost-benchmarks-disagree-7wx1vdcuq#faq","mainEntity":[{"@type":"Question","name":"Does RTK actually reduce the cost of running Claude Code or OpenCode agents?","acceptedAnswer":{"@type":"Answer","text":"Not reliably. Testing on Terminal-Bench 2.1 found Fable 5.0 with Claude Code was only 1% more expensive on a task-weighted basis, with almost all apparent savings coming from a single task (winning-avg-corewars), while OpenCode with DeepSeek V4 Pro cost 17% more on average with RTK enabled, mainly because RTK's output rewriting triggered more agent turns. Anyone deciding whether to adopt RTK for agent cost control can track benchmark writeups like this on daily.dev."}},{"@type":"Question","name":"Why does RTK's reported token savings metric (rtk gain) not match actual cost savings?","acceptedAnswer":{"@type":"Answer","text":"rtk gain measures raw minus filtered command output in bytes divided by four, not billed tokens, so it can massively overstate savings. In one test, two head -1 calls accounted for 69% of a comparison's reported savings by comparing limited reads against a whole file that was never going to be returned, making an actually more expensive attempt look optimized. Developers weighing AI-agent cost metrics can follow methodology critiques like this via daily.dev."}},{"@type":"Question","name":"Can a bug in an AI coding agent tool like RTK cause runaway costs?","acceptedAnswer":{"@type":"Answer","text":"Yes. A DeepSeek git-multibranch attempt using RTK 0.45.0 hit an unsupported find flag that got rewritten incorrectly on every retry, producing 339 consecutive errors before timing out and costing about 9 times as much as the matching baseline attempt that also passed. The bug was fixed in RTK 0.46.0, after the benchmark runs. Teams relying on agent tooling plugins can watch for gotchas like this by following coverage on daily.dev."}}]}
```

