I burned all my tokens researching how to save tokens

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

After burning an entire Claude team plan token limit in 30 minutes using /deep-research, the author built a custom multi-agent research pipeline that runs roughly 10x longer without extra cost. The approach assigns specific roles to different models (Sonnet for finding, Opus for verifying, Fable only for planning), uses shared memory across Claude, Codex, and Antigravity subscriptions via a small bash wrapper, and places /deep-research at the end of the pipeline rather than the beginning. To reduce hallucinations, strict verification rules require every claim to have a URL and primary source quote, with the finder and verifier always being different agents. Key findings from the resulting knowledge base include: harness choice can cause a 66x token spread for the same model, context compaction can double costs through re-compaction loops, mid-session tool schema changes silently invalidate prompt caches, and real invoices can be 7-11x higher than naive token estimates.

10m read timeFrom quesma.com
Post cover image
Table of contents
Using every subscription I already pay forCheaper models as subagentsReducing hallucinations/deep-research goes lastThe result: a knowledge base I can trustA few findings from the knowledge baseBefore you burn your next limit
13 Impressions