Prompt Caching Savings Estimator

Per-request and per-run cache savings on Claude models, computed in your browser from verified rates

Uncached cost, one request-
Cached cost, one request-
One-time cache write, amortized-
Savings across all requests-
Savings as share of uncached total-
-

How the estimate is computed

All rates below are per token, in USD. The cache write rate bills once on the first request. The tool spreads that one-time write across the request count you enter, so treat the result as a ceiling when your cache expires and rewrites often.

uncached(req) = (cached + fresh) * input_rate + out * output_rate cached(req) = cached * cache_read_rate + fresh * input_rate + out * output_rate + cached * cache_write_rate / requests savings = (uncached(req) - cached(req)) * requests

Reference rates, verified 2026-09-29 against the OpenRouter models API. You can check them yourself with one command.

curl -s https://openrouter.ai/api/v1/models | jq '.data[] | select(.id=="claude-sonnet-5") | .pricing'
ModelInputOutputCache readCache write
claude-opus-50.0000050.0000255e-70.00000625
claude-opus-4-60.0000050.0000255e-70.00000625
claude-sonnet-50.0000020.000012e-70.0000025
claude-haiku-4-50.0000010.0000051e-70.00000125

Why the cache read rate moves the total

The cache read rate is the lever. For claude-sonnet-5 it is 2e-7 USD per token against a fresh input rate of 0.000002 USD per token, a 10x gap. The estimator computes the exact per-request delta from those two rates rather than stating a fixed ratio, because the savings depend on your token mix and request count.

Keep the cached prefix stable. If your system prompt or retrieved context changes between calls, the cache misses, you pay the write rate again, and you never earn the read discount. The estimator assumes a stable prefix, so its output is a ceiling when your prompts vary.

Frequently asked

Does caching help with short prompts?

Usually not much. Savings scale with cached token volume times the gap between the input and cache read rates. A few thousand cached tokens produce cents per thousand requests, while a long stable system prompt with retrieved context produces noticeable sums. Run your own counts rather than trusting a rule of thumb.

What rate applies on the first call?

The cache write rate. The estimator includes it once, amortized across the request count you enter. A low request count with an expiring cache will overstate savings.

Are these prices live?

No. They are reference rates, verified 2026-09-29 against the OpenRouter models API. Pricing may not reflect current provider rates, so verify against official pricing pages before making budget decisions. The jq command above lets you check in seconds.

About the author

I build and ship the ClaudFlow workflow builder, including its eight block types (input, output, prompt, transform, condition, file, button, submit) and the JSON export format, so the agent loop patterns this estimator prices are patterns implemented in shipped code. The sibling tool KickLLM sources pricing from provider pricing pages directly, so per-token cost math is operating territory here. The whole network runs client side by design, which is why nothing you type leaves this page. More on the about page.

ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts, no tracking, no data leaves your browser. Part of the Zovo Tools network by Michael Lip. Token counts are estimates, and actual usage depends on encoding, model version, and prompt structure.

Agent Loop Cost Calculator Prompt Caching Savings Opus vs Sonnet Per Turn All Tools