Per-request and per-run cache savings on Claude models, computed in your browser from verified rates
All rates below are per token, in USD. The cache write rate bills once on the first request. The tool spreads that one-time write across the request count you enter, so treat the result as a ceiling when your cache expires and rewrites often.
Reference rates, verified 2026-09-29 against the OpenRouter models API. You can check them yourself with one command.
| Model | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| claude-opus-5 | 0.000005 | 0.000025 | 5e-7 | 0.00000625 |
| claude-opus-4-6 | 0.000005 | 0.000025 | 5e-7 | 0.00000625 |
| claude-sonnet-5 | 0.000002 | 0.00001 | 2e-7 | 0.0000025 |
| claude-haiku-4-5 | 0.000001 | 0.000005 | 1e-7 | 0.00000125 |
The cache read rate is the lever. For claude-sonnet-5 it is 2e-7 USD per token against a fresh input rate of 0.000002 USD per token, a 10x gap. The estimator computes the exact per-request delta from those two rates rather than stating a fixed ratio, because the savings depend on your token mix and request count.
Keep the cached prefix stable. If your system prompt or retrieved context changes between calls, the cache misses, you pay the write rate again, and you never earn the read discount. The estimator assumes a stable prefix, so its output is a ceiling when your prompts vary.
Usually not much. Savings scale with cached token volume times the gap between the input and cache read rates. A few thousand cached tokens produce cents per thousand requests, while a long stable system prompt with retrieved context produces noticeable sums. Run your own counts rather than trusting a rule of thumb.
The cache write rate. The estimator includes it once, amortized across the request count you enter. A low request count with an expiring cache will overstate savings.
No. They are reference rates, verified 2026-09-29 against the OpenRouter models API. Pricing may not reflect current provider rates, so verify against official pricing pages before making budget decisions. The jq command above lets you check in seconds.
I build and ship the ClaudFlow workflow builder, including its eight block types (input, output, prompt, transform, condition, file, button, submit) and the JSON export format, so the agent loop patterns this estimator prices are patterns implemented in shipped code. The sibling tool KickLLM sources pricing from provider pricing pages directly, so per-token cost math is operating territory here. The whole network runs client side by design, which is why nothing you type leaves this page. More on the about page.
ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts, no tracking, no data leaves your browser. Part of the Zovo Tools network by Michael Lip. Token counts are estimates, and actual usage depends on encoding, model version, and prompt structure.