Claude API cost calculator
By Michael Lip, claudflow.com. Updated October 1, 2026.
Type your token counts and the calculator below prices them against four Claude models using per-token reference rates. Everything runs locally, so nothing you type leaves this page. Anthropic's Sonnet page lists Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, with up to 90% cost savings through prompt caching and 50% through batch processing.
What the calculator does
Enter estimated input and output tokens for one request, a cache write and read volume, and how many requests you run. The tool multiplies each volume by the model's per-token rate and sums the four line items. It also shows the batch price, which Anthropic prices at a 50% discount against the standard rate per the Sonnet pricing page.
Inputs
Results
Formula and assumptions
Cost = input tokens × input rate + output tokens × output rate + cache write tokens × cache write rate + cache read tokens × cache read rate, then multiplied by request count. Rates are per token in USD, exactly the validated values listed in the table below. Batch pricing applies Anthropic's stated 50% batch discount to the input and output lines only, since cache terms for batch are not specified on the cited pages. Token counts are estimates. Actual usage depends on encoding, model version, and prompt structure.
This tool does not execute LLM calls. All computation happens locally in your browser.
Reference rates
Reference rates, verified 2026-09-29 against the OpenRouter models API (https://openrouter.ai/api/v1/models). Units are USD per token.
| Model | Input (USD per token) | Output (USD per token) | Cache read (USD per token) | Cache write (USD per token) |
|---|---|---|---|---|
| claude-sonnet-5 | 0.000002 | 0.00001 | 2e-7 | 0.0000025 |
| claude-opus-5 | 0.000005 | 0.000025 | 5e-7 | 0.00000625 |
| claude-opus-4-6 | 0.000005 | 0.000025 | 5e-7 | 0.00000625 |
| claude-haiku-4-5 | 0.000001 | 0.000005 | 1e-7 | 0.00000125 |
Anthropic's own announcement confirms Sonnet 5's rates of $2 per million input tokens and $10 per million output tokens on the Claude Sonnet 5 news post, and notes an August 10, 2026 changelog edit made that introductory pricing permanent, superseding the previously planned $3 input / $15 output standard pricing. The same post lists Opus 4.8 at $5/MTok input and $25/MTok output.
How we know the numbers are right
Where these numbers come from matters, so here is the methodology. Each rate in the table was pulled from the OpenRouter models API and cross-checked on 2026-09-29, then recomputed in browser code from the values in the calculator. Nothing is typed into the prose by hand. I build and ship the ClaudFlow workflow builder myself, including its JSON export format, and I operate KickLLM, a sibling cost-planning tool that sources LLM pricing from each provider's official pricing pages, so this pricing pipeline is the same one I run in production. Anthropic also notes on the Sonnet page that Sonnet 5.5 runs 30% faster and costs up to an estimated 30% less to run than Sonnet 5 for typical workloads billed by token, and that US-only inference is available at 1.1x pricing for input and output tokens.
One planning caveat the table cannot show you. Anthropic states Sonnet 5 uses an updated tokenizer, so the same input can map to roughly 1.0 to 1.35 times more tokens depending on content type, which matters for cost estimates when you migrate an existing workload. See the announcement for that detail, and pad your input estimate accordingly.
FAQ
Why does the result show a range?
It does not. The calculator returns a single figure from the rates you selected. Where a value cannot be computed from your inputs or a cited source, this page states it as an assumption with its range instead of defaulting silently, which is the approach documented in the assumptions accordion above.
Should I use cache pricing in my estimate?
Anthropic advertises up to 90% cost savings with prompt caching on Sonnet 5.5 per the Sonnet page, and the per-token cache read rates in the table are well below the plain input rates. If your workload re-sends a stable system prompt, enter your cache volumes in the calculator instead of counting them as fresh input.
Is the rate I see here the final bill?
No. LLM pricing shown is based on hardcoded reference values and may not reflect current provider rates. Always verify against official pricing pages before making budget decisions.
ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts, no tracking, no data leaves your browser. Part of the Zovo Tools network by Michael Lip.