Claude Haiku 5.5 pricing calculator
Which prices apply to your prompt, what one request and one month cost, and where the price changes
Anthropic prices Claude Haiku 5.5 (model ID claude-haiku-5-5) by prompt length. Its pricing page says that "a request whose prompt is over 100,000 tokens pays higher prices" and that the prompt length "counts all of its input tokens, including cache reads and cache writes". Enter the token counts of one request. The calculator adds up the prompt, shows which set of prices applies, and prices the request and the month.
Formula and assumptions
- The cache write rate is the 5 minute rate or the 1 hour rate that you select.
- One set of prices applies to the whole request. A request over the limit pays the higher rate on every token, output tokens included.
- A prompt of exactly 100,000 tokens gets the base prices. Anthropic's table says "for prompts up to 100,000 tokens" and "for prompts over 100,000 tokens".
- Each request in the month has the same token counts. If your requests differ, run the calculator once for each group.
- The cache write tokens that you enter are billed at the write rate on each request. If only the first request writes the cache, enter 0 here and price that first request alone.
- The Batch API option halves all four line items. Anthropic's batch documentation says "All usage is charged at 50% of the standard API prices."
- The rates are the Claude API rates with global routing. US-only inference (a 1.1x multiplier in Anthropic's pricing page) and partner platforms are not included.
- The calculator checks the prompt against the 1M token context window that Anthropic lists. It does not check the output limit.
- All sums use whole-number arithmetic on the per-token rates. Amounts below 1 USD are shown exactly. Amounts of 1 USD or more are rounded to the cent.
- Token counts are estimates. Actual usage depends on encoding, model version, and prompt structure.
This tool does not execute LLM calls. All computation happens locally in your browser.
Reference rates
Reference rates, verified 2026-10-10 against the OpenRouter models API (https://openrouter.ai/api/v1/models) and the Anthropic pricing page (https://platform.claude.com/docs/en/about-claude/pricing). Units are USD per token. The two sources give the same value for each rate.
| Token type | Prompt up to 100,000 tokens | Prompt over 100,000 tokens |
|---|---|---|
| Input | 0.0000001 | 0.0000005 |
| Output | 0.0000005 | 0.0000025 |
| Cache read | 0.00000001 | 0.00000005 |
| Cache write, 5 minutes | 0.000000125 | 0.000000625 |
| Cache write, 1 hour | 0.0000002 | 0.000001 |
| Token type | Prompt up to 100,000 tokens | Prompt over 100,000 tokens |
|---|---|---|
| Input | 0.00000005 | 0.00000025 |
| Output | 0.00000025 | 0.00000125 |
| Cache read | 0.000000005 | 0.000000025 |
| Cache write, 5 minutes | 0.0000000625 | 0.0000003125 |
| Cache write, 1 hour | 0.0000001 | 0.0000005 |
LLM pricing shown is based on hardcoded reference values and may not reflect current provider rates. Always verify against official pricing pages before making budget decisions.
What the sources say
Each rule in the calculator comes from one of these passages. All were read on 2026-10-10.
Claude Haiku 5.5 is priced by prompt length: a request whose prompt is over 100,000 tokens pays higher prices. A request's prompt length counts all of its input tokens, including cache reads and cache writes. Each request is priced on its own: a request over the threshold pays the higher prices even when part of its prompt is a cache hit, and earlier requests keep the prices they were billed at.Anthropic pricing page, Long context pricing, platform.claude.com/docs/en/about-claude/pricing
Claude Haiku 5.5 (for prompts up to 100,000 tokens): base input $0.10 / MTok, 5m cache writes $0.125 / MTok, 1h cache writes $0.20 / MTok, cache hits and refreshes $0.01 / MTok, output $0.50 / MTok. Claude Haiku 5.5 (for prompts over 100,000 tokens): base input $0.50 / MTok, 5m cache writes $0.625 / MTok, 1h cache writes $1 / MTok, cache hits and refreshes $0.05 / MTok, output $2.50 / MTok.Anthropic pricing page, Model pricing table, two rows written out with their column names. MTok is Anthropic's unit for one million tokens.
Claude Haiku 5.5 (for prompts up to 100,000 tokens): batch input $0.05 / MTok, batch output $0.25 / MTok. Claude Haiku 5.5 (for prompts over 100,000 tokens): batch input $0.25 / MTok, batch output $1.25 / MTok.Anthropic pricing page, Batch processing table, two rows written out with their column names
The Batches API offers significant cost savings. All usage is charged at 50% of the standard API prices.Anthropic batch processing guide, platform.claude.com/docs/en/build-with-claude/batch-processing
The OpenRouter models API has two entries, anthropic/claude-haiku-5.5 and anthropic/claude-haiku-5.5:batch. Each entry gives the five base rates in USD per token and an override with "min_prompt_tokens": 100000 that carries the five higher rates. The values in the two tables above are those strings. Note one difference in wording. The OpenRouter field name reads as a minimum of 100,000 tokens, and Anthropic says "over 100,000 tokens". This calculator follows Anthropic, the provider that bills the request.
How the price changes at 100,000 tokens
The change is a step, not a slope. Below the limit, each added uncached input token costs one input rate. The token that takes the prompt past 100,000 changes the rate of every token in the request. The result panel shows the size of that step for your own request: it prices the same request at a prompt of 100,001 tokens, or at exactly 100,000 tokens if you are already over.
Cached tokens do not help you stay under the limit. A cache hit is cheaper than uncached input, but it has the same weight in the prompt length. A request with a small uncached part and a large cached prefix can be over the limit, and then its cache reads, its cache writes and its output all use the higher rates.
Count tokens with the model itself. Anthropic's Claude Haiku 5.5 overview says it "uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5". A prompt that you measured on an older model can be longer on this one.
Questions
Does the higher price apply only to the tokens above 100,000?
No. The pricing page says "a request over the threshold pays the higher prices". The calculator applies the higher rates to the whole request.
Do cache reads and cache writes count toward the 100,000 tokens?
Yes. The pricing page says the prompt length "counts all of its input tokens, including cache reads and cache writes".
Does one long request change the price of other requests?
No. The pricing page says "Each request is priced on its own". Only the requests that are over the limit pay the higher prices.
Does the Batch API discount apply at both price levels?
Yes. Anthropic's batch table has a row for prompts up to 100,000 tokens and a row for prompts over 100,000 tokens.
Are these prices live?
No. They are reference rates, verified 2026-10-10. Check the two sources above before you set a budget.
Related tools
This page does not model retries or failed calls. Use the retry calculators for that. The other tools cover cache savings and batch savings on other Claude models.
ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts. Tool inputs stay in your browser. Page and product-interest analytics is optional and requires consent. Part of the Zovo Tools network by Michael Lip.