Claude Haiku 5.5 pricing calculator

Which prices apply to your prompt, what one request and one month cost, and where the price changes

By Michael Lip, claudflow.com. Updated October 10, 2026.

Anthropic prices Claude Haiku 5.5 (model ID claude-haiku-5-5) by prompt length. Its pricing page says that "a request whose prompt is over 100,000 tokens pays higher prices" and that the prompt length "counts all of its input tokens, including cache reads and cache writes". Enter the token counts of one request. The calculator adds up the prompt, shows which set of prices applies, and prices the request and the month.

Input tokens that are not read from or written to the cache.
Cache hits. They count toward the prompt length.
Tokens written to the cache. They count toward the prompt length.
Sets the price of the cache write tokens.
Output does not count toward the prompt length, but its price follows it.
Each request is assumed to have the token counts above.
Prompt length (uncached input + cache read + cache write)-
Prices that apply-
Distance to the 100,000 token limit-
Batch API-
Uncached input, one request-
Cache read, one request-
Cache write, one request-
Output, one request-
Cost of one request-
Cost per month-
-
Formula and assumptions
prompt = uncached input + cache read + cache write prices = base if prompt <= 100,000 higher if prompt > 100,000 request cost = uncached input * input rate + cache read * cache read rate + cache write * cache write rate + output * output rate batch: each rate * 0.5 month cost = request cost * requests

This tool does not execute LLM calls. All computation happens locally in your browser.

Reference rates

Reference rates, verified 2026-10-10 against the OpenRouter models API (https://openrouter.ai/api/v1/models) and the Anthropic pricing page (https://platform.claude.com/docs/en/about-claude/pricing). Units are USD per token. The two sources give the same value for each rate.

claude-haiku-5-5, standard prices, USD per token
Token typePrompt up to 100,000 tokensPrompt over 100,000 tokens
Input0.00000010.0000005
Output0.00000050.0000025
Cache read0.000000010.00000005
Cache write, 5 minutes0.0000001250.000000625
Cache write, 1 hour0.00000020.000001
claude-haiku-5-5, Batch API prices, USD per token
Token typePrompt up to 100,000 tokensPrompt over 100,000 tokens
Input0.000000050.00000025
Output0.000000250.00000125
Cache read0.0000000050.000000025
Cache write, 5 minutes0.00000006250.0000003125
Cache write, 1 hour0.00000010.0000005

LLM pricing shown is based on hardcoded reference values and may not reflect current provider rates. Always verify against official pricing pages before making budget decisions.

What the sources say

Each rule in the calculator comes from one of these passages. All were read on 2026-10-10.

Claude Haiku 5.5 is priced by prompt length: a request whose prompt is over 100,000 tokens pays higher prices. A request's prompt length counts all of its input tokens, including cache reads and cache writes. Each request is priced on its own: a request over the threshold pays the higher prices even when part of its prompt is a cache hit, and earlier requests keep the prices they were billed at.Anthropic pricing page, Long context pricing, platform.claude.com/docs/en/about-claude/pricing
Claude Haiku 5.5 (for prompts up to 100,000 tokens): base input $0.10 / MTok, 5m cache writes $0.125 / MTok, 1h cache writes $0.20 / MTok, cache hits and refreshes $0.01 / MTok, output $0.50 / MTok. Claude Haiku 5.5 (for prompts over 100,000 tokens): base input $0.50 / MTok, 5m cache writes $0.625 / MTok, 1h cache writes $1 / MTok, cache hits and refreshes $0.05 / MTok, output $2.50 / MTok.Anthropic pricing page, Model pricing table, two rows written out with their column names. MTok is Anthropic's unit for one million tokens.
Claude Haiku 5.5 (for prompts up to 100,000 tokens): batch input $0.05 / MTok, batch output $0.25 / MTok. Claude Haiku 5.5 (for prompts over 100,000 tokens): batch input $0.25 / MTok, batch output $1.25 / MTok.Anthropic pricing page, Batch processing table, two rows written out with their column names
The Batches API offers significant cost savings. All usage is charged at 50% of the standard API prices.Anthropic batch processing guide, platform.claude.com/docs/en/build-with-claude/batch-processing

The OpenRouter models API has two entries, anthropic/claude-haiku-5.5 and anthropic/claude-haiku-5.5:batch. Each entry gives the five base rates in USD per token and an override with "min_prompt_tokens": 100000 that carries the five higher rates. The values in the two tables above are those strings. Note one difference in wording. The OpenRouter field name reads as a minimum of 100,000 tokens, and Anthropic says "over 100,000 tokens". This calculator follows Anthropic, the provider that bills the request.

How the price changes at 100,000 tokens

The change is a step, not a slope. Below the limit, each added uncached input token costs one input rate. The token that takes the prompt past 100,000 changes the rate of every token in the request. The result panel shows the size of that step for your own request: it prices the same request at a prompt of 100,001 tokens, or at exactly 100,000 tokens if you are already over.

Cached tokens do not help you stay under the limit. A cache hit is cheaper than uncached input, but it has the same weight in the prompt length. A request with a small uncached part and a large cached prefix can be over the limit, and then its cache reads, its cache writes and its output all use the higher rates.

Count tokens with the model itself. Anthropic's Claude Haiku 5.5 overview says it "uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5". A prompt that you measured on an older model can be longer on this one.

Questions

Does the higher price apply only to the tokens above 100,000?

No. The pricing page says "a request over the threshold pays the higher prices". The calculator applies the higher rates to the whole request.

Do cache reads and cache writes count toward the 100,000 tokens?

Yes. The pricing page says the prompt length "counts all of its input tokens, including cache reads and cache writes".

Does one long request change the price of other requests?

No. The pricing page says "Each request is priced on its own". Only the requests that are over the limit pay the higher prices.

Does the Batch API discount apply at both price levels?

Yes. Anthropic's batch table has a row for prompts up to 100,000 tokens and a row for prompts over 100,000 tokens.

Are these prices live?

No. They are reference rates, verified 2026-10-10. Check the two sources above before you set a budget.

Related tools

This page does not model retries or failed calls. Use the retry calculators for that. The other tools cover cache savings and batch savings on other Claude models.

ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts. Tool inputs stay in your browser. Page and product-interest analytics is optional and requires consent. Part of the Zovo Tools network by Michael Lip.