What a timestamp or unstable prefix costs you when every request pays the 12.5x cache-write rate instead of the read rate.
Writes pay the 5-minute cache-write rate; reads pay roughly a tenth of that. The first request always writes. Invalidated requests rewrite the entire prefix, so the multiplier is the blend of read and write prices across your traffic.
On Anthropic pricing a 5-minute cache write costs 12.5x a cache read for Sonnet-class models ($2.50 vs $0.20 per million tokens). A 50k-token prefix costs $0.01 per request when cached and $0.125 per request when invalidated. Opus is worse: $5.00 per million to write, 25x the read price.
Anthropic prompt caches expire after 5 minutes of inactivity. Any prefix change also invalidates the cache: one new token at the start forces a rewrite of everything downstream, even if the rest is byte-identical.
Yes. Agent stacks routinely inject timestamps, session IDs, or freshly serialized tool lists into the system prompt. A framework that prepends a per-request trace ID silently disables caching for everything below it.
After the static prefix. Put system prompt, tool definitions, and large stable context first; put timestamps, user data, and retrieved documents last. Sessions with a fixed prefix hold cache across turns; sessions that mutate it pay full price every turn.
I build and ship the ClaudFlow workflow builder, including its eight block types and JSON export format. The agent loop patterns these calculators price are patterns implemented in shipped code, not invented for a blog post. Everything runs client side by design: nothing you type leaves this page. More on the about page.
ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts, no tracking, no data leaves your browser. Part of the Zovo Tools network by Michael Lip. Token counts are estimates, and actual usage depends on encoding, model version, and prompt structure.