Prompt Caching Savings Calculator
Prompt caching reuses repeated prompt prefixes instead of reprocessing them — Anthropic charges cache reads at 0.1× and OpenAI discounts cached input to 0.5×. Enter your usage to see your monthly bill with vs. without caching.
Rates last updated: October 2026. Providers change prices — re-verify before budgeting.
Share of each request's input that is a repeatable prefix (system prompt, docs, examples).
How often the cacheable prefix is actually served from cache.
$0
Monthly cost — no caching
$0
Monthly cost — with caching
$0
Saved per month
| Token bucket | Tokens / month | Rate applied | Cost |
|---|
Assumptions used
- Anthropic: cache writes cost 1.25× the base input price; cache reads cost 0.10×. A missed (uncached) prefix token is modeled as one cache write.
- OpenAI: cached input costs 0.50× the base input price; cache writes carry no surcharge.
- Output tokens are excluded — caching only affects input pricing.
- Real-world costs also depend on cache TTLs, minimum cacheable lengths, and refresh patterns; this is a planning estimate.
Rates change. The per-1M-token prices below were last updated in October 2026 and are hardcoded for offline use. Always re-verify current pricing on the provider's site before making budget decisions.