Prompt Caching Savings Calculator
Calculate the cost savings from prompt caching when a large shared prompt prefix is reused across many requests at a discounted cached rate.
Last reviewed by the Radiatus Cloud team
Calculate cost savings from caching a reused prompt prefix.
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Calculate prompt caching savings
Prompt caching lets a model reuse a previously processed prompt prefix, such as a long system prompt or document context, at a steeply discounted rate on subsequent requests. This calculator estimates the savings when a shared prefix of a given token count is reused across many requests, comparing the full price with the discounted cached price. If eight thousand cached tokens are reused across ten thousand requests, the savings can be substantial because the cached rate is often a small fraction of the full rate.
The savings scale with the size of the cached prefix and the number of requests that reuse it.
When caching pays off
Prompt caching is most valuable for applications that send the same large context repeatedly, such as chatbots with long system prompts, document question answering, and tools that include extensive instructions on every call. Because the cached rate is typically far below the full rate, the savings on the reused portion can dominate the bill. The first request usually pays to populate the cache, after which reuses are cheap.
Only the shared prefix is cached, so the variable part of each request is still billed at the normal rate. All calculation happens locally in your browser.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
What is prompt caching?
It reuses a previously processed prompt prefix at a discounted rate on later requests, avoiding reprocessing the same shared context each time.
When is it worth using?
When many requests share a large common prefix, such as a long system prompt or document, so the discounted reuse saves on most of the input.
Does the whole prompt get cached?
Only the shared prefix. The variable part of each request is billed at the normal rate, so savings apply to the reused portion.
Is there a cost to populate the cache?
Usually the first request pays the normal rate to create the cache, after which subsequent reuses are charged at the lower cached rate.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter the cached tokens, requests, and the full and cached prices.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.