AI Security

Prompt Cost Optimizer

Compare prompt variants by input and output token cost at your actual request volume, including the effect of prompt caching, and see what each change is worth per month.

Last reviewed by the Radiatus Cloud team

Comparison appears here.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Prompt cost is dominated by volume, not by length

A prompt that is two hundred tokens longer than it needs to be costs almost nothing on one request and a substantial amount at ten million. The decision about whether a longer prompt is worth its cost cannot be made by looking at the prompt, only by multiplying it by the request volume, and teams reliably underestimate this because they develop against a handful of test calls. The same arithmetic run against production volume frequently changes which variant is preferred.

Output tokens usually cost several times more than input

Most providers price output tokens at three to five times input tokens, so a change that shortens the response is worth several times as much per token as one that shortens the prompt. Teams optimise the prompt because it is the part they wrote and can see, while the larger saving sits in instructing the model to be more concise or in constraining the output format. Checking the ratio in your own pricing before optimising anything is a minute of work that redirects the effort.

Caching changes the arithmetic entirely

Where a provider offers prompt caching, the stable prefix of a prompt is charged at a large discount on every request after the first, which inverts the usual advice: a long fixed system prompt followed by a short variable part becomes far cheaper than a shorter prompt that varies throughout. Restructuring a prompt so the invariant material sits at the front, unchanged between requests, is often a larger saving than deleting anything from it.

Related tools

Frequently Asked Questions

Should I optimise the prompt or the response?

Check your ratio first. Most providers price output at three to five times input, so shortening the response is usually worth several times as much per token as shortening the prompt.

Does prompt caching change what I should do?

Substantially. A cached stable prefix is charged at a large discount after the first request, so a long fixed prefix followed by a short variable part can be cheaper than a shorter prompt that varies throughout.

How accurate are the token estimates?

They are approximations from characters per token, which varies by tokeniser, language and content. Code and non-Latin scripts use more tokens per character, so use your provider’s counter for anything you will commit to.

Is the cheapest variant the right one?

Only if quality is equal, and this measures cost alone. A variant that is cheaper and worse is a false economy, so pair this with an evaluation that measures whether the shorter prompt still works.

Why compare at monthly volume?

Because that is where the difference is visible. A change worth a fraction of a penny per request is worth thousands a month at scale, and the same change is invisible in development.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter your prompt variants and pricing to compare them at scale.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.