AI Security

Completion Token Calculator

Calculate how many completion tokens remain for a model response given the context window size, prompt tokens and a reserved margin.

Last reviewed by the Radiatus Cloud team

Calculate the maximum completion tokens left after your prompt.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Calculate available completion tokens

Large language models have a fixed context window that must hold both the prompt and the generated response. This calculator works out how many tokens remain for the completion after subtracting the prompt and an optional reserve from the total context window. If a model has a context window of a hundred and twenty-eight thousand tokens and your prompt uses four thousand, with five hundred reserved, you have about a hundred and twenty-three thousand five hundred tokens for the response.

Reserving a margin is wise because token counts are estimates and some tokens are consumed by formatting and special markers.

Avoiding truncation

If the prompt plus the requested completion exceeds the context window, the model will error or silently truncate, losing information. Calculating the available completion budget in advance lets you set the maximum output length safely and decide whether a long prompt needs trimming or summarising. It also helps when designing retrieval-augmented systems that pack context, where leaving room for the answer is essential.

The percentage of context used highlights how much headroom remains. All calculation happens locally in your browser.

Notes on these estimates

Because the completion token calculator runs entirely in your browser, nothing you enter is uploaded, so you can use it with private data safely. The figures are estimates based on the values you provide and common rules of thumb, so treat them as planning guidance rather than exact measurements, and run the tool as often as you need for free.

Related tools

Frequently Asked Questions

What is a context window?

It is the maximum number of tokens a model can process at once, shared between the prompt and the generated completion.

Why reserve extra tokens?

Token counts are estimates and some tokens go to formatting and special markers, so a reserve prevents accidentally exceeding the window.

What happens if the prompt is too long?

If the prompt plus the response exceeds the window, the model errors or truncates, so you should trim or summarise the prompt first.

How do I use the result?

Set the maximum output length to the available completion tokens, or lower, so the response fits comfortably within the context window.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the context window, prompt tokens and any reserved tokens.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.