LLM Context Window Packer
Work out which retrieved documents fit inside a context budget after the system prompt and reserved output, which get dropped, and where each one lands relative to the positions models attend to worst.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
The context window is not the budget
A model advertising a 128,000-token window does not offer 128,000 tokens for retrieved documents. The system prompt, the conversation so far, the tool definitions and the space reserved for the response all come out first, and tool definitions in particular are easy to forget because nobody writes them by hand. Teams routinely discover the real budget when a request fails in production, at which point the failure is a hard error rather than a degradation, because exceeding the window is not something the model can partially accommodate.
Position within the context matters measurably
Models attend unevenly across a long context. The well-replicated finding is a U-shaped curve: material at the beginning and at the end is used substantially more reliably than material in the middle, and the effect grows with context length. This has a direct practical consequence for retrieval-augmented generation. Placing the most relevant retrieved document in the middle of twenty others is a way of including it while reducing the chance it is used, and reordering so the best matches sit at the edges costs nothing.
More context is not free even when it fits
Filling a window because it is available increases cost linearly, increases latency, and past some point reduces answer quality by diluting the relevant material among the irrelevant. The number of documents worth retrieving is an empirical question with a maximum, and it is usually far lower than the window permits. A system that retrieves twenty passages because twenty fit is usually worse than one that retrieves five and ranks them well.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Why is the usable budget smaller than the window?
Because the system prompt, conversation history, tool definitions and reserved output space all come out of the same window first. Tool definitions are the component most often forgotten, since nobody writes them by hand.
What is the lost-in-the-middle effect?
The well-replicated finding that models use material at the start and end of a long context substantially more reliably than material in the middle, with the effect growing as the context lengthens.
Should I fill the window if the documents fit?
Usually not. Filling it increases cost and latency and past some point dilutes the relevant material among the irrelevant. The number of passages worth retrieving has a maximum, and it is normally far below what the window allows.
How accurate is the token estimate?
It is an approximation based on characters per token, which varies by tokeniser, by language and by content type. Code and non-Latin scripts use more tokens per character, so leave headroom rather than packing to the limit.
What should I reserve for the response?
At least the longest response you expect, plus margin. Running out of output space truncates the answer mid-sentence, which is a worse failure than dropping a document because it looks like a model problem.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter your context limit and document sizes to see what fits.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.