LLM Latency Budget Planner
Model the latency of an LLM feature from retrieval, prefill, decode and network stages, see where the time actually goes, and check the result against what users tolerate.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Time to first token and total time are different products
A streaming interface is judged on when text starts appearing; a batch call that returns a JSON object is judged on when it finishes. These have different bottlenecks. Time to first token is dominated by retrieval and prefill, which scale with input length, while total time is dominated by decode, which scales with output length and is essentially unaffected by how long the prompt was. Optimising the wrong one is common: teams shorten prompts to make a slow generation faster and find that nothing changes, because the generation was slow for a reason prompt length does not touch.
Decode is sequential and that is the whole constraint
Prefill processes the entire input in parallel, so doubling the prompt adds relatively little. Decode produces one token at a time, each depending on the last, so output length translates almost linearly into wall-clock time. A response of 800 tokens at 50 tokens per second takes 16 seconds no matter what hardware sits behind it, and the only levers are generating fewer tokens, generating them faster, or starting to show them sooner.
Retrieval before the model is latency the model gets blamed for
In a retrieval-augmented feature the embedding call, the vector search and any reranking all happen before the first model token can be produced, and they are serial with it. A reranker adding 400 milliseconds is 400 milliseconds of blank screen. Running retrieval concurrently with anything else that must happen, and starting the stream as early as possible, buys more perceived speed than most model-level tuning.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Why does prompt length barely affect total time?
Because prefill processes the input in parallel while decode produces one token at a time. Input length affects time to first token; output length dominates total time.
What is a tolerable time to first token?
Under about one second feels immediate for a streaming interface. Beyond roughly two to three seconds users start checking whether anything is happening, which is why a visible stream matters more than raw speed.
How do I make a long generation faster?
Generate fewer tokens, generate them faster, or start showing them sooner. There is no fourth option, because decode is sequential and each token depends on the one before it.
Does a reranker hurt latency?
Yes, and directly. It sits before the first model token, so its time is blank screen. Whether it is worth it depends on whether the retrieval improvement is larger than the perceived cost.
Are these figures accurate for my setup?
They are a model built from the numbers you enter. Real throughput varies with load, batching, context length and provider, so measure your own percentiles rather than trusting a planning figure.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter your token counts and throughput to model the latency.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.