LLM Rate Limit Calculator
Calculate how many requests per minute you can make within token-per-minute and request-per-minute rate limits based on tokens per request.
Last reviewed by the Radiatus Cloud team
Calculate your effective request throughput under LLM rate limits.
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Calculate your LLM rate limits
Language model APIs usually impose two rate limits: a cap on tokens per minute and a cap on requests per minute. Your effective throughput is whichever of these you hit first, given how many tokens each request uses. This calculator works out how many requests per minute the token limit allows, compares it with the request limit, and reports the lower of the two as your real ceiling, along with which limit is the bottleneck.
If each request uses many tokens, the token-per-minute limit usually binds first; if requests are small, the request-per-minute limit dominates.
Planning for throughput
Understanding which limit constrains you is essential for sizing a workload and avoiding rate-limit errors. If the token limit is the bottleneck, shortening prompts or responses increases throughput, whereas if the request limit binds, batching work into fewer, larger requests helps. Knowing the ceiling also informs whether you need a higher-tier plan or to spread load across time.
The requests-per-hour figure scales the per-minute ceiling for longer-horizon planning. All calculation happens locally in your browser.
Notes on these estimates
Because the llm rate limit calculator runs entirely in your browser, nothing you enter is uploaded, so you can use it with private data safely. The figures are estimates based on the values you provide and common rules of thumb, so treat them as planning guidance rather than exact measurements, and run the tool as often as you need for free.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
What limits my LLM throughput?
Whichever rate limit you hit first: the tokens-per-minute cap or the requests-per-minute cap, given your tokens per request.
How do I find the bottleneck?
Divide the token limit by tokens per request to get the token-bound requests per minute, then compare with the request limit. The lower one binds.
How do I increase throughput if tokens bind?
Reduce the tokens per request by shortening prompts or responses, so more requests fit within the tokens-per-minute cap.
How do I increase throughput if requests bind?
Batch more work into each request so fewer, larger requests use your token budget without hitting the requests-per-minute cap.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter the rate limits and the tokens per request.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.