AI Security

LLM Rate Limit Calculator

Calculate how many requests per minute you can make within token-per-minute and request-per-minute rate limits based on tokens per request.

Last reviewed by the Radiatus Cloud team

Calculate your effective request throughput under LLM rate limits.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Calculate your LLM rate limits

Language model APIs usually impose two rate limits: a cap on tokens per minute and a cap on requests per minute. Your effective throughput is whichever of these you hit first, given how many tokens each request uses. This calculator works out how many requests per minute the token limit allows, compares it with the request limit, and reports the lower of the two as your real ceiling, along with which limit is the bottleneck.

If each request uses many tokens, the token-per-minute limit usually binds first; if requests are small, the request-per-minute limit dominates.

Planning for throughput

Understanding which limit constrains you is essential for sizing a workload and avoiding rate-limit errors. If the token limit is the bottleneck, shortening prompts or responses increases throughput, whereas if the request limit binds, batching work into fewer, larger requests helps. Knowing the ceiling also informs whether you need a higher-tier plan or to spread load across time.

The requests-per-hour figure scales the per-minute ceiling for longer-horizon planning. All calculation happens locally in your browser.

Notes on these estimates

Because the llm rate limit calculator runs entirely in your browser, nothing you enter is uploaded, so you can use it with private data safely. The figures are estimates based on the values you provide and common rules of thumb, so treat them as planning guidance rather than exact measurements, and run the tool as often as you need for free.

Related tools

Frequently Asked Questions

What limits my LLM throughput?

Whichever rate limit you hit first: the tokens-per-minute cap or the requests-per-minute cap, given your tokens per request.

How do I find the bottleneck?

Divide the token limit by tokens per request to get the token-bound requests per minute, then compare with the request limit. The lower one binds.

How do I increase throughput if tokens bind?

Reduce the tokens per request by shortening prompts or responses, so more requests fit within the tokens-per-minute cap.

How do I increase throughput if requests bind?

Batch more work into each request so fewer, larger requests use your token budget without hitting the requests-per-minute cap.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the rate limits and the tokens per request.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.