AI Security

LLM FLOPs Calculator

Estimate the floating-point operations for a language model forward pass from its parameter count and the number of tokens processed.

Last reviewed by the Radiatus Cloud team

Estimate the compute (FLOPs) for a language model forward pass.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Estimate LLM inference compute

The compute cost of running a language model can be estimated with a simple and widely used rule of thumb: a forward pass costs about two floating-point operations per parameter per token. This calculator applies that formula to estimate the operations for processing a given number of tokens through a model of a given size. A seven-billion-parameter model processing a thousand tokens performs roughly fourteen trillion floating-point operations.

The factor of two comes from each parameter being involved in a multiply and an add during the matrix operations that dominate the model.

Why FLOPs estimates help

Estimating the floating-point operations helps you reason about inference speed and hardware requirements, since a GPU rated for so many operations per second gives a rough upper bound on throughput. It also makes the cost of larger models and longer inputs concrete, as compute scales linearly with both parameters and tokens. The per-token figure shows the marginal cost of each additional token.

This is an approximation that captures the dominant matrix multiplications and ignores lower-order terms like attention overhead, which grows with context length. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

How are inference FLOPs estimated?

With the rule of thumb of about two floating-point operations per parameter per token, multiplied across the tokens processed.

Why the factor of two?

Each parameter takes part in a multiply and an add during the matrix operations that dominate the forward pass, giving roughly two operations each.

What does this tell me about speed?

Dividing the FLOPs by a GPU operations-per-second rating gives a rough upper bound on how fast the model can process the tokens.

What does the estimate ignore?

Lower-order costs such as attention, which grows with context length, and overheads, so it captures the dominant compute rather than the exact total.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the parameter count in billions and the number of tokens.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.