AI Security

Training Compute Calculator

Estimate the total compute to train a language model using the approximation of six FLOPs per parameter per token, and the GPU-days needed.

Last reviewed by the Radiatus Cloud team

Estimate the total compute to train a model, and the GPU-days required.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Estimate training compute

Training a language model from scratch requires a vast amount of compute, commonly estimated with the rule of thumb of about six floating-point operations per parameter per token. This calculator applies that formula to estimate the total operations, then divides by a GPU effective throughput, accounting for real-world utilization, to estimate how many GPU-days the run would take on a single accelerator. Training a seven-billion-parameter model on a trillion tokens requires an enormous number of operations.

The factor of six, compared with two for inference, reflects the forward pass plus the roughly twice-as-expensive backward pass used in training.

The scale of training

Estimating training compute makes the scale of modern models tangible and helps with planning cluster size, budget and schedule. Because real GPUs achieve only a fraction of their peak throughput on large training runs, the utilization factor is important for a realistic estimate. Dividing the single-GPU time by the number of GPUs in a cluster gives the wall-clock duration, assuming perfect scaling.

This is a high-level approximation; actual training also involves data loading, communication overhead and imperfect scaling across many GPUs. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

How is training compute estimated?

With the rule of thumb of about six floating-point operations per parameter per token, multiplied across all training tokens.

Why six FLOPs, not two?

Training includes the forward pass and a backward pass that is roughly twice as costly, so the combined estimate is about six operations per parameter per token.

Why apply a utilization factor?

Real GPUs achieve only part of their peak throughput on large training runs, so utilization gives a more realistic effective speed and time estimate.

How do I get wall-clock time?

Divide the single-GPU time by the number of GPUs in your cluster, assuming near-perfect scaling, to approximate the real training duration.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the parameter count and the training tokens.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.