AI Security

LLM VRAM Calculator

Estimate the GPU memory needed to run a large language model from its parameter count and quantization level, with an overhead allowance.

Last reviewed by the Radiatus Cloud team

Estimate the GPU memory (VRAM) needed to run an LLM.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Estimate LLM VRAM requirements

Running a large language model requires enough GPU memory to hold its weights, and this calculator estimates how much from the parameter count and the precision used. Each parameter takes a number of bytes set by the precision: four for 32-bit, two for 16-bit, one for 8-bit and half a byte for 4-bit quantization. A seven-billion-parameter model in 16-bit precision needs about fourteen gigabytes for its weights, plus roughly twenty percent overhead for activations and runtime.

Quantizing to lower precision dramatically reduces memory, which is how large models are made to fit on consumer hardware.

Fitting models on your hardware

The dominant memory cost for inference is the model weights, so quantization is the main lever for fitting a model onto a given GPU. Dropping from 16-bit to 4-bit cuts the weight memory to a quarter, often with modest quality loss, enabling much larger models on the same card. The overhead allowance accounts for activations, buffers and framework usage that add to the raw weight size.

For long contexts, the key-value cache adds further memory that grows with sequence length and should be estimated separately. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

How much VRAM does a 7B model need?

In 16-bit precision, about fourteen gigabytes for the weights, plus roughly twenty percent overhead, so around seventeen gigabytes in total for inference.

How does quantization reduce memory?

It stores each parameter in fewer bits. Moving from 16-bit to 4-bit cuts the weight memory to a quarter, letting larger models fit on smaller GPUs.

What is the overhead for?

It accounts for activations, buffers and framework memory that add to the raw weight size during inference, estimated here at about twenty percent.

Does this include the KV cache?

No. The key-value cache grows with context length and should be estimated separately, especially for long-context inference.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the parameter count in billions and choose the precision.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.