LLM VRAM Calculator
Estimate the GPU memory needed to run a large language model from its parameter count and quantization level, with an overhead allowance.
Last reviewed by the Radiatus Cloud team
Estimate the GPU memory (VRAM) needed to run an LLM.
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Estimate LLM VRAM requirements
Running a large language model requires enough GPU memory to hold its weights, and this calculator estimates how much from the parameter count and the precision used. Each parameter takes a number of bytes set by the precision: four for 32-bit, two for 16-bit, one for 8-bit and half a byte for 4-bit quantization. A seven-billion-parameter model in 16-bit precision needs about fourteen gigabytes for its weights, plus roughly twenty percent overhead for activations and runtime.
Quantizing to lower precision dramatically reduces memory, which is how large models are made to fit on consumer hardware.
Fitting models on your hardware
The dominant memory cost for inference is the model weights, so quantization is the main lever for fitting a model onto a given GPU. Dropping from 16-bit to 4-bit cuts the weight memory to a quarter, often with modest quality loss, enabling much larger models on the same card. The overhead allowance accounts for activations, buffers and framework usage that add to the raw weight size.
For long contexts, the key-value cache adds further memory that grows with sequence length and should be estimated separately. All calculation happens locally in your browser.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
How much VRAM does a 7B model need?
In 16-bit precision, about fourteen gigabytes for the weights, plus roughly twenty percent overhead, so around seventeen gigabytes in total for inference.
How does quantization reduce memory?
It stores each parameter in fewer bits. Moving from 16-bit to 4-bit cuts the weight memory to a quarter, letting larger models fit on smaller GPUs.
What is the overhead for?
It accounts for activations, buffers and framework memory that add to the raw weight size during inference, estimated here at about twenty percent.
Does this include the KV cache?
No. The key-value cache grows with context length and should be estimated separately, especially for long-context inference.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter the parameter count in billions and choose the precision.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.