AI Security

KV Cache Calculator

Estimate the key-value cache memory a language model uses for a given context length, from the number of layers, hidden size and precision.

Last reviewed by the Radiatus Cloud team

Estimate the key-value cache memory for a given context length.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Estimate key-value cache memory

When a language model generates text, it caches the key and value vectors for every token processed so it does not recompute them, and this key-value cache can consume significant GPU memory for long contexts. This calculator estimates the cache size from the number of layers, the hidden size, the context length, the batch size and the precision. The cache stores two tensors, keys and values, for each layer, so the memory is twice the product of these factors.

For a thirty-two-layer model with a hidden size of four thousand and ninety-six, an eight-thousand-token context in 16-bit precision uses several gigabytes of cache.

Why the KV cache matters

The key-value cache grows linearly with the context length, so long-context inference can require more memory for the cache than for the model weights themselves. This is a key constraint when serving many simultaneous users, since the batch size multiplies the cache, or when supporting very long prompts. Techniques like cache quantization and grouped attention reduce this cost.

This estimate uses a standard formula and assumes the full hidden size per layer; architectures with grouped or multi-query attention use less. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

What is the KV cache?

It is the stored key and value vectors for every processed token, kept so the model avoids recomputing them when generating each new token.

Why does it grow with context length?

The cache holds entries for every token in the context, so doubling the context length doubles the cache memory.

Can the KV cache exceed the model size?

Yes. For very long contexts or large batches, the key-value cache can require more memory than the model weights themselves.

How can the KV cache be reduced?

Through lower-precision cache storage, and architectures with grouped or multi-query attention that share key-value vectors across attention heads.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the layers, hidden size, context length and batch size.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.