AI Eval Sample Size Calculator
Work out how many evaluation examples are needed to detect a real difference between two models, using the paired McNemar test for shared test sets and the two-proportion test for independent ones.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Most model comparisons are run on far too few examples
A team compares two models on fifty prompts, sees 82 percent against 78 percent, and ships the winner. That difference is well inside what random variation produces at that sample size: the 95 percent confidence interval on 82 percent from fifty examples runs from roughly 69 to 91 percent. The comparison was not close to conclusive, and the decision was made anyway because the number looked like evidence. Working out the required size before running the evaluation is what prevents this, and it usually reveals that the honest answer needs several hundred examples rather than fifty.
A shared test set needs a paired test, and that is good news
When both models are evaluated on the same examples, the results are paired, and the right test is McNemar's, which looks only at the cases where the two models disagree. This is far more powerful than treating the two accuracy figures as independent samples, because it removes the variation caused by some examples simply being harder than others. The practical effect is large: a paired design frequently needs a third or a quarter of the examples an unpaired one would, and evaluations on a shared test set are paired by construction whether or not anyone analysed them that way.
The number you need depends on the difference you care about
There is no sample size that detects any difference. The calculation needs a minimum effect worth detecting, and choosing it is a product decision rather than a statistical one: a one-point accuracy gain may be worth a great deal or nothing at all depending on what the model does. Picking the effect size after seeing the data is the standard way to arrive at a significant result that does not replicate.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Why is fifty examples not enough?
Because the confidence interval is wide. At 82 percent accuracy on fifty examples the 95 percent interval runs from roughly 69 to 91 percent, so a four-point difference against another model is well inside random variation.
What is McNemar’s test?
A paired test for comparing two classifiers on the same examples. It counts only the cases where they disagree, which removes the variation caused by some examples being harder than others and makes it much more powerful than an unpaired comparison.
When should I use the unpaired calculation?
Only when the two models genuinely see different examples. If they share a test set the results are paired by construction, and treating them as independent throws away most of the power you already have.
How do I choose the minimum detectable effect?
As a product decision about what difference would change what you do. Choosing it after seeing the data is the standard route to a significant result that does not replicate.
Does this apply to generative evaluation?
It applies wherever each example produces a pass or fail judgement, including a rubric or a model-graded score thresholded into pass and fail. It does not apply to continuous scores without adaptation.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter your baseline accuracy and the improvement you want to detect.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.