Training Data Memorisation Risk Estimator
Estimate the risk that a fine-tuned model reproduces training examples verbatim, from duplication, dataset size, model capacity, epochs and the mitigations applied.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Duplication drives memorisation more than anything else
The most robust finding in the memorisation literature is that a sequence appearing many times in a training corpus is far more likely to be reproduced verbatim than one appearing once, and the relationship is steep rather than gradual. This matters because duplication is usually accidental: the same support article appears in a hundred exported tickets, the same boilerplate paragraph sits at the foot of every document, the same customer record was exported three times during a migration. Deduplication is consequently the cheapest and most effective single intervention available, and it is regularly skipped because the duplicates are not visible in a sample.
Capacity relative to data is the second factor
A large model fine-tuned on a small dataset for many epochs has both the capacity to store the examples and the repetition to encode them, which is exactly the configuration most organisations use when they fine-tune. The generic advice to train longer for better results pushes directly against memorisation risk, and there is no setting that gives both. Where the training data contains anything sensitive, that trade-off should be made deliberately rather than by accepting a default.
Extraction is testable and almost nobody tests it
Whether a specific model memorised specific data is an empirical question with a straightforward test: prompt it with prefixes drawn from the training set and measure how often the continuation matches. That takes an afternoon and gives a real number, and it is skipped in favour of an assumption in nearly every deployment. An estimate like this one is a reason to run the test, not a substitute for having run it.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
What is the single most effective mitigation?
Deduplication. A sequence appearing many times in the corpus is far more likely to be reproduced verbatim, and the relationship is steep. Duplicates are usually accidental and invisible in a sample.
Does a bigger model memorise more?
Yes, all else equal, because capacity is what stores the examples. A large model fine-tuned on a small dataset for many epochs is the highest-risk configuration and also the most common one.
Do more epochs increase the risk?
Substantially. The advice to train longer for better results pushes directly against memorisation risk, and there is no setting that gives both.
How do I actually test for memorisation?
Prompt the model with prefixes drawn from the training set and measure how often the continuation matches the original. It takes an afternoon and gives a real number rather than an estimate.
Does differential privacy solve this?
It bounds it in a provable way, at a cost in accuracy that is often significant for a fine-tune. It is the strongest available answer and it is not free, which is why the trade-off has to be a decision rather than a default.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter your training configuration to estimate memorisation risk.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.