Training Data Provenance Checker
Audit the provenance of a training corpus across rights, personal data, contamination and lineage, and identify which questions about the model you would currently be unable to answer.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Nearly every question about a model reduces to its training data
Is the output infringing, is there personal data in it, are the benchmark scores real, why does it fail on this group, can it be deployed in this jurisdiction: all of these resolve to what the model was trained on. An organisation that cannot enumerate its training corpus cannot answer any of them, and the questions arrive together, usually in response to a complaint rather than in advance. Provenance work is unusually front-loaded because after training the only remedy for a bad dataset is retraining.
Contamination makes performance claims void, not merely optimistic
If an evaluation set appears in the training data, the resulting score measures recall of memorised examples rather than capability, and every decision made on the strength of that score rests on nothing. This is common enough in publicly assembled corpora that silence on the question is itself informative, and the check is not expensive: hashing evaluation examples against the corpus catches the bulk of it.
Lineage from checkpoint to corpus is what makes any of it usable
Documentation that describes datasets in general, without recording which exact version of which dataset produced a given checkpoint, cannot answer a question about the model that is actually deployed. That is the only model anyone asks about. A manifest per training run costs almost nothing at the time and is the difference between answering a question and reconstructing a guess.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Why does provenance need to be documented before training?
Because after training the only remedy for an unapproved or unlicensed dataset is retraining. Every other question can be researched later; this one cannot be fixed later.
What is benchmark contamination?
Evaluation data present in the training set, which makes the resulting score a measure of recall rather than capability. It is common in publicly assembled corpora and cheap to check by hashing.
Do erasure requests apply to training data?
Personal data in a training set engages data protection law regardless of how it was collected. If removal requires retraining, that is a position to decide and document deliberately rather than to discover when the first request arrives.
Why does deduplication matter?
It drives memorisation and inflates benchmark scores, and removing duplicates is the cheapest single intervention available on both problems.
Is buying a dataset enough?
Not by itself. Purchase does not transfer risk unless the contract says so, and standard terms usually do not include a provenance warranty.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Tick what you have documented to score provenance.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.