AI Security

Training Data Provenance Checker

Audit the provenance of a training corpus across rights, personal data, contamination and lineage, and identify which questions about the model you would currently be unable to answer.

Last reviewed by the Radiatus Cloud team

Results appear here.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Nearly every question about a model reduces to its training data

Is the output infringing, is there personal data in it, are the benchmark scores real, why does it fail on this group, can it be deployed in this jurisdiction: all of these resolve to what the model was trained on. An organisation that cannot enumerate its training corpus cannot answer any of them, and the questions arrive together, usually in response to a complaint rather than in advance. Provenance work is unusually front-loaded because after training the only remedy for a bad dataset is retraining.

Contamination makes performance claims void, not merely optimistic

If an evaluation set appears in the training data, the resulting score measures recall of memorised examples rather than capability, and every decision made on the strength of that score rests on nothing. This is common enough in publicly assembled corpora that silence on the question is itself informative, and the check is not expensive: hashing evaluation examples against the corpus catches the bulk of it.

Lineage from checkpoint to corpus is what makes any of it usable

Documentation that describes datasets in general, without recording which exact version of which dataset produced a given checkpoint, cannot answer a question about the model that is actually deployed. That is the only model anyone asks about. A manifest per training run costs almost nothing at the time and is the difference between answering a question and reconstructing a guess.

Related tools

Frequently Asked Questions

Why does provenance need to be documented before training?

Because after training the only remedy for an unapproved or unlicensed dataset is retraining. Every other question can be researched later; this one cannot be fixed later.

What is benchmark contamination?

Evaluation data present in the training set, which makes the resulting score a measure of recall rather than capability. It is common in publicly assembled corpora and cheap to check by hashing.

Do erasure requests apply to training data?

Personal data in a training set engages data protection law regardless of how it was collected. If removal requires retraining, that is a position to decide and document deliberately rather than to discover when the first request arrives.

Why does deduplication matter?

It drives memorisation and inflates benchmark scores, and removing duplicates is the cheapest single intervention available on both problems.

Is buying a dataset enough?

Not by itself. Purchase does not transfer risk unless the contract says so, and standard terms usually do not include a provenance warranty.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Tick what you have documented to score provenance.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.