Compliance

K-Anonymity Calculator

Paste a CSV, choose the quasi-identifier columns, and measure k-anonymity, l-diversity and the number of unique records that could be re-identified, entirely in your browser.

Last reviewed by the Radiatus Cloud team

Results appear here.

Going for ISO 27001, SOC 2, HIPAA or GDPR?

Radiatus runs end-to-end compliance & GRC programs.

Get a free readiness review

Removing the name is not anonymisation

A dataset with the direct identifiers stripped can still be re-identified from the combination of attributes that remain. The classic demonstration is that a large share of the United States population is uniquely identified by the triple of postcode, date of birth and sex, none of which is a name. Those combining attributes are called quasi-identifiers, and the measure of how much protection remains is k-anonymity: the size of the smallest group of records sharing an identical combination.

What k actually tells you

If k is 1, at least one record is unique on its quasi-identifiers, and anyone who knows those attributes about a person can pick that person out of the file with certainty. If k is 5, every record hides among at least four others. Under the GDPR this matters directly, because data that can still be attributed to an identifiable person by any means reasonably likely to be used remains personal data, and Recital 26 is explicit that pseudonymised data is still in scope. A dataset with k of 1 has not left that scope.

k alone is not enough

A group of five records that all share the same sensitive value gives away that value without anyone needing to be identified. That failure is what l-diversity measures: the number of distinct sensitive values within each equivalence class. A dataset can be 5-anonymous and 1-diverse simultaneously, which is why this tool reports both. Neither measure accounts for outside knowledge, so both are upper bounds on privacy rather than guarantees.

Related tools

Frequently Asked Questions

What is a quasi-identifier?

An attribute that is not identifying alone but becomes identifying in combination, such as postcode, birth date, job title or employer. The classic result is that postcode, date of birth and sex together uniquely identify a large share of a population.

What value of k is enough?

There is no universal threshold. k of 5 is a common floor for research data releases and k of 11 appears in some health guidance, but the right value depends on how sensitive the data is and who might realistically try to re-identify it.

Why does l-diversity matter if k is high?

Because a group of five records that all share the same sensitive value reveals that value without identifying anyone. A dataset can be 5-anonymous and 1-diverse at the same time.

Does k-anonymity make data non-personal under the GDPR?

Not by itself. Recital 26 keeps pseudonymised data in scope, and the test is whether attribution is possible by means reasonably likely to be used. A high k reduces that likelihood; it does not settle the legal question.

Is my data uploaded anywhere?

No. The parsing and all counting happen in your browser, and nothing is transmitted. You can confirm that by loading the page and disconnecting from the network before pasting.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Paste a CSV and pick the quasi-identifier columns to measure.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.