What is the difference between anonymisation and pseudonymisation?
Whether the data can still be linked back to a person. Anonymised data cannot, and falls outside data protection law entirely. Pseudonymised data can, with additional information, and remains fully regulated — a distinction organisations get wrong expensively.
Pseudonymisation. Replacing identifying details with a reference — a code, token or hashed value — while keeping the means to reverse it separately. The classic example is replacing names with participant numbers and keeping the key in a different, restricted system.
It remains personal data. The legislation is explicit: pseudonymised data that could be attributed to an individual by use of additional information is still personal data, and all obligations continue to apply. It is a security measure that reduces risk, not an exit from the regime — and it is explicitly encouraged as such.
Anonymisation. Rendering data such that the individual is no longer identifiable by anyone, by any reasonably likely means. Genuinely anonymised data is outside the scope of data protection law, so it can be used and shared freely.
Why true anonymisation is much harder than it appears. The test considers all means reasonably likely to be used, by anyone, accounting for available technology and other datasets. Research has repeatedly demonstrated that datasets believed anonymous were re-identifiable by combining them with other sources — a small number of ordinary attributes such as postcode, birth date and sex is frequently enough to isolate an individual in a population.
High-profile re-identification cases involving supposedly anonymised medical, location and viewing data have made regulators considerably more sceptical of anonymisation claims.
What genuine anonymisation involves: removing direct identifiers, plus techniques addressing indirect ones — aggregation, generalisation of values into ranges, suppression of rare combinations, and adding statistical noise.
k-anonymity requires each record to be indistinguishable from at least k−1 others, with l-diversity and differential privacy addressing weaknesses in that approach.
The practical test: if you retain the ability to re-identify, it is pseudonymised. Destroying the key is what converts one into the other, and organisations are frequently reluctant to do so — which tells you it was never anonymised.