What is differential privacy?
A mathematical guarantee that the output of an analysis reveals almost nothing about any single individual in the data — achieved by adding carefully calibrated random noise, and notable for being one of the few privacy techniques with a provable definition rather than a best-effort promise.
The problem it solves. Removing names does not make data anonymous. Repeated famous results have shown that individuals can be re-identified by combining "anonymised" datasets with other information — a handful of data points about location, ratings or dates is frequently enough to single someone out. Anonymisation as usually practised does not survive contact with auxiliary data.
The core idea. An analysis is differentially private if its output is almost exactly as likely whether or not any particular person's record was included. If your presence cannot meaningfully change the result, the result cannot meaningfully reveal your presence. This holds regardless of what else an attacker knows — which is the property that makes it strong.
How it is achieved. Noise drawn from a specific distribution is added, scaled to how much one individual could affect the answer. A count is easy; an average over a small group needs more noise; an outlier-sensitive statistic needs more still.
The privacy budget, epsilon. A parameter controlling the trade-off: smaller epsilon means more noise, more privacy, less accuracy. Crucially, repeated queries consume budget cumulatively, because each answer leaks a little — so systems must track and cap total queries, which is the practical difficulty.
Where it is genuinely used: national statistical agencies including in census publication; telemetry collection by major operating system and browser vendors; and local differential privacy, where noise is added on your device before anything is sent, so the collector never holds the true value.
Its limits, honestly stated: it protects individuals, not groups or aggregate facts; it reduces accuracy, particularly for small subpopulations, which has real equity consequences; epsilon values used in practice are sometimes far larger than theorists consider meaningful; and it is no use where the individual record itself is the product.