What does a p-value actually mean?
The probability of observing results at least as extreme as yours, assuming the null hypothesis is true. It is a statement about the data given a hypothesis — not about the hypothesis given the data, and reversing those is the most consequential misunderstanding in applied statistics.
What it is not:
Not the probability that the null hypothesis is true. A p-value of 0.03 does not mean a 3% chance there is no effect.
Not the probability that your hypothesis is true.
Not the probability that the result was chance.
Not a measure of effect size. A tiny, meaningless effect will produce a very small p-value in a large enough sample. Statistical significance and practical importance are entirely different questions, and conflating them is how trivial findings become headlines.
Not a measure of how likely the result is to replicate.
Where 0.05 came from. It is a convention, attributable largely to Fisher, who suggested it as a convenient threshold and explicitly did not intend it as a rule. There is nothing special about 0.05, and treating 0.049 and 0.051 as categorically different is indefensible — the underlying evidence is nearly identical.
Why it is so easily abused:
Dichotomising a continuous measure into significant and not significant discards information.
Multiple comparisons. Testing many hypotheses at 0.05 means roughly one in twenty produces a false positive by chance. Without correction, testing enough things guarantees a finding.
Optional stopping, where data collection continues until significance appears.
"Approaching significance" and similar phrasing, which is either significant or it is not.
What the profession has said. The American Statistical Association issued a formal statement on p-values, warning against these misinterpretations and against basing conclusions solely on whether a threshold is crossed. Some journals have restricted or banned significance testing.
What to report alongside, or instead: the effect size, with a confidence interval, which conveys both magnitude and precision; the sample size; and the pre-specified analysis plan.
The useful reframing: a p-value is a measure of surprise given an assumption, not a verdict.