Question

What is an effect size, and why does it matter more than significance?

Vault Verified
Curated Intelligence
Definitive Source
Answer

A measure of how large a difference or relationship is, expressed in a standardised or interpretable form — and it answers the question significance testing does not: not whether an effect exists, but whether it matters.

The distinction that makes it essential. Statistical significance tells you whether an observed result is unlikely under the null hypothesis. With a large enough sample, almost any difference becomes significant, including one far too small to be of any practical consequence. Effect size tells you the magnitude, independently of sample size.

This is why a headline reporting a "significant" finding conveys almost nothing on its own.

The common measures:

Cohen's d — the difference between two means, expressed in standard deviations. Conventionally described as small, medium and large at around 0.2, 0.5 and 0.8, though Cohen himself intended these as rough guidance and warned against mechanical application. What counts as large depends entirely on the field.

Correlation coefficient (r) — the strength of a linear relationship, from −1 to 1. r² gives the proportion of variance explained, which is frequently more sobering: a correlation of 0.3 explains 9% of the variance.

Odds ratio and risk ratio, in medicine and epidemiology.

Absolute versus relative risk, which is the most consequential distinction in health reporting. "Doubles your risk" sounds alarming; if the baseline risk is 1 in 100,000, the absolute increase is negligible. Relative figures without the baseline are uninformative and frequently misleading, and they dominate coverage because they produce better headlines.

Number needed to treat — how many patients must receive a treatment for one to benefit. Directly interpretable and rarely reported outside clinical literature.

Why effect sizes should always be reported: they allow comparison across studies, are required for meta-analysis, permit power calculations for future work, and let readers judge practical importance for themselves.

The caution: effect sizes from small studies are imprecise and, where selected for significance, systematically inflated — so a large reported effect from a small study should be treated as an upper bound rather than an estimate.

Related Questions