What is statistical power, and why do underpowered studies mislead?
Power is the probability that a study will detect an effect if one genuinely exists. Low power means a real effect is likely to be missed — and, far less intuitively, it also makes the findings that are significant less trustworthy.
What determines it:
Sample size, the main controllable factor.
Effect size — larger effects are easier to detect.
Variability in the data.
The significance threshold used.
The conventional target is 80% power, meaning a 20% chance of missing a genuine effect — itself a fairly permissive standard that persists by convention.
The obvious problem: false negatives. An underpowered study that finds nothing has not shown there is no effect. "No significant difference" is not evidence of no difference, and reporting it as such is among the most common errors in both research and its coverage. Absence of evidence is not evidence of absence, and distinguishing them requires knowing the power.
The non-obvious problem, which matters more. In an underpowered study, an effect can only reach significance if the observed estimate is unusually large — because a modest estimate will not clear the threshold with so little data. So the significant findings that emerge from underpowered research are systematically inflated.
This is the winner's curse, and it explains a great deal about why striking early findings shrink or vanish on replication. The first small study reports a large effect; larger replications find something much smaller or nothing. Nothing dishonest has occurred — it is a mechanical consequence of the design.
Combined with publication bias, it compounds. If underpowered studies only get published when they reach significance, and reaching significance requires an inflated estimate, the literature fills with overstated effects.
Where it is worst: fields where data collection is expensive or participants are scarce, subgroup analyses within larger studies, and neuroimaging and genetics, where power problems have been documented extensively.
What helps: power calculations before collecting data; larger and multi-site studies; pre-registration; and treating a single small significant study as a reason to investigate rather than as a finding.