Question

What makes a sample representative?

Vault Verified
Curated Intelligence
Definitive Source
Answer

That every member of the population has a known, non-zero chance of being included — which is a far stronger requirement than the sample merely being large, and it is why enormous samples can be badly wrong.

The crucial point: size does not fix bias. A biased sampling method produces a biased estimate no matter how many people are surveyed. Increasing the sample makes the wrong answer more precise, not more correct.

The historical demonstration. The 1936 Literary Digest poll surveyed a very large number of Americans — millions of responses — and confidently predicted the wrong presidential winner, while George Gallup predicted correctly with a far smaller but better-constructed sample. The Digest drew its sample from telephone directories, magazine subscribers and vehicle registrations, which over-represented wealthier households. The enormous sample made the bias precise rather than correcting it.

The forms of sampling bias:

Coverage bias, where the sampling frame excludes part of the population — the Literary Digest problem, and the reason telephone polling now has difficulties.

Non-response bias, the dominant modern problem. When response rates are low, the people who respond differ systematically from those who do not, and no amount of effort reaches the rest.

Self-selection, where participants opt in. Online polls, product reviews and voluntary surveys all suffer badly, because people with strong views participate disproportionately.

Survivorship bias, where the sample includes only those who reached the point of being measured.

Convenience sampling, including the heavy reliance in psychology on university students — the WEIRD problem, where samples are Western, Educated, Industrialised, Rich and Democratic and are generalised to humanity.

How good sampling is done: probability sampling — simple random, stratified to ensure subgroups are proportionally represented, cluster for practicality, and systematic.

Weighting adjusts for known differences between sample and population, and is standard practice — but it can only correct for characteristics that are measured, so it does not fix unknown biases.

Margin of error assumes random sampling, so quoting one for a self-selected online poll is meaningless.

Related Questions