Question

Why do marketing tests so often produce misleading results?

Vault Verified
Curated Intelligence
Definitive Source
Answer

Because most tests are underpowered, stopped early, or measuring the wrong thing — and each of those produces a confident number that does not survive contact with reality.

Stopping early is the largest problem. Watching a test and stopping when it reaches significance dramatically inflates the false positive rate. Random variation crosses the threshold frequently on the way to converging, so a test peeked at repeatedly will eventually show a "winner" that does not exist. Decide the sample size and duration in advance, and do not look at significance until then — or use a method designed for continuous monitoring.

Underpowering. Detecting a small effect requires a large sample. Most businesses do not have the traffic to detect a 2% improvement reliably, so they either run tests that cannot answer the question or accept noise as signal. Calculate the required sample before running, and if it exceeds what you can gather in a reasonable period, the test is not worth running.

Running for too short a period. A test must cover complete business cycles — at minimum full weeks, since weekday and weekend behaviour differ — and ideally longer for anything with a monthly rhythm.

Measuring a proxy. Optimising clicks frequently reduces revenue, because the change that attracts more clicks attracts less qualified ones. Measure the outcome you actually want, even though it is further down the funnel and noisier.

Ignoring novelty and primacy effects. Existing users react to change itself, which fades — so an early result can reverse.

Running many tests and reporting the winners. Testing twenty things at a 5% threshold produces one false winner by chance alone.

Segment mining after the fact, where a flat result is rescued by finding a subgroup where it worked. This is nearly always chance unless the segment was specified in advance.

What good practice looks like: a written hypothesis, a primary metric chosen beforehand, a predetermined sample and duration, guardrail metrics to catch harm elsewhere, and a record of tests that failed — since organisations that only remember their wins systematically overestimate the value of testing.

Related Questions