Multiple Comparisons & P-hacking
Test enough hypotheses and one turns “significant” by chance alone — about 1 in 20 at α = 0.05.
Every hypothesis test carries a false-positive risk of α (usually 0.05). Run many independent tests and the chance that at least one is “significant” by chance balloons — this is the multiple-comparisons problem. P-hacking (data-dredging) exploits it: trying many analyses, subgroups or outcomes and reporting only the ones that crossed p<0.05. Defences are multiplicity corrections (Bonferroni; false-discovery rate) and — more powerfully — pre-registration with a single pre-specified primary outcome.
Test 20 truly-null effects and on average one comes up “significant” — roughly a 64% chance of at least one false positive.
Twenty tests, every one of an effect that is genuinely nothing. Nineteen correctly come back non-significant (grey); one lights up red at p<0.05 purely by chance. The bar underneath makes the cumulative point: across 20 independent null tests the probability of at least one false positive is about 1 − 0.9520 ≈ 64% — far above the 5% you assumed per test.
Count the tests, not just the wins. A single red square is the expected yield of testing 20 null hypotheses — so a lone “significant” result among many is uninformative until you correct for how many shots were taken. Bonferroni divides α by the number of tests (a strict guard against any false positive); false-discovery rate controls the proportion of significant findings that are false (less strict, better powered). Best of all, a pre-specified primary outcome means only one square was ever in play.
Trials and reviews routinely test dozens of secondary outcomes and subgroups. Celebrating one “positive” secondary or subgroup finding from a trial that quietly ran many comparisons is how spurious results enter practice. The defence the examiner wants is structural: was the outcome pre-specified, and was multiplicity accounted for? If not, the finding is hypothesis-generating, not practice-changing.
- At
α = 0.05, ≈ 1 in 20 null tests is “significant” by chance P(≥1 false +) = 1 − (0.95)kfor k independent testsBonferroni= use α ÷ k |FDR= control the false-discovery proportion- Strongest defence =
pre-registration+ one pre-specified primary outcome
ISIS-2 (Lancet, 13 Aug 1988) — 17,187 patients with suspected acute MI; aspirin clearly reduced vascular death overall. To teach the danger of subgroups, the investigators deliberately split patients by astrological birth sign: aspirin appeared not to help those born under Gemini or Libra, while helping every other sign. The signal was pure chance — an honest demonstration that a subgroup can “abolish” a real, large effect when you carve the data finely enough. The lesson the examiners love: if a manifestly absurd subgroup (star sign) can reach “significance,” so can a plausible-sounding clinical one — subgroup and secondary findings need pre-specification and multiplicity correction before you believe them.
- Unadjusted multiplicity — many outcomes/tests reported with no Bonferroni or FDR correction.
- Post-hoc subgroups — effects “found” after the fact, not pre-specified, are hypothesis-generating only.
- Outcome switching — the primary outcome in the protocol differs from the one highlighted in the publication.
Quick check
A trial reports 1 significant result among 25 tested — convincing?
Answer: Not by itself — at α = 0.05 about one false positive is expected by chance across 25 tests. It needs multiplicity correction (or pre-specification) and ideally independent replication before you trust it.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.