Home/Critical Appraisal
Critical Appraisal

Multiple Comparisons & P-hacking

EM FINAL EXAMS Critical Appraisal · Statistics Multiple Comparisons & P-hacking Test enough hypotheses and one turns “significant” by chance alone — about 1 in 20 at α = 0.05. Definition Every hypothesis test carries a false-positive risk of α (usually 0.05). Run many independent tests and the chance that at least one is “significant” […]

EM FINAL EXAMS Critical Appraisal · Statistics

Multiple Comparisons & P-hacking

Test enough hypotheses and one turns “significant” by chance alone — about 1 in 20 at α = 0.05.

Definition

Every hypothesis test carries a false-positive risk of α (usually 0.05). Run many independent tests and the chance that at least one is “significant” by chance balloons — this is the multiple-comparisons problem. P-hacking (data-dredging) exploits it: trying many analyses, subgroups or outcomes and reporting only the ones that crossed p<0.05. Defences are multiplicity corrections (Bonferroni; false-discovery rate) and — more powerfully — pre-registration with a single pre-specified primary outcome.

The picture
20 tests of effects that are truly NULL 1 “significant” by chance (p < 0.05, false positive) P(≥1 false positive in 20 tests) ≈ 64%

Test 20 truly-null effects and on average one comes up “significant” — roughly a 64% chance of at least one false positive.

What it shows

Twenty tests, every one of an effect that is genuinely nothing. Nineteen correctly come back non-significant (grey); one lights up red at p<0.05 purely by chance. The bar underneath makes the cumulative point: across 20 independent null tests the probability of at least one false positive is about 1 − 0.9520 ≈ 64% — far above the 5% you assumed per test.

How to read it

Count the tests, not just the wins. A single red square is the expected yield of testing 20 null hypotheses — so a lone “significant” result among many is uninformative until you correct for how many shots were taken. Bonferroni divides α by the number of tests (a strict guard against any false positive); false-discovery rate controls the proportion of significant findings that are false (less strict, better powered). Best of all, a pre-specified primary outcome means only one square was ever in play.

Why it matters

Trials and reviews routinely test dozens of secondary outcomes and subgroups. Celebrating one “positive” secondary or subgroup finding from a trial that quietly ran many comparisons is how spurious results enter practice. The defence the examiner wants is structural: was the outcome pre-specified, and was multiplicity accounted for? If not, the finding is hypothesis-generating, not practice-changing.

Key
  • At α = 0.05, ≈ 1 in 20 null tests is “significant” by chance
  • P(≥1 false +) = 1 − (0.95)k for k independent tests
  • Bonferroni = use α ÷ k  |  FDR = control the false-discovery proportion
  • Strongest defence = pre-registration + one pre-specified primary outcome
Pitfall
Pitfall Celebrating a “significant” secondary or subgroup finding from a trial that quietly tested dozens of comparisons. Without multiplicity correction or pre-specification, that lone p<0.05 is exactly what chance predicts — not evidence of a real effect.
emfinalexams.com · FRCEM / MRCEM revision
EM trial in the wild

ISIS-2 (Lancet, 13 Aug 1988) — 17,187 patients with suspected acute MI; aspirin clearly reduced vascular death overall. To teach the danger of subgroups, the investigators deliberately split patients by astrological birth sign: aspirin appeared not to help those born under Gemini or Libra, while helping every other sign. The signal was pure chance — an honest demonstration that a subgroup can “abolish” a real, large effect when you carve the data finely enough. The lesson the examiners love: if a manifestly absurd subgroup (star sign) can reach “significance,” so can a plausible-sounding clinical one — subgroup and secondary findings need pre-specification and multiplicity correction before you believe them.

Examiner traps
  • Unadjusted multiplicity — many outcomes/tests reported with no Bonferroni or FDR correction.
  • Post-hoc subgroups — effects “found” after the fact, not pre-specified, are hypothesis-generating only.
  • Outcome switching — the primary outcome in the protocol differs from the one highlighted in the publication.
Quick check

A trial reports 1 significant result among 25 tested — convincing?
Answer: Not by itself — at α = 0.05 about one false positive is expected by chance across 25 tests. It needs multiplicity correction (or pre-specification) and ideally independent replication before you trust it.

Related cards

Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.

Share
0
    0
    Your Cart
    Your cart is emptyReturn to Shop