Home/Critical Appraisal
Critical Appraisal

Type I & Type II Errors

EM FINAL EXAMS Critical Appraisal · Statistics Type I & Type II Errors Concluding there’s a difference when there isn’t (Type I) versus missing a real one (Type II). Definition A Type I error (α) is rejecting a true null hypothesis — a false positive, conventionally accepted at α = 0.05. A Type II error […]

EM FINAL EXAMS Critical Appraisal · Statistics

Type I & Type II Errors

Concluding there’s a difference when there isn’t (Type I) versus missing a real one (Type II).

Definition

A Type I error (α) is rejecting a true null hypothesis — a false positive, conventionally accepted at α = 0.05. A Type II error (β) is failing to reject a false null — a false negative. Power = 1 − β: the probability of detecting a true effect of a given size, conventionally set at ≥ 80%.

The picture
critical value H0 H1 α = false positive (Type I) β = false negative (Type II) H0 true H1 true Correct Type II (β) Type I (α) Correct

Two distributions split by the critical value: the red tail (α) is the false positive, the amber region (β) the false negative.

What it shows

Two overlapping distributions — the null hypothesis (H0) and the alternative (H1) — split by the critical value, with the α tail and β region shaded, plus a truth × decision 2×2. Each shaded area is one of the two ways a study can reach the wrong conclusion.

How to read it

The red tail (α) is the false-positive rate; the amber region (β) the false-negative rate. Lowering α without increasing the sample size pushes the cut-off to the right and increases β — the two errors trade off against each other unless you add more patients.

Why it matters

Many “negative” EM trials are simply underpowered (Type II) — the effect was real but the study was too small to catch it. Conversely, testing many outcomes inflates Type I. Recognising both stops you over-believing a lone significant p-value and under-believing a non-significant one.

Key
  • Type I (α) = false positive (usually 0.05)
  • Type II (β) = false negative
  • Power = 1 − β (aim ≥ 80%)
Pitfall
Pitfall Equating a non-significant result with “no effect” — often a Type II error from low power. And testing many outcomes without adjusting for multiple comparisons inflates Type I until a “significant” finding is likely by chance alone.
emfinalexams.com · FRCEM / MRCEM revision
EM trial in the wild

ATACH-2 (Qureshi et al., NEJM 2016) — 1,000 patients with acute intracerebral haemorrhage, intensive vs standard BP lowering. Death or severe disability (mRS 4–6) at 3 months was 38.7% vs 37.7% → adjusted RR 1.04 (95% CI 0.85–1.27). The power calculation assumed a 60% control event rate, but only ~38% of patients reached the outcome — far fewer events than planned, so the trial was underpowered. A “no significant difference” with a CI spanning 0.85–1.27 is inconclusive, not proof of no effect — a worthwhile benefit or harm can’t be excluded.

Examiner traps
  • Non-significance ≠ no effect — a wide CI usually means the trial was underpowered (Type II).
  • Multiple comparisons inflate Type I — test enough outcomes and one will “reach” p<0.05 by chance.
  • Confusing α (the pre-set acceptable false-positive rate) with the p-value of one specific result.
Quick check

A trial reports p=0.20 for its primary outcome — proof the treatment doesn’t work?
Answer: No — it may be a Type II error from low power; check the confidence interval and the trial’s power calculation before concluding “no effect”.

Related cards

Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.

Share
0
    0
    Your Cart
    Your cart is emptyReturn to Shop