Statistical Power
The probability a study will detect a real effect — if one truly exists.
Power = 1 − β — the probability of detecting a true effect of a specified size, if it really exists. Conventionally set at ≥ 80%. It is driven by the effect size, the variability of the data, the chosen α, and above all the sample size. An underpowered study risks a Type II error (a false negative).
Power is the slice of the H1 curve beyond the critical value — everything that isn’t the β region.
The alternative hypothesis (H1) curve is where the truth lies when a real effect exists. The portion to the right of the critical value is the study’s power — the chance it lands in “significant” territory and detects the effect. The amber β slice to the left is the chance it misses (Type II error).
A larger sample narrows both curves, shrinking the β region and growing power. A bigger true effect slides H1 further right, also raising power. Conversely, an underpowered study leaves a fat β slice — so a real effect is quite likely to be reported as “non-significant”, with a wide, uninformative confidence interval.
Power is decided before a trial recruits — it is the sample-size calculation. An underpowered ED trial that returns “no significant difference” hasn’t shown the treatment doesn’t work; it may simply have been too small to tell. Reading that as proof of “no effect” is one of the commonest appraisal errors.
Power = 1 − β(aim ≥ 80%)- Driven most by the
sample size - Low power →
Type IIerror + wide CI
ATACH-2 (Qureshi et al., NEJM 2016) — 1,000 patients with acute intracerebral haemorrhage, intensive vs standard BP lowering. The power calculation assumed a 60% control event rate with a 10% absolute reduction, but only ~38% reached the primary outcome (death/disability, mRS 4–6: 38.7% vs 37.7%) → far fewer events than planned. Adjusted RR 1.04 (95% CI 0.85–1.27); the trial stopped early for futility. A CI running 0.85–1.27 cannot exclude a worthwhile benefit or harm — that is “inconclusive”, not “no effect”.
- Non-significance ≠ no effect — inspect the confidence interval before concluding a treatment doesn’t work.
- Post-hoc (“observed”) power is circular and uninformative — power is set a priori, in the sample-size calculation.
- An over-optimistic assumed effect size shrinks the required sample — the study is then underpowered for the real, smaller effect.
Quick check
An underpowered trial reports p=0.3 — does that prove there’s no effect?
Answer: No — it may be a Type II error from low power. Inspect the confidence interval: if it’s wide and includes a clinically important effect, the result is inconclusive, not negative.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.