Hypothesis Testing & the P-value
The p-value is the probability of data this extreme IF the null were true — not the probability the null is true.
Hypothesis testing pits a null hypothesis (H0: no difference) against an alternative (H1). The p-value is the probability of observing data as extreme as — or more extreme than — yours if the null were true. By convention p < 0.05 is called “statistically significant” and we reject H0 — but 0.05 is an arbitrary line, not a law of nature.
The red tail is the p-value — how often chance alone would produce a result this extreme if the null held, not the chance the null is true.
The bell curve is the spread of results you’d expect if H0 were true. Your observed result sits out in the tail; the shaded area beyond it — in both directions for a two-sided test — is the p-value. The further out the result, the smaller that tail, the smaller the p.
Read the p as a measure of compatibility with the null, not of truth. A small p means the data would be surprising under H0, so we doubt it. It does not tell you the probability the null is true, the chance the result is “due to chance”, or how big or important the effect is.
A trial can be statistically significant yet clinically trivial, or non-significant yet promising but underpowered. Worshipping the 0.05 cut-off — “dichotomania” — collapses a continuous measure of evidence into a misleading yes/no and invites mis-statement of what p means.
p = P(data this extreme | H₀ true)- p is not P(H₀ true) and not effect size
- Statistical significance ≠ clinical importance
Same p, different stories — imagine two ED analgesia trials that both report p = 0.04. Trial A finds a 12 mm drop on a 100 mm pain scale; Trial B, a huge study, finds a 1 mm drop. Identical p-values, yet only one is clinically meaningful — the p-value alone cannot tell them apart. A small p flags that an effect is unlikely to be pure chance; it says nothing about how large or useful the effect is. Always read the effect size and its confidence interval next to the p.
- Reading p as P(null true) or as “the probability the result is due to chance” — it is neither.
- Confusing statistical significance with clinical significance.
- Dichotomania at 0.05, and ignoring multiplicity — test enough outcomes and one crosses 0.05 by chance.
Quick check
p = 0.04 — does it mean a 96% probability the drug works?
Answer: No — p is P(data | null), not P(null | data). It’s the chance of data this extreme if there were truly no effect, not the chance the treatment works.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.