Heterogeneity (I²)
How much the results of studies in a meta-analysis differ beyond what chance alone would produce.
I² is the percentage of the total variation across studies that is due to real (between-study) heterogeneity rather than chance. It complements Cochran’s Q (χ²) test and τ² (the between-study variance).
Four study squares whose 95% CIs overlap only partly — that inconsistency is summarised as I² = 64% (moderate).
A forest plot: each study’s effect is a square (sized by its weight) sitting on its 95% CI, with a vertical line of no effect and a pooled diamond at the bottom. The I² statistic beneath it summarises how inconsistent those study results are.
Use rough bands: <25% low, ~50% moderate, >75% high heterogeneity. Visually, the more the individual study CIs fail to overlap, the higher the I² — the studies are not telling the same story.
Pooling clinically or statistically heterogeneous trials yields a meaningless average. A high I² should push you to a random-effects model and, more importantly, to explain the heterogeneity through subgroups and sensitivity analyses — not just report a single number and move on.
I² = % of total variation due to heterogeneity(not chance)<25% low · ~50% moderate · >75% high- High I² → random-effects + investigate, don’t just pool
Li et al., Cochrane Review 2007 (intravenous magnesium for acute MI) — pooling early mortality across 22 trials (72,476 participants) gave marked heterogeneity, I² = 64%. That inconsistency flipped the answer depending on the model: a fixed-effect analysis showed no benefit (OR 0.99, 95% CI 0.94–1.04) while a random-effects analysis showed an apparent large benefit (OR 0.66, 95% CI 0.53–0.82). The heterogeneity was driven by small early positive trials sitting against the huge neutral mega-trials (ISIS-4 dominated), with likely publication bias — so the authors urged caution rather than trusting the pooled number. High I² means explain it, don’t just average.
- Confusing statistical heterogeneity (I²) with clinical/methodological diversity of the trials.
- Using a fixed-effect model despite a high I².
- Treating I² as precise/absolute — it has a CI and rises with study size and number.
Quick check
I² = 0% — does that prove the trials are clinically homogeneous?
Answer: No — it indicates low statistical inconsistency, but clinical and methodological diversity (different populations, doses, comparators) can still make pooling inappropriate.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.