Effect Modification & Interaction
When a treatment genuinely works differently across subgroups — a real signal you REPORT, not a nuisance you adjust away.
Effect modification (statistical interaction) occurs when the size — or even direction — of a treatment or exposure effect genuinely differs across levels of a third variable: it helps in one subgroup and less (or not at all, or harms) in another. Unlike confounding, it is a true feature of the biology. You do not adjust it away; you describe it with stratified / subgroup estimates and test it with a formal interaction term, ideally pre-specified.
Same drug, opposite effects by timing — the effect is modified. You report each subgroup; you do not average them away.
Two subgroups of the same trial, plotted separately. One estimate sits clearly left of the null (benefit); the other sits right of it (harm). The point estimates land on opposite sides and the intervals barely overlap — that separation is the visual signature of effect modification, confirmed by a significant test for interaction.
Don’t read the gap between two subgroup p-values (“significant in one, not the other”) — that is a classic error. Read the formal test for interaction: it asks whether the two effect estimates genuinely differ from each other. A small interaction p-value, especially if pre-specified and biologically plausible, says the effect is truly modified and the subgroups must be reported separately.
Effect modification is real, actionable information — it tells you which patients to treat. Confounding is the opposite: a distortion to be removed. Mishandle the distinction and you either bury a genuine subgroup signal by pooling, or you chase a phantom one created by data-dredging. The treatment of the two is exactly reversed.
- Effect modification = effect
genuinely differsby subgroup → report it - Confounding = distortion → remove it (opposite handling)
- Detect with a
test for interaction, not two separate p-values - Trust it most when
pre-specified& biologically plausible
CRASH-2 — tranexamic acid by time-to-treatment (Lancet 2011) — the effect of TXA on death due to bleeding varied strongly with how early it was given (test for interaction p<0.0001). Given ≤1 h: RR 0.68 (0.57–0.82); 1–3 h: RR 0.79 (0.64–0.97); given >3 h it increased bleeding deaths. Treatment time is a genuine effect modifier — so practice is “give TXA early”, reported per stratum rather than as a single average. This is the right kind of subgroup: pre-specified, biologically plausible, and confirmed by a formal interaction test — not a post-hoc fishing trip.
- Post-hoc subgroups — the ISIS-2 authors famously showed aspirin “didn’t work” for patients born under Gemini or Libra to ridicule unplanned subgroup analysis (Lancet 1988). Chance alone manufactures such splits.
- Multiplicity — many subgroups mean many tests; without correction a spurious “significant” one is near-inevitable.
- Confusing effect modification (report it) with confounding (adjust it away) — opposite responses.
Quick check
A drug helps men but harms women — do you adjust this away?
Answer: No — that is effect modification, not confounding. You report the subgroup-specific effects (and confirm with a test for interaction); you do not average or adjust them into a single misleading estimate.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.