Reliability & Reproducibility
Whether a measurement gives the same answer on repeat — consistency, not correctness.
Reliability (reproducibility) is the consistency of a measurement: repeat it and you get the same result. Validity is accuracy: it gives the true value. The two are independent — a measure can be reliable but invalid (consistently wrong). Reliability has three flavours: inter-observer (different raters agree), intra-observer (the same rater agrees with themselves on re-reading) and test–retest (the same instrument agrees over time).
| B says + | B says − | |
|---|---|---|
| A says + | 80Both agree + | 30Disagree |
| A says − | 20Disagree | 870Both agree − |
Agreement reads down the diagonal · the off-diagonal cells are disagreements
Two clinicians independently reading the same 1,000 films. They agree on 950 (80 + 870) and disagree on 50 — raw agreement 95%. The table captures consistency between observers; it says nothing about whether either reading is correct, because there is no “true disease” column here.
Read down the diagonal for agreement and across the off-diagonal for disagreement. High agreement = reliable. But note: if almost every film is normal, two raters could agree most of the time just by both saying “normal” — so raw agreement flatters reliability when one category dominates. That is why agreement is corrected for chance using kappa.
An unreliable measure injects noise into every study that uses it, biasing effects towards the null and making findings hard to reproduce. But reliability is necessary, not sufficient: a miscalibrated D-dimer assay or a mis-set scale is wonderfully reliable and uniformly wrong. Examiners love this gap — precise is not the same as accurate.
- Reliable = consistent / repeatable · Valid = accurate / true
- Types:
inter-observer·intra-observer·test–retest - Quantify agreement with
kappa(categorical) orICC(continuous), not raw %
Inter-observer reliability of ECG interpretation — agreement between emergency physicians reading the same ECGs is only modest for subjective calls. For ST-elevation MI, studies repeatedly show meaningful disagreement on whether STEMI criteria are met (kappa often only fair–moderate, ~0.3–0.6), which is why second reads and serial ECGs are recommended. High reliability between readers still would not prove the ECG call is correct — that needs a true reference standard (angiography / troponin), i.e. validity.
- Reliable ≠ valid — consistency does not imply accuracy; a measure can be precisely wrong.
- Precision vs accuracy — tight scatter (precise) is not the same as hitting the bullseye (accurate).
- Raw % agreement ignores chance — when one category dominates it overstates reliability; use kappa.
Quick check
A test gives the same wrong answer every single time. Is it reliable? Is it valid?
Answer: Reliable but not valid — it is perfectly consistent (reproducible) yet systematically inaccurate, so it is reliably wrong.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.