Kappa Statistic (κ)
Agreement between two raters, after stripping out the agreement you would expect by chance.
Cohen’s kappa (κ) measures agreement between two raters — or a test against itself on retest — corrected for chance. κ = (Po − Pe) / (1 − Pe), where Po is observed agreement and Pe is the agreement expected by chance alone. It runs from ~0 (no better than chance) to 1 (perfect); 0 means chance-level, negatives mean worse than chance.
| B says + | B says − | |
|---|---|---|
| A says + | 45Both + | 5Disagree |
| A says − | 10Disagree | 40Both − |
Observed agreement 85% → kappa subtracts chance → κ ≈ 0.70
100 cases scored by two raters. They agree on 85 (45 + 40), so observed agreement Po = 0.85. By chance, with each rater calling ~half the cases positive, you would expect ~0.50 agreement anyway, so Pe ≈ 0.50. Kappa = (0.85 − 0.50) / (1 − 0.50) ≈ 0.70 — substantial agreement that already discounts luck.
Bigger is better. Landis & Koch (1977): <0.20 poor/slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, 0.81–1.00 almost perfect. The bands are conventions, not laws — for a high-stakes call (e.g. STEMI), even “substantial” may be uncomfortable. Use weighted kappa for ordered categories so a near-miss counts more than a wild miss.
Raw percent agreement is seductive but inflated — when one answer is common, two raters agree most of the time just by both guessing it. Kappa is the honest version: it asks how much agreement there is over and above chance, which is what you actually care about when you ask “can I trust a second reader?”
κ = (Po − Pe) / (1 − Pe)- 0 = chance · 1 = perfect · <0 = worse than chance
Weighted κfor ordered categories ·ICCfor continuous data
Inter-observer agreement on imaging findings — kappa is the standard metric when papers report how well two radiologists or EPs agree on a sign (e.g. a fracture line, free fluid on FAST, or a CT-head bleed). A study quoting “94% agreement” can still report only a “moderate” kappa (~0.4–0.6) once the dominant “normal” calls are discounted. Always read the kappa, not the headline percent — for rare findings, percent agreement is dominated by everyone agreeing the film is normal.
- Ignoring chance agreement — quoting raw percent agreement instead of kappa.
- Prevalence / base-rate effects — the kappa paradox: high agreement, low kappa when one category dominates.
- Weighted vs unweighted — unweighted kappa treats a near-miss the same as a gross disagreement on ordered scales.
Quick check
Why not just report the percent agreement between two raters?
Answer: Because raw percent agreement counts agreement that would happen by chance anyway; kappa subtracts that expected-by-chance agreement, giving the agreement attributable to genuine concordance.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.