Home/Critical Appraisal
Critical Appraisal

Kappa Statistic

EM FINAL EXAMS Critical Appraisal · Measurement Kappa Statistic (κ) Agreement between two raters, after stripping out the agreement you would expect by chance. Definition Cohen’s kappa (κ) measures agreement between two raters — or a test against itself on retest — corrected for chance. κ = (Po − Pe) / (1 − Pe), where […]

EM FINAL EXAMS Critical Appraisal · Measurement

Kappa Statistic (κ)

Agreement between two raters, after stripping out the agreement you would expect by chance.

Definition

Cohen’s kappa (κ) measures agreement between two raters — or a test against itself on retest — corrected for chance. κ = (Po − Pe) / (1 − Pe), where Po is observed agreement and Pe is the agreement expected by chance alone. It runs from ~0 (no better than chance) to 1 (perfect); 0 means chance-level, negatives mean worse than chance.

The picture

Observed agreement 85% → kappa subtracts chance → κ ≈ 0.70

What it shows

100 cases scored by two raters. They agree on 85 (45 + 40), so observed agreement Po = 0.85. By chance, with each rater calling ~half the cases positive, you would expect ~0.50 agreement anyway, so Pe ≈ 0.50. Kappa = (0.85 − 0.50) / (1 − 0.50) ≈ 0.70 — substantial agreement that already discounts luck.

How to read it

Bigger is better. Landis & Koch (1977): <0.20 poor/slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, 0.81–1.00 almost perfect. The bands are conventions, not laws — for a high-stakes call (e.g. STEMI), even “substantial” may be uncomfortable. Use weighted kappa for ordered categories so a near-miss counts more than a wild miss.

Why it matters

Raw percent agreement is seductive but inflated — when one answer is common, two raters agree most of the time just by both guessing it. Kappa is the honest version: it asks how much agreement there is over and above chance, which is what you actually care about when you ask “can I trust a second reader?”

Key
  • κ = (Po − Pe) / (1 − Pe)
  • 0 = chance · 1 = perfect · <0 = worse than chance
  • Weighted κ for ordered categories · ICC for continuous data
Pitfall
Pitfall The kappa paradox: a high raw percent agreement can sit alongside a poor kappa when one category dominates (very high or very low prevalence of the finding). Most chance agreement is “built in”, so kappa is harshly penalised — report kappa and the marginal prevalences.
emfinalexams.com · FRCEM / MRCEM revision
EM trial in the wild

Inter-observer agreement on imaging findings — kappa is the standard metric when papers report how well two radiologists or EPs agree on a sign (e.g. a fracture line, free fluid on FAST, or a CT-head bleed). A study quoting “94% agreement” can still report only a “moderate” kappa (~0.4–0.6) once the dominant “normal” calls are discounted. Always read the kappa, not the headline percent — for rare findings, percent agreement is dominated by everyone agreeing the film is normal.

Examiner traps
  • Ignoring chance agreement — quoting raw percent agreement instead of kappa.
  • Prevalence / base-rate effects — the kappa paradox: high agreement, low kappa when one category dominates.
  • Weighted vs unweighted — unweighted kappa treats a near-miss the same as a gross disagreement on ordered scales.
Quick check

Why not just report the percent agreement between two raters?
Answer: Because raw percent agreement counts agreement that would happen by chance anyway; kappa subtracts that expected-by-chance agreement, giving the agreement attributable to genuine concordance.

Related cards

Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.

Share
0
    0
    Your Cart
    Your cart is emptyReturn to Shop