ROC Curve & AUC
How well a test separates disease from no-disease across every possible threshold.
A receiver operating characteristic (ROC) curve plots sensitivity against 1 − specificity across all possible cut-offs of a continuous test. The area under it (AUC, or c-statistic) summarises discrimination — the probability the test ranks a random diseased patient above a random non-diseased one. 0.5 = useless (a coin toss), 1.0 = perfect.
Curve hugging the top-left → AUC 0.85; the gold dot is one chosen cut-off
Every point on the curve is one possible threshold and its sensitivity / false-positive-rate pair. The further the curve bows toward the top-left corner, the better the test discriminates; the dashed diagonal is a coin toss (AUC 0.5). The gold dot is a single chosen operating point.
Read the whole curve for overall discrimination (the AUC), then read a single point for the sensitivity and specificity you would actually get at that cut-off. Sliding the operating point up-and-right buys sensitivity at the cost of specificity; down-and-left does the reverse.
AUC is the single most-quoted headline for a diagnostic or prognostic score — but it is a summary of ranking only. It tells you the test can separate cases from non-cases on average; it does not tell you which cut-off to use, nor whether the predicted probabilities are actually correct.
AUC 0.5 = useless·1.0 = perfect- Rough guide: 0.7–0.8 acceptable, 0.8–0.9 good, >0.9 excellent
- AUC = discrimination, not calibration or the right threshold
National Early Warning Score (NEWS) — in a cohort of 35,585 acute medical admissions, NEWS discriminated death within 24 h with an AUROC of 0.894 (95% CI 0.887–0.902), beating 33 other early-warning scores (Smith et al., Resuscitation 2013). Excellent discrimination — yet the trigger thresholds for escalation still have to be chosen separately, balancing missed deteriorations against alarm fatigue. A high AUC does not pick the cut-off for you: NEWS still needs defined trigger points, and its predicted risk must be calibrated to the local population to be trusted.
- Equating AUC with calibration — discrimination and calibration are different things.
- Forgetting you still must choose a threshold for the clinical purpose.
- Assuming a higher AUC is clinically better — it can miss differences that matter at the relevant cut-off.
Quick check
A score has AUC 0.9 — is it ready to use at any cut-off?
Answer: No — you must choose a threshold appropriate to the clinical purpose and check calibration. A high AUC means good ranking overall, not that any given cut-off performs well or that its predicted risks are accurate.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.