Appraising a Diagnostic Study
Was the test compared blindly against a proper reference standard, in all patients, in the right population — and do the numbers transfer to my setting?
Structured appraisal of a diagnostic-accuracy study (CASP diagnostic / QUADAS-2 framework). Validity: was the index test compared independently and blindly against an appropriate reference (gold) standard, applied to all patients regardless of the index result, in an appropriate spectrum of patients — the real diagnostic-dilemma population? Results: sensitivity and specificity, likelihood ratios, and predictive values. Applicability: does the study population’s disease prevalence match my own setting?
Valid? → Results? → Applicable? — blind, complete, right-spectrum, right-prevalence.
The appraisal order for an accuracy study. Validity hinges on three design features: the reference standard must be applied blind to the index result, applied to everyone (not just those who tested positive), and the patients must span the realistic clinical spectrum rather than obvious cases and healthy controls. Only then do sensitivity, specificity and likelihood ratios mean what they claim — and only then can you ask whether they transfer to your prevalence.
At gate 1, hunt for the classic biases: did anyone interpreting the index test know the reference result (and vice versa)? Did test-negatives also get the gold standard, or were they assumed disease-free? Were the patients a true diagnostic dilemma or a stacked case-control sample? At gate 2, read sensitivity/specificity and likelihood ratios. At gate 3, remember predictive values are prevalence-dependent — do not import a PPV from a high-prevalence cohort into low-prevalence ED practice.
A test can look brilliant purely because of how it was studied. Spectrum bias inflates sensitivity when only florid cases are included; verification bias distorts accuracy when only test-positives reach the gold standard; and a predictive value lifted from a different-prevalence population can be dangerously wrong at the bedside. Appraising the design, not just the headline accuracy, is what protects your patient.
VALID? → RESULTS? → APPLICABLE?— tools: QUADAS-2, CASP diagnostic- Validity = blind · independent · proper reference standard · all patients · right spectrum
PPV/NPVmove with prevalence; sensitivity & specificity are (broadly) prevalence-stable
Ottawa Ankle Rules (Bachmann, BMJ 2003) — a systematic review pooling 27 studies (15,581 patients) reported a pooled sensitivity around 97.6% (roughly 96–99%) for excluding clinically significant ankle/mid-foot fracture, with modest specificity. Appraise it: the reference standard is radiography, applied across a realistic ED injury spectrum, with a high sensitivity that makes the rule a sound rule-out tool. Specificity is low, so a “positive” rule does not confirm a fracture — it only flags who needs an X-ray. Beware studies where test-negative patients never got radiography (verification bias) or where children/atypical injuries fall outside the validated spectrum.
- Verification / spectrum bias — only test-positives get the gold standard, or only obvious cases are studied.
- A weak or differential reference standard (an imperfect “gold” standard, or a different one for positives vs negatives).
- Prevalence transfer — importing a PPV/NPV from a population whose prevalence differs from yours.
Quick check
Why must ALL patients receive the reference standard?
Answer: To avoid verification (work-up) bias — if only test-positive patients are confirmed against the gold standard, the missed false-negatives distort sensitivity and specificity.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.