QUADAS-2
The standard tool for appraising risk of bias and applicability in diagnostic accuracy studies — not RoB 2, which is for RCTs.
QUADAS-2 (Whiting et al., Ann Intern Med 2011) is the recommended instrument for assessing the quality of diagnostic test accuracy studies. It works across four domains: (1) patient selection, (2) index test, (3) reference standard and (4) flow & timing. Every domain is rated for risk of bias; the first three are also rated for concerns about applicability (does this study match my review question?).
Four domains rated for risk of bias; the first three also for applicability.
The journey of a patient through a test-accuracy study — selection → index test and reference standard → the 2×2 table that yields sensitivity and specificity — with each QUADAS-2 domain mapped onto the step it polices. Flow & timing sits across the whole path: it captures whether everyone got the same reference standard, and whether the two tests were close enough in time to be comparing like with like.
Walk down the diagram and ask one question per domain. Was the spectrum of patients representative, or a convenient case–control sample? (patient selection). Was the index test interpreted blind to the reference result? (index test). Is the reference standard a credible truth, read blind to the index? (reference standard). Did everyone get the same reference standard, at an appropriate interval? (flow & timing). Then, for the first three, ask the separate applicability question: do these patients, this test and this standard match the question I actually care about?
Diagnostic studies have their own bias structure — partial verification, differential verification, an imperfect reference standard, spectrum bias — that an RCT tool simply cannot see. Reach for RoB 2 or a Jadad score and you appraise the wrong things entirely. QUADAS-2 is what GRADE and Cochrane diagnostic reviews use, and it is what the exam expects when the paper in front of you reports sensitivity and specificity rather than a relative risk.
4 domains= patient selection · index test · reference standard · flow & timing- All four rated for risk of bias; first three also for applicability
Diagnostic accuracy → QUADAS-2(RCT → RoB 2 / Jadad)
Applying QUADAS-2 in practice — take an ED study of point-of-care lung ultrasound for pneumonia (the index test) against CT or a clinical reference standard. Patient selection: a convenience sample of stable, scanned patients risks spectrum bias and poor applicability to the breathless undifferentiated ED arrival. Index test: was the sonographer blind to the CT result? Reference standard: is the chosen truth credible, and read blind? Flow & timing: did everyone get the reference standard, or only the ultrasound-positives (partial verification), and were the two done close enough in time? Each domain earns a low / high / unclear risk-of-bias rating, and the first three a separate applicability rating. QUADAS-2 gives a structured risk-of-bias judgement per domain — it deliberately does not roll up into a single numeric quality score.
- Reaching for the wrong tool — using RoB 2 or Jadad (RCT instruments) on a diagnostic accuracy study instead of QUADAS-2.
- Ignoring flow & timing — missing partial or differential verification, the diagnostic biases most often overlooked.
- Conflating risk of bias with applicability — a study can be internally sound yet answer a different question to yours.
Quick check
Which tool appraises a diagnostic accuracy study’s risk of bias?
Answer: QUADAS-2 — not RoB 2. RoB 2 and Jadad are for randomised trials; diagnostic accuracy studies have a distinct bias structure (e.g. partial verification under flow & timing) that only QUADAS-2 captures.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.