Home/Critical Appraisal
Critical Appraisal

Hierarchy of Evidence

EM FINAL EXAMS Critical Appraisal · Study design Hierarchy of Evidence A ranking of study designs by how susceptible they are to bias. Definition A traditional pyramid ordering designs by internal validity / risk of bias for therapy questions — from systematic reviews & meta-analyses of RCTs at the top, down to expert opinion at […]

EM FINAL EXAMS Critical Appraisal · Study design

Hierarchy of Evidence

A ranking of study designs by how susceptible they are to bias.

Definition

A traditional pyramid ordering designs by internal validity / risk of bias for therapy questions — from systematic reviews & meta-analyses of RCTs at the top, down to expert opinion at the base.

The picture
1 2 3 4 5 6 Systematic reviews & meta-analyses RCTs Cohort studies Case-control studies Case series & reports Expert opinion higher = lower risk of bias

Narrow & stronger at the top → wide & weaker at the base.

What it shows

The pyramid tiers: systematic reviews & meta-analyses → RCTs → cohort → case-control → case series & reports → expert opinion. Narrow and stronger at the top; wide and weaker at the base.

How to read it

Higher generally means lower risk of bias — but the best design depends on the question (RCT for therapy, cohort for prognosis/harm, cross-sectional for diagnosis), and quality within a tier varies enormously.

Why it matters

It is a fast triage of evidence strength — but applied blindly it misleads. A flawed RCT can be weaker than a strong cohort, and the wrong design for the question ranks poorly regardless of tier.

Key
  • Higher tier = lower risk of bias (in general)
  • Best design depends on the question
  • GRADE rates certainty of evidence, not just design
Pitfall
Pitfall Treating the hierarchy as absolute: ignoring study quality within a tier, or using the wrong design for the question type, can rank good evidence below bad.
emfinalexams.com · FRCEM / MRCEM revision
EM trial in the wild

Therapeutic hypothermia after cardiac arrest — watch one question climb the pyramid. Two small RCTs (HACA and Bernard, both NEJM 2002) cooled comatose survivors to ~33°C and reported better neurological outcomes vs no temperature control → cooling entered guidelines. The larger TTM trial (NEJM 2013) found 33°C no better than 36°C, and TTM2 (NEJM 2021, ~1900 patients) found targeted hypothermia no better than normothermia for death or function. The early “positive” signal sat near the top of the pyramid (RCTs) yet was overturned by bigger, better RCTs and pooled analyses — tier alone never settles a question; size, quality and the comparator do.

Examiner traps
  • Treating the hierarchy as absolute — tier rank is not a verdict.
  • Ignoring within-tier quality — a small, biased RCT can sit below a rigorous cohort.
  • Using the wrong design for the question (e.g. an RCT where a cohort answers prognosis/harm).
Quick check

Is a large, well-conducted cohort study always weaker than any RCT?
Answer: No — a rigorous cohort can outrank a small, biased RCT; fitness-for-question and quality matter, not just tier.

Related cards

Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.

Share
0
    0
    Your Cart
    Your cart is emptyReturn to Shop