How to Pass the FRCEM Critical Appraisal Section
TL;DR — FRCEM CA section needs 6 things: study design recognition, p-values + CIs, NNT/NNH, sensitivity/specificity, common biases, and 1–2 statistical test names.
Last updated: 30 May 2026
Process at a glance
4 weeks out
p-values, CIs, NNT and NNH
RCTs, cohorts, case-control, systematic reviews
6 common types
with full debrief
30 SBA paper, 60 minutes
Critical appraisal is no longer examined as a standalone FRCEM paper or viva, but it remains a high-yield part of emergency medicine exams. In current UK EM exams, evidence interpretation appears mainly in SBA questions and may also influence OSCE reasoning, justification of management, and safe application of guidelines. The skill being tested is practical: identify the study design, spot the main threat to validity, interpret the result correctly, and decide whether the evidence should change UK emergency practice.
For MRCEM SBA, the level is usually more basic: common study designs, simple statistics, and safe interpretation. For FRCEM SBA, the questions are broader and more nuanced: confidence intervals, diagnostic studies, non-inferiority, meta-analysis, and applicability to NHS practice. For FRCEM OSCE, you are unlikely to face a formal appraisal station, but you may need to justify decisions using evidence, guidelines, and risk-benefit reasoning.
Why the FRCEM Critical Appraisal Section Matters
Emergency medicine depends on rapid decisions made under uncertainty. Many of those decisions are driven by evidence: whether a troponin pathway is safe, whether age-adjusted D-dimer reduces unnecessary imaging, whether a risk score is ready for use, or whether a new intervention offers meaningful benefit.
In real practice, poor appraisal leads to two common errors:
- Adopting weak evidence too early
- Rejecting useful evidence because the statistics are misunderstood
In the exam, the same errors lose marks. Candidates often memorise isolated definitions but struggle when shown an abstract, a 2 x 2 table, a confidence interval, or a forest plot. The exam rewards structured judgement, not detached statistical trivia.
The core consultant skill is to answer five questions quickly:
- What is the clinical question?
- What study design is this?
- What is the biggest threat to validity?
- What do the results actually mean?
- Would this change UK emergency practice?
Key Definitions
| Term | Meaning | Exam relevance |
|---|---|---|
| Internal validity | Whether the study result is likely to be true for the patients studied | Usually the first thing to judge before applicability |
| External validity | Whether the findings apply to your patients and setting | Important for NHS and ED applicability |
| Bias | Systematic error that distorts the result | Questions often ask for the single most important bias |
| Confounding | A third factor associated with both exposure and outcome, creating a misleading association | Key issue in observational studies |
| Randomisation | Allocation by chance to reduce selection bias and confounding | Core feature of treatment trials |
| Allocation concealment | Preventing recruiters knowing the next treatment assignment | Protects randomisation; often more important than blinding |
| Blinding | Keeping participants, clinicians, or assessors unaware of allocation | Reduces performance and detection bias |
| Intention-to-treat analysis | Analysing patients in the groups to which they were randomised | Preserves benefits of randomisation |
| Confidence interval | Range of values compatible with the data, reflecting precision | More useful than p value alone |
| p value | Probability of observing the data, or more extreme, if the null hypothesis were true | Does not measure effect size or clinical importance |
| Absolute risk reduction | Difference in event rate between groups | More clinically useful than relative risk reduction |
| Number needed to treat | Number of patients needing treatment for one additional beneficial outcome | Derived from absolute risk reduction |
| Sensitivity | Proportion of people with disease who test positive | Useful for rule-out when high and combined with appropriate context |
| Specificity | Proportion of people without disease who test negative | Useful for rule-in when high and combined with appropriate context |
| PPV / NPV | Probability of disease given a positive or negative test | Depend on prevalence; do not transfer unchanged between settings |
| Likelihood ratio | How much a test result changes disease probability | High-yield for diagnostic SBA questions |
| Heterogeneity | Variation in results between studies in a meta-analysis | Often summarised by I² |
| Non-inferiority trial | Trial designed to show a new treatment is not unacceptably worse than standard treatment by a pre-specified margin | Common exam trap |
Essential Pathophysiology
Critical appraisal has no biological pathophysiology, but there is an underlying logic to how evidence becomes trustworthy or misleading.
- Bias distorts the estimate away from the truth
- Confounding creates false associations in non-randomised studies
- Random error causes imprecision, often seen as wide confidence intervals
- Poor outcome choice can make a study statistically positive but clinically unhelpful
- Poor applicability means a valid study may still not fit UK ED practice
Think of evidence failure in three layers:
- Was the study designed appropriately for the question?
- Was it conducted well enough to trust the result?
- Even if valid, does it apply to my patients, pathways, and resources?
This is the framework examiners expect.
Clinical Presentation
In exam terms, critical appraisal usually presents in one of a small number of ways:
- A short abstract asking for the study design
- A results table asking for the best interpretation
- A confidence interval asking whether the result is significant, precise, or clinically important
- A diagnostic 2 x 2 table asking for sensitivity, specificity, predictive values, or likelihood ratios
- A forest plot asking about pooled effect or heterogeneity
- A paper summary asking whether the authors’ conclusion is justified
- A clinical scenario asking whether evidence should change UK practice
The commonest paper types are:
- Randomised controlled trials
- Diagnostic accuracy studies
- Cohort studies
- Case-control studies
- Systematic reviews and meta-analyses
- Clinical prediction rule studies
- Before-and-after service evaluations
- Non-inferiority trials
Red Flags and High-Risk Features
These are the features that should make you cautious in both exams and practice.
- Observational study making causal claims
- Single-centre study with highly selected patients
- Surrogate outcomes instead of patient-important outcomes
- Large relative effect but tiny absolute benefit
- Wide confidence interval compatible with both benefit and harm
- Non-significant superiority trial described as proving equivalence
- Positive subgroup analysis despite neutral primary outcome
- Diagnostic study with poor or inconsistent reference standard
- Verification bias: not all patients receive the reference standard
- Spectrum bias: study population unlike real ED case mix
- Derivation study presented as if ready for implementation
- Systematic review pooling poor-quality or very heterogeneous studies
- Service redesign paper ignoring secular trends or co-interventions
- Single paper contradicting established NICE guidance without sufficient supporting evidence
Differential Diagnosis
When a question asks for the “best criticism” or “most likely interpretation”, your differential diagnosis is the list of possible flaws. The key is to choose the dominant one.
| If the paper is about | Main alternative explanations to consider |
|---|---|
| Treatment | Poor randomisation, lack of allocation concealment, baseline imbalance, loss to follow-up, no intention-to-treat, unblinded outcome assessment |
| Diagnosis | Bad reference standard, verification bias, incorporation bias, spectrum bias, lack of blinding, case-control design for test accuracy |
| Association / prognosis | Confounding, selection bias, reverse causation, incomplete follow-up, poor adjustment |
| Meta-analysis | Poor included studies, publication bias, major heterogeneity, inappropriate pooling |
| Clinical prediction rule | Derivation only, no external validation, poor calibration, unclear impact on management |
| Service evaluation | Secular trends, co-interventions, case-mix change, regression to the mean |
Initial ED Assessment
The fastest safe approach in the exam is a structured five-step appraisal.
Step 1: Define the clinical question
Decide whether the paper is about treatment, diagnosis, prognosis, risk prediction, harm, or service delivery. If you misclassify the question, you will often misjudge the design.
Step 2: Identify the study design
This is the highest-yield first move.
| Question type | Best usual design | Common exam clue |
|---|---|---|
| Treatment | Randomised controlled trial | Patients allocated to intervention or control |
| Diagnosis | Diagnostic accuracy study | Index test compared with reference standard |
| Prognosis | Cohort study | Patients followed over time for outcomes |
| Risk factor / association | Cohort or case-control study | Exposure compared with outcome occurrence |
| Evidence summary | Systematic review / meta-analysis | Multiple studies pooled |
| Risk tool | Derivation / validation study | Score or rule created or tested |
Step 3: Judge internal validity first
Ask what single flaw most threatens the result. In SBA questions, the best answer is usually the biggest problem, not a list of minor imperfections.
Step 4: Interpret the result in plain English
Translate statistics into a clinical statement.
- What is the direction of effect?
- How large is it?
- How precise is it?
- Is it clinically important?
Step 5: Assess applicability to UK emergency practice
Even a valid study may not fit your setting.
- Does the population resemble NHS ED patients?
- Does the comparator reflect current UK practice?
- Does the pathway depend on a specific assay or resource?
- Is it consistent with NICE, RCEM, BTS, SIGN, or local governance?
Investigations
For exam purposes, your “investigations” are the tools used to interrogate a paper.
High-yield statistics you must know
| Statistic | What it tells you | High-yield interpretation |
|---|---|---|
| Risk ratio / relative risk | Relative event rate between groups | RR 1 means no difference |
| Odds ratio | Odds of outcome in one group versus another | OR approximates RR when outcomes are rare; common in case-control studies |
| Hazard ratio | Relative event rate over time | Used in time-to-event analyses; HR 1 means no difference |
| Absolute risk reduction | Absolute difference in event rates | More clinically meaningful than relative reduction |
| Number needed to treat | 1 / absolute risk reduction | Lower NNT usually means larger benefit |
| Confidence interval | Precision and plausible range of effect | Wide CI means imprecision; may include benefit and harm |
| p value | Compatibility with null hypothesis | Does not tell you clinical importance |
| Sensitivity | True positive rate | High sensitivity helps rule out only in the right context |
| Specificity | True negative rate | High specificity helps rule in only in the right context |
| PPV / NPV | Probability of disease after test result | Depend on prevalence |
| LR+ | How much a positive test increases disease probability | Higher is better for ruling in |
| LR- | How much a negative test decreases disease probability | Lower is better for ruling out |
| I² | Proportion of variation due to heterogeneity | Higher values suggest more inconsistency between studies |
Confidence intervals: the exam-safe approach
- For ratio measures such as RR, OR, or HR, a CI crossing 1 is compatible with no effect
- For difference measures, a CI crossing 0 is compatible with no effect
- A narrow CI suggests precision
- A wide CI suggests imprecision and may be compatible with both clinically important benefit and harm
- A non-significant result does not prove no difference
- A superiority trial that fails to show significance does not prove equivalence or non-inferiority
Absolute versus relative effects
Examiners like papers that overstate benefit using relative measures.
Example:
- Control event rate 2%
- Treatment event rate 1%
- Relative risk reduction 50%
- Absolute risk reduction 1%
- NNT 100
The relative effect sounds dramatic. The absolute benefit is modest. In exam questions, the best answer often highlights this distinction.
Diagnostic test interpretation
| Concept | Key point |
|---|---|
| Sensitivity / specificity | More stable test characteristics than predictive values |
| PPV / NPV | Change with prevalence and case mix |
| Likelihood ratios | Useful for moving from pre-test to post-test probability |
| Rule-out strategy | Needs low enough post-test risk, not just a “good” sensitivity |
| Reference standard | Must be appropriate and applied consistently |
Forest plots and meta-analysis basics
- Each study has a point estimate and confidence interval
- The pooled estimate is usually shown as a diamond
- If the pooled CI crosses the line of no effect, the meta-analysis is compatible with no overall effect
- Look for heterogeneity before trusting the pooled result
- Major clinical or methodological heterogeneity may make pooling inappropriate
- A meta-analysis of poor studies does not become strong evidence simply because it is pooled
Management in the Emergency Department
In practice and in the exam, management means how you handle evidence safely.
Immediate approach in the exam
- Read the question stem before the abstract or table if possible
- Identify the clinical question and study design
- Look for the primary outcome, not just any positive result
- Find the main effect estimate and confidence interval
- Ask what the biggest validity problem is
- Decide whether the conclusion is proportionate
- Check UK applicability before choosing a practice-changing answer
Paper-type-specific appraisal
Randomised controlled trials
Use RCTs mainly for treatment questions.
- Was randomisation truly random?
- Was allocation concealed?
- Were groups similar at baseline?
- Was blinding adequate where feasible?
- Was follow-up complete enough?
- Were patients analysed by intention to treat?
- Was the primary outcome clinically important?
- Was the effect size meaningful, not just statistically significant?
Common traps:
- Focusing on a positive secondary outcome when the primary outcome was neutral
- Treating a non-significant result as proof of no difference
- Ignoring crossover or major loss to follow-up
Diagnostic accuracy studies
These are extremely common in emergency medicine.
- Was there an appropriate reference standard?
- Did all or nearly all patients receive both index test and reference standard?
- Were index test and reference standard interpreted independently?
- Was the population representative of real ED patients?
- Was the threshold pre-specified?
- Were indeterminate results handled properly?
Common biases:
- Verification bias
- Incorporation bias
- Spectrum bias
- Review bias from lack of blinding
Common traps:
- Assuming PPV and NPV will be the same in a different setting
- Accepting a case-control diagnostic study as if it reflects real-world test performance
- Ignoring whether the rule-out miss rate is acceptable in UK practice
Cohort studies
Useful for prognosis, risk factors, and some harms questions.
- Were exposed and unexposed groups comparable?
- Was follow-up adequate and complete?
- Were outcomes measured objectively?
- Were important confounders measured and adjusted for?
- Is the association biologically and clinically plausible?
Common trap: concluding causation from association.
Case-control studies
Often used for rare outcomes.
- Were cases and controls selected appropriately?
- Were they drawn from the same underlying population?
- Was exposure measured similarly in both groups?
- Is recall bias likely?
- Were important confounders addressed?
Common trap: interpreting odds ratio as if it were always the same as relative risk.
Systematic reviews and meta-analyses
- Was the review question focused?
- Was the search strategy adequate?
- Were inclusion criteria appropriate?
- Was study quality assessed?
- Were the included studies sufficiently similar to pool?
- Was heterogeneity explored?
- Is publication bias likely?
Common trap: assuming a pooled result is reliable despite poor included studies or marked heterogeneity.
Clinical prediction rules
These are highly relevant in ED practice.
- Was the rule derived in an appropriate cohort?
- Has it been externally validated?
- Does it show acceptable discrimination and calibration?
- Has it been tested in an impact analysis to show it improves care?
- Does it outperform or safely complement current practice?
Exam rule: derivation is not enough. Validation is better. Impact analysis is strongest.
Before-and-after service studies
Common in operational and pathway papers.
- Were there secular trends over time?
- Were there other changes occurring at the same time?
- Did case mix change?
- Were outcomes measured consistently before and after?
Common trap: attributing all improvement to the intervention.
Non-inferiority and equivalence trials
These are common exam traps.
- Was there a pre-specified non-inferiority margin?
- Was the margin clinically justified?
- Was the analysis appropriate?
- Did the confidence interval stay within the non-inferiority boundary?
Key exam rule: a superiority trial with a non-significant result does not prove non-inferiority or equivalence.
Later consolidation for revision
- Practise one treatment paper, one diagnostic paper, and one meta-analysis each week
- Summarise each in six lines: question, design, main bias, main result, applicability, bottom line
- Use question banks to convert theory into SBA decisions
Disposition, Referral and Follow-Up
For evidence questions, disposition means deciding what the evidence supports.
| Evidence pattern | Safest conclusion |
|---|---|
| Well-designed study with clinically important effect and good applicability | May support change in practice, especially if consistent with guidance |
| Valid study but narrow benefit in selected population | May support cautious use in similar patients only |
| Promising derivation study or single-centre pathway paper | Supports further validation, not immediate widespread adoption |
| Neutral or imprecise study | No clear evidence of benefit; may be underpowered |
| Study conflicts with established NICE or local pathway | Interesting evidence, but usually insufficient alone to change UK practice |
In the NHS, single studies rarely override established guidance on their own. In exam questions, the safest answer is often the balanced one: the evidence is interesting, but implementation depends on validation, guideline context, assay availability, governance, and acceptable risk.
Special Groups
Critical appraisal is not disease-specific, but some groups matter because applicability changes.
Paediatrics
- Do not assume adult evidence applies to children
- Check age range, outcomes, and whether paediatric pathways differ from NICE or local guidance
Pregnancy
- Diagnostic pathways and acceptable investigations may differ
- Applicability of non-pregnant adult studies is often limited
Older adults
- Frailty, multimorbidity, and atypical presentation may reduce applicability of highly selected trial populations
- Risk scores often perform differently in older cohorts
Immunosuppressed patients
- Often under-represented in trials
- Diagnostic test performance and baseline risk may differ substantially
Exam point: if a study excludes the group in front of you, external validity is limited even if internal validity is good.
Common Pitfalls
- Memorising definitions without being able to apply them
- Failing to identify the study design
- Listing many minor biases instead of choosing the main one
- Equating statistical significance with clinical importance
- Ignoring confidence intervals and focusing only on p values
- Confusing association with causation
- Assuming non-significant means no effect
- Assuming predictive values transfer unchanged between settings
- Accepting subgroup analyses too readily
- Treating derivation studies as practice-changing
- Ignoring whether the comparator reflects current UK care
- Ignoring NICE, RCEM, BTS, SIGN, local assay-specific pathways, or governance constraints
FRCEM and MRCEM Exam Tips
For MRCEM SBA:
- Know the common study designs
- Be able to interpret sensitivity, specificity, predictive values, and simple confidence intervals
- Recognise basic bias and overclaiming
For FRCEM SBA:
- Expect more nuance around confidence intervals, absolute versus relative effect, non-inferiority, forest plots, and applicability
- Be able to identify the single most important flaw
- Be cautious with pathway papers and risk tools
For FRCEM OSCE:
- You are unlikely to appraise a paper formally
- You may need to justify a decision using evidence and guidelines
- Safe answers acknowledge uncertainty, current standards, and patient context
High-yield revision plan for the final 8 weeks:
| Time before exam | Priority |
|---|---|
| 8 to 6 weeks | Revise study designs, core bias, confidence intervals, diagnostic statistics |
| 6 to 4 weeks | Practise RCTs, diagnostic studies, cohort studies, meta-analyses |
| 4 to 2 weeks | Timed SBA practice; focus on common traps and UK applicability |
| Final 2 weeks | Rapid abstract interpretation, forest plots, non-inferiority, decision rules, high-yield guideline context |
Useful revision habits:
- Read one abstract a day and identify design, main flaw, and bottom line
- Practise converting statistics into plain English
- Revise common UK examples: troponin pathways, age-adjusted D-dimer, head injury imaging, syncope tools
How This Appears in SBA Questions
Typical question stems
- Which study design is described?
- What is the most important limitation of this study?
- What does this confidence interval imply?
- Which statement best interprets the results?
- Which conclusion is most justified?
- Should this evidence change current UK emergency practice?
- What is the sensitivity / specificity / PPV / NPV / likelihood ratio?
- What does this forest plot show?
- Does this study demonstrate superiority, non-inferiority, or neither?
Key discriminator clues
| Clue in stem | What to think |
|---|---|
| Patients randomly assigned | RCT; check allocation concealment and intention to treat |
| Index test compared with gold standard | Diagnostic study; check reference standard and verification bias |
| Patients followed over time | Cohort study; think confounding and follow-up |
| Cases with disease compared with controls | Case-control study; think selection and recall bias |
| Pooled studies with diamond on plot | Meta-analysis; think heterogeneity and study quality |
| New score derived from predictors | Clinical prediction rule; ask whether externally validated |
| Before and after a pathway change | Service evaluation; think secular trends and co-interventions |
Common wrong answer traps
- “No significant difference” therefore “the treatments are equivalent”
- “The p value is below 0.05” therefore “the effect is clinically important”
- “The subgroup result was positive” therefore “the treatment works”
- “The test had a high NPV in one study” therefore “it will rule out disease everywhere”
- “The observational study showed an association” therefore “the exposure causes the outcome”
- “The meta-analysis is positive” therefore “the evidence is strong”, despite poor included studies
- “The new score performed well in derivation” therefore “it should replace current pathways”
Worked mini-examples
Example 1: Confidence interval
An RCT reports RR 0.82, 95% CI 0.60 to 1.12 for the primary outcome.
- Correct interpretation: the study does not show a statistically significant difference; the estimate is imprecise and remains compatible with benefit or no effect
- Wrong interpretation: the treatments are equivalent
Example 2: Diagnostic study
A D-dimer study reports NPV 99.5% in a low-risk ambulatory cohort.
- Correct interpretation: the result may not transfer unchanged to a higher-risk or different-prevalence population
- Wrong interpretation: the same NPV applies in all ED settings
Example 3: Prediction rule
A syncope score is derived in one centre and shows excellent discrimination.
- Correct interpretation: promising derivation study; external validation is needed before routine implementation
- Wrong interpretation: ready to replace current assessment
Example 4: Non-inferiority
A superiority trial comparing two analgesic strategies finds p = 0.18.
- Correct interpretation: no evidence of superiority; this does not prove non-inferiority
- Wrong interpretation: the treatments are non-inferior
Example 5: Meta-analysis
A pooled estimate favours treatment, but I² is high and included studies vary markedly in population and protocol.
- Correct interpretation: heterogeneity limits confidence in the pooled result
- Wrong interpretation: the pooled estimate alone is sufficient to change practice
Key Takeaways
- There is no standalone current FRCEM critical appraisal paper, but evidence interpretation remains heavily testable in SBA format.
- Start by identifying the clinical question and study design.
- Choose the biggest threat to validity, not a long list of minor flaws.
- Interpret confidence intervals, not just p values.
- Distinguish statistical significance from clinical importance.
- Use absolute effects to judge real benefit.
- In diagnostic studies, check the reference standard, blinding, verification, and prevalence effects.
- Derivation studies do not usually justify immediate practice change.
- A non-significant superiority trial does not prove equivalence or non-inferiority.
- Single studies rarely override NICE guidance or established NHS pathways on their own.
- The best SBA answer is usually balanced, accurate, and proportionate.
Further Reading
- RCEM curriculum and current RCEM examination information
- RCEM Learning
- NICE guidance relevant to emergency medicine pathways, including chest pain, VTE, head injury, stroke and sepsis
- Resuscitation Council UK guidelines
- British Thoracic Society guidance where relevant to respiratory and pleural topics
- SIGN guidelines where relevant to UK practice
- CASP checklists for RCTs, cohort studies, case-control studies, diagnostic studies and systematic reviews
- Users’ Guides to the Medical Literature for practical appraisal methods
Related on EM Final Exams
- P Values Confidence Intervals and Bias Explained Simply
- Most Common FRCEM Critical Appraisal SBA Questions
- How Hard is the FRCEM Exam
- SBA Question Dissection How to Break Down Any Question in 30 Seconds
Authoritative Sources
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.
