Most Common FRCEM Critical Appraisal SBA Questions
TL;DR — The recurring FRCEM CA question types: identify the bias, calculate NNT, interpret a CI, pick the right test, judge external validity. 5 patterns dominate.
Last updated: 30 May 2026
Critical appraisal is no longer examined as a standalone written paper, but it remains heavily tested in MRCEM SBA, FRCEM SBA and often influences FRCEM OSCE discussion. In practice, the exam rarely announces itself as “statistics”. Instead, it hides appraisal inside ordinary emergency medicine stems: a trial result in sepsis, a diagnostic pathway for chest pain, a confidence interval in trauma, or a question asking whether new evidence should change UK practice. The aim is not advanced epidemiology. The aim is safe interpretation of evidence in the ED.
A reliable approach is to ask three questions:
- Is the study likely to be true: internal validity
- How certain is the estimate: precision
- Does it apply to my patient and my department: external validity
Most appraisal SBAs can be solved by identifying which of those domains is being tested.
Why These Most Common FRCEM Critical Appraisal Questions Matter
Emergency medicine depends on rapid decisions under uncertainty. We use evidence to decide who needs imaging, who can be discharged, whether a biomarker is useful, and whether a new intervention should alter practice. Poor appraisal leads to predictable errors:
- adopting statistically significant but clinically trivial interventions
- overinterpreting underpowered negative trials
- misusing diagnostic tests outside validated pathways
- confusing relative benefit with absolute benefit
- assuming a study from a very different healthcare system applies directly to a UK ED
In exam terms, this is high-yield because the same concepts recur repeatedly. In clinical terms, it matters because NICE, RCEM, BTS, SIGN and Resuscitation Council UK guidance all depend on correct interpretation of evidence rather than isolated headline results.
Key Definitions
| Term | Meaning | Exam point |
|---|---|---|
| Internal validity | Whether the study result is likely to be true for the participants studied | Think bias, confounding, randomisation, blinding, follow-up |
| External validity | Whether the result applies to your patients and setting | Think UK ED population, pathway, resources, case-mix |
| Precision | How narrow or wide the estimate is | Usually judged from the confidence interval |
| P-value | Probability of observing data this extreme, or more extreme, if the null hypothesis were true and assumptions held | Does not tell you effect size or clinical importance |
| Confidence interval | Range of values compatible with the data under the model used | Shows both significance and precision |
| Absolute risk reduction | Control event rate minus treatment event rate | Usually more clinically useful than relative risk reduction |
| Relative risk reduction | Absolute risk reduction divided by control event rate | Can exaggerate apparent benefit |
| Number needed to treat | 1 divided by ARR as a proportion | Needs a clear outcome and time frame |
| Sensitivity | Proportion of people with disease who test positive | High sensitivity helps rule out when negative, but only in the right pathway |
| Specificity | Proportion of people without disease who test negative | High specificity helps rule in when positive, but only with appropriate pre-test probability |
| Positive predictive value | Probability of disease given a positive test | Depends on prevalence |
| Negative predictive value | Probability of no disease given a negative test | Depends on prevalence |
| Likelihood ratio | How much a test result changes probability | More robust than slogans such as SnNout and SpPin |
| Type I error | False positive | Usually linked to alpha 0.05 |
| Type II error | False negative | Often due to low power |
| Power | Probability of detecting a true difference if one exists | Power = 1 minus beta |
| Allocation concealment | Preventing recruiters knowing the next treatment assignment before enrolment | Protects against selection bias |
| Blinding | Keeping participants, clinicians or assessors unaware of allocation after randomisation | Reduces performance and detection bias |
| Intention-to-treat analysis | Analysing patients in the groups to which they were randomised | Usually best for superiority trials |
Essential Pathophysiology
Critical appraisal has no biological pathophysiology, but there is a practical “decision pathophysiology” that matters in emergency medicine. Evidence moves from study design to bedside action through a series of filters:
- Was the study designed well enough to minimise bias?
- Was the estimate precise enough to be useful?
- Was the outcome clinically meaningful?
- Does the result apply to the patient in front of you?
- Does it fit current UK guidance and validated pathways?
Most exam questions probe failure at one of these steps. A trial may be randomised but underpowered. A diagnostic test may have excellent sensitivity but only in a selected population. A statistically significant result may be based on a surrogate outcome that does not improve patient-centred outcomes. A positive meta-analysis may still be unreliable if heterogeneity is high or the included studies are poor.
Clinical Presentation
In the exam, critical appraisal usually presents in one of a few recognisable formats:
- a trial summary embedded in a clinical stem
- a confidence interval, odds ratio, relative risk or mean difference to interpret
- a question asking for the main source of bias
- event rates requiring ARR, RRR or NNT calculation
- a diagnostic test result requiring interpretation in context
- a forest plot or meta-analysis summary
- a question asking whether evidence should change UK emergency practice
Typical ED contexts include:
- high-sensitivity troponin pathways in chest pain
- D-dimer in suspected VTE
- CT decision rules in head injury
- biomarkers in sepsis
- analgesia or sedation trials
- airway, trauma or resuscitation interventions
Red Flags and High-Risk Features
These are the features that should make you cautious about accepting a study result or choosing an enthusiastic answer in an SBA.
- confidence interval crosses the line of no effect
- confidence interval is very wide
- primary outcome negative but secondary outcomes positive
- subgroup benefit claimed despite neutral overall trial
- surrogate outcome used instead of patient-centred outcome
- composite outcome driven by a weak or less important component
- small sample size with low event rate
- large loss to follow-up, especially if unequal between groups
- post-randomisation exclusions
- unblinded trial with subjective outcomes
- diagnostic test studied in a very different population from your patient
- predictive values quoted without prevalence context
- negative superiority trial interpreted as equivalence
- single-centre study from a non-UK system being used to justify immediate practice change
Differential Diagnosis
When a question looks like “critical appraisal”, the real diagnosis is usually one of the following archetypes.
- P-value versus clinical significance
- Confidence interval interpretation
- ARR, RRR and NNT
- Type I error, Type II error and power
- Bias in randomised trials
- Diagnostic test accuracy and likelihood ratios
- Observational studies, confounding and causation
- Systematic reviews and meta-analysis
- Applicability and external validity
- Superiority, non-inferiority, equivalence, surrogate outcomes and subgroup traps
Initial ED Assessment
Use a rapid appraisal sequence when faced with an evidence-based SBA.
- Identify the study type.
- RCT, cohort, case-control, diagnostic study, systematic review, meta-analysis
- Identify the main outcome.
- mortality, admission, symptom score, biomarker change, imaging finding
- Check whether the result is statistically significant.
- ratio measure: does CI cross 1?
- difference measure: does CI cross 0?
- Check precision.
- narrow CI suggests more precision; wide CI suggests uncertainty
- Check clinical importance.
- how large is the effect, and does it matter to patients?
- Check validity.
- bias, confounding, follow-up, blinding, analysis method
- Check applicability.
- similar patients, similar pathway, similar healthcare setting, consistent with NICE/RCEM practice
If short of time, ask: true, precise, applicable.
Investigations
The “investigations” in critical appraisal are the core tools you need to interpret results quickly.
Quick formula box
| Measure | Formula |
|---|---|
| ARR | Control event rate minus treatment event rate |
| RRR | ARR divided by control event rate |
| NNT | 1 divided by ARR as a proportion |
| NNH | 1 divided by absolute risk increase as a proportion |
| Power | 1 minus beta |
| LR+ | Sensitivity divided by (1 minus specificity) |
| LR- | (1 minus sensitivity) divided by specificity |
Line of no effect
| Statistic | No-effect value |
|---|---|
| Relative risk | 1 |
| Odds ratio | 1 |
| Hazard ratio | 1 |
| Mean difference | 0 |
| Absolute risk difference | 0 |
Likelihood ratio interpretation
| Likelihood ratio | Usual interpretation |
|---|---|
| LR+ greater than 10 | Large increase in probability; useful for ruling in |
| LR+ 5 to 10 | Moderate increase |
| LR+ 2 to 5 | Small increase |
| LR- less than 0.1 | Large decrease in probability; useful for ruling out |
| LR- 0.1 to 0.2 | Moderate decrease |
| LR- 0.2 to 0.5 | Small decrease |
These are rough guides, not absolute rules. Test utility still depends on pre-test probability and whether the test is being used within a validated pathway.
Management in the Emergency Department
The practical management of critical appraisal in the ED is not about bedside treatment. It is about safe evidence use.
Immediate approach in the exam
- Read the final line first.
- Know whether you are being asked about significance, bias, applicability or management change.
- Identify the metric.
- OR, RR, HR, mean difference, sensitivity, specificity, NNT.
- Check the confidence interval.
- This often answers the question faster than the p-value.
- Look for the main flaw.
- Do not list every imperfection. Choose the one most likely to distort the result.
- Prefer cautious interpretations.
- Examiners usually reward safe, methodologically sound answers rather than overclaiming.
Immediate approach in clinical practice
- Do not change practice on the basis of a single statistically significant study alone.
- Check whether the finding is consistent with NICE, RCEM, BTS, SIGN or Resuscitation Council UK guidance.
- Use diagnostic tests only within validated pathways where appropriate.
- D-dimer for VTE assessment depends on pre-test probability and pathway use.
- High-sensitivity troponin depends on timing, assay and validated rule-out strategy.
- Prioritise patient-centred outcomes over surrogate markers.
- Balance benefit against harms, resource implications and feasibility in a UK ED.
The 10 most common critical appraisal SBA archetypes
1. P-value versus clinical significance
A p-value below 0.05 does not prove that an intervention matters clinically. Large studies can detect very small differences. The key question is whether the effect size is meaningful to patients.
High-yield points:
- p-value does not tell you effect size
- p-value does not tell you whether the null hypothesis is true
- clinical significance depends on outcome importance, magnitude of benefit, harms, cost and feasibility
- absolute effects are usually more informative than relative effects
Classic trap:
- ED length of stay reduced by 18 minutes, p<0.001
- Best answer: statistically significant, but uncertain clinical importance
Common distractors:
- “This proves the treatment should become standard care”
- “P<0.05 means the treatment is definitely superior”
2. Confidence interval interpretation
Confidence intervals tell you both significance and precision.
- For RR, OR and HR, if the 95% CI crosses 1, the result is not statistically significant at the 5% level in a standard two-sided analysis.
- For mean difference or risk difference, if the 95% CI crosses 0, the result is not statistically significant.
- Wide intervals mean imprecision.
Examples:
- OR 0.72, 95% CI 0.51 to 0.98: statistically significant
- Mean difference 1.8 days, 95% CI -0.4 to 4.0: not statistically significant
- RR 0.85, 95% CI 0.40 to 1.70: very imprecise
Exam trap:
- Choosing “no effect” when the correct interpretation is “study is inconclusive because the CI is wide and includes both benefit and no important benefit”.
3. ARR, RRR and NNT
These are common calculation questions.
Example:
- Control mortality 10%
- Treatment mortality 6%
- ARR = 4%
- RRR = 40%
- NNT = 1/0.04 = 25
High-yield points:
- ARR and NNT are usually more clinically useful than RRR
- NNT must relate to a specific outcome over a specific time period
- NNT changes with baseline risk
- Benefit should be weighed against NNH
Common trap:
- Using 4 instead of 0.04 when calculating NNT
4. Type I error, Type II error and power
Negative trials are often used to test whether candidates understand power.
- Type I error = false positive
- Type II error = false negative
- Power = 1 – beta
High-yield interpretation:
- a small underpowered study may miss a real difference
- a non-significant result does not prove no difference exists
- failure to show superiority is not the same as proving equivalence
Classic trap:
- “The treatments are equivalent because p>0.05”
- Best answer: no evidence of superiority shown; study may be underpowered
5. Bias in randomised controlled trials
Know where the distortion occurs.
| Clue in stem | Likely issue |
|---|---|
| Recruiters knew the next assignment | Poor allocation concealment; selection bias |
| Patients and clinicians knew treatment allocation | Performance bias |
| Outcome assessors knew allocation | Detection bias |
| Large unequal loss to follow-up | Attrition bias |
| Patients excluded after randomisation | Compromised internal validity |
| Per-protocol only in a superiority trial | Loss of randomisation benefits |
High-yield distinctions:
- allocation concealment happens before assignment
- blinding happens after assignment
- intention-to-treat usually best preserves randomisation in superiority trials
6. Diagnostic test accuracy
This is one of the commonest emergency medicine appraisal themes.
Core principles:
- sensitivity and specificity are test characteristics
- predictive values depend on prevalence
- likelihood ratios tell you how much a result changes probability
- test performance only matters if used in the right population and pathway
Useful heuristics:
- a highly sensitive test can help rule out when negative
- a highly specific test can help rule in when positive
But these are only heuristics. In real practice and in better SBA questions, pre-test probability matters more.
ED examples:
- D-dimer is useful for ruling out VTE only in an appropriate low or intermediate-risk population within a validated pathway
- high-sensitivity troponin cannot be interpreted safely without considering symptom timing, assay and pathway
- imaging studies may perform differently in tertiary referral populations than in general ED populations
Diagnostic study biases worth knowing:
- spectrum bias: test studied in a narrow or unrepresentative population
- verification bias: not all patients receive the reference standard
- incorporation bias: the index test forms part of the reference standard
- review bias: test interpretation influenced by knowledge of the reference standard or vice versa
Common trap:
- choosing an answer based on predictive value without noticing that prevalence differs markedly from the study population
7. Observational studies, confounding and causation
Not all important EM evidence comes from RCTs. Cohort and case-control studies are common, especially for prognosis, harms and rare outcomes.
High-yield points:
- association does not prove causation
- confounding occurs when a third factor is associated with both exposure and outcome
- confounding by indication is common in treatment studies outside randomisation
- case-control studies are efficient for rare outcomes but prone to recall and selection bias
- cohort studies are useful for prognosis and incidence but remain vulnerable to confounding
Exam trap:
- Patients receiving a treatment had worse outcomes, therefore the treatment is harmful.
- Best interpretation: this may reflect confounding by indication, because sicker patients were more likely to receive the treatment.
8. Systematic reviews and meta-analysis
These are high-yield because they look authoritative, but the exam often tests whether you can spot limitations.
Key concepts:
- a meta-analysis is only as good as the studies included
- forest plots show individual study estimates and the pooled estimate
- heterogeneity means study results differ more than expected by chance
- I² estimates inconsistency; higher values suggest more heterogeneity
- publication bias can exaggerate apparent benefit
Practical interpretation:
- if the pooled CI crosses the line of no effect, the meta-analysis is not statistically significant
- high heterogeneity reduces confidence in a single pooled estimate
- clinical heterogeneity matters as much as statistical heterogeneity
Common trap:
- “Meta-analysis proves treatment works” despite poor-quality included studies and marked heterogeneity
9. Applicability and external validity
A valid study may still not apply to your patient or your department.
Ask:
- Were the patients similar to a UK ED population?
- Was the intervention feasible in NHS emergency care?
- Was the comparator relevant to current UK practice?
- Were outcomes patient-centred and clinically relevant?
- Does the result fit with NICE or RCEM pathways?
Examples of limited applicability:
- single-centre tertiary trauma study applied to a district general hospital ED
- US admission-threshold study applied directly to NHS practice
- diagnostic pathway using a different assay from the one used locally
Exam trap:
- assuming a statistically positive study should immediately change UK practice even when it conflicts with established NICE guidance
10. Superiority, non-inferiority, equivalence, surrogate outcomes and subgroup traps
This cluster generates many SBA distractors.
Superiority versus non-inferiority versus equivalence
| Design | Question asked | Exam trap |
|---|---|---|
| Superiority | Is treatment A better than treatment B? | Negative result does not prove equivalence |
| Non-inferiority | Is treatment A not unacceptably worse than treatment B by a prespecified margin? | Need prespecified margin and appropriate analysis |
| Equivalence | Are treatments sufficiently similar within prespecified margins? | Cannot infer from an ordinary negative superiority trial |
Surrogate outcomes
Surrogate outcomes are laboratory or physiological markers used instead of patient-centred outcomes. They may be useful, but they do not guarantee improved survival, symptoms or function.
Examples:
- biomarker reduction without mortality benefit
- improved imaging appearance without better functional outcome
Composite outcomes
Composite outcomes combine several endpoints. They can be misleading if the apparent benefit is driven by a less important component.
Subgroup analyses
Subgroup findings are often underpowered and prone to false positives. They are more credible if prespecified, biologically plausible and consistent with the overall result.
Classic trap:
- overall trial neutral, but one subgroup positive
- best answer: hypothesis-generating only, not practice-changing
Disposition, Referral and Follow-Up
For exam purposes, disposition means deciding what to do with evidence after interpreting it.
When evidence should not immediately change practice
- single small study with wide confidence intervals
- primary outcome negative but secondary outcomes positive
- benefit shown only in post hoc subgroup analysis
- surrogate outcome only
- study population very different from UK ED patients
- result conflicts with established NICE or RCEM guidance without broader supporting evidence
When evidence is more likely to be practice-relevant
- methodologically sound multicentre study or high-quality systematic review
- patient-centred outcomes
- precise estimates
- intervention feasible in NHS emergency care
- consistent with or incorporated into UK guidance
Special Groups
Critical appraisal principles are the same, but applicability becomes even more important in special populations.
Paediatrics
- adult evidence may not generalise to children
- diagnostic thresholds and disease prevalence differ
- small paediatric studies are often underpowered
Pregnancy
- many trials exclude pregnant patients, limiting applicability
- diagnostic pathways may differ because of imaging and physiological changes
- be cautious about extrapolating non-pregnant adult data
Older adults
- frailty, multimorbidity and polypharmacy may reduce generalisability
- trial populations are often younger and fitter than real ED patients
- outcomes such as function and delirium may matter more than short-term physiological endpoints
Immunosuppressed patients
- baseline risk and disease spectrum differ
- diagnostic test performance may change in populations with atypical presentations
- external validity is often limited if such patients were excluded from the original studies
Common Pitfalls
- thinking p<0.05 means clinically important
- ignoring the confidence interval
- confusing allocation concealment with blinding
- calling a negative superiority trial “equivalent”
- using predictive values without considering prevalence
- treating odds ratio as if it were risk ratio when outcomes are common
- forgetting to convert percentages to proportions for NNT
- overvaluing secondary outcomes or subgroup analyses
- accepting composite outcomes without checking which component drove the result
- using diagnostic tests outside validated pathways
- assuming a statistically interesting paper should override NICE or RCEM guidance
FRCEM and MRCEM Exam Tips
For MRCEM, know the core definitions and calculations. For FRCEM, expect more interpretation and more traps around methodology, applicability and safe practice change.
Top one-line rules
- If a ratio CI crosses 1, it is not statistically significant in a standard two-sided 95% analysis.
- If a difference CI crosses 0, it is not statistically significant in a standard two-sided 95% analysis.
- Wide confidence intervals mean imprecision.
- P-value does not tell you effect size.
- ARR and NNT usually describe patient benefit better than RRR.
- Negative superiority trial does not prove equivalence.
- Allocation concealment prevents selection bias before randomisation.
- Blinding reduces bias after randomisation.
- Intention-to-treat usually best preserves randomisation in superiority trials.
- Predictive values depend on prevalence.
- Likelihood ratios are more useful than sensitivity and specificity alone for changing probability.
- Association does not prove causation.
- Confounding by indication is common in observational treatment studies.
- A meta-analysis cannot rescue poor primary studies.
- Subgroup findings are usually hypothesis-generating unless strongly prespecified and credible.
- Surrogate outcomes are weaker than patient-centred outcomes.
- Composite outcomes may exaggerate importance if driven by minor components.
- External validity matters: ask whether the study fits a UK ED population and pathway.
- Do not change practice on one positive paper if it conflicts with established UK guidance.
- In SBAs, the safest answer is often the most methodologically cautious one.
How This Appears in SBA Questions
Typical question stems
- “A randomised trial found a relative risk of 0.82 with 95% CI 0.61 to 1.11. What is the best interpretation?”
- “A diagnostic study of D-dimer reports sensitivity 98% and specificity 42%. Which statement is most accurate?”
- “In a superiority trial, the primary outcome was not significant, but two secondary outcomes were positive. What is the best conclusion?”
- “Investigators were aware of the next treatment allocation before enrolment. Which bias is most likely?”
- “Mortality was 12% in the control group and 9% in the intervention group. What is the NNT?”
- “A meta-analysis shows substantial heterogeneity with I² 78%. What is the main concern?”
- “A cohort study found lower mortality in patients receiving treatment X. What is the main limitation in inferring causation?”
Key discriminator clues
| Clue | Think |
|---|---|
| CI crosses 1 or 0 | Not statistically significant |
| Very wide CI | Imprecision, often low power |
| Small study, negative result | Possible Type II error |
| Recruiter knew next allocation | Allocation concealment problem |
| Outcome assessor unblinded | Detection bias |
| Positive subgroup only | Hypothesis-generating, not definitive |
| Predictive value quoted | Ask about prevalence and study population |
| Observational treatment study | Think confounding by indication |
| Meta-analysis with high I² | Heterogeneity limits confidence in pooled estimate |
| US single-centre pathway study | Question external validity to UK ED practice |
Common wrong answer traps
- “Statistically significant, therefore clinically important”
- “No significant difference, therefore treatments are equivalent”
- “High sensitivity means the test rules out disease in any patient”
- “Positive predictive value is an intrinsic property of the test”
- “Blinding and allocation concealment are the same thing”
- “Odds ratio can be read as risk ratio regardless of event rate”
- “Meta-analysis automatically provides the highest quality answer”
Mini worked examples
Example 1
A trial reports hazard ratio 0.90, 95% CI 0.68 to 1.19 for mortality. What is the best interpretation?
- Best answer: no statistically significant mortality benefit shown; the estimate is compatible with benefit or no important effect.
Example 2
A treatment reduces admission from 20% to 15%. What is the ARR and NNT?
- ARR = 5%
- NNT = 1/0.05 = 20
Example 3
A superiority trial is neutral overall, but a post hoc subgroup of patients under 50 appears to benefit. What is the safest interpretation?
- Best answer: subgroup finding is hypothesis-generating and should not alone change practice.
Example 4
A diagnostic study of a biomarker in sepsis includes only ICU patients with advanced disease. What is the main concern when applying it to ED patients?
- Best answer: spectrum bias and limited external validity.
Key Takeaways
- Critical appraisal remains heavily tested in MRCEM and FRCEM, even without a standalone paper.
- Most questions can be solved by asking: is it true, how sure are we, and does it apply here?
- Confidence intervals are often more useful than p-values because they show both significance and precision.
- ARR and NNT usually describe clinical benefit better than RRR.
- Negative superiority trials do not prove equivalence.
- Allocation concealment and blinding are different and commonly confused.
- Diagnostic tests must be interpreted in the context of pre-test probability and validated pathways.
- Predictive values depend on prevalence.
- Observational studies are vulnerable to confounding, especially confounding by indication.
- Meta-analyses can mislead if included studies are poor or heterogeneity is high.
- External validity matters: a good study may still not apply to a UK ED.
- In the exam, cautious, methodologically sound interpretations usually beat overconfident claims.
Further Reading
- NICE guideline NG158: Venous thromboembolic diseases: diagnosis, management and thrombophilia testing
- NICE guideline NG185: Acute coronary syndromes
- NICE guideline NG232: Head injury: assessment and early management
- RCEM curriculum and current examination regulations
- Resuscitation Council UK guidelines
- SIGN critical appraisal checklists
- CASP checklists for randomised trials, cohort studies, case-control studies and systematic reviews
- BTS guidelines relevant to acute respiratory presentations where diagnostic evidence is commonly tested
Related on EM Final Exams
- How to Pass the FRCEM Critical Appraisal Section
- P Values Confidence Intervals and Bias Explained Simply
- How Hard is the FRCEM Exam
- SBA Question Dissection How to Break Down Any Question in 30 Seconds
Authoritative Sources
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.
