Derivation & Validation of Prediction Models
Building a model on one dataset, then proving it still works on patients it has never seen.
Derivation = building the model on a development dataset. Then internal validation (bootstrapping / cross-validation, correcting for optimism and overfitting); then external validation on a separate population; then an impact / implementation study. Overfitting makes a model look excellent on its own data yet fail elsewhere.
Derivation looks perfect; on external data the same model over-predicts and falls away
The same model judged two ways. On its own development data (gold) it tracks the perfect-calibration line and looks excellent — because it was fitted to exactly that data. Tested on a separate population (plum) it sags below the line: it over-predicts risk and its real performance is lower. The gap is the optimism baked into derivation.
Treat any derivation-set figure (AUC, calibration) as a best-case. Internal validation (bootstrap / cross-validation) shrinks the estimate towards honesty by correcting for optimism, but only external validation — a different time, place or case-mix — tests whether the model travels. Implementation is proven last, by an impact study.
Most published models never get externally validated, and many that do lose performance. Small datasets with too many candidate predictors overfit badly: the model memorises noise. Knowing where a model sits in the derive→validate→impact chain tells you how much to trust its numbers in your department.
derive → internal → external → impact- Internal validation corrects optimism / overfitting
- Aim for ≥~10 events per variable (a rough floor, not a law)
HEART score — a model that did survive the journey. Derived in a single cohort (Six et al., Neth Heart J 2008), it was then externally validated prospectively in 2,440 ED chest-pain patients across 10 hospitals (Backus et al., Int J Cardiol 2013), holding a c-statistic of 0.83 for 6-week MACE before going on to impact trials. The point of contrast: a model that looks strong only at derivation has not yet earned that trust. Many scores post equally bright derivation numbers and then lose performance — sometimes substantially — the first time they meet a new population.
- Quoting optimism / overfitting-inflated derivation performance as if it were real-world.
- Accepting a model with no external validation in a separate population.
- Too few events per variable — over-fitted on a small, predictor-heavy dataset.
Quick check
A model has AUC 0.9 in its derivation sample — ready for practice?
Answer: No — derivation performance is optimistic (the model was fitted to that data). It needs external validation on a separate population, and ideally an impact study, before you trust it.
Ready to build your plan? EMF Premium gives you all 40,000+ questions, 20 mocks and 1,215 OSCE stations from £29/month — or a one-off 3- or 6-month pass.