Identification first. Then estimation. Then stress-testing.
An external control is only as good as the argument that it answers your trial's question. At Variacle we write that argument down, estimate under it, and measure how far it would have to fail before the conclusion changes.
- 01Estimand
What exactly is the question, in which population?
- 02Identification
Under which assumptions does the external data answer it?
- 03Sensitivity
How wrong can those assumptions be before the answer changes?
The estimand
Five attributes, every time: population, treatment condition, variable, intercurrent events and population-level summary. For a single-arm study with an external control, the usual target is the effect in the treated — the population defined by the trial's own eligibility criteria.
We name intercurrent-event strategies — hypothetical, principal stratum, treatment policy — rather than writing “per-protocol”, the label ICH E9(R1) discourages and reviewers notice.
Identification
We state the assumptions under which the external source answers the trial's question: that the populations are comparable once measured covariates are accounted for (exchangeability, or transportability between settings), that every trial patient has external counterparts (overlap), and that the endpoint means the same thing in both sources.
When an important confounder is not measured but a proxy for it is, proximal methods can recover the effect under different, explicit conditions. Sources do not need identical covariates: each is modelled on its own terms, and the assumption linking them stays visible instead of being hidden in a pooled model.
E[Y(0) | X, S = trial] = E[Y(0) | X, S = external] + δ
δ = 0 is the identifying assumption. Every other value of δ is a specific way it could fail, and the estimate is reported along all of them.
Estimation
Doubly robust and targeted estimators (AIPW, TMLE), which stay consistent if either the outcome model or the weighting model is right. Separate models per source rather than one pooled model with the source as a covariate. Bootstrap resampling by site or study, because that is the unit that varies.
Every external source is reported with its effective sample size — what its patients are worth to the comparison after between-study variation — not its headcount.
Report how far each assumption would have to fail.
Each identifying assumption gets a sensitivity parameter δ. We report the effect as a curve ψ(δ), relaxing one assumption at a time, and the tipping point where the conclusion would change. Reviewers then compare that distance with clinical knowledge, instead of taking our word for it.
E-values for unmeasured confounding sit alongside, in the form regulators already recognise.
Illustrative. The conclusion survives until the assumption is wrong by δ ≈ 0.21, about two and a half times the largest value clinical experts consider plausible. Reviewers judge that distance, not our confidence.
Process
The difference between an accepted and a rejected external control is usually process, not method. Our checklist:
- Write and date the analysis plan before any external outcome data is touched
- Define time zero explicitly, and how it is matched in each external source
- Justify pooling sources, or keep them separate and report each one
- Check, before any comparison, whether every eligibility criterion can be reproduced in the external data
A real-world external comparator for a single-arm trial, with no analysis plan agreed in advance and an eligibility criterion that could not be applied at time zero.
Comparator protocol and statistical analysis plan fixed before external outcomes were seen; two data sources kept separate instead of pooled.
Delivery
The code travels; the data stays. Where patient-level data cannot leave a sponsor, hospital or data partner, the analysis is delivered as a versioned, reproducible package that runs inside their environment and returns aggregates only — the model Europe's Health Data Space will require for secondary use of health data from 2029.
Every reported number traces to versioned code and checksummed inputs, derived quantities are double-programmed independently, and each deliverable passes a documented review before release.
Selected literature
- ICH E9(R1). Addendum on estimands and sensitivity analysis in clinical trials, 2019.
- FDA. Considerations for the design and conduct of externally controlled trials for drug and biological products. Draft guidance, February 2023.
Want to see confounding hide a treatment effect across study environments? Try the interactive simulation.
The three questions at the top of this page follow the causal roadmap of Petersen and van der Laan (2014).
Want the method applied to your study?
Put your design through the checklist.
A 30-minute call: whether an external control is viable for your question, what data could serve, and what regulators have accepted in similar settings.