Skip to content
Method

Identification first. Then estimation. Then stress-testing.

An external control is only as good as the argument that it answers your trial's question. At Variacle we write that argument down, estimate under it, and measure how far it would have to fail before the conclusion changes.

Three questions, in this order
  1. 01
    Estimand

    What exactly is the question, in which population?

  2. 02
    Identification

    Under which assumptions does the external data answer it?

  3. 03
    Sensitivity

    How wrong can those assumptions be before the answer changes?

01

The estimand

Five attributes, every time: population, treatment condition, variable, intercurrent events and population-level summary. For a single-arm study with an external control, the usual target is the effect in the treated — the population defined by the trial's own eligibility criteria.

We name intercurrent-event strategies — hypothetical, principal stratum, treatment policy — rather than writing “per-protocol”, the label ICH E9(R1) discourages and reviewers notice.

02

Identification

We state the assumptions under which the external source answers the trial's question: that the populations are comparable once measured covariates are accounted for (exchangeability, or transportability between settings), that every trial patient has external counterparts (overlap), and that the endpoint means the same thing in both sources.

When an important confounder is not measured but a proxy for it is, proximal methods can recover the effect under different, explicit conditions. Sources do not need identical covariates: each is modelled on its own terms, and the assumption linking them stays visible instead of being hidden in a pooled model.

The sensitivity model, in one line

E[Y(0) | X, S = trial] = E[Y(0) | X, S = external] + δ

δ = 0 is the identifying assumption. Every other value of δ is a specific way it could fail, and the estimate is reported along all of them.

03

Estimation

Doubly robust and targeted estimators (AIPW, TMLE), which stay consistent if either the outcome model or the weighting model is right. Separate models per source rather than one pooled model with the source as a covariate. Bootstrap resampling by site or study, because that is the unit that varies.

Every external source is reported with its effective sample size — what its patients are worth to the comparison after between-study variation — not its headcount.

04 · Sensitivity

Report how far each assumption would have to fail.

Each identifying assumption gets a sensitivity parameter δ. We report the effect as a curve ψ(δ), relaxing one assumption at a time, and the tipping point where the conclusion would change. Reviewers then compare that distance with clinical knowledge, instead of taking our word for it.

E-values for unmeasured confounding sit alongside, in the form regulators already recognise.

Effect estimate ψ(δ) as one assumption is relaxed
ψ(δ) with 95% band Tipping point Plausible range
Assumed bias δ in the identifying assumption

Illustrative. The conclusion survives until the assumption is wrong by δ ≈ 0.21, about two and a half times the largest value clinical experts consider plausible. Reviewers judge that distance, not our confidence.

05

Process

The difference between an accepted and a rejected external control is usually process, not method. Our checklist:

  • Write and date the analysis plan before any external outcome data is touched
  • Define time zero explicitly, and how it is matched in each external source
  • Justify pooling sources, or keep them separate and report each one
  • Check, before any comparison, whether every eligibility criterion can be reproduced in the external data
Criticised · FDA review, 2019

A real-world external comparator for a single-arm trial, with no analysis plan agreed in advance and an eligibility criterion that could not be applied at time zero.

Pre-specified · cell therapy programme

Comparator protocol and statistical analysis plan fixed before external outcomes were seen; two data sources kept separate instead of pooled.

06

Delivery

The code travels; the data stays. Where patient-level data cannot leave a sponsor, hospital or data partner, the analysis is delivered as a versioned, reproducible package that runs inside their environment and returns aggregates only — the model Europe's Health Data Space will require for secondary use of health data from 2029.

Every reported number traces to versioned code and checksummed inputs, derived quantities are double-programmed independently, and each deliverable passes a documented review before release.

07

Selected literature

  1. ICH E9(R1). Addendum on estimands and sensitivity analysis in clinical trials, 2019.
  2. FDA. Considerations for the design and conduct of externally controlled trials for drug and biological products. Draft guidance, February 2023.

Want to see confounding hide a treatment effect across study environments? Try the interactive simulation.

The three questions at the top of this page follow the causal roadmap of Petersen and van der Laan (2014).

Want the method applied to your study?

Put your design through the checklist.

A 30-minute call: whether an external control is viable for your question, what data could serve, and what regulators have accepted in similar settings.

Book a feasibility call