Tutorials › Biostatistics › Umbrella Trials Explained

Precision Oncology & Master Protocols

Umbrella Trials Explained

A practical guide to umbrella trials in precision oncology, including biomarker-driven treatment assignment, master protocols, multiple treatment substudies, statistical design, multiplicity, adaptive decision-making, operating characteristics, and a complete worked example.

Advanced 18 min read

What You'll Learn

  • What an umbrella trial is and how it differs from basket and platform trials
  • How molecular biomarkers can determine treatment assignment
  • How a master protocol can contain multiple disease-specific treatment substudies
  • How statistical inference changes when several hypotheses are evaluated within one protocol
  • How to calculate sample size and operating characteristics for individual umbrella substudies
  • How adaptive enrollment, arm dropping, and biomarker evolution can be incorporated

Introduction

Precision oncology attempts to match patients with treatments according to the biological characteristics of their tumors. Instead of assuming that every patient with a particular cancer will respond to the same treatment, investigators can use genomic, molecular, or other biomarker information to identify subgroups that may benefit from specific therapies.

An umbrella trial is a master-protocol approach designed to evaluate multiple treatments or treatment strategies within a single disease or disease subtype, with patients assigned to different treatment substudies according to biomarkers or other predefined characteristics.

For example, imagine a study enrolling patients with advanced non-small-cell lung cancer. A molecular screening panel identifies several alterations:

  • EGFR alteration
  • ALK rearrangement
  • ROS1 rearrangement
  • MET alteration
  • KRAS G12C mutation

Rather than conducting five completely independent screening studies, a single master protocol could screen patients and route eligible patients into different treatment substudies.

Key idea: An umbrella trial generally starts with one disease or disease population and then divides that population into multiple biomarker-defined or otherwise biologically defined treatment paths. The umbrella structure is therefore primarily about multiple treatment questions within one disease framework.

What Makes a Trial an Umbrella Trial?

The defining structure can be summarized as:

1
One disease framework: patients enter the master protocol because they have the disease or disease subtype of interest.
2
Molecular screening: tumors or other biological samples are characterized using predefined biomarkers.
3
Biomarker-defined subgroups: patients are classified according to molecular or clinical characteristics.
4
Multiple treatment paths: different subgroups can receive different investigational treatments.
5
Master-protocol infrastructure: screening, eligibility, data collection, governance, and operational processes can be shared.

A Simple Umbrella Trial Diagram

Conceptually, an umbrella trial can be represented as:

Screened Population Biomarker Treatment Substudy
Patients with the same disease Biomarker A Experimental Treatment A
Biomarker B Experimental Treatment B
Biomarker C Experimental Treatment C
Biomarker D Experimental Treatment D
No actionable biomarker Standard treatment / control / other predefined pathway

The important point is that the trial does not necessarily test one treatment against every patient. Instead, the treatment question can depend on the patient's biological subgroup.

Umbrella vs. Basket vs. Platform Trials

Three terms are frequently used together in precision oncology: umbrella, basket, and platform trials. They overlap, but they describe different structural concepts.

Feature Umbrella Basket Platform
Core population structure Usually one disease Usually one biomarker across multiple diseases Can span multiple arms or substudies
Multiple treatments Yes May be yes or no Yes
Biomarker-driven assignment Common Common Common but not required
Shared master protocol Often Often Central feature
Arms can change over time Sometimes Sometimes Often
Example conceptual question Which treatment works for each subgroup of one cancer? Does this treatment work across cancers sharing a biomarker? Which treatments should remain in an evolving trial platform?
Important: The terms are not mutually exclusive. A trial can have an umbrella-like disease structure while also operating as a platform with arms that are added, dropped, or modified over time.

The Statistical Unit Is Often the Substudy

One of the most important statistical concepts is that an umbrella trial does not necessarily represent one homogeneous hypothesis test.

Suppose there are four biomarker-defined treatment groups. The scientific questions might be:

$$ H_{01}: \theta_1\le\theta_{01} \qquad\text{vs.}\qquad H_{A1}: \theta_1>\theta_{01} $$
$$ H_{02}: \theta_2\le\theta_{02} \qquad\text{vs.}\qquad H_{A2}: \theta_2>\theta_{02} $$
$$ H_{03}: \theta_3\le\theta_{03} \qquad\text{vs.}\qquad H_{A3}: \theta_3>\theta_{03} $$
$$ H_{04}: \theta_4\le\theta_{04} \qquad\text{vs.}\qquad H_{A4}: \theta_4>\theta_{04} $$

Each hypothesis can correspond to a different treatment-biomarker pair.

The resulting statistical architecture is therefore more complicated than a conventional single-arm Phase II trial.

Why Umbrella Trials Are Attractive

Traditional clinical development can require separate protocols for each biomarker-defined subgroup.

That can create:

  • Repeated screening infrastructure
  • Separate contracts and site activation
  • Separate molecular testing workflows
  • Separate regulatory submissions
  • Competition for the same patient population
  • Longer startup times

A master protocol can centralize many of these activities.

Patients can potentially be screened once and then routed to the appropriate substudy.

Operational principle: The major efficiency is not necessarily that every patient contributes to every hypothesis. Instead, the master protocol can make it substantially easier to identify the appropriate treatment pathway without repeatedly screening patients in separate trials.

Biomarker Screening

Suppose a study evaluates four actionable alterations:

Biomarker Prevalence Assigned Treatment
A 20% Treatment A
B 15% Treatment B
C 10% Treatment C
D 5% Treatment D
No actionable alteration 50% Predefined non-targeted pathway

If 1,000 patients are screened, the expected number entering each biomarker substudy is approximately:

$$ E(n_A)=1000(0.20)=200 $$
$$ E(n_B)=1000(0.15)=150 $$
$$ E(n_C)=1000(0.10)=100 $$
$$ E(n_D)=1000(0.05)=50 $$

This illustrates a major design problem. Biomarker prevalence directly affects recruitment into each substudy.

The Rare-Biomarker Problem

Suppose Treatment D targets a biomarker present in only 5% of screened patients. If the treatment substudy requires 40 evaluable patients, the investigators may need approximately:

$$ \frac{40}{0.05}=800 $$

screened patients just to obtain the required number of biomarker-positive patients, assuming perfect screening and enrollment.

Real studies will require more because of:

  • Screening failures
  • Ineligible patients
  • Assay failures
  • Patient refusal
  • Competing studies
  • Loss before treatment
  • Unevaluable outcomes
Planning implication: Sample size within an umbrella substudy is not sufficient for recruitment planning. Investigators must also model the screening prevalence and attrition pathway required to generate that number of evaluable patients.

Biomarker Assignment Is Part of the Design

The treatment assignment process must be specified before enrollment. For example:

1
Obtain tumor tissue or another prespecified biological sample.
2
Run the prespecified molecular assay.
3
Determine biomarker status according to validated criteria.
4
Apply the eligibility and assignment algorithm.
5
Enroll the patient into the appropriate treatment substudy.

What If a Patient Has More Than One Biomarker?

Real tumors may contain multiple potentially actionable alterations. For example, a patient might test positive for both Biomarker A and Biomarker B.

The protocol therefore needs an explicit hierarchy or assignment strategy. Possible approaches include:

  • A prespecified biomarker priority hierarchy
  • Eligibility for one specific treatment based on mutation type
  • Randomization among eligible treatment paths
  • Assignment based on treatment availability
  • Assignment using a molecular tumor board

The statistical consequences depend on the rule.

Do not leave overlapping biomarkers to ad hoc clinical judgment. If biomarker overlap can affect treatment assignment, the assignment algorithm should be defined prospectively.

A Simple Single-Arm Umbrella Substudy

Consider a hypothetical umbrella trial containing three treatment substudies. Each substudy evaluates objective response rate. Suppose the design assumptions are:

Substudy Biomarker \(p_0\) \(p_1\)
A Biomarker A 10% 30%
B Biomarker B 15% 35%
C Biomarker C 20% 40%

Each substudy can have its own statistical design because the clinical expectations differ.

For example:

$$ H_{0A}:p_A\le0.10 \qquad\text{vs.}\qquad H_{AA}:p_A\ge0.30 $$

while:

$$ H_{0C}:p_C\le0.20 \qquad\text{vs.}\qquad H_{AC}:p_C\ge0.40 $$

The fact that the substudies are contained within one umbrella protocol does not require every substudy to have identical statistical assumptions.

Shared Protocol, Separate Statistical Questions

An umbrella trial often has two levels of structure.

Level Examples
Master protocol Eligibility framework, screening, biomarker testing, safety reporting, governance, common procedures
Substudy Treatment, biomarker definition, endpoint, sample size, analysis, stopping rules

This distinction is critical. A common protocol does not automatically mean that all statistical analyses can be pooled.

Randomized Umbrella Trials

Not every umbrella trial is single-arm. A substudy can use randomization. For example, among patients with Biomarker A:

Arm Treatment Allocation
A1 Experimental Treatment A 1:1
A2 Control 1:1

A different biomarker subgroup might have its own randomized comparison.

The statistical model could then compare treatment effects within each biomarker-defined population.

$$ H_{0k}:\theta_k=0 \qquad\text{vs.}\qquad H_{Ak}:\theta_k\ne0 $$

where \(\theta_k\) represents the treatment effect in substudy \(k\).

Why Randomization Can Be Important

Single-arm umbrella substudies can be efficient for screening highly targeted therapies, but historical or external controls may introduce substantial uncertainty. Differences in:

  • Patient selection
  • Supportive care
  • Response assessment
  • Imaging schedules
  • Prior therapies
  • Calendar time

can make historical response rates difficult to interpret.

A randomized control arm can provide contemporaneous estimation of the counterfactual outcome.

Common Controls Across Substudies

A particularly important design issue occurs when multiple treatment arms share a common control group. Suppose:

Substudy Experimental Arm Control
A A1 Common Control
B B1 Common Control
C C1 Common Control

This can reduce the number of control patients required compared with running three completely separate trials.

However, shared controls also create statistical dependence among treatment comparisons.

Important: If the same control patients contribute to multiple comparisons, those comparisons are not statistically independent. This dependence must be considered when calculating multiplicity-adjusted inference and operating characteristics.

Multiplicity in Umbrella Trials

Suppose an umbrella trial contains \(K\) treatment hypotheses:

$$ H_{01},H_{02},\ldots,H_{0K} $$

If each hypothesis is tested at level \(\alpha\), the probability of making at least one false-positive declaration can exceed \(\alpha\).

If the tests were independent, a simple illustration would be:

$$ P(\text{at least one false positive}) = 1-(1-\alpha)^K $$

For \(K=5\) and \(\alpha=0.05\):

$$ 1-(0.95)^5 \approx0.226 $$

or approximately 22.6%.

That does not mean every umbrella trial has a 22.6% family-wise type I error. It is simply an illustration of why multiple hypothesis testing must be considered.

Do All Umbrella Trials Need Family-Wise Error Control?

Not necessarily. The appropriate multiplicity strategy depends on the scientific objectives and regulatory context. Possible approaches include:

  • Controlling family-wise type I error
  • False discovery rate control
  • Hierarchical testing
  • Gatekeeping procedures
  • Separate confirmatory hypotheses
  • Exploratory treatment-specific analyses
  • Bayesian decision criteria

The key question is: What inferential claim is the study intended to support?

If the study is a screening platform intended to identify promising treatment signals, the multiplicity strategy may differ from that of a confirmatory trial intended to support several formal efficacy claims.

A Bonferroni Illustration

Suppose five hypotheses must jointly control family-wise type I error at 5%. A simple Bonferroni allocation would use:

$$ \alpha_k=\frac{0.05}{5}=0.01 $$

for each hypothesis.

Bonferroni is simple and robust, but it can be conservative. More efficient multiplicity procedures may be appropriate when the hypotheses have a prespecified structure.

Multiplicity and Shared Controls

The previous Bonferroni calculation illustrates the basic concept but does not capture every dependency created by a shared control group. Suppose two treatment effects are estimated as:

$$ \hat\theta_A=\bar Y_A-\bar Y_C $$
and:

$$ \hat\theta_B=\bar Y_B-\bar Y_C $$

Both estimates contain the same control estimate \(\bar Y_C\). Therefore:

$$ \operatorname{Cov}(\hat\theta_A,\hat\theta_B) = \operatorname{Var}(\bar Y_C) $$

under the usual independence assumptions for the treatment groups and common control.

This covariance affects joint statistical calculations.

Adaptive Treatment Arm Decisions

Umbrella trials are often particularly attractive when the treatment landscape is evolving. A master protocol can potentially allow:

  • New treatment arms to be added
  • Underperforming arms to be dropped
  • Biomarker definitions to evolve
  • Recruitment to be redirected
  • Substudies to close independently

These changes can dramatically improve operational efficiency. But they also make statistical planning more complicated.

Adaptive does not mean unplanned. The rules governing adaptation must be prospectively specified, or the statistical consequences of the adaptation must be addressed using an appropriate inferential framework.

Arm Dropping

Suppose three treatment arms begin enrollment:

A
Treatment A → continue if interim evidence remains promising.
B
Treatment B → stop if the predictive probability of success becomes too low.
C
Treatment C → continue because accumulating data remain encouraging.

Treatment B may therefore stop accruing while A and C continue. This is one of the major operational advantages of a flexible master protocol.

Interim Analysis in an Umbrella Trial

An interim analysis may be conducted after a predefined number of events or patients. For each treatment-biomarker subgroup, investigators may examine:

  • Response rate
  • Treatment effect
  • Progression-free survival
  • Overall survival
  • Safety
  • Predictive probability of success

The decision rules may be different for different substudies.

For example:

Interim Result Decision
Very low activity Drop treatment arm
Intermediate activity Continue enrollment
Strong activity Continue, graduate, or expand according to prespecified rules

Frequentist vs. Bayesian Decision-Making

Umbrella and platform trials can use either frequentist or Bayesian methods. A frequentist design might define:

$$ P(\text{reject }H_0\mid H_0)\le\alpha $$

while a Bayesian design might use a posterior probability:

$$ P(\theta>\theta_0\mid\text{data}) \ge c $$

for a prespecified decision threshold \(c\).

Bayesian approaches can be particularly natural when investigators want to make explicit probability-based decisions about whether a treatment is likely to meet a clinically meaningful target.

A Bayesian Example

Suppose the clinically meaningful response rate is 30%. An interim rule might be:

$$ P(p>0.30\mid\text{data})>0.90 \quad\Rightarrow\quad \text{continue} $$

while:

$$ P(p>0.30\mid\text{data})<0.10 \quad\Rightarrow\quad \text{drop} $$

The region between those thresholds represents an uncertain state.

This is only an illustrative framework; actual decision thresholds require careful calibration to the desired operating characteristics.

Complete Worked Example

Consider a hypothetical umbrella trial in advanced disease X. The master protocol screens all patients for three actionable biomarkers.

Substudy Biomarker Treatment Expected Prevalence
A A-positive Drug A 25%
B B-positive Drug B 15%
C C-positive Drug C 10%

Assume each substudy is a single-arm Phase II study using objective response rate. For simplicity, use:

Parameter Value
Null response rate \(p_0\) 10%
Target response rate \(p_1\) 30%
One-sided \(\alpha\) 5%
Power 80%

Assume a hypothetical single-arm design for each substudy requiring approximately 30 evaluable patients. The exact sample size should be obtained using the selected statistical design and operating-characteristic calculation rather than by simply using the illustrative number below.

Step 1: Calculate the Number of Patients Needed in Each Substudy

Suppose the design requires:

$$ n_A=n_B=n_C=30 $$

evaluable biomarker-positive patients. The biomarker prevalence determines the approximate screening requirement. For Biomarker A:

$$ N_{\text{screen},A} \approx \frac{30}{0.25} = 120 $$

For Biomarker B:

$$ N_{\text{screen},B} \approx \frac{30}{0.15} = 200 $$

For Biomarker C:

$$ N_{\text{screen},C} \approx \frac{30}{0.10} = 300 $$

These are simplified calculations assuming no losses.

Step 2: Add Screening Inefficiency

Suppose only 80% of biomarker-positive patients who are identified actually enter treatment because of eligibility and operational losses. Then the required number of biomarker-positive patients becomes:

$$ \frac{30}{0.80}=37.5 $$

or approximately 38 biomarker-positive patients. For Biomarker C, with a prevalence of 10%:

$$ N_{\text{screen},C} \approx \frac{38}{0.10} = 380 $$

Thus the rarest biomarker may determine the overall screening burden.

Step 3: Build the Treatment Assignment Algorithm

The master protocol might specify:

1
Screen all eligible patients for Biomarkers A, B, and C.
2
If exactly one actionable biomarker is present, assign the patient to the corresponding substudy.
3
If multiple actionable biomarkers are present, apply the prespecified biomarker hierarchy.
4
If no actionable biomarker is present, assign the patient to the predefined non-targeted pathway or screen failure category.

Step 4: Define the Statistical Hypotheses

For Treatment A:

$$ H_{0A}:p_A\le0.10 \qquad H_{AA}:p_A\ge0.30 $$

For Treatment B:

$$ H_{0B}:p_B\le0.10 \qquad H_{AB}:p_B\ge0.30 $$

For Treatment C:

$$ H_{0C}:p_C\le0.10 \qquad H_{AC}:p_C\ge0.30 $$

These are three distinct treatment questions even though they are contained within the same master protocol.

Step 5: Consider the Multiplicity Strategy

Suppose the three hypotheses are intended to support separate confirmatory claims. The statistical team might choose to control family-wise type I error at 5%. A simple Bonferroni illustration would allocate:

$$ \alpha_k=\frac{0.05}{3}=0.0167 $$

per hypothesis.

Alternatively, the study might define the three treatment hypotheses as separate screening objectives and report each according to a prespecified exploratory framework.

The statistical objective comes first. There is no universal multiplicity correction that defines an umbrella trial. The appropriate method depends on whether the hypotheses are confirmatory, co-primary, hierarchical, exploratory, or otherwise structured.

Step 6: Evaluate Each Substudy Independently

For each substudy, calculate:

  • Type I error
  • Power
  • Expected sample size
  • Probability of early stopping
  • Probability of success
  • Recruitment feasibility

For example, if Treatment C has a low biomarker prevalence, its statistical sample size may be reasonable while its operational recruitment burden is substantial.

Step 7: Evaluate the Entire Master Protocol

The umbrella trial should also be evaluated globally. Relevant quantities include:

  • Total number of patients screened
  • Probability of entering each substudy
  • Expected enrollment in each arm
  • Probability each arm reaches its target sample size
  • Probability of at least one successful treatment
  • Probability of false-positive declarations
  • Overall study duration
  • Site burden

Probability of Entering a Substudy

Suppose the prevalence of Biomarker A is \(q_A\). If \(N\) eligible patients are screened, the number who are biomarker-positive can be represented as:

$$ X_A\sim\operatorname{Binomial}(N,q_A) $$

The probability of obtaining at least \(n_A\) biomarker-positive patients is:

$$ P(X_A\ge n_A) = 1- P(X_A\le n_A-1) $$

This probability can be extremely important for rare biomarkers.

Example: Probability of Recruiting a Rare Biomarker

Suppose Biomarker C has prevalence:

$$ q_C=0.10 $$

and the trial needs 30 biomarker-positive patients. If 300 patients are screened, then:

$$ X_C\sim\operatorname{Binomial}(300,0.10) $$

The expected number of biomarker-positive patients is:

$$ E(X_C)=300(0.10)=30 $$

But the probability of obtaining at least 30 is substantially less than 50% because the expected value is exactly 30 and the distribution is variable.

Recruitment lesson: Planning to screen exactly \(n/q\) patients for a biomarker does not guarantee that \(n\) biomarker-positive patients will be obtained. Recruitment feasibility should therefore incorporate the stochastic nature of biomarker prevalence.

Screening Sample Size vs. Treatment Sample Size

Umbrella trials therefore have at least two related sample-size problems.

Problem Question
Substudy sample size How many biomarker-positive patients are needed to evaluate the treatment?
Screening requirement How many patients must be screened to generate those biomarker-positive patients?

These calculations should not be confused. A study can have excellent statistical power conditional on enrollment but still be operationally infeasible because the biomarker is too rare.

Umbrella Trials and Enrichment

Umbrella trials often use biomarker enrichment. If Treatment A is hypothesized to work specifically in Biomarker A-positive patients, enrolling only those patients into Treatment A increases the biological relevance of the treatment comparison.

Instead of estimating:

$$ E[Y\mid\text{all patients}] $$

the study focuses on:

$$ E[Y\mid\text{Biomarker A positive}] $$

This can increase the expected treatment effect if the biomarker truly identifies a responsive population.

But Enrichment Has a Cost

Restricting treatment assignment to biomarker-positive patients reduces the available population. If the biomarker prevalence is low, recruitment becomes slower. Thus:

$$ \text{Biological enrichment} \quad\leftrightarrow\quad \text{Recruitment feasibility} $$

is a fundamental design trade-off.

Biomarker Misclassification

Umbrella trials depend heavily on the biomarker assay. Suppose the true biomarker status is \(B\), but the observed assay result is \(B^*\). The treatment effect may then be diluted because some patients are assigned to a treatment despite not having the intended biological feature.

For example, if:

$$ P(B^*=1\mid B=1)=Se $$

and:

$$ P(B^*=0\mid B=0)=Sp $$

then sensitivity and specificity influence the actual treatment population.

Statistical consequence: Biomarker measurement error is not merely a laboratory problem. It can directly affect treatment effect estimation, power, subgroup definitions, and interpretation of the trial.

Dynamic Biomarkers

Another challenge is that the relevance of a biomarker can change during development. New molecular alterations may become clinically actionable while an existing treatment becomes obsolete. A platform-like umbrella protocol can potentially accommodate these changes more efficiently than a collection of independent protocols.

For example:

1
Initial protocol contains Treatments A, B, and C.
2
Treatment B shows insufficient activity and stops enrollment.
3
A new biomarker becomes biologically compelling.
4
Treatment D is added under the master protocol according to the predefined amendment and statistical framework.

Adding a New Arm Changes the Statistical Problem

Suppose an umbrella platform initially contains two hypotheses:

$$ H_1,\ H_2 $$

and later adds:

$$ H_3 $$

The final multiplicity structure depends on the inferential objective and the rules governing the addition.

If the new hypothesis is intended to contribute to a common confirmatory family, the overall type I error strategy must account for it.

If it is a new exploratory question with a separately defined objective, the statistical treatment may differ.

Protocol amendments can have statistical consequences. Adding a treatment arm is not simply an operational decision if the new arm changes the family of formal hypotheses or the shared-control structure.

Adaptive Enrichment

Some precision-oncology designs go beyond fixed biomarker assignment and allow the evidence about treatment benefit to influence the population being studied. Suppose a treatment appears effective in two related biomarker subgroups. The protocol might prospectively allow enrollment to continue in the more responsive population.

Conceptually:

$$ \text{Initial population} \rightarrow \text{Interim evidence} \rightarrow \text{Refined population} $$

This is known as an adaptive enrichment strategy.

The statistical validity of such a strategy depends on how the enrichment rule was designed and how inference is performed after adaptation.

Survival Endpoints in Umbrella Trials

Umbrella trials are not limited to response rates. Suppose the primary endpoint is progression-free survival. For each substudy \(k\), the treatment effect may be represented by a hazard ratio:

$$ HR_k = \frac{\lambda_{k,T}}{\lambda_{k,C}} $$

A typical hypothesis might be:

$$ H_{0k}:HR_k\ge1 \qquad\text{vs.}\qquad H_{Ak}:HR_k<1 $$

Sample size then depends on the number of events rather than simply the number of enrolled patients.

Events Can Be Uneven Across Substudies

Suppose four umbrella substudies enroll the same number of patients but have different progression rates. The number of observed events can differ substantially. Therefore, equal enrollment does not imply equal statistical information.

Substudy Patients Event Rate Approximate Events
A 100 80% 80
B 100 60% 60
C 100 40% 40
D 100 25% 25

This can make operational planning particularly challenging in a multi-substudy master protocol.

Shared Data Infrastructure

One of the practical advantages of a master protocol is that many data structures can be standardized. For example:

  • Common screening forms
  • Common demographic variables
  • Common adverse-event collection
  • Common biomarker metadata
  • Common imaging standards
  • Common data transfer specifications
  • Common data review procedures

The treatment-specific variables can then be layered onto the common data structure.

Statistical Programming Considerations

An umbrella trial can generate a large number of related analysis datasets. The programming architecture should therefore distinguish between:

  • Master-protocol data
  • Biomarker data
  • Treatment assignment
  • Substudy-specific analysis populations
  • Shared control data
  • Substudy-specific endpoints

A useful programming structure might include a treatment-assignment variable:

substudy <- c(
  "A",
  "A",
  "B",
  "C",
  "A",
  "B"
)

and a biomarker variable:

biomarker <- c(
  "A+",
  "A+",
  "B+",
  "C+",
  "A+",
  "B+"
)

These variables can then be used to construct substudy-specific analysis populations.

Example R Data Structure

data <- data.frame(
  id = 1:8,
  biomarker = c(
    "A+","A+","B+","C+",
    "A+","B+","C+","A+"
  ),
  substudy = c(
    "A","A","B","C",
    "A","B","C","A"
  ),
  response = c(
    1,0,1,0,
    1,1,0,1
  )
)

A treatment-specific analysis can then be performed using a subset of the master dataset.

data_A <- subset(
  data,
  substudy == "A"
)

table(data_A$response)

Estimating Response Rates by Substudy

For a single-arm response endpoint, the observed response rate in substudy \(k\) is:

$$ \hat p_k = \frac{X_k}{n_k} $$

Suppose Treatment A has 18 responses among 50 evaluable patients:

$$ \hat p_A = \frac{18}{50} = 0.36 $$

or 36%.

The corresponding confidence interval should be constructed using an appropriate method for a binomial proportion rather than relying automatically on the Wald approximation, especially when sample sizes are small.

Comparing Treatment Effects Across Biomarkers

A tempting question is: Which treatment worked best?

But comparing observed response rates across different biomarker-defined populations can be misleading. Suppose:

Substudy Observed Response
A 40%
B 30%
C 20%

The 40% response rate in A does not automatically mean Treatment A is superior to Treatment B. The underlying patient populations, prognostic characteristics, biomarker biology, control response, and treatment effects may all differ.

Interpretation principle: Umbrella substudies often answer different biological questions. Observed response rates across unrelated biomarker populations should not be treated as though they came from one randomized comparison.

Common Control vs. Separate Controls

Consider two designs.

Design Structure
Separate controls A vs. Control A; B vs. Control B
Shared control A and B vs. one common control

A shared control can improve efficiency, but its validity depends on the comparability of the patient populations and the statistical design.

For example, if treatment A is only available to Biomarker A-positive patients and treatment B is only available to Biomarker B-positive patients, a single unselected control population may not provide an appropriate counterfactual for both groups.

This is why the control strategy must be considered jointly with the biomarker architecture.

Umbrella Trial Operating Characteristics

Operating characteristics should be evaluated at both the substudy level and the master-protocol level. At the substudy level:

  • Type I error
  • Power
  • Expected sample size
  • Probability of success
  • Probability of stopping

At the master-protocol level:

  • Probability each arm recruits adequately
  • Probability at least one arm succeeds
  • Expected total enrollment
  • Expected screening burden
  • Probability of false discoveries
  • Expected number of arms remaining at the end

Probability of at Least One Successful Arm

Suppose three independent treatment hypotheses each have probability \(q\) of being declared successful. Then:

$$ P(\text{at least one success}) = 1-(1-q)^3 $$

For example, if:

$$ q=0.80 $$

then:

$$ P(\text{at least one success}) = 1-(0.20)^3 = 0.992 $$

or 99.2%.

However, this quantity should not be interpreted as evidence that the master protocol has a 99.2% chance of identifying a clinically useful therapy. It is only a mathematical illustration under highly simplified assumptions.

In a real umbrella trial, the probability of at least one success depends on the number of truly active treatments, the correlations among tests, the multiplicity strategy, biomarker prevalence, and the adaptive decision rules.

Probability of False Successes

If multiple inactive treatments are tested, the probability of at least one false-positive signal must also be considered. Suppose \(K\) independent null hypotheses each have type I error \(\alpha\). Then:

$$ P(\text{at least one false positive}) = 1-(1-\alpha)^K $$

This is one reason why the statistical architecture of the umbrella protocol should be established before the results are observed.

Adaptive Arm Dropping and Simulation

Once arm dropping, shared controls, multiple biomarkers, delayed endpoints, and adaptive additions are introduced, exact closed-form calculations can become difficult. Simulation is often useful.

A simulation can:

  1. Generate a virtual patient.
  2. Generate biomarker status.
  3. Assign the patient to a substudy.
  4. Generate treatment outcomes.
  5. Perform the interim analysis.
  6. Apply the arm-dropping rules.
  7. Continue enrollment in surviving arms.
  8. Apply the final decision rules.
  9. Repeat thousands of times.

A Simple Simulation Skeleton in R

set.seed(123)

nsim <- 5000

results <- data.frame(
  sim = 1:nsim,
  success_A = NA,
  success_B = NA,
  success_C = NA
)

for (i in 1:nsim) {

  response_A <- rbinom(
    1,
    size = 30,
    prob = 0.30
  )

  response_B <- rbinom(
    1,
    size = 30,
    prob = 0.20
  )

  response_C <- rbinom(
    1,
    size = 30,
    prob = 0.10
  )

  results$success_A[i] <-
    response_A >= 9

  results$success_B[i] <-
    response_B >= 9

  results$success_C[i] <-
    response_C >= 9
}

colMeans(
  results[,2:4]
)

The numerical thresholds in this illustrative code are not intended to represent a validated umbrella-trial design. In practice, the simulation must reproduce the exact proposed protocol.

What Should Be Simulated?

A serious simulation study can vary:

  • Biomarker prevalence
  • True treatment effects
  • Control response rates
  • Patient accrual rates
  • Dropout rates
  • Biomarker misclassification
  • Interim timing
  • Arm-dropping thresholds
  • New-arm addition rules
  • Multiplicity procedures
  • Correlation among treatment effects

The objective is to determine whether the proposed design behaves as intended under both favorable and unfavorable scenarios.

Simulation Under the Global Null

One particularly important scenario is the global null, in which none of the experimental treatments is effective.

For example:

$$ H_{01}\cap H_{02}\cap\cdots\cap H_{0K} $$

The simulation estimates the probability that the trial nevertheless declares one or more treatments successful. This is a direct way to evaluate the overall false-positive behavior of a complex adaptive umbrella design.

Simulation Under Mixed Scenarios

The global null is not the only scenario of interest. Suppose:

Substudy True Response Rate Interpretation
A 35% Active
B 10% Inactive
C 25% Intermediate

A useful simulation should determine whether the design:

  • Retains A
  • Stops B efficiently
  • Handles C appropriately
  • Maintains the intended error properties

Recruitment Rate Matters

An umbrella trial may be statistically efficient but operationally slow. Suppose the overall disease population accrues at 20 patients per month. If Biomarker C has prevalence 5%, only approximately:

$$ 20(0.05)=1 $$

patient per month is expected to qualify for that treatment pathway.

If the treatment requires 40 evaluable patients, the biomarker prevalence alone could imply approximately:

$$ \frac{40}{1}=40 $$

months of recruitment before accounting for screening failures and dropout.

Operational lesson: The rarest biomarker-defined subgroup can become the bottleneck for the entire development program even when the overall disease population is large.

Umbrella Trial Site Selection

Site selection should consider molecular screening capability in addition to ordinary patient volume. A site may have many patients with the disease but few patients with a particular actionable alteration.

Relevant site-level characteristics include:

  • Historical disease volume
  • Biomarker prevalence
  • Availability of tumor tissue
  • Turnaround time for molecular testing
  • Ability to administer study treatments
  • Imaging capacity
  • Clinical trial experience

Turnaround Time Can Become a Statistical Issue

Suppose biomarker testing takes four weeks. Patients may progress clinically while waiting for assignment. This can create:

  • Screen failures after testing
  • Delayed treatment initiation
  • Selection effects
  • Unequal treatment availability

Therefore, the biomarker workflow should be included in operational simulations when it can affect enrollment or treatment assignment.

Common Mistakes

  1. Confusing umbrella and basket trials. An umbrella trial generally evaluates multiple treatments within one disease framework, whereas basket designs typically evaluate a treatment or strategy across multiple diseases sharing a biological feature.
  2. Assuming the master protocol means one hypothesis. An umbrella protocol can contain several distinct statistical hypotheses.
  3. Ignoring biomarker prevalence. A statistically reasonable substudy may be impossible to recruit if its biomarker is rare.
  4. Ignoring biomarker misclassification. Assay sensitivity and specificity can affect the treatment population and therefore power.
  5. Using a common control without examining comparability. A shared control is useful only when its relationship to the treatment populations is scientifically and statistically defensible.
  6. Ignoring multiplicity. Multiple treatment hypotheses can create a substantial false-positive problem if the inferential objective requires joint error control.
  7. Adding treatment arms without considering statistical consequences. New arms can change the multiplicity and shared-control structure.
  8. Comparing raw response rates across unrelated biomarkers. Different biomarker populations can have different prognoses and baseline response probabilities.
  9. Assuming every arm needs the same design. Different biomarkers can justify different clinical effect assumptions and sample sizes.
  10. Confusing screening efficiency with statistical efficiency. A master protocol can simplify screening while still requiring substantial sample sizes within individual substudies.

A Practical Umbrella Trial Design Workflow

1
Define the disease population covered by the master protocol.
2
Define the biomarkers and assay strategy.
3
Define how overlapping biomarker results are handled.
4
Define treatment assignment for each biomarker subgroup.
5
Determine whether each substudy is single-arm or randomized.
6
Define the endpoint and estimand for each substudy.
7
Specify the null and alternative treatment effects.
8
Determine the sample size for each treatment-biomarker comparison.
9
Calculate the screening requirement using biomarker prevalence.
10
Specify multiplicity and error-control strategy where applicable.
11
Define interim analyses and arm-dropping rules.
12
Simulate the complete master protocol under null, alternative, and mixed scenarios.

What Should Be in the Statistical Analysis Plan?

The statistical documentation for an umbrella trial should be sufficiently detailed that every treatment pathway can be reconstructed. At minimum, specify:

  • Master-protocol population
  • Biomarker definitions
  • Assay methodology
  • Biomarker assignment algorithm
  • Substudy eligibility criteria
  • Treatment assignment
  • Randomization methodology where applicable
  • Primary endpoint for each substudy
  • Estimand for each treatment question
  • Analysis population
  • Null and alternative hypotheses
  • Sample-size assumptions
  • Multiplicity strategy
  • Interim analysis timing
  • Arm-dropping criteria
  • Arm-addition rules
  • Handling of missing data
  • Handling of unevaluable patients
  • Safety monitoring
  • Final analysis rules

Umbrella Trials and Estimands

An umbrella trial can contain different estimands. For example, a randomized substudy might target the treatment effect among patients with a particular biomarker:

$$ \theta_A = E[Y(1)-Y(0)\mid B=A] $$

where \(B=A\) denotes the relevant biomarker-defined population.

A different substudy may target a response probability:

$$ p_B=P(\text{response}\mid B=B,\text{Treatment B}) $$

These are not interchangeable statistical questions.

Estimand principle: Before choosing the statistical test, define precisely what treatment effect is supposed to be estimated, in which patients, under what treatment strategy, and using which endpoint.

Master Protocol Does Not Mean Master Analysis

The master protocol can standardize operations without forcing every substudy to use the same analysis. For example:

Substudy Endpoint Design
A Objective response rate Single-arm
B Progression-free survival Randomized
C Major pathological response Single-arm
D Overall survival Randomized

A master protocol can therefore contain substantial statistical heterogeneity.

Why Simulation Becomes Increasingly Important

For a simple single-arm trial, exact binomial calculations may be sufficient. For a complex umbrella platform, the combination of:

  • Multiple biomarkers
  • Multiple treatment arms
  • Shared controls
  • Adaptive enrollment
  • Arm dropping
  • Arm addition
  • Interim analyses
  • Delayed outcomes
  • Biomarker prevalence
  • Missing data

can make simulation the most practical method for evaluating the complete design.

A Master-Protocol Simulation Structure

A useful simulation can be organized as:

for (sim in 1:nsim) {

  initialize_master_protocol()

  while (study_is_active) {

    generate_patient()

    generate_biomarker()

    assign_substudy()

    generate_outcome()

    update_substudy_data()

    if (interim_analysis_due()) {

      evaluate_stopping_rules()

      drop_inactive_arms()

      add_prespecified_arms_if_allowed()

    }
  }

  store_final_results()
}

The exact implementation depends on the protocol. The important principle is that the simulation should reproduce the entire decision process, not merely simulate final treatment effects.

Global Operating Characteristics

After many simulations, investigators can estimate:

$$ \widehat{P}(\text{false positive}) = \frac{\#\text{simulations with false success}} {N_{\text{sim}}} $$

Similarly:

$$ \widehat{P}(\text{success}) = \frac{\#\text{simulations with correct success}} {N_{\text{sim}}} $$

and:

$$ \widehat{E}(N) = \frac{1}{N_{\text{sim}}} \sum_{i=1}^{N_{\text{sim}}}N_i $$

These quantities allow investigators to compare the proposed design against alternative operating strategies.

Umbrella Trial vs. Conventional Parallel Trials

Feature Separate Trials Umbrella Master Protocol
Screening Potentially repeated Can be centralized
Biomarker testing Protocol-specific Can be shared
Treatment pathways Separate protocols Multiple substudies
Arm changes Requires new trial Can potentially occur within platform framework
Shared controls Usually unavailable across independent trials Potentially available
Statistical complexity Usually lower per trial Often substantially higher
Operational infrastructure Duplicated Potentially shared

The Trade-Off: Efficiency vs. Complexity

Umbrella trials can reduce operational duplication, but they do not make statistical planning simpler. The master protocol must coordinate:

  • Biomarker testing
  • Treatment assignment
  • Multiple hypotheses
  • Different treatment effects
  • Potentially different endpoints
  • Multiplicity
  • Adaptive decisions
  • Shared controls
  • Regulatory considerations

The statistical team therefore needs to think about both individual substudy validity and global master-protocol behavior.

When an Umbrella Trial May Be Useful

An umbrella approach may be particularly useful when:

  • A disease contains several biologically distinct molecular subgroups.
  • Multiple targeted treatments are being developed simultaneously.
  • Biomarker testing can identify relevant treatment populations.
  • Patient populations overlap operationally.
  • A centralized screening infrastructure is valuable.
  • The treatment landscape is expected to evolve.
  • Individual biomarker populations are too small for separate large programs.

When a Conventional Trial May Be Simpler

A conventional trial may be preferable when:

  • There is only one treatment question.
  • There is no meaningful biomarker stratification.
  • The disease population is sufficiently homogeneous.
  • There is little need for adaptive modification.
  • The statistical objectives are straightforward.
  • Shared infrastructure provides little practical benefit.

Umbrella Trials in Oncology Development

Umbrella designs have become particularly relevant to precision oncology because many cancers are now understood as collections of molecularly distinct subgroups rather than single homogeneous diseases.

Examples of master-protocol approaches in oncology include studies such as Lung-MAP and other biomarker-driven programs in which molecular testing is integrated into treatment assignment. The exact statistical architecture differs across studies, and the terms umbrella, platform, and master protocol are not always used identically.

Terminology matters: A trial should be classified according to its actual design architecture, not merely because the words "umbrella," "basket," or "platform" appear in its description.

Common Statistical Questions in an Umbrella Trial

Before the protocol is finalized, the statistical team should be able to answer:

  1. What is the primary hypothesis for each substudy?
  2. Which patients contribute to each hypothesis?
  3. How are biomarker-positive patients identified?
  4. What happens when a patient has multiple biomarkers?
  5. How is the control population defined?
  6. Are treatment comparisons independent?
  7. What multiplicity strategy applies?
  8. When are interim analyses performed?
  9. What causes an arm to stop?
  10. What causes an arm to expand?
  11. Can new arms be added?
  12. How is type I error controlled after adaptation?
  13. What happens if a biomarker becomes rare or obsolete?
  14. How many patients must be screened?
  15. What happens when an assay fails?

A Compact Statistical Framework

An umbrella trial can be represented mathematically as a collection of substudies:

$$ \mathcal{T} = \{T_1,T_2,\ldots,T_K\} $$

where each \(T_k\) corresponds to a treatment-biomarker question. For each substudy:

$$ T_k: \quad H_{0k} \quad\text{vs.}\quad H_{Ak} $$

with a corresponding sample size \(n_k\), treatment effect \(\theta_k\), and decision rule \(D_k\).

The master protocol defines how patients are routed into the \(T_k\)'s and how the individual decision rules interact.

The Most Important Concept

The most important conceptual point is that an umbrella trial is not simply one large clinical trial with several treatment arms.

Its defining feature is the relationship between:

  • A common disease framework
  • Biomarker or subgroup classification
  • Multiple treatment pathways
  • A master protocol
  • Potentially distinct statistical hypotheses

The statistical design therefore has to operate at two levels.

At the first level, each treatment-biomarker substudy must have an appropriate design.

At the second level, the master protocol must account for the interactions created by multiple hypotheses, shared controls, adaptive decisions, and biomarker-dependent recruitment.

Bottom line: An umbrella trial evaluates multiple treatment strategies within a common disease framework, often using biomarker information to route patients to different treatment substudies. Its major operational advantage is the ability to centralize screening, biomarker testing, infrastructure, and governance while evaluating several treatment questions. Its major statistical challenge is that the resulting master protocol can contain multiple hypotheses, different patient populations, shared controls, adaptive decisions, and unequal biomarker prevalence. A rigorous umbrella-trial design therefore requires both substudy-specific statistical planning and master-protocol-level evaluation, frequently supported by simulation.

Umbrella Trial Checklist

Design Element Question to Resolve
Disease population What population enters the master protocol?
Biomarkers Which molecular characteristics determine treatment assignment?
Assay How are biomarker statuses established?
Assignment How is each patient routed to a substudy?
Overlap What happens when multiple biomarkers are present?
Treatment What treatment is evaluated in each subgroup?
Control Is a concurrent control used, and is it shared?
Endpoint What is the primary endpoint for each substudy?
Estimand What treatment effect is being estimated?
Sample size How many evaluable patients are needed?
Screening How many patients must be screened?
Multiplicity How are multiple hypotheses handled?
Interim analysis When are treatment arms evaluated?
Arm dropping What evidence causes an arm to stop?
Arm addition Can new treatments or biomarkers be introduced?
Simulation Does the complete design achieve its intended operating characteristics?

References

Woodcock, J. & LaVange, L.M. (2017). Master protocols to study multiple therapies, multiple diseases, or both. New England Journal of Medicine, 377, 62–70.
Park, J.W., Liu, M.C., Yee, D., et al. (2016). Adaptive randomization of neratinib in early breast cancer. New England Journal of Medicine.
Redig, A.J. & Jänne, P.A. (2015). Basket trials and the evolution of clinical trial design in an era of genomic medicine. Journal of Clinical Oncology, 33, 975–977.
Berry, D.A. (2015). The brave new world of clinical cancer research: adaptive biomarker-driven trials integrating clinical practice with clinical research. Clinical Trials, 12, 92–95.
Renfro, L.A. & Sargent, D.J. (2017). Statistical controversies in clinical research: basket trials, umbrella trials, and platform trials. Journal of Clinical Oncology.
Park, J.W., Liu, M.C., Yee, D., et al. (2016). Adaptive randomization of neratinib in early breast cancer. New England Journal of Medicine.
Mullard, A. (2017). Multicancer master protocols expand clinical-trial options. Nature Reviews Drug Discovery.
Redig, A.J. & Jänne, P.A. (2015). Basket trials and the evolution of clinical trial design in an era of genomic medicine. Journal of Clinical Oncology.