Introduction
Traditional clinical development often evaluates one experimental treatment against a control in a standalone trial.
Once that trial is completed, another trial may begin to evaluate a different treatment, frequently using a new protocol, new control group, new infrastructure, and a new set of investigators and patients.
A platform trial takes a fundamentally different approach.
Instead of treating each treatment comparison as a completely independent trial, a platform can evaluate multiple interventions within a common master protocol.
Treatment arms can be added when new interventions become available and removed when they demonstrate insufficient activity, meet a success criterion, or otherwise reach a prespecified decision point.
What Is a Platform Trial?
A platform trial is a clinical trial framework designed to evaluate multiple interventions, often against a common control, with the possibility of adding or dropping treatment arms over time.
The term "platform" emphasizes that the trial is not necessarily a single fixed comparison.
Instead, the platform provides an infrastructure within which multiple treatment questions can be addressed.
A simplified representation is:
Platform Trials vs. Conventional Trials
Consider a conventional development strategy in which three experimental treatments are each evaluated in separate randomized trials.
| Feature | Conventional Approach | Platform Approach |
|---|---|---|
| Protocol | Separate protocol for each trial | Common master protocol |
| Control group | Typically separate for each trial | May be shared |
| Infrastructure | Repeated | Can be shared |
| Treatment arms | Usually fixed | Can enter or leave |
| Interim adaptation | Usually limited to prespecified trial | Can govern multiple active arms |
| Trial duration | Separate development timelines | Platform can remain active over time |
The important point is that a platform trial is not simply "three trials combined."
The statistical dependencies created by a shared control, common eligibility criteria, shared calendar time, and adaptive decisions must be explicitly considered.
The Master Protocol
The master protocol is the central document defining the overall platform.
It can specify:
- Eligibility criteria
- Study population
- Control treatment
- Active treatment arms
- Randomization procedures
- Primary and secondary endpoints
- Interim analyses
- Stopping rules
- Rules for adding treatment arms
- Rules for dropping treatment arms
- Statistical analysis methods
- Safety monitoring
- Data collection procedures
A master protocol can therefore provide the common framework while individual treatment comparisons address distinct scientific questions.
Multi-Arm Trials
The simplest platform structure is a multi-arm trial.
Suppose three experimental treatments are being evaluated:
against a common control:
The randomization might therefore be:
For equal allocation, each patient could initially have probability:
If treatment C is dropped, the randomization distribution can be revised for the remaining active arms.
For example:
could then use:
provided that such a change was prespecified and statistically justified.
Multi-Stage Trials
A platform becomes multi-stage when treatment decisions are made using accumulating data at interim analyses.
For example, suppose each treatment is evaluated after approximately 100 evaluable patients.
At an interim analysis, an experimental treatment might:
- Continue because the evidence is promising.
- Stop for futility because the evidence is insufficient.
- Stop for efficacy if a sufficiently strong prespecified criterion is met.
- Continue with a modified randomization probability under an adaptive randomization strategy.
The combination of multiple treatment arms and sequential decisions gives the platform its multi-arm, multi-stage character.
Why Shared Controls Matter
One of the most important features of many platform trials is the ability to use a common control group.
Suppose three experimental treatments each require 100 control patients if evaluated separately.
Three independent trials would therefore require:
control patients.
A platform might instead recruit a shared control cohort.
For illustration, suppose 150 control patients can provide sufficient information for the three comparisons under the prespecified design.
| Strategy | Experimental Patients | Control Patients | Total |
|---|---|---|---|
| Three separate trials | 300 | 300 | 600 |
| Illustrative platform | 300 | 150 | 450 |
The numerical values above are illustrative rather than a universal sample-size rule.
The potential efficiency comes from not requiring a completely new control cohort for every experimental treatment.
The Statistical Structure of a Shared-Control Platform
Suppose the endpoint is continuous and larger values indicate better outcomes. Let:
and:
The treatment effect for arm \(j\) is:
The platform may simultaneously evaluate:
for:
The comparisons are statistically related because the same control patients can contribute to multiple treatment-control comparisons.
Correlation Between Treatment Comparisons
This is one of the most important statistical concepts in platform trials.
Suppose treatments A and B are each compared with the same control.
The estimated treatment effects are:
Because both estimates contain \(\bar Y_0\), they are not generally independent.
Indeed, under simple independent observations:
when the treatment groups themselves are independent.
This covariance is one reason that platform-trial statistical methods must account for the shared control rather than treating every comparison as an independent standalone test.
Why the Correlation Can Be Useful
The shared control creates statistical dependence, but it can also provide efficiency.
Instead of estimating a separate control mean for every treatment, the platform obtains information about the control population from a common control cohort.
This can be especially attractive when:
- The control treatment is stable over the relevant enrollment period.
- The patient population is sufficiently comparable across treatment comparisons.
- The endpoint and assessment procedures are common.
- The platform has adequate control information.
When Sharing a Control Becomes More Complicated
A control group is not automatically exchangeable across every treatment arm.
For example, suppose treatment A enters the platform in January while treatment B enters in December.
Patient characteristics, standard of care, diagnostic procedures, background therapy, or other aspects of the clinical environment may have changed.
Therefore, platform designs may need to account for:
- Calendar time
- Site
- Patient characteristics
- Changes in standard of care
- Changes in eligibility criteria
- Changes in endpoint assessment
Adding a New Treatment Arm
One of the defining characteristics of a platform trial is that new treatment arms can potentially be introduced after the platform has already started.
Suppose the original platform contains:
A new treatment D becomes available.
The platform could transition to:
New patients can then be randomized according to the updated allocation scheme.
The statistical analysis must distinguish patients enrolled under different treatment-arm configurations and account for the information actually available for each comparison.
Dropping a Futile Treatment
Suppose treatment B performs poorly at an interim analysis.
The platform might specify a rule such as:
under a Bayesian framework, or a corresponding frequentist predictive-power or conditional-power criterion.
If the criterion is met, treatment B can be dropped for futility.
The platform then continues with:
rather than continuing to allocate patients to B.
Dropping a Successful Treatment
A treatment may also leave the platform because it has met a prespecified success criterion.
For example:
might constitute a Bayesian success rule.
Alternatively, a frequentist platform might use a prespecified test statistic crossing a multiplicity-adjusted efficacy boundary.
Once the success criterion has been met, further randomized patients may no longer be necessary for that comparison.
Platform Lifecycle
A platform can therefore evolve over time.
The platform is therefore a dynamic clinical-development system rather than a single fixed comparison.
Platform Trials and Master Protocols
Several trial concepts are sometimes discussed together under the master protocol umbrella.
| Design Type | Basic Question |
|---|---|
| Umbrella trial | Which treatments are effective for different subtypes within one disease? |
| Basket trial | Does one treatment work across multiple diseases or molecular subgroups? |
| Platform trial | How can multiple treatments be evaluated within an ongoing common framework? |
These concepts can overlap.
For example, a platform could also contain biomarker-defined subgroups, making the design simultaneously multi-arm, multi-stage, and stratified by disease subtype.
Platform Trial vs. Umbrella Trial
An umbrella trial typically evaluates multiple treatments within a single disease, often with treatment assignment determined by molecular or clinical subtype.
For example:
A platform trial is defined more by its ability to maintain a common, adaptable infrastructure for multiple interventions over time.
Therefore, an umbrella trial can be a platform, but the concepts are not synonymous.
Platform Trial vs. Basket Trial
A basket trial typically evaluates the same intervention across multiple diseases, histologies, or molecular subgroups.
For example:
The scientific question is different from the typical multi-treatment platform question.
A Platform Is Not Necessarily Bayesian
Platform trials are often associated with Bayesian statistics, particularly because Bayesian methods provide a convenient framework for making repeated probability-based decisions.
However, a platform trial does not require Bayesian statistics.
Frequentist, Bayesian, and hybrid approaches can all be used.
| Approach | Possible Decision Quantity |
|---|---|
| Frequentist | Test statistic, p-value, confidence interval, conditional power |
| Bayesian | Posterior probability of clinically meaningful benefit |
| Predictive Bayesian | Predictive probability of future success |
| Hybrid | Bayesian decision rules with frequentist operating-characteristic evaluation |
Bayesian Decision Rules
Suppose the treatment effect for arm \(j\) is:
A Bayesian success rule could be:
where:
- \(D\) is the accumulated data.
- \(\delta\) is the minimum clinically meaningful effect.
A futility rule could instead be:
The exact thresholds must be established during design development.
Predictive Probability of Success
Another useful Bayesian quantity is the probability that a treatment will ultimately succeed if additional patients are enrolled.
This can be expressed conceptually as:
A treatment could therefore be dropped if:
and continued if:
This can be more informative than simply asking whether the current estimate has crossed a fixed significance threshold.
Frequentist Interim Decisions
Frequentist platform designs can use interim test statistics and prespecified boundaries.
For treatment \(j\), define:
An efficacy boundary could take the form:
while a futility boundary might use:
The boundaries can vary by analysis time and can be constructed to preserve the desired overall error properties.
Multiplicity Becomes Important
Suppose a platform evaluates three treatments:
If each hypothesis is tested at an unadjusted one-sided level of 0.05, the probability of at least one false positive can exceed 5%.
If the tests were independent, the probability of at least one false positive would be:
or approximately 14.3%.
The actual familywise error for a shared-control platform is not necessarily this value because treatment comparisons may be correlated.
Nevertheless, the example illustrates why multiplicity must be addressed.
Multiplicity Strategies
Potential approaches include:
- Bonferroni-type adjustments
- Holm procedures
- Hierarchical testing
- Graphical multiple-testing procedures
- Closed-testing approaches
- Joint modeling of correlated treatment comparisons
- Bayesian decision rules with prespecified operating-characteristic evaluation
The appropriate strategy depends on the estimand, confirmatory objectives, number and relationship of hypotheses, and regulatory framework.
Familywise Error vs. Per-Arm Error
A platform protocol should clearly distinguish between:
and:
If the platform has a confirmatory objective involving several treatment comparisons, the relevant error-control strategy should be established before the confirmatory analysis.
Common Control and Calendar Time
Consider a platform that begins with treatments A and B.
Six months later, treatment C enters.
The control patients enrolled during the first six months may differ from those enrolled later.
Therefore, a treatment effect can potentially be influenced by temporal changes unrelated to the treatment itself.
A statistical model may therefore include calendar time or other relevant covariates.
For example:
where \(Z_i\) could represent a prespecified baseline or temporal adjustment variable.
The exact model depends on the endpoint and estimand.
Concurrent vs. Non-Concurrent Controls
This distinction is particularly important in long-running platforms.
A concurrent control patient is enrolled during the period in which the corresponding treatment arm is active.
A non-concurrent control patient may have been enrolled during a different period.
Using non-concurrent controls can create additional assumptions because the clinical environment may have changed.
Randomization in Platform Trials
Randomization can be fixed or adaptive.
Under fixed randomization, suppose four active groups exist:
An equal allocation ratio is:
An alternative could be:
where the control receives twice the allocation probability of each experimental treatment.
The choice of allocation affects both information and patient exposure.
Adaptive Randomization
In adaptive randomization, allocation probabilities can change in response to accumulating information.
For example, after an interim analysis the probabilities might change from:
to an allocation such as:
These numbers are illustrative only.
The new probabilities would need to arise from a prespecified adaptive randomization algorithm.
Platform Trials and Response-Adaptive Randomization
Response-adaptive randomization attempts to allocate more future patients to treatments that appear more promising.
Conceptually:
The exact function must be defined by the design.
Response-adaptive randomization can introduce substantial statistical complexity.
It can affect:
- Allocation probabilities
- Estimator properties
- Variance
- Power
- Type I error
- Interpretation of treatment comparisons
- Operational predictability
Therefore, it should not be added to a platform merely because it appears intuitively attractive.
A Complete Worked Example
Consider a hypothetical Phase II platform evaluating treatments for a disease with a continuous efficacy endpoint.
The platform begins with:
where:
- \(C_0\) = common control
- \(A\) = experimental treatment A
- \(B\) = experimental treatment B
A third treatment, C, may be introduced later.
Suppose the clinically meaningful treatment effect is:
units of the endpoint.
The initial platform is designed with an interim analysis after approximately 120 patients have contributed evaluable data.
| Design Component | Illustrative Value |
|---|---|
| Initial treatment arms | A, B |
| Common control | C0 |
| Interim information | Approximately 120 evaluable patients |
| Clinically meaningful effect | \(\delta=5\) |
| Futility criterion | Prespecified low conditional/predictive probability of success |
| Efficacy criterion | Prespecified success boundary |
| Potential new arm | C |
The numbers above are deliberately illustrative. Actual platform-trial sample sizes and boundaries must be derived from the specific endpoint, estimand, error-control strategy, and operating characteristics.
Stage 1: Initial Platform
Suppose the initial randomization is:
Patients are enrolled under the master protocol.
At the interim analysis, suppose treatment A has accumulated strong evidence of benefit, while treatment B has weak evidence.
For illustration, assume the decision rules are:
What Happens to the Control Group?
The platform does not necessarily stop merely because A succeeds or B fails.
If A stops for success and B stops for futility, the platform might temporarily have no active experimental arm.
If treatment C is ready to enter, the platform can transition to:
The control therefore becomes part of the continuing platform infrastructure.
Introducing Treatment C
Suppose treatment C becomes available after A and B have reached their decisions.
The platform can add C according to a prespecified amendment process.
The new active comparison is:
The statistical analysis must specify:
- Which control observations contribute to the C comparison.
- Whether historical or non-concurrent controls are eligible.
- How calendar time is handled.
- How the new comparison enters the multiplicity framework.
- How its type I error or posterior decision criterion is established.
A Simple Timeline
| Time | Active Arms | Decision |
|---|---|---|
| Month 0 | C0, A, B | Platform launches |
| Month 6 | C0, A, B | Interim analysis |
| Month 6 | C0, A | B dropped for futility |
| Month 10 | C0, A | C becomes available |
| Month 10 | C0, A, C | C added |
| Month 14 | C0, C | A reaches success criterion |
| Month 20 | C0, C | Final C analysis |
This illustrates why the term "platform" is useful: the infrastructure persists while the treatment portfolio changes.
Statistical Information Accumulates Differently for Each Arm
A common misconception is that every treatment arm necessarily has the same amount of information at every interim analysis.
That is generally not true.
Suppose:
- A entered at Month 0.
- B entered at Month 0.
- C entered at Month 10.
At Month 12, A and B may have substantial information while C has relatively little.
Therefore, platform designs often use information fractions or arm-specific information times.
Information Fraction
For treatment \(j\), define an information fraction:
where \(I_j(t)\) is the information available at interim time \(t\).
The value \(t_j\) might range from:
This allows interim boundaries to be defined according to statistical information rather than simply calendar time.
Interim Analyses Need Prespecified Rules
A platform can have many possible interim analyses.
For example:
At each interim analysis, a treatment can be evaluated according to the prespecified decision rules.
The statistical design must account for repeated opportunities to stop for success.
Group Sequential Methods Within a Platform
A platform can incorporate group-sequential methods.
For example, an efficacy boundary may be specified as:
where \(c(t)\) changes according to the information fraction.
An O'Brien-Fleming-like approach may use a more stringent early boundary and a less stringent final boundary.
A Pocock-like approach may use more similar boundaries across analyses.
The exact boundary construction depends on the platform's multiplicity and interim-analysis structure.
Multiplicity Exists in Two Dimensions
Platform trials can face multiplicity across both:
- Treatments
- Interim analyses
For example, three treatment arms evaluated at three interim analyses create multiple opportunities to make efficacy decisions.
The design therefore needs to consider the joint probability of false positive decisions across the relevant treatment and time dimensions.
Familywise Error in a Platform
Suppose the confirmatory family contains three treatment hypotheses:
The familywise error rate can be written as:
A confirmatory design may target:
The exact calculation can require multivariate distributions, simulation, closed testing, graphical procedures, or other methods depending on the design.
Why Simulation Is Often Essential
For a simple two-stage single-arm trial, exact binomial enumeration may be enough.
Platform trials can be substantially more complicated.
For example, a platform may include:
- Multiple treatment arms
- Multiple interim analyses
- Shared controls
- Unequal randomization
- Adaptive randomization
- Arms entering at different times
- Arms dropping at different times
- Covariate adjustment
- Missing outcomes
- Delayed outcomes
- Different treatment effects
In such settings, Monte Carlo simulation is often the most practical way to evaluate operating characteristics.
Simulation-Based Operating Characteristics
Important quantities include:
- Familywise type I error
- Power for each treatment
- Probability of correctly dropping futile treatments
- Probability of incorrectly dropping active treatments
- Probability of early success
- Expected sample size
- Maximum sample size
- Probability each arm enters the platform
- Probability each arm reaches final analysis
- Bias and precision of treatment-effect estimates
- Average trial duration
A platform design should be evaluated under many plausible scenarios rather than only one assumed treatment-effect configuration.
Scenario-Based Simulation
Suppose a platform has three experimental treatments.
A useful simulation program might include:
| Scenario | A | B | C |
|---|---|---|---|
| Global null | 0 | 0 | 0 |
| A active | δ | 0 | 0 |
| B active | 0 | δ | 0 |
| C active | 0 | 0 | δ |
| A and B active | δ | δ | 0 |
| All active | δ | δ | δ |
| Small effects | δ/2 | δ/2 | δ/2 |
These scenarios reveal how the platform behaves when different combinations of treatments are actually effective.
Global Null Scenario
The global null is particularly important for evaluating familywise type I error.
Suppose:
The simulation asks:
This should satisfy the relevant error-control requirement.
Power Is Not a Single Number in a Platform
In a conventional two-arm trial, people often refer to "the power" of the study.
In a platform, several power quantities may be relevant.
For treatment A:
Similarly:
may be evaluated separately.
Investigators may also consider:
- Marginal power
- Joint power
- Power to identify at least one effective treatment
- Power to identify all effective treatments
- Power under different combinations of active and inactive treatments
Marginal vs. Joint Power
Suppose A and B are both truly effective.
Marginal power asks whether A is detected:
while joint power asks whether both are detected:
Joint power is generally lower than either marginal probability.
This distinction becomes important when the scientific objective is to identify a complete set of effective interventions rather than simply detect individual effects.
Conditional Power
Frequentist platform designs can also use conditional power.
Conceptually:
A treatment may be dropped if the conditional power falls below a prespecified futility threshold.
Conditional power can therefore provide an evidence-based futility decision without requiring the interim estimate itself to cross a fixed efficacy boundary.
Predictive Probability vs. Conditional Power
These concepts are related but not identical.
| Quantity | General Interpretation |
|---|---|
| Conditional power | Probability of eventual success conditional on current data and specified assumptions about the future effect |
| Predictive probability | Probability of eventual success integrating uncertainty about future outcomes and model parameters |
| Posterior probability | Probability that the current treatment effect exceeds a clinically relevant threshold given the observed data |
These quantities should not be treated as interchangeable.
Delayed Outcomes
Platform trials frequently encounter delayed endpoints.
For example, progression-free survival or overall survival may take months or years to mature.
The platform may therefore use:
- Interim information based on partially mature data
- Time-to-event methods
- Information-based stopping rules
- Predictive probabilities
- Event-driven interim analyses
The timing of decisions should be aligned with the actual information available rather than simply the number of patients enrolled.
Time-to-Event Platform Trials
For a time-to-event endpoint, treatment effects may be expressed through a hazard ratio:
where \(HR<1\) may indicate benefit when lower hazard is favorable.
A platform might test:
Interim decisions can then be based on information measured through the number of observed events.
Sample Size in Platform Trials
Sample-size planning is more complicated than simply multiplying the sample size for one two-arm trial by the number of treatments.
The required information depends on:
- Number of treatment arms
- Control allocation
- Treatment allocation
- Endpoint variance or event rate
- Target treatment effect
- Multiplicity adjustment
- Interim analyses
- Shared-control correlations
- Expected treatment-arm entry and dropout
- Accrual rate
- Information timing
A Simple Continuous-Endpoint Approximation
For a single treatment-control comparison with equal allocation and a continuous endpoint, a rough two-sided sample-size relationship is:
A platform with multiple treatment arms cannot simply use this equation without modification.
The allocation ratio, multiplicity strategy, shared control, and adaptive structure all matter.
Control Allocation in a Multi-Arm Trial
Suppose there are \(K\) experimental arms and a control.
If each experimental arm receives \(n_T\) patients and the control receives \(n_C\), then the allocation ratio is:
A larger control group improves precision for every treatment-control comparison, but it also means fewer patients are allocated to experimental treatments for a fixed total sample size.
Thus, control allocation is a design optimization problem rather than a purely statistical afterthought.
Operational Efficiency
Platform trials can potentially reduce more than patient numbers.
A common platform can also share:
- Clinical sites
- Data-management infrastructure
- Central laboratories
- Imaging procedures
- Randomization systems
- Statistical programming
- Safety monitoring infrastructure
- Vendor relationships
- Protocol training
These efficiencies can be particularly important when multiple treatments address the same disease population.
Operational Complexity
Efficiency does not mean simplicity.
Platform trials can be operationally complex because:
- Treatment arms may have different dosing schedules.
- Arms may have different safety profiles.
- Arms may enter at different times.
- Randomization probabilities can change.
- Eligibility criteria may evolve.
- Control treatment may change over time.
- Different arms may have different follow-up requirements.
The platform therefore requires strong governance and careful version control.
Protocol Amendments
Adding or removing an arm generally requires a prespecified amendment and appropriate regulatory and operational processes.
The statistical framework should specify:
- How a new treatment becomes eligible for entry
- Which data contribute to its analysis
- How randomization is modified
- How multiplicity is handled
- How its sample size is determined
- How success and futility are defined
Blinding
Platform trials can be blinded or open-label depending on the interventions and endpoints.
Blinding can become more complicated as multiple treatments enter and leave the platform.
For example, if treatments have visibly different administration procedures, maintaining blinding may require separate matching procedures.
The statistical design should be developed together with the operational blinding strategy.
Safety Monitoring in a Platform
Each treatment arm can have its own safety profile.
The platform therefore needs rules for:
- Arm-specific safety monitoring
- Shared safety oversight
- Stopping an individual treatment for toxicity
- Stopping the platform for broader safety concerns
- Reviewing cumulative safety data
An arm-specific safety stop does not necessarily imply that other treatments must stop.
A Simple Platform Decision Matrix
| Efficacy Evidence | Safety Evidence | Potential Decision |
|---|---|---|
| Strong | Acceptable | Continue or declare success according to protocol |
| Weak | Acceptable | Continue or stop for futility |
| Strong | Unacceptable | Stop treatment for safety |
| Weak | Unacceptable | Stop treatment |
Estimands in Platform Trials
A platform trial should define its estimands explicitly.
For treatment \(j\), an estimand might specify:
- Population
- Treatment condition
- Endpoint
- Summary measure
- Handling of intercurrent events
For example, a continuous endpoint estimand could target:
The exact estimand depends on the scientific question and treatment context.
Different Arms Can Address Different Questions
A platform should not be assumed to have a single universal estimand.
Treatment A might be evaluated in one biomarker-defined population while treatment B is evaluated in another.
In that situation, the platform can share infrastructure without pretending that all treatment comparisons answer exactly the same scientific question.
Common Statistical Models
Depending on the endpoint, platform analyses may use:
- Linear models
- Generalized linear models
- Logistic regression
- Cox proportional-hazards models
- Mixed-effects models
- Repeated-measures models
- Bayesian hierarchical models
- Joint models
The platform structure determines how the treatment indicators, shared control, time, site, and other factors enter the analysis.
Hierarchical Borrowing
Bayesian platform trials sometimes consider hierarchical models that allow information to be partially shared across treatment arms or populations.
Conceptually, treatment effects might follow:
where \(\tau^2\) represents between-treatment heterogeneity.
This can allow some information sharing while retaining differences between treatments.
Platform Trials With Biomarkers
Platform trials can be especially useful when treatments are associated with different biomarkers.
For example:
The platform can therefore combine treatment and subgroup adaptation.
This creates additional multiplicity and sample-size considerations.
What Happens When a Treatment Is Dropped?
Dropping a treatment has several consequences.
First, future patients are no longer randomized to that treatment.
Second, the allocation probabilities among the remaining arms may change.
Third, the analysis population for the dropped treatment is effectively closed.
Fourth, the platform's future control observations may or may not contribute to the final analysis depending on the prespecified analysis plan.
Therefore, "drop an arm" is a statistical operation as well as an operational one.
What Happens When a Treatment Enters?
Adding an arm introduces a new treatment-control comparison.
The platform should therefore specify:
- Entry criteria
- Required sample size
- Initial randomization probability
- Interim timing
- Futility rule
- Efficacy rule
- Final analysis rule
- Multiplicity framework
- Control eligibility
The fact that a new treatment entered later does not eliminate the need to define its statistical operating characteristics.
Platform Trial vs. Standard Multi-Arm Trial
| Feature | Fixed Multi-Arm Trial | Platform Trial |
|---|---|---|
| Number of treatments | Usually fixed | Can change |
| Master protocol | Possible | Central feature |
| Shared control | Possible | Commonly important |
| Arms added over time | Uncommon | Potentially planned |
| Arms dropped over time | Possible | Central adaptive feature |
| Long-running infrastructure | Not necessarily | Typical |
Platform Trial vs. Group Sequential Trial
A group sequential trial generally has one primary treatment comparison and multiple interim analyses.
A platform can have:
Therefore, a platform can incorporate group-sequential methods without being equivalent to a conventional group sequential trial.
Platform Trial vs. Adaptive Trial
The terms are also not synonymous.
An adaptive trial is any trial that prospectively plans changes based on accumulating data while maintaining the specified validity of inference.
A platform trial is a particular type of adaptive infrastructure designed to evaluate multiple interventions over time.
Thus:
conceptually, although the exact terminology used in the literature can vary.
Simulation in R
A simplified platform simulation can illustrate the basic mechanics.
set.seed(123)
nsim <- 10000
n_control <- 100
n_A <- 100
n_B <- 100
mu_control <- 0
mu_A <- 0.5
mu_B <- 0
sigma <- 1
results <- data.frame(
A_success = logical(nsim),
B_success = logical(nsim)
)
for (i in seq_len(nsim)) {
y_control <- rnorm(
n_control,
mean = mu_control,
sd = sigma
)
y_A <- rnorm(
n_A,
mean = mu_A,
sd = sigma
)
y_B <- rnorm(
n_B,
mean = mu_B,
sd = sigma
)
z_A <- (
mean(y_A) - mean(y_control)
) / sqrt(
sigma^2/n_A +
sigma^2/n_control
)
z_B <- (
mean(y_B) - mean(y_control)
) / sqrt(
sigma^2/n_B +
sigma^2/n_control
)
results$A_success[i] <- z_A > 2.5
results$B_success[i] <- z_B > 2.5
}
This is intentionally a simplified illustration.
A production platform-trial simulation should reproduce the actual randomization algorithm, interim analyses, stopping rules, enrollment timing, missingness, endpoint distribution, multiplicity procedure, and analysis model.
Estimating Operating Characteristics
The simulated probability of success for treatment A is:
mean(results$A_success)
Similarly:
mean(results$B_success)
If A is truly effective and B is not, these quantities help characterize the marginal behavior of the platform.
The simulation can be extended to calculate the probability of at least one false positive:
mean( results$A_success | results$B_success )
The actual expression should use the intended logical structure carefully; for example:
mean( results$A_success | results$B_success )
would estimate the proportion of simulations in which at least one of the two decisions was positive.
Adding an Interim Analysis to the Simulation
A more realistic simulation would divide enrollment into stages.
n_stage1 <- 50
n_stage2 <- 50
for (i in seq_len(nsim)) {
# Stage 1 enrollment
# Calculate interim treatment effect
# Apply futility/efficacy boundaries
# If treatment continues:
# enroll Stage 2 patients
# Perform final analysis
}
This structure allows the simulation to estimate:
- Probability of early success
- Probability of early futility
- Probability of reaching final analysis
- Expected sample size
- Final power
- Familywise type I error
Expected Sample Size in a Platform
For a fixed two-stage trial, expected sample size can often be written simply.
For a platform, the calculation can become substantially more complicated.
Suppose there are \(K\) treatment arms.
The total sample size can be represented as:
where:
- \(N_C\) is the number of control patients.
- \(N_j\) is the number of patients assigned to treatment \(j\).
But \(N_j\) is itself random when treatment arms can enter, stop, or change allocation.
Therefore:
with each expectation depending on the platform's adaptive rules.
Expected Number of Patients by Treatment
A useful operating characteristic is:
for every treatment arm.
This can reveal how much exposure is expected under different treatment-effect scenarios.
For example, a simulation might produce:
| Arm | Mean N | Maximum N |
|---|---|---|
| Control | 180 | 250 |
| A | 92 | 120 |
| B | 61 | 120 |
| C | 48 | 100 |
These values are illustrative.
They demonstrate why reporting only total sample size can hide important features of an adaptive platform.
Probability of Dropping an Effective Treatment
Futility rules create a potential risk:
This is sometimes called a false-futility probability.
A good platform evaluation should examine this probability under clinically meaningful effect sizes.
For example:
| True Effect | Probability of Futility Stop |
|---|---|
| 0 | High |
| Small positive effect | Intermediate |
| Clinically meaningful effect | Should be acceptably low |
| Large effect | Very low |
Probability of Early Success
Similarly, investigators can evaluate:
This helps determine how often highly effective treatments can leave the platform before reaching the maximum planned information.
Trial Duration
Patient numbers are not the only efficiency metric.
Platform simulations should often estimate:
- Calendar time to first treatment decision
- Calendar time to final treatment decision
- Time required to introduce a new treatment
- Time spent waiting for endpoint maturation
- Time to platform completion
A design that saves patients but requires substantially longer follow-up may have a different operational profile from one that uses more patients but produces faster decisions.
Accrual Rate Matters
Suppose a platform enrolls 10 patients per month.
If treatment A requires 100 evaluable patients, recruitment alone requires approximately:
months, before accounting for endpoint maturation.
If three arms are recruiting simultaneously, the total platform accrual rate may be considerably higher, depending on eligibility and randomization.
Eligibility and Arm-Specific Eligibility
A platform can have common eligibility criteria plus treatment-specific criteria.
For example:
This means a patient can be eligible for the platform but not eligible for every treatment arm.
Randomization must therefore occur only among the treatments for which that patient is eligible.
Platform Randomization Is Not Always a Simple \(1:1:1:1\)
Suppose a patient is eligible for A and B but not C.
A randomization algorithm could assign among:
rather than:
The platform therefore needs an eligibility-aware randomization system.
Master Protocol Governance
Long-running platforms require governance structures that can manage:
- Treatment-arm entry
- Treatment-arm closure
- Protocol amendments
- Statistical changes
- Safety signals
- Data-access rules
- Blinding
- Independent decision-making
An independent data monitoring or decision-making body may be involved depending on the trial.
Operational Blinding and Information Leakage
Adaptive platforms can create opportunities for information leakage.
For example, investigators may infer that a treatment is performing well if its enrollment probability suddenly increases.
The platform should therefore determine which adaptation information is visible to:
- Investigators
- Patients
- Sponsors
- Statistical teams
- Operational teams
and which information remains restricted.
Statistical Team Structure
Complex platforms often benefit from clearly separated responsibilities for:
- Trial statisticians
- Independent statistical monitoring
- Data management
- Randomization programming
- Simulation programming
- Final analysis programming
The exact structure depends on the platform's design and blinding requirements.
Common Mistakes
- Thinking a platform is simply several trials sharing a website. The shared-control and adaptive statistical structure must be explicitly designed.
- Ignoring calendar time. A treatment entering years later may not have a directly exchangeable control population.
- Using an unadjusted 0.05 threshold repeatedly. Multiple treatment arms and interim looks can inflate false-positive error.
- Assuming all arms have equal information. Arms can enter and leave at different times.
- Automatically using response-adaptive randomization. Adaptive randomization adds substantial statistical and operational complexity and is not required for a platform.
- Failing to simulate the actual platform. Simple closed-form calculations often do not capture the full adaptive structure.
- Ignoring treatment-arm entry. Adding an arm creates a new statistical comparison.
- Ignoring treatment-arm exit. Dropping an arm changes future randomization and information accumulation.
- Failing to distinguish concurrent from non-concurrent controls. This can create important interpretational and statistical issues.
- Reporting only overall power. Treatment-specific marginal and joint operating characteristics may also be important.
- Confusing posterior probability with frequentist type I error. These are different statistical concepts.
- Assuming a platform must be Bayesian. Platform trials can use Bayesian, frequentist, or hybrid methods.
A Practical Platform-Trial Design Workflow
What Should Be Specified in the Statistical Analysis Plan?
The statistical documentation should be sufficiently detailed that the platform's decisions can be reproduced.
At minimum, specify:
- Estimands
- Primary endpoints
- Analysis populations
- Treatment-arm definitions
- Control definition
- Randomization procedure
- Shared-control rules
- Concurrent-control rules
- Interim-analysis timing
- Futility boundaries
- Efficacy boundaries
- Multiplicity strategy
- Type I error control
- Power definitions
- Sample-size assumptions
- Arm-entry criteria
- Arm-removal criteria
- Missing-data handling
- Safety stopping procedures
- Simulation assumptions
Example Operating-Characteristic Table
A platform design might summarize its simulated operating characteristics as follows.
| Scenario | Power A | Power B | Power C | Mean Total N |
|---|---|---|---|---|
| Global null | Type I error metric | Type I error metric | Type I error metric | Depends on stopping rules |
| A active | High | Low false-positive probability | Low false-positive probability | Scenario-dependent |
| B active | Low false-positive probability | High | Low false-positive probability | Scenario-dependent |
| C active | Low false-positive probability | Low false-positive probability | High | Scenario-dependent |
| All active | High | High | High | Scenario-dependent |
Actual numerical values should come from the final platform simulation rather than being inferred from a simple two-arm sample-size calculation.
Platform Trials and Regulatory Considerations
A platform trial can be attractive from a development perspective, but its statistical complexity means that the design should be justified carefully.
Important topics include:
- Definition of each treatment comparison
- Control comparability
- Multiplicity
- Interim analyses
- Type I error
- Statistical adaptations
- Integrity of treatment comparisons
- Data access and blinding
- Protocol amendments
For confirmatory development, the platform's statistical operating characteristics should be demonstrated under realistic scenarios before the trial begins.
Confirmatory vs. Exploratory Platform Objectives
Not every platform comparison has the same evidentiary purpose.
Some treatment arms may be intended for exploratory screening.
Others may have confirmatory objectives.
The distinction matters because the required error control, multiplicity strategy, and interpretation of results can differ substantially.
Platform Trials Can Be Perpetual
One of the most interesting possibilities is the perpetual platform.
Instead of designing the platform to end after treatment A, B, and C have been evaluated, the infrastructure can continue as new interventions become available.
Conceptually:
The platform becomes a continuing mechanism for evaluating treatments within a disease area.
Why Perpetual Platforms Are Attractive
A successful perpetual platform can potentially avoid repeatedly rebuilding:
- Clinical infrastructure
- Site contracts
- Randomization systems
- Data standards
- Central laboratory procedures
- Statistical programming infrastructure
- Safety monitoring processes
This can make the development ecosystem more responsive to new candidate treatments.
Why Perpetual Platforms Are Difficult
A long-running platform also creates accumulating complexity.
Over time:
- Standard of care can change.
- Eligibility criteria can evolve.
- New treatments can have different mechanisms.
- Control outcomes can change.
- Statistical assumptions can become outdated.
- Operational processes can change.
The longer a platform runs, the more important careful governance and statistical monitoring become.
A Mental Model for Platform Trials
A useful way to remember the concept is:
The Most Important Statistical Insight
The defining statistical challenge of a platform trial is not simply that there are many treatments.
It is that the treatment comparisons are embedded in a shared, adaptive system.
The system creates dependencies through:
- Shared controls
- Shared patients and eligibility criteria
- Common calendar time
- Repeated interim analyses
- Adaptive randomization
- Arm entry and exit
- Common decision rules
Consequently, the platform should be designed as one statistical system rather than as a collection of independent analyses.
Platform Trial Checklist
| Question | Design Requirement |
|---|---|
| What is the scientific question? | Define the estimand |
| Which treatments are active? | Define initial arms |
| Can new treatments enter? | Define entry rules |
| Can treatments leave? | Define efficacy and futility rules |
| Is the control shared? | Justify exchangeability and define control rules |
| When are interim analyses performed? | Define information-based timing |
| How are treatment effects analyzed? | Specify statistical model |
| How is multiplicity handled? | Define error-control strategy |
| How is adaptive randomization performed? | Specify algorithm or use fixed allocation |
| How is performance evaluated? | Simulate operating characteristics |
| What happens if the control changes? | Prespecify transition rules |
| How is safety monitored? | Define arm-specific and platform-level rules |
Summary of the Worked Platform
The hypothetical platform began with:
Treatment B was dropped for futility.
Treatment C subsequently entered.
Treatment A eventually met its success criterion and left the platform.
The resulting active platform became:
The key feature is that the platform infrastructure remained operational throughout these transitions.
| Platform Feature | Illustration |
|---|---|
| Master protocol | Common clinical and statistical framework |
| Multiple arms | A, B, C |
| Shared control | C0 |
| Interim analysis | Used for treatment decisions |
| Futility | B removed |
| New-arm entry | C added |
| Success | A removed after meeting success criterion |
| Continuing platform | C remains active |
What a Platform Trial Is Really Buying You
The central value proposition of a platform is not simply "fewer patients."
It is the ability to create a reusable statistical and operational infrastructure for answering multiple treatment questions.
Potential advantages include:
- Shared controls
- Shared infrastructure
- Faster evaluation of new treatments
- Early discontinuation of futile treatments
- Potentially earlier identification of successful treatments
- Efficient use of accumulated control information
- Continuous learning within a disease area
These advantages come with corresponding costs in statistical planning, governance, simulation, operational complexity, and regulatory coordination.
Key Takeaways
- A platform trial is a persistent trial framework, not merely a multi-arm trial.
- A master protocol defines the overarching operational and statistical framework.
- Multiple treatment arms can be evaluated simultaneously.
- Shared controls can improve efficiency but create statistical dependence and require careful consideration of concurrent control and calendar-time effects.
- Treatment arms can be added or dropped according to prespecified rules.
- Interim analyses can be used to identify futile or successful treatments before maximum enrollment.
- Adaptive randomization is optional.
- Bayesian statistics are optional. Platform trials can be frequentist, Bayesian, or hybrid.
- Multiplicity must be addressed when multiple confirmatory hypotheses are evaluated.
- Simulation is often essential for demonstrating type I error, power, sample size, arm-selection properties, and trial duration.
References
Berry, S.M., Carlin, B.P., Lee, J.J. & Muller, P. (2010).
Bayesian Adaptive Methods for Clinical Trials.
CRC Press.
Park, J.W., Liu, M.C., Yee, D., et al. (2016).
Adaptive randomization of neratinib in early breast cancer.
Clinical Cancer Research.
Berry, D.A. (2015).
The brave new world of clinical cancer trials.
New England Journal of Medicine, 373, 1753–1755.
Saville, B.R. & Berry, S.M. (2016).
Efficiencies of platform clinical trials: a vision of the future.
Clinical Trials.
Woodcock, J. & LaVange, L.M. (2017).
Master protocols to study multiple therapies, multiple diseases, or both.
New England Journal of Medicine, 377, 62–70.
Angus, D.C. (2019).
Fewer patients, more trials: a new paradigm for clinical trials.
Critical Care Medicine.
Park, J.W., Liu, M.C., Yee, D., et al. (2020).
Adaptive randomization of neratinib in early breast cancer.
Clinical Cancer Research.
Hobbs, B.P., Landin, R. & Wang, L. (2021).
Statistical considerations for platform trials.
Clinical Trials.
Greenstreet, A., et al. (2023).
Platform trial designs and their statistical considerations.
Contemporary Clinical Trials.