Introduction
Clinical trials do not always need to wait until the planned final analysis before evaluating whether the accumulating evidence is sufficient to continue. An interim analysis is a prespecified statistical evaluation performed before the final analysis, usually while patients are still being followed or enrolled.
Interim analyses can serve several purposes. A study may stop early because the experimental treatment has demonstrated compelling evidence of benefit, because the treatment is unlikely to provide the desired benefit, or because an important safety concern has emerged.
The challenge is that repeatedly examining accumulating data can change the probability of a false-positive conclusion if the statistical procedure is not properly designed.
What Is an Interim Analysis?
Suppose a randomized clinical trial is ultimately expected to accumulate 400 primary endpoint events. Instead of waiting for all 400 events, the study might perform analyses after approximately 100, 200, 300, and 400 events.
The first three are interim analyses and the final analysis occurs at the fourth look.
| Analysis | Events | Information Fraction |
|---|---|---|
| Interim 1 | 100 | 25% |
| Interim 2 | 200 | 50% |
| Interim 3 | 300 | 75% |
| Final | 400 | 100% |
The important quantity is not necessarily the percentage of patients enrolled. It is usually the proportion of the statistical information that has accumulated for the primary analysis.
Information Fraction
Let \(I_k\) denote the information available at analysis \(k\), and let \(I_K\) denote the information planned for the final analysis. The information fraction is:
At the final analysis:
For a time-to-event endpoint in which information is approximately proportional to the number of events, the information fraction is often approximated by:
where \(D_k\) is the number of events observed at analysis \(k\).
Why Perform Interim Analyses?
There are several fundamentally different reasons for introducing interim analyses.
Interim Analyses and Type I Error
Suppose a conventional one-sided hypothesis test uses:
If the investigator performs one analysis at the end, a critical value can be chosen to provide the desired type I error.
Now suppose the investigator examines the data three times and declares success if any analysis produces a nominal \(p<0.025\).
The probability of obtaining at least one apparently significant result under the null hypothesis can exceed 0.025.
The analyses are correlated, so the inflation is not simply the sum of the nominal alpha levels, but the principle is straightforward: repeated unadjusted looks at accumulating data can inflate the overall type I error.
Group Sequential Designs
A group sequential design divides the trial into groups or information periods and permits formal statistical decisions at interim analyses.
Suppose there are \(K\) analyses. Let:
- \(Z_k\) = test statistic at analysis \(k\)
- \(t_k\) = information fraction at analysis \(k\)
- \(u_k\) = upper efficacy boundary
- \(l_k\) = lower futility boundary
An efficacy decision can be represented as:
A futility decision can be represented as:
If the statistic lies between the boundaries, the trial continues.
The Three-Way Interim Decision
For a two-sided or efficacy-and-futility group sequential design, each interim analysis can produce three broad outcomes.
| Observed Statistic | Decision |
|---|---|
| \(Z_k\ge u_k\) | Stop for efficacy |
\(l_k| Continue |
|
| \(Z_k\le l_k\) | Stop for futility |
This simple structure is the foundation of many group sequential designs.
Why Early Efficacy Boundaries Are Usually Stringent
At an early interim analysis, relatively little information has accumulated. Therefore, an extreme result is required before the trial can conclude that the treatment is effective.
As more information accumulates, the efficacy boundary generally becomes less extreme.
Conceptually:
for many common efficacy-boundary constructions.
The final analysis therefore generally has the least stringent standardized boundary.
O'Brien-Fleming-Type Boundaries
One of the best-known group sequential approaches is the O'Brien-Fleming framework. A simplified representation of an O'Brien-Fleming-type boundary is:
where \(c\) is selected so that the overall type I error has the desired value.
Because \(t_k\) is small early in the study, the boundary can be very large at the first interim analysis. As \(t_k\) approaches 1, the boundary approaches its final value.
| Information Fraction | Qualitative Boundary |
|---|---|
| 20% | Very stringent |
| 50% | Moderately stringent |
| 75% | Closer to final boundary |
| 100% | Least stringent |
The exact critical values depend on the number and timing of analyses and the desired overall type I error.
Pocock-Type Boundaries
A contrasting approach is the Pocock family of boundaries. The conceptual feature is that the standardized critical value is much more similar across analyses.
The price is that the final boundary can be somewhat more stringent than the corresponding O'Brien-Fleming-type final boundary.
| Feature | O'Brien-Fleming-Type | Pocock-Type |
|---|---|---|
| Early boundary | Very stringent | Less stringent |
| Final boundary | Close to conventional threshold | More stringent |
| Alpha allocation | Heavily concentrated late | More evenly distributed |
| Early stopping for efficacy | More difficult | More accessible |
Neither approach is universally correct. The choice reflects how the investigators want the statistical evidence distributed across the information times.
Alpha Spending
An alternative way to construct group sequential boundaries is to specify how the overall type I error is spent as information accumulates. Let:
be the overall one-sided type I error. An alpha-spending function \(A(t)\) satisfies:
At information fraction \(t_k\), the cumulative alpha spent is:
The incremental alpha spent between analyses \(k-1\) and \(k\) is:
The corresponding efficacy boundary is then determined from the joint distribution of the sequential test statistics.
Why Alpha Spending Is Useful
A prespecified fixed-number group sequential design assumes that the interim analyses occur at specific information fractions. In practice, however, event-driven trials may not reach exactly those planned information times.
Alpha-spending approaches allow the boundary to be recalculated according to the information fraction actually available at the analysis.
O'Brien-Fleming-Type Alpha Spending
A commonly used Lan-DeMets approximation to an O'Brien-Fleming spending philosophy can be represented as:
for a two-sided setting, with the exact formulation adjusted according to the one-sided or two-sided testing framework.
The important conceptual feature is that very little alpha is spent early.
This produces an O'Brien-Fleming-like design even when the interim analysis does not occur at exactly the originally anticipated information fraction.
Pocock-Type Alpha Spending
A Pocock-like alpha-spending function can be constructed so that alpha is spent more evenly over information time. One commonly used formulation has the general form:
where \(\theta\) is selected to reproduce the desired spending behavior.
Again, the precise function should be specified consistently with the chosen one-sided or two-sided design and statistical software implementation.
Information Fraction Versus Calendar Time
Interim analyses are sometimes described as occurring at six months, 12 months, and 18 months. That description is operational rather than statistical.
For many confirmatory trials, the more relevant quantity is the amount of information accumulated.
For example, a study might plan analyses at:
If event accrual is faster than expected, the calendar dates can occur earlier. If event accrual is slower, they can occur later.
The statistical design is therefore tied to information rather than a fixed calendar date.
A Worked Group Sequential Example
Consider a randomized Phase III trial comparing an experimental treatment with control. The primary endpoint is time-to-event, and the final analysis is planned after 400 events.
Suppose the trial is designed with two interim analyses and one final analysis:
| Analysis | Target Events | Information Fraction |
|---|---|---|
| Interim 1 | 200 | 0.50 |
| Interim 2 | 300 | 0.75 |
| Final | 400 | 1.00 |
The primary efficacy hypothesis is tested using a one-sided type I error of:
The trial will use an O'Brien-Fleming-type efficacy boundary.
Step 1: Define the Test Statistic
Let \(Z_k\) denote the standardized test statistic at analysis \(k\), with larger values favoring the experimental treatment.
Under the null hypothesis, the statistic is approximately standard normal at each analysis:
The statistics are correlated across analyses because the data observed at an earlier analysis are contained within the later analysis.
Step 2: Specify the Efficacy Boundaries
Suppose the resulting efficacy boundaries are approximately:
| Analysis | Information Fraction | Efficacy Boundary |
|---|---|---|
| Interim 1 | 0.50 | Approximately 2.96 |
| Interim 2 | 0.75 | Approximately 2.45 |
| Final | 1.00 | Approximately 2.02 |
These values are illustrative of the shape of an O'Brien-Fleming-type design. The exact values in an actual trial must be calculated from the specified design and overall type I error.
The corresponding decisions are:
| Analysis | Decision Rule |
|---|---|
| Interim 1 | \(Z_1\ge2.96\) → stop for efficacy |
| Interim 2 | \(Z_2\ge2.45\) → stop for efficacy |
| Final | \(Z_3\ge2.02\) → declare efficacy |
Step 3: Interpret the First Interim Analysis
Suppose 200 events have occurred and the observed test statistic is:
Because:
the efficacy boundary is not crossed.
The trial therefore continues.
Step 4: The Second Interim Analysis
Suppose the trial reaches 300 events and:
Now:
The efficacy boundary is crossed. The trial can therefore stop early for efficacy, subject to the prespecified decision process and governance structure.
There is no need to wait for the final 400 events to make the formal efficacy decision.
Why the Early Boundary Is Higher
At 200 events, the study has only half of the planned information. A result of \(Z=2.20\) is encouraging, but the design requires a more extreme result at this early look.
By 300 events, the information fraction is larger and the required boundary is lower.
This is the fundamental logic of group sequential testing:
and:
Futility Boundaries
Efficacy stopping asks: Is there sufficiently strong evidence that the treatment works?
Futility stopping asks a different question: Is it unlikely that continuing the trial will produce a successful result?
A futility boundary can therefore be specified as:
Futility boundaries are conceptually different from efficacy boundaries because their principal purpose is not necessarily to preserve type I error.
Binding and Nonbinding Futility
A particularly important distinction is whether a futility boundary is binding or nonbinding.
Binding Futility
A binding futility rule means that if the boundary is crossed, the trial must stop according to the design. The futility rule is therefore part of the formal decision procedure.
Nonbinding Futility
A nonbinding futility rule means that crossing the boundary recommends stopping but does not necessarily invalidate the type I error control if the trial continues.
Futility Is Not the Same as Lack of Statistical Significance
A common mistake is to define futility simply as:
That is generally not an appropriate group sequential futility rule.
At an interim analysis, failure to demonstrate efficacy may simply mean that insufficient information has accumulated.
A futility rule should instead be linked to a prespecified criterion such as conditional power, predictive probability, or a suitable test-statistic boundary.
Conditional Power
Conditional power asks: Given what has been observed so far, what is the probability of eventually rejecting the null if the remaining data follow a specified assumption?
Let \(X_k\) denote the information accumulated at the interim analysis and let \(I_K\) denote the final information. The conditional power can be written generally as:
The calculation requires assumptions about the future treatment effect.
For example, investigators might calculate conditional power assuming that the originally planned alternative effect remains true.
A Simplified Conditional Power Concept
Suppose a trial has completed half of its planned information. The observed treatment effect is smaller than anticipated. The conditional power under the originally assumed alternative may therefore be substantially lower than the original planned power.
A prespecified futility rule might state that:
The exact threshold is a design choice and should be established before the interim analysis.
Predictive Probability
A Bayesian alternative to conditional power is the predictive probability of success. Rather than conditioning on one fixed future treatment effect, predictive probability averages over uncertainty about the underlying parameter using the posterior distribution.
Conceptually:
This can be useful when investigators want to incorporate uncertainty about the treatment effect rather than conditioning on one assumed value.
Conditional Power vs. Predictive Probability
| Feature | Conditional Power | Predictive Probability |
|---|---|---|
| Framework | Usually frequentist | Bayesian |
| Future effect | Specified assumption | Averaged over posterior uncertainty |
| Output | Probability of eventual success | Predictive probability of success |
| Dependence on assumptions | Strong | Depends on prior and posterior model |
These quantities should not be treated as interchangeable simply because both produce a number between 0 and 1.
Alpha Spending and Futility Are Different Concepts
Alpha spending controls the probability of a false-positive efficacy conclusion.
Futility stopping controls whether the trial continues when success appears unlikely.
They answer different questions.
| Concept | Primary Purpose |
|---|---|
| Alpha spending | Control overall type I error |
| Efficacy boundary | Permit early declaration of benefit |
| Futility boundary | Permit early stopping when success is unlikely |
| Conditional power | Quantify chance of eventual success under an assumption |
| Predictive probability | Quantify predictive chance of future success |
Multiplicity Across Multiple Interim Analyses
If a trial has several interim analyses, the individual tests cannot generally be treated as independent.
The same accumulating data contribute to multiple analyses.
The joint distribution of the test statistics therefore matters. For many group sequential designs, the canonical joint distribution has correlations approximately determined by the information fractions.
For \(t_i
This correlation structure is what allows the boundaries to be calibrated
while accounting for repeated looks at the accumulating data.
Under standard group sequential assumptions, the sequential statistics can be
represented approximately as correlated normal random variables.
For analyses at information fractions \(t_1,\ldots,t_K\):
where the covariance structure is determined by the information times.
This is why simply applying a standard normal critical value independently at
each interim analysis is generally inappropriate.
Some confirmatory trials require a two-sided hypothesis:
In that case there may be both an upper and lower efficacy boundary:
and:
The overall two-sided type I error must be allocated appropriately across the
two tails.
The choice must be consistent with the scientific question, regulatory
strategy, estimand, and statistical analysis plan.
Safety monitoring is often performed alongside efficacy monitoring but uses
different criteria.
For example, a study might specify stopping rules for:
Safety monitoring may involve different statistical methods than efficacy
monitoring.
In many confirmatory trials, interim unblinded efficacy and safety results are
reviewed by an independent data monitoring committee (DMC).
The DMC may receive unblinded information while the sponsor's broader study
team remains blinded to treatment-comparative results.
This structure can help preserve trial integrity while allowing the prespecified
stopping rules to be evaluated.
Not every interim analysis requires unblinded comparative treatment results.
For example, a study may conduct a blinded sample-size reassessment based on
the overall event rate or variance without examining the treatment effect.
An unblinded efficacy interim analysis, by contrast, examines comparative
treatment information.
A trial can reassess sample size without examining comparative treatment
effects.
For example, investigators may discover that the pooled event rate is lower
than anticipated.
A blinded reassessment may then determine that more events are required to
maintain the desired power.
This is conceptually different from looking at the treatment effect and
deciding whether the study should stop for success or futility.
A particularly important principle is that interim analyses should be
prespecified whenever possible.
If investigators unexpectedly examine accumulating treatment-effect data and
then modify the trial based on what they see, the original operating
characteristics may no longer apply.
Therefore, the protocol and statistical analysis plan should specify:
Real trials rarely behave exactly as planned.
Suppose the intended first interim analysis was at:
but the trial reaches the analysis after only:
or:
An alpha-spending design can calculate the boundary based on the information
fraction actually available.
This flexibility is one of the major practical advantages of alpha-spending
approaches.
For time-to-event trials, it is often preferable to define interim analyses
using event counts rather than calendar dates.
For example:
This approach aligns the statistical timing of the analyses with the amount of
information available.
Suppose a Phase III study uses three analyses at information fractions:
Suppose the efficacy boundaries are:
Now suppose a nonbinding futility rule is based on conditional power:
At the first interim analysis, suppose the observed statistic is:
The efficacy boundary is clearly not crossed.
Suppose the conditional power under the original design alternative is
calculated to be 12%.
The trial would therefore meet the prespecified futility criterion.
The DMC could recommend stopping for futility according to the protocol and
the governance framework.
Consider the same \(Z_1=0.10\), but imagine that the first analysis occurs
very early, at only 20% information.
There may still be substantial information yet to come.
A low interim statistic alone therefore does not tell us whether the trial is
futile.
The appropriate question is:
which is why conditional power, predictive probability, or a prespecified
futility boundary can be useful.
Some designs specify futility using an estimated treatment effect.
For example, investigators might consider whether the observed effect remains
compatible with a clinically meaningful benefit.
However, an effect estimate alone should not be interpreted without considering
its uncertainty and the amount of information accumulated.
An observed hazard ratio of 0.90 after very few events means something
different from the same hazard ratio after hundreds of events.
Confidence intervals can be useful for describing the accumulating treatment
effect.
Suppose the estimated hazard ratio is:
with a confidence interval:
The interval describes the current statistical uncertainty.
However, the confidence interval should not automatically be used as the
stopping rule unless the group sequential design was constructed around that
criterion.
A common misunderstanding is that investigators can simply perform ordinary
tests at each interim analysis and "adjust the final p-value" after observing
the results.
A proper group sequential design instead defines the sequential testing
procedure in advance.
The critical boundaries and final inference are mathematically linked to the
entire sequence of analyses.
The final reported p-value or confidence interval should therefore be
consistent with the sequential design.
Ordinary 95% confidence intervals are not automatically valid as simultaneous
confidence intervals across repeated interim looks.
If repeated interim estimates are used for formal inference, the interval
procedure should account for the group sequential design.
This distinction is important when interpreting a treatment effect after
early stopping.
Stopping early for efficacy is statistically efficient for decision-making,
but it can introduce bias into naive treatment-effect estimates.
If a trial stops because the observed treatment effect is unusually favorable,
the observed estimate at stopping may tend to overestimate the true effect.
This phenomenon is sometimes described as
estimation bias following early stopping.
Suppose a trial stops only when:
This means that the observed treatment effect has already passed an unusually
high threshold.
Conditional on crossing that threshold, the observed effect is therefore
selected from a favorable portion of its sampling distribution.
The treatment effect observed at stopping is consequently not a random draw
from the same distribution as it would have been under fixed sample size.
Some adaptive designs use interim information to modify the planned sample
size.
For example, the sample size may be increased if the nuisance variance or
event rate differs from the planning assumption.
This should not be confused with simply increasing the sample size because the
observed treatment effect is disappointing or promising.
Effect-driven sample-size adaptation requires careful statistical planning to
preserve type I error and avoid introducing operational bias.
An interim analysis may also trigger decisions about enrollment or treatment
allocation.
For example, a study could stop enrollment to one arm after an efficacy or
safety boundary is crossed.
However, such adaptations create additional statistical and operational
complexity.
Any treatment-allocation adaptation should therefore be included in the
prespecified design framework.
The same principles apply to single-arm trials, although the statistical
model may differ.
Suppose the primary endpoint is response rate.
The investigators might evaluate response after a prespecified number of
patients and stop if the response count is either sufficiently high or
insufficiently promising.
For a binary endpoint:
The stopping boundaries are then constructed using the relevant exact or
asymptotic distribution.
For survival endpoints, the number of events is often more informative than
the number of enrolled patients.
A log-rank statistic or another standardized statistic can be used for
group sequential monitoring.
The test statistic may be expressed approximately as:
where \(\widehat{\theta}\) represents an estimated treatment effect on the
appropriate scale.
For example, in a proportional hazards framework:
so that:
The sign convention should be defined so that larger values consistently
favor the prespecified beneficial direction.
For continuous endpoints, the test statistic may be based on a treatment
difference:
where \(\widehat{\Delta}\) is the estimated treatment difference.
The same group sequential principles apply:
A high-quality interim analysis plan should be detailed enough that another
statistician could reproduce the stopping boundaries without needing to know
the interim results.
At minimum, document:
Access to unblinded interim results should be restricted according to the
trial's governance plan.
In many confirmatory studies, an independent DMC reviews the unblinded
comparative data while the sponsor's operational team remains blinded to
treatment-effect results.
This separation helps reduce the risk that knowledge of interim results
influences trial conduct.
Interim analyses create a potential source of operational bias if investigators
learn how the treatment is performing.
For example, knowledge that the treatment appears highly effective could
influence:
This is one reason why unblinded interim results are often restricted to an
independent monitoring group.
Stopping for efficacy does not mean that every statistical question disappears.
The final report may still need to address:
The protocol and SAP should specify how these analyses will be handled.
If a trial stops for futility, the study may still need to complete predefined
follow-up for safety and other clinical outcomes.
Stopping efficacy enrollment does not necessarily mean that every participant
immediately stops treatment.
The operational consequences of stopping should therefore be explicitly
defined.
A simple representation of information fractions in R is:
which gives:
For an event-driven design, the actual information fraction at each analysis
can be calculated from the observed information rather than assuming that
calendar time or enrollment provides an adequate approximation.
A simple illustrative O'Brien-Fleming-type calculation uses the general
relationship:
This demonstrates the characteristic shape of the boundary.
The exact critical values used in a formal group sequential design should be
calculated using a procedure that accounts for the joint distribution of the
sequential statistics and the desired overall type I error.
An alpha-spending function can be represented conceptually as a function of
information time.
For example:
This illustrates how cumulative alpha spending changes with information.
For a production analysis, use a validated group sequential implementation
rather than treating this simplified calculation as the final boundary
derivation.
A generic conditional-power workflow might look like:
The important point is not the code itself but the definition of the
assumption used for future data.
Formal group sequential designs are commonly implemented using specialized
statistical software or validated statistical packages.
For example, R users may work with packages designed specifically for group
sequential boundaries and alpha spending.
The software should be used to obtain:
An interim analysis does not automatically make a trial an adaptive design.
A conventional group sequential trial may have fully prespecified interim
rules and no flexibility to alter other aspects of the trial.
An adaptive design generally permits one or more prospectively defined
modifications based on accumulating data.
The interim analysis should be aligned with the same clinical question as the
final analysis.
If the final primary estimand concerns a specific treatment effect in a
defined population, the interim analysis should not casually switch to a
different population or endpoint because the interim data are inconvenient.
This is particularly important when handling:
At an interim analysis, some patients may not yet have complete endpoint
information.
This is particularly common when the endpoint requires long follow-up.
The analysis plan should therefore specify:
The nominal information fraction should not be confused with data maturity.
For example, a trial may have enrolled enough patients to reach the planned
sample size while a substantial proportion of patients have not yet reached
the primary endpoint assessment.
The interim analysis should therefore be based on the prespecified definition
of statistical information and endpoint maturity.
There is no universal optimal number.
More interim analyses provide more opportunities to stop early but can also
complicate trial operations and increase the burden of statistical monitoring.
Common confirmatory designs may have one or two interim efficacy analyses
before the final analysis.
The choice should consider:
A trial could theoretically be monitored very frequently.
But frequent looks may provide little practical advantage if the treatment
effect would not be sufficiently identifiable at very early information times.
The monitoring schedule should therefore be driven by the decision problem,
not by a desire to maximize the number of statistical looks.
A protocol might state, in simplified form:
The SAP should provide enough detail to reproduce the complete interim
monitoring procedure.
At minimum, include:
The central relationship in group sequential testing can be summarized as:
At early information times, the efficacy boundary is generally more stringent.
As information accumulates, the boundary becomes less stringent.
The final analysis therefore completes the sequence rather than representing an
independent test performed after several unrelated interim tests.
Imagine the trial moving along an information scale:
At each point, the accumulating evidence is compared with a prespecified
boundary.
The trial can leave the sequence early through an efficacy or futility
decision.
If it does not cross either boundary, it continues toward the final analysis.
Several statements sound reasonable but are statistically incorrect.
Not if the look is being used for formal efficacy inference without an
appropriate sequential procedure.
Not necessarily.
The critical value at an interim analysis may be substantially more stringent
than the conventional 0.05 threshold.
Not necessarily.
The study may simply need more information.
No.
O'Brien-Fleming-type procedures generally spend very little alpha early and
more as the trial approaches its final analysis.
Not necessarily.
Equal division is only one possible spending philosophy and is not the
defining feature of group sequential testing.
For confirmatory clinical trials, interim analyses should be integrated into
the overall statistical design rather than treated as an informal monitoring
exercise.
The protocol, SAP, DMC charter, and relevant statistical documentation should
be consistent about:
The exact regulatory expectations depend on the development program and trial
context, so the statistical design should be reviewed with the appropriate
regulatory and clinical teams.
The numerical boundaries above illustrate the shape of the design rather than
serving as a universal set of critical values. In an actual trial, the
boundaries must be generated from the exact planned design, testing direction,
number of analyses, information fractions, and overall type I error.
There are several concepts that should remain clear when planning an interim
analysis.
O'Brien, P.C. & Fleming, T.R. (1979).
A multiple testing procedure for clinical trials.
Biometrics, 35(3), 549–556.The Canonical Joint Distribution
A Two-Sided Trial
One-Sided vs. Two-Sided Testing
Design
Typical Efficacy Decision
One-sided
Only the prespecified beneficial direction triggers success
Two-sided
Either sufficiently positive or sufficiently negative result can cross an efficacy boundary
Safety Stopping Rules
Independent Data Monitoring Committees
Blinded vs. Unblinded Interim Analyses
Analysis Type
Typical Purpose
Blinded
Assess nuisance parameters or planning assumptions
Unblinded
Formal efficacy, futility, or safety decision
Sample Size Reassessment Is Not Automatically an Interim Efficacy Analysis
Unplanned Interim Analyses
What Happens If the Interim Analysis Occurs at the Wrong Information Fraction?
Event-Driven Interim Analyses
Worked Example: Efficacy and Futility Together
Analysis
Efficacy Boundary
50%
\(Z\ge2.96\)
75%
\(Z\ge2.45\)
100%
\(Z\ge2.02\)
Why a Low Interim Z-Statistic Does Not Automatically Mean Futility
Futility Can Be Based on Effect Estimates
Confidence Intervals at Interim Analyses
Alpha Spending vs. Adjusting the P-Value Afterward
Repeated Confidence Intervals
Early Stopping Changes Estimation
Why the Naive Point Estimate Can Be Optimistic
Sample Size Re-estimation and Interim Analyses
Interim Analysis and Treatment Allocation
Interim Analysis in a Single-Arm Study
Interim Analysis for Time-to-Event Endpoints
Interim Analysis for Continuous Endpoints
Interim Analysis Planning Checklist
What Should Be Prespecified?
Who Should See the Interim Results?
Operational Bias
What Happens After Early Efficacy Stopping?
What Happens After Futility Stopping?
Common Mistakes
R Implementation: Information Fractions
events <- c(200, 300, 400)
information_fraction <-
events / max(events)
information_fraction
# 0.50 0.75 1.00
R Implementation: O'Brien-Fleming-Type Boundaries
t <- c(0.50, 0.75, 1.00)
z_final <- qnorm(1 - 0.025)
z_approx <- z_final / sqrt(t)
z_approx
R Implementation: Alpha-Spending Concept
alpha <- 0.025
t <- c(0.50, 0.75, 1.00)
z_alpha <- qnorm(1 - alpha)
of_spend <- function(t) {
1 - pnorm(z_alpha / sqrt(t))
}
of_spend(t)
R Implementation: Conditional Power Concept
conditional_power <- function(
z_current,
information_current,
information_final,
critical_value,
assumed_effect
) {
# Illustrative framework.
# Production calculations should use
# the exact design-specific formula.
remaining_information <-
information_final - information_current
# Future information and the assumed
# treatment effect determine the
# distribution of the final statistic.
# Calculate P(final statistic exceeds
# the prespecified critical value).
# Return the resulting probability.
}
Design Software
Interim Analysis vs. Adaptive Design
Feature
Group Sequential Trial
Adaptive Trial
Interim analysis
Yes
Usually
Early stopping
Common
Common
Design modification
Usually limited
May be permitted
Prespecification
Extensive
Extensive
Statistical complexity
Moderate
Potentially high
Interim Analyses and Estimands
Interim Analyses and Missing Data
Interim Analysis Maturity
How Many Interim Analyses Should a Trial Have?
More Interim Analyses Are Not Automatically Better
Interim Analysis Planning Workflow
Protocol Example
What Should Be Reported in the Statistical Analysis Plan?
Interim Analysis Decision Table
Finding
Possible Decision
Statistical Framework
Strong evidence of benefit
Stop for efficacy
Efficacy boundary
Low likelihood of eventual success
Stop for futility
Futility boundary or conditional power
Unacceptable safety
Stop or modify trial
Safety monitoring rule
Evidence remains intermediate
Continue
Prespecified continuation region
The Relationship Between Information and Boundary
A Mental Model for Group Sequential Testing
Key Differences Among Common Stopping Approaches
Approach
Main Purpose
Characteristic
O'Brien-Fleming
Efficacy monitoring
Very stringent early boundaries
Pocock
Efficacy monitoring
More similar boundaries across looks
Alpha spending
Flexible sequential testing
Spends cumulative alpha according to information
Conditional power
Futility assessment
Probability of eventual success under an assumption
Predictive probability
Futility assessment
Predicts success while incorporating parameter uncertainty
Common Misinterpretations
"We can look at the data whenever we want."
"If the interim p-value is below 0.05, we can stop."
"If the interim p-value is above 0.05, the trial is futile."
"O'Brien-Fleming spends the same alpha at every analysis."
"Alpha spending means we divide 0.025 equally among three analyses."
Interim Analysis and Regulatory Interpretation
Summary of the Worked Example
Component
Illustrative Value
Study type
Randomized Phase III
Primary endpoint
Time-to-event
Final information
400 events
Interim 1
200 events / 50%
Interim 2
300 events / 75%
Final
400 events / 100%
Overall one-sided alpha
0.025
Efficacy framework
O'Brien-Fleming-type
Illustrative boundary at 50%
Approximately 2.96
Illustrative boundary at 75%
Approximately 2.45
Illustrative final boundary
Approximately 2.02
Futility
Conditional power example
The Most Important Concepts
References
Pocock, S.J. (1977).
Group sequential methods in the design and analysis of clinical
trials.
Biometrika, 64(2), 191–199.
Lan, K.K.K. & DeMets, D.L. (1983).
Discrete sequential boundaries for clinical trials.
Biometrika, 70(3), 659–663.
Jennison, C. & Turnbull, B.W. (2000).
Group Sequential Methods with Applications to Clinical Trials.
Chapman & Hall/CRC.
Wassmer, G. & Brannath, W. (2016).
Group Sequential Designs in Clinical Trials.
Springer.
Proschan, M.A., Lan, K.K.K. & Wittes, J.T. (2006).
Statistical Monitoring of Clinical Trials: A Unified Approach.
Springer.
Whitehead, J. (1997).
The Design and Analysis of Sequential Clinical Trials.
Wiley.