← Clinical Trials
Triple-Negative Breast Cancer Phase 3 Randomized NCT02425891

IMpassion130: Complete Statistical Analysis of Atezolizumab in Triple-Negative Breast Cancer

An independent statistical review of the randomized phase 3 IMpassion130 trial comparing atezolizumab plus nab-paclitaxel with placebo plus nab-paclitaxel in participants with previously untreated metastatic triple-negative breast cancer.

Phase 3  ·  Completed  ·  Enrollment 902  ·  Primary completion April 14, 2020
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record for IMpassion130 and the posted statistical analyses summarized below.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

IMpassion130 was a randomized, double-blind, parallel-group phase 3 trial evaluating atezolizumab in combination with nab-paclitaxel versus placebo with nab-paclitaxel in participants with previously untreated metastatic triple-negative breast cancer.

902
Enrolled
2 treatment arms
0.80
All-participant PFS HR
95% CI 0.69–0.92
0.62
PD-L1 PFS HR
95% CI 0.49–0.78
0.67
PD-L1 OS HR
95% CI 0.53–0.86
FeatureIMpassion130
PhasePhase 3
ConditionTriple Negative Breast Cancer
DesignRandomized, double-blind, parallel
AllocationRandomized
Enrollment902
Primary purposeTreatment
Primary endpointsFour registered primary endpoints: PFS in all randomized participants; PFS in participants with detectable PD-L1; OS in all randomized participants; OS in participants with detectable PD-L1
Results postedYes
Outcome measures posted15
Statistical analyses posted10
ClinicalTrials.govNCT02425891
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry
Trial datesStart: June 23, 2015  ·  Primary completion: April 14, 2020

2. Clinical Question

The central statistical question was whether adding atezolizumab to nab-paclitaxel produced a superior time-to-event outcome compared with placebo plus nab-paclitaxel in participants with previously untreated metastatic triple-negative breast cancer.

Population

Participants with previously untreated metastatic triple-negative breast cancer enrolled in the randomized phase 3 study.

Intervention

Atezolizumab (MPDL3280A), an engineered anti-PDL1 antibody, in combination with nab-paclitaxel.

Comparator

Placebo in combination with nab-paclitaxel.

Primary question

Does the atezolizumab plus nab-paclitaxel regimen improve PFS and OS relative to placebo plus nab-paclitaxel?

3. Trial Design

01
Randomize902 participants enrolled
02
Two armsAtezolizumab or placebo plus nab-paclitaxel
03
Double-blindRandomized treatment comparison
04
AssessPFS, OS, response, DOR and PRO endpoints
05
AnalyzeLog-rank and Cochran-Mantel-Haenszel methods
PLACEBO ARM · 430 AT RISK FOR SERIOUS AE ANALYSIS

Placebo + Nab-Paclitaxel

  • Placebo (q2w)
  • Nab-paclitaxel
  • Double-blind randomized comparator regimen
ATEZOLIZUMAB ARM · 460 AT RISK FOR SERIOUS AE ANALYSIS

Atezolizumab + Nab-Paclitaxel

  • Atezolizumab (MPDL3280A), an engineered anti-PDL1 antibody
  • Atezolizumab (q2w)
  • Nab-paclitaxel
  • Double-blind randomized treatment regimen
Allocation
Randomized allocation to two parallel treatment groups.
Masking
Double-blind.
Primary purpose
Treatment.
Hypothesis type
Superiority.

4. Endpoints

The registry lists four primary endpoints and a mixture of time-to-event and binary endpoint types. The primary efficacy analyses reported in the ClinicalTrials.gov record uses log-rank testing and hazard ratios for the four primary endpoints.

Primary endpointRegistered definitionTime frameType
Progression Free Survival (PFS) According to RECIST Version 1.1 (v1.1) in All Randomized Participants PFS was defined as the time from randomization to the occurrence of disease progression, as determined by investigators from tumor assessments per RECIST v1.1, or death from any cause, whichever occurred first. Baseline up to approximately 34 months Time-to-event
PFS According to RECIST v1.1 in Participants With Detectable PD-L1 PFS was defined as the time from randomization to the occurrence of disease progression, as determined by investigators from tumor assessments per RECIST v1.1, or death from any cause, whichever occurred first. Baseline up to approximately 34 months Time-to-event
Overall Survival (OS) in All Randomized Participants OS was defined as the time from the date of randomization to the date of death from any cause. Baseline until death due to any cause (up to approximately 58 months) Time-to-event
OS in Participants With Detectable PD-L1 OS was defined as the time from the date of randomization to the date of death from any cause. Baseline until death due to any cause (up to approximately 58 months) Time-to-event

Secondary endpoints with posted formal analyses

EndpointTime frameAnalysis methodEffect measure
Objective response rate in all randomized participants Baseline up to approximately 34 months Cochran-Mantel-Haenszel test Difference in overall response rates
Objective response rate in participants with detectable PD-L1 Baseline up to approximately 34 months Cochran-Mantel-Haenszel test Difference in overall response rates
Duration of response in all randomized participants Baseline up to approximately 34 months Log-rank test Hazard ratio
Duration of response in participants with detectable PD-L1 Baseline up to approximately 34 months Log-rank test Hazard ratio
Time to deterioration in global health status/health-related quality of life in all randomized participants Baseline up to approximately 58 months Log-rank test Hazard ratio
Time to deterioration in global health status/health-related quality of life in participants with detectable PD-L1 Baseline up to approximately 58 months Log-rank test Hazard ratio

5. Statistical Methodology

Intention-to-treat analysis

For the primary all-randomized analyses, the registry defines the ITT population as all randomized patients, whether or not the assigned study treatment was received. This preserves the randomized treatment comparison: once randomization occurs, the efficacy analysis remains tied to the assigned group rather than being redefined according to subsequent treatment exposure.

ITT principle
Analyze randomized participants according to assigned treatment

The principal statistical advantage is that the treatment groups retain the allocation created by randomization. This is especially important when treatment discontinuation, protocol deviations, or other post-randomization events occur.

Log-rank test

The registry reports the log-rank test for PFS, OS, DOR, and time to deterioration. The log-rank test is designed for comparing time-to-event distributions between randomized groups while incorporating the timing of events and the presence of right-censored observations.

Conceptual comparison
H0: the event-time distributions are equivalent between groups

The resulting P-value addresses evidence against the null comparison under the specified test. It does not itself quantify how large or clinically important the treatment difference is.

Hazard ratio

The principal effect measure for the time-to-event analyses was the hazard ratio. A hazard ratio compares the estimated instantaneous event rate between treatment groups within the analysis framework.

Interpretation
HR < 1  →  lower estimated instantaneous event rate in the atezolizumab group

For example, an HR of 0.80 corresponds to an estimated instantaneous event rate that is 80% of the comparator group's rate under the fitted analysis framework. Equivalently, 0.80 corresponds to a 20% lower estimated hazard; it does not mean 20% of patients avoided the event.

Stratified analysis

The registry identifies the primary time-to-event analyses as stratified analyses. The ClinicalTrials.gov record does not specify the individual stratification variables, so they are not reproduced here. The important statistical point is that the treatment comparison was not described as an unstratified log-rank analysis.

Cochran-Mantel-Haenszel test

The objective-response analyses used the Cochran-Mantel-Haenszel test. This is a categorical-data method that can compare treatment groups while accounting for stratification. The associated effect measure reported in the registry is the difference in overall response rates.

Superiority framework

All ten registry-reported formal analyses are labeled with a superiority hypothesis type. Thus, the reported hazard ratios and response-rate differences are being interpreted in a framework intended to test whether the treatment groups differ in the favorable direction specified by the trial's hypotheses. The ClinicalTrials.gov record does not provide a non-inferiority margin or equivalence margin.

6. Results: Progression-Free Survival in All Randomized Participants

The first primary endpoint was PFS according to RECIST v1.1 in all randomized participants, assessed from baseline up to approximately 34 months. The analysis population was the ITT population, defined as all randomized patients whether or not the assigned study treatment was received.

Hazard ratio for progression or death

0.80

95% CI: 0.69–0.92   ·   P = 0.0025

Log-rank test  ·  Stratified analysis  ·  Superiority hypothesis

Primary PFS analysisValue
EndpointPFS according to RECIST v1.1 in all randomized participants
Analysis populationITT
ComparisonPlacebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel
MethodLog-rank test
Effect measureHazard ratio
Estimate0.80
95% CI0.69–0.92
P-value0.0025
Clinical Biostats interpretation

An HR of 0.80 means that the estimated instantaneous rate of progression or death in the atezolizumab-plus-nab-paclitaxel group was approximately 80% of that in the placebo-plus-nab-paclitaxel group under the reported time-to-event analysis. Put another way, 0.80 corresponds to a 20% lower estimated hazard.

The HR does not mean that 20% of participants avoided progression or death, that every participant experienced exactly a 20% reduction, or that the median PFS differed by a particular number of months. Those would require different reported quantities.

The 95% CI of 0.69–0.92 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient outcomes. The interval also gives more information about precision than the point estimate alone.

The P-value of 0.0025 measures the evidence against the relevant null hypothesis under the specified test; it is not a measure of effect size. A small P-value can coexist with a modest effect, while a larger effect can be estimated imprecisely in a small dataset.

Because this is a hazard ratio, interpretation also depends on the behavior of the hazards over time and on censoring. The ClinicalTrials.gov record identifies a stratified log-rank analysis but do not provide enough information here to independently assess the proportional-hazards assumption.

7. Results: Progression-Free Survival in Participants With Detectable PD-L1

The second primary endpoint evaluated PFS according to RECIST v1.1 in participants with detectable PD-L1. The registry defines this PD-L1-selected subpopulation as patients in the ITT population whose PD-L1 status was IC1/2/3 at the time of randomization.

Hazard ratio for progression or death

0.62

95% CI: 0.49–0.78   ·   P < 0.0001

Log-rank test  ·  Stratified analysis  ·  Superiority hypothesis

Primary PFS analysisValue
EndpointPFS according to RECIST v1.1 in participants with detectable PD-L1
PD-L1-selected populationPatients in the ITT population whose PD-L1 status was IC1/2/3 at randomization
ComparisonPlacebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel
MethodLog-rank test
Effect measureHazard ratio
Estimate0.62
95% CI0.49–0.78
P-value<0.0001
Clinical Biostats interpretation

An HR of 0.62 corresponds to an estimated instantaneous rate of progression or death approximately 62% of the comparator rate within this PD-L1-selected analysis. Expressed as a relative hazard interpretation, this is approximately a 38% lower estimated hazard.

The estimate applies to the defined PD-L1-selected population. It should not automatically be generalized to participants outside that population, and it does not establish that the treatment effect differs from the all-randomized estimate simply because the numerical HRs are different.

The 95% CI of 0.49–0.78 communicates the statistical uncertainty around 0.62. It is narrower than a statement such as "most patients had a 38% reduction"; individual patients do not have fixed hazard ratios assigned to them by the trial.

The reported P < 0.0001 indicates strong statistical evidence against the corresponding null hypothesis under the reported test. It does not mean the probability that the treatment has an effect is less than 0.0001, and it does not measure the magnitude of clinical benefit.

Because this is a prespecified primary endpoint within a superiority framework, its interpretation must be understood in the context of the trial's overall multiplicity and testing strategy. The ClinicalTrials.gov record identifies four primary endpoints but do not provide a complete alpha-allocation scheme, so no additional multiplicity adjustment is inferred here.

8. Results: Overall Survival in All Randomized Participants

The third primary endpoint was OS in all randomized participants. OS was defined as the time from randomization to death from any cause, with the registered time frame extending from baseline until death due to any cause, up to approximately 58 months.

Hazard ratio for death

0.87

95% CI: 0.75–1.02   ·   P = 0.0770

Log-rank test  ·  Stratified analysis  ·  Superiority hypothesis

Primary OS analysisValue
EndpointOverall survival in all randomized participants
DefinitionTime from date of randomization to date of death from any cause
Analysis populationITT
ComparisonPlacebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel
MethodLog-rank test
Effect measureHazard ratio
Estimate0.87
95% CI0.75–1.02
P-value0.0770
Clinical Biostats interpretation

An HR of 0.87 corresponds to an estimated instantaneous rate of death approximately 87% of the comparator rate under the reported analysis, or approximately a 13% lower estimated hazard.

The estimate is not an absolute survival difference. It does not mean that 13% of patients survived because of treatment, nor does it specify how long individual patients lived. Median OS, survival probabilities at particular time points, and absolute risk differences are separate quantities and are not reported in the ClinicalTrials.gov record.

The 95% CI of 0.75–1.02 spans 1.00. That means the reported interval includes the null hazard-ratio value. The P-value of 0.0770 likewise does not provide conventional evidence against the null at a 0.05 threshold if that threshold were used. Importantly, the P-value itself is not an effect-size measure.

The correct statistical reading is therefore more nuanced than simply describing the HR as "good" or "bad": the point estimate is below 1, but the confidence interval reflects uncertainty that includes a hazard ratio of 1. The conclusion also depends on the trial's prespecified hypothesis-testing framework and multiplicity handling, which are not fully specified in the ClinicalTrials.gov record.

As with the PFS analyses, a Cox-type hazard-ratio interpretation should not be extended beyond what the registry analysis supports. Censoring and the behavior of hazards over time remain important considerations for any time-to-event analysis.

9. Results: Overall Survival in Participants With Detectable PD-L1

The fourth primary endpoint evaluated OS in the PD-L1-selected subpopulation. This population consisted of patients in the ITT population whose PD-L1 status was IC1/2/3 at the time of randomization.

Hazard ratio for death

0.67

95% CI: 0.53–0.86   ·   P = 0.0016

Log-rank test  ·  Stratified analysis  ·  Superiority hypothesis

Primary OS analysisValue
EndpointOS in participants with detectable PD-L1
DefinitionTime from date of randomization to date of death from any cause
PD-L1-selected populationPatients in the ITT population whose PD-L1 status was IC1/2/3 at randomization
ComparisonPlacebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel
MethodLog-rank test
Effect measureHazard ratio
Estimate0.67
95% CI0.53–0.86
P-value0.0016
Clinical Biostats interpretation

An HR of 0.67 means that the estimated instantaneous rate of death was approximately 67% of the comparator rate in the PD-L1-selected analysis, corresponding to approximately a 33% lower estimated hazard.

The estimate is specific to the PD-L1-selected population defined by the registry. It does not imply that the same hazard ratio applies to every participant in the overall ITT population.

The 95% CI of 0.53–0.86 indicates uncertainty around the point estimate while remaining below 1.00. The P-value of 0.0016 indicates statistical evidence against the relevant null hypothesis under the reported analysis, but it does not quantify the magnitude of the treatment effect.

It is also important not to compare P-values as though they measure the strength of treatment benefit across endpoints. The PFS and OS analyses answer different clinical questions, and the PD-L1-selected and all-randomized populations are different analysis populations.

Finally, the ClinicalTrials.gov record does not provide enough detail to reconstruct the full multiplicity hierarchy among the four primary endpoints. The four results should therefore be interpreted within the prespecified statistical design rather than treated as four unrelated hypothesis tests.

10. Secondary Results: Objective Response Rate

Objective response was defined as the percentage of participants with an objective response of complete response (CR) or partial response (PR) according to RECIST v1.1. The registry reports formal Cochran-Mantel-Haenszel analyses in both the all-randomized response-evaluable population and the PD-L1-selected response-evaluable population.

Objective Response in All Randomized Participants

Difference in overall response rates

10.12

95% CI: 3.40–16.84   ·   P = 0.0021

Cochran-Mantel-Haenszel test  ·  Stratified analysis

Clinical Biostats interpretation

The reported effect measure is a difference in overall response rates, not a hazard ratio. A value of 10.12 represents a 10.12-unit difference in the response-rate scale used by the registry analysis, comparing placebo plus nab-paclitaxel with atezolizumab plus nab-paclitaxel as listed in the analysis. It should not be converted into a hazard ratio or interpreted as a 10.12-fold effect.

The 95% CI of 3.40–16.84 quantifies uncertainty around the estimated response-rate difference. The P-value of 0.0021 addresses the hypothesis test and does not describe the size of the response difference.

The response analysis uses an evaluable population defined as patients in the ITT population with measurable disease at baseline. That differs conceptually from simply assuming that every randomized participant contributes a binary response observation.

Objective Response in Participants With Detectable PD-L1

Difference in overall response rates

16.30

95% CI: 5.67–26.92   ·   P = 0.0016

Cochran-Mantel-Haenszel test  ·  Stratified analysis

Clinical Biostats interpretation

The reported difference in overall response rates is 16.30 in the PD-L1-selected response-evaluable analysis. Its 95% CI is 5.67–26.92, giving the statistical uncertainty around that estimate.

The P-value of 0.0016 indicates evidence against the corresponding null hypothesis under the Cochran-Mantel-Haenszel analysis. It does not mean there is a 0.16% probability that the observed difference occurred by chance, because a frequentist P-value is not the posterior probability of a hypothesis.

The difference in response rates also answers a different question from PFS or OS. Response is a categorical tumor-assessment outcome; PFS and OS incorporate the timing of events and censoring. A complete statistical interpretation should therefore retain these endpoints as complementary rather than interchangeable measures.

11. Secondary Results: Duration of Response

Duration of response (DOR) was analyzed among participants with an objective response. The registry reports log-rank tests and hazard ratios for both the all-participant and detectable-PD-L1 analyses.

AnalysisHR95% CIP-valueMethod
DOR in all randomized participants 0.78 0.63–0.98 0.0285 Log-rank; unstratified analysis
DOR in participants with detectable PD-L1 0.60 0.43–0.86 0.0047 Log-rank; unstratified analysis

All randomized participants

An HR of 0.78 corresponds to an estimated 22% lower instantaneous rate of loss of response under the reported analysis. The 95% CI is 0.63–0.98 and the P-value is 0.0285.

Detectable PD-L1

An HR of 0.60 corresponds to an estimated 40% lower instantaneous rate of loss of response under the reported analysis. The 95% CI is 0.43–0.86 and the P-value is 0.0047.

DOR is inherently conditional on achieving an objective response. That makes its analysis population different from the ITT population used for the primary all-randomized PFS and OS analyses. A DOR hazard ratio therefore should not be interpreted as though it describes the event experience of every randomized participant.

Population distinction: the registry defines the DOR-evaluable population as patients with an objective response. This conditioning is important statistically: DOR asks how long responses persist among responders, whereas PFS asks about progression or death from randomization onward.

12. Secondary Results: Time to Deterioration in Global Health Status / Quality of Life

The registry also reports time to deterioration (TTD) in global health status/health-related quality of life according to the EORTC QLQ-C30 v3.0. The analysis population was the PRO-evaluable population: patients in the ITT population with a baseline and at least one post-baseline PRO assessment.

AnalysisHR95% CIP-valueMethod
TTD in all randomized participants 0.98 0.81–1.18 0.8078 Log-rank; stratified analysis
TTD in participants with detectable PD-L1 0.98 0.73–1.31 0.8879 Log-rank; stratified analysis
Clinical Biostats interpretation

For both TTD analyses, the hazard ratio is 0.98. This is very close to the null value of 1.00, meaning that the estimated instantaneous rate of deterioration was approximately 98% of the comparator rate under each reported analysis.

For all randomized participants, the 95% CI is 0.81–1.18 and the P-value is 0.8078. For participants with detectable PD-L1, the 95% CI is 0.73–1.31 and the P-value is 0.8879. Both intervals include 1.00.

These results should not be translated into a claim that the two groups had identical quality-of-life experiences. Rather, the reported analyses do not provide evidence of a detectable difference in the time-to-deterioration hazard under the specified statistical framework. The confidence intervals also show that a range of underlying effects remains statistically compatible with the data.

13. Summary of Posted Statistical Analyses

EndpointRoleEstimate95% CIP-valueMethod
PFS, all randomizedPrimaryHR 0.800.69–0.920.0025Log-rank
PFS, detectable PD-L1PrimaryHR 0.620.49–0.78<0.0001Log-rank
OS, all randomizedPrimaryHR 0.870.75–1.020.0770Log-rank
OS, detectable PD-L1PrimaryHR 0.670.53–0.860.0016Log-rank
ORR, all randomizedSecondaryDifference 10.123.40–16.840.0021Cochran-Mantel-Haenszel
ORR, detectable PD-L1SecondaryDifference 16.305.67–26.920.0016Cochran-Mantel-Haenszel
DOR, all randomizedSecondaryHR 0.780.63–0.980.0285Log-rank
DOR, detectable PD-L1SecondaryHR 0.600.43–0.860.0047Log-rank
TTD, all randomizedSecondaryHR 0.980.81–1.180.8078Log-rank
TTD, detectable PD-L1SecondaryHR 0.980.73–1.310.8879Log-rank

The pattern of estimates illustrates why clinical-trial interpretation should not be reduced to a single P-value. The four primary analyses include two PFS hazard ratios below 1, an all-randomized OS hazard ratio of 0.87 with a confidence interval crossing 1, and a PD-L1-selected OS hazard ratio of 0.67 with a confidence interval below 1. The secondary analyses additionally cover tumor response, duration of response, and patient-reported deterioration.

14. Serious Adverse Events

The ClinicalTrials.gov record provides serious adverse event counts by treatment arm. These are presented as affected participants divided by participants at risk.

Treatment groupSerious adverse eventsAffected / at risk
Placebo (q2w) + Nab-Paclitaxel Serious adverse events 80 / 430
Atezolizumab (q2w) + Nab-Paclitaxel Serious adverse events 110 / 460

Placebo arm

80 of 430 participants at risk were affected by a serious adverse event.

Atezolizumab arm

110 of 460 participants at risk were affected by a serious adverse event.

The ClinicalTrials.gov record does not provide a formal between-group statistical test for serious adverse events. Accordingly, this page reports the affected and at-risk counts without constructing an unreported P-value, risk ratio, confidence interval, or hypothesis test.

Safety interpretation: efficacy and safety use different statistical populations and questions. A serious adverse event count should not be treated as the inverse of an efficacy endpoint, and the available counts alone do not establish a causal difference in serious-adverse-event risk without a prespecified safety analysis framework.

15. Statistical Methods Explained

Why was a log-rank test used for PFS and OS?

PFS and OS are time-to-event endpoints. A simple comparison of proportions would discard the timing of progression or death and would handle censoring poorly. The log-rank test instead compares the observed and expected pattern of events across the follow-up period while accommodating right-censored observations. That makes it a natural method for the registered PFS and OS analyses.

What does an HR of 0.80 mean?

An HR of 0.80 means that the estimated instantaneous event rate in the treatment group is 80% of that in the comparator group under the fitted time-to-event analysis. The corresponding relative hazard reduction is 20%. It does not mean a 20% absolute improvement, a 20-percentage-point increase in survival, or that every patient receives the same benefit.

Why is the confidence interval important?

The point estimate is only one summary of the treatment comparison. A 95% confidence interval shows the uncertainty surrounding that estimate under the statistical model and sampling framework. For example, the all-randomized OS estimate of 0.87 has a 95% CI of 0.75–1.02. That interval is more informative than the number 0.87 alone because it communicates the range of values compatible with the analysis at the stated confidence level.

Why does the P-value not measure effect size?

A P-value quantifies how compatible the observed data are with a specified null hypothesis under the test procedure. It is not a percentage benefit, probability of treatment success, or measure of clinical importance. Effect magnitude is described by measures such as the hazard ratio or response-rate difference, while the confidence interval describes uncertainty around that effect estimate.

Why are the PD-L1 analyses separate from the all-randomized analyses?

The registry explicitly defines a PD-L1-selected subpopulation: patients in the ITT population whose PD-L1 status was IC1/2/3 at randomization. Restricting an analysis population changes the question being answered. A treatment effect estimated in this subgroup cannot automatically be assumed to be identical to the effect in the full randomized population.

Why is DOR different from PFS?

DOR begins conceptually among participants who have achieved an objective response and asks how long that response persists. PFS begins at randomization and counts progression or death as the event. Because DOR conditions on response, its analysis population is different from the ITT population and its hazard ratio answers a different question.

Why use the Cochran-Mantel-Haenszel test for objective response?

Objective response is a binary categorical endpoint, unlike PFS and OS. The Cochran-Mantel-Haenszel test provides a way to compare treatment groups while accounting for stratification. The reported effect measure is the difference in overall response rates, so the test and effect estimate are aligned with the categorical nature of the endpoint.

16. Confidence Intervals and the Null Value

EndpointEstimate95% CINull valueDoes CI include null?
PFS, all randomizedHR 0.800.69–0.921.00No
PFS, detectable PD-L1HR 0.620.49–0.781.00No
OS, all randomizedHR 0.870.75–1.021.00Yes
OS, detectable PD-L1HR 0.670.53–0.861.00No
ORR, all randomizedDifference 10.123.40–16.840No
ORR, detectable PD-L1Difference 16.305.67–26.920No
DOR, all randomizedHR 0.780.63–0.981.00No
DOR, detectable PD-L1HR 0.600.43–0.861.00No
TTD, all randomizedHR 0.980.81–1.181.00Yes
TTD, detectable PD-L1HR 0.980.73–1.311.00Yes

For hazard ratios, the null value is 1.00 because a ratio of 1 represents equal hazards between groups. For a difference in response rates, the null value is 0 because a difference of zero represents equal response rates.

Important: whether a confidence interval crosses its null value is a useful descriptive feature, but formal confirmatory interpretation should follow the trial's prespecified hypothesis-testing and multiplicity framework rather than applying an isolated 0.05 rule to every reported endpoint.

17. Primary Endpoint Structure and Multiplicity

IMpassion130 has four registered primary endpoints: PFS in all randomized participants, PFS in participants with detectable PD-L1, OS in all randomized participants, and OS in participants with detectable PD-L1. This creates an important multiplicity issue because multiple primary hypotheses are being evaluated.

Primary endpointPopulationEffect measureP-value
PFSAll randomizedHR 0.800.0025
PFSDetectable PD-L1HR 0.62<0.0001
OSAll randomizedHR 0.870.0770
OSDetectable PD-L1HR 0.670.0016

Multiplicity is the statistical problem created when several hypotheses are tested within the same confirmatory trial. Without a prespecified strategy, repeatedly testing hypotheses can increase the familywise probability of a false-positive conclusion.

The ClinicalTrials.gov record identifies all four endpoints as primary and identify each analysis as a superiority hypothesis, but they do not provide the complete alpha-allocation or hierarchical testing procedure. Therefore, this page does not infer an unreported multiplicity adjustment or assign a new interpretation to the individual P-values beyond their reported values.

18. Comparing the Overall and PD-L1-Selected Analyses

The trial provides a useful statistical lesson in how an overall population and a biomarker-defined population can produce different estimates without necessarily demonstrating that the treatment effect itself is statistically different between populations.

Primary hazard-ratio estimates
All randomized PFS
HR 0.80
PD-L1 PFS
HR 0.62
All randomized OS
HR 0.87
PD-L1 OS
HR 0.67

The numerical differences between the all-randomized and PD-L1-selected estimates are descriptive. A formal claim that PD-L1 modifies the treatment effect would require an appropriate interaction or heterogeneity analysis. A smaller P-value in one subgroup is not, by itself, evidence that the treatment effect is statistically different from the effect in another subgroup.

19. Time-to-Event Analysis: What Is Being Compared?

For PFS, the event is the first occurrence of disease progression according to investigator tumor assessments using RECIST v1.1 or death from any cause, whichever occurs first. For OS, the event is death from any cause. These definitions mean that PFS and OS are not simply percentages measured at a single point in time.

Conceptual survival function
S(t) = P(T > t)

A survival function describes the probability of remaining event-free beyond time t. In clinical-trial analysis, Kaplan-Meier estimation is commonly used to estimate this function when observations may be right-censored.

The registry-reported statistical-method field specifically reports log-rank testing rather than a Kaplan-Meier method. Kaplan-Meier estimation is nevertheless the standard descriptive framework associated with these time-to-event endpoints because it estimates the event-free distribution while retaining information from censored participants up to their censoring times.

Educational note: this page does not draw a fabricated Kaplan-Meier curve from summary hazard ratios and P-values. A valid KM reconstruction requires appropriate event and censoring information or sufficiently detailed source data.

20. P-Values in Context

The ten formal analyses span a wide range of P-values. The appropriate interpretation is endpoint-specific and must distinguish statistical evidence from effect magnitude.

EndpointP-valueEffect estimateInterpretive point
PFS, all randomized0.0025HR 0.80Evidence against the null under the reported test; effect magnitude is summarized by HR and CI.
PFS, detectable PD-L1<0.0001HR 0.62Strong statistical evidence under the reported test; the HR describes relative hazard.
OS, all randomized0.0770HR 0.87CI includes 1.00; the point estimate remains below 1 but is uncertain.
OS, detectable PD-L10.0016HR 0.67Evidence against the null under the reported test in the selected population.
ORR, all randomized0.0021Difference 10.12Categorical response difference, not a time-to-event effect.
ORR, detectable PD-L10.0016Difference 16.30Response-rate difference in the selected population.
DOR, all randomized0.0285HR 0.78Duration among responders, not the randomized population's PFS.
DOR, detectable PD-L10.0047HR 0.60Duration-of-response comparison among responders.
TTD, all randomized0.8078HR 0.98Estimate close to the null; CI includes 1.00.
TTD, detectable PD-L10.8879HR 0.98Estimate close to the null; CI includes 1.00.

This illustrates why statistical significance should never substitute for reporting the effect estimate and confidence interval. A complete result contains at least the endpoint definition, analysis population, effect measure, point estimate, confidence interval, and P-value.

21. Limitations

22. Why This Trial Matters Statistically

IMpassion130 is a useful teaching case because the registry record contains several layers of clinical-trial statistics in a single randomized study: time-to-event endpoints, a biomarker-defined analysis population, categorical response endpoints, duration-of-response analysis, patient-reported time-to-deterioration analysis, stratified testing, and ITT analysis.

ConceptHow it appears in IMpassion130
RandomizationTwo-arm randomized parallel-group phase 3 design.
Double blindingBoth treatment strategies were evaluated in a double-blind framework.
ITT analysisPrimary all-randomized analyses use an ITT population consisting of all randomized patients.
Time-to-event endpointsPFS, OS, DOR, and TTD are analyzed using log-rank methods.
Hazard ratioPrimary and secondary time-to-event effects are reported as HRs.
Confidence intervalsAll ten registry-reported formal analyses include 95% confidence intervals.
Log-rank testUsed for all four primary endpoints and the DOR and TTD secondary analyses.
Cochran-Mantel-Haenszel testUsed for the two reported objective-response analyses.
Biomarker-selected populationSeparate PFS and OS primary analyses use the detectable-PD-L1 population.
Different analysis populationsORR, DOR, and PRO analyses use endpoint-specific evaluable populations.
MultiplicityFour registered primary endpoints require interpretation within a multiple-testing framework.
Safety analysisSerious adverse events are reported as affected participants over participants at risk by arm.

The statistical lesson is broader than any individual result. A randomized trial does not generate one universal "treatment effect." It generates a collection of estimates tied to specific endpoints, populations, time frames, estimands, and statistical methods. Understanding those connections is essential for interpreting the evidence correctly.

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

Apply the same statistical concepts with Clinical Biostats calculators:

25. Sources

Continue with the statistical methods behind this trial

Explore the underlying biostatistical concepts, then apply them with Clinical Biostats calculators and related clinical-trial analyses.

26. Record Summary

IMpassion130 provides a compact example of how modern randomized-trial evidence is assembled from multiple statistical layers. The four registered primary endpoints include PFS and OS in both the all-randomized and detectable-PD-L1 populations. The reported primary hazard ratios were 0.80 and 0.62 for PFS and 0.87 and 0.67 for OS, with their corresponding confidence intervals and P-values. Secondary analyses extended the statistical picture to objective response, duration of response, and patient-reported time to deterioration.

The most important interpretive principle is to keep the endpoint, analysis population, effect measure, confidence interval, and hypothesis test connected. An HR of 0.62 for PD-L1-selected PFS cannot be substituted for the HR of 0.87 for all-randomized OS, just as a response-rate difference cannot be interpreted as a survival effect. Each statistic describes a particular comparison under a particular analysis framework.

Clinical Biostats methodology: A trial-results page should not merely reproduce a list of P-values. The goal is to reconstruct the statistical story of the trial, distinguish randomized efficacy analyses from endpoint-specific evaluable populations, explain what each effect measure means, and identify the uncertainty and design features that determine how the results should be interpreted.