This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record for IMpassion130 and the posted statistical analyses summarized below.
1. Trial at a Glance
IMpassion130 was a randomized, double-blind, parallel-group phase 3 trial evaluating atezolizumab in combination with nab-paclitaxel versus placebo with nab-paclitaxel in participants with previously untreated metastatic triple-negative breast cancer.
| Feature | IMpassion130 |
|---|---|
| Phase | Phase 3 |
| Condition | Triple Negative Breast Cancer |
| Design | Randomized, double-blind, parallel |
| Allocation | Randomized |
| Enrollment | 902 |
| Primary purpose | Treatment |
| Primary endpoints | Four registered primary endpoints: PFS in all randomized participants; PFS in participants with detectable PD-L1; OS in all randomized participants; OS in participants with detectable PD-L1 |
| Results posted | Yes |
| Outcome measures posted | 15 |
| Statistical analyses posted | 10 |
| ClinicalTrials.gov | NCT02425891 |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | Industry |
| Trial dates | Start: June 23, 2015 · Primary completion: April 14, 2020 |
2. Clinical Question
The central statistical question was whether adding atezolizumab to nab-paclitaxel produced a superior time-to-event outcome compared with placebo plus nab-paclitaxel in participants with previously untreated metastatic triple-negative breast cancer.
Population
Participants with previously untreated metastatic triple-negative breast cancer enrolled in the randomized phase 3 study.
Intervention
Atezolizumab (MPDL3280A), an engineered anti-PDL1 antibody, in combination with nab-paclitaxel.
Comparator
Placebo in combination with nab-paclitaxel.
Primary question
Does the atezolizumab plus nab-paclitaxel regimen improve PFS and OS relative to placebo plus nab-paclitaxel?
3. Trial Design
Placebo + Nab-Paclitaxel
- Placebo (q2w)
- Nab-paclitaxel
- Double-blind randomized comparator regimen
Atezolizumab + Nab-Paclitaxel
- Atezolizumab (MPDL3280A), an engineered anti-PDL1 antibody
- Atezolizumab (q2w)
- Nab-paclitaxel
- Double-blind randomized treatment regimen
4. Endpoints
The registry lists four primary endpoints and a mixture of time-to-event and binary endpoint types. The primary efficacy analyses reported in the ClinicalTrials.gov record uses log-rank testing and hazard ratios for the four primary endpoints.
| Primary endpoint | Registered definition | Time frame | Type |
|---|---|---|---|
| Progression Free Survival (PFS) According to RECIST Version 1.1 (v1.1) in All Randomized Participants | PFS was defined as the time from randomization to the occurrence of disease progression, as determined by investigators from tumor assessments per RECIST v1.1, or death from any cause, whichever occurred first. | Baseline up to approximately 34 months | Time-to-event |
| PFS According to RECIST v1.1 in Participants With Detectable PD-L1 | PFS was defined as the time from randomization to the occurrence of disease progression, as determined by investigators from tumor assessments per RECIST v1.1, or death from any cause, whichever occurred first. | Baseline up to approximately 34 months | Time-to-event |
| Overall Survival (OS) in All Randomized Participants | OS was defined as the time from the date of randomization to the date of death from any cause. | Baseline until death due to any cause (up to approximately 58 months) | Time-to-event |
| OS in Participants With Detectable PD-L1 | OS was defined as the time from the date of randomization to the date of death from any cause. | Baseline until death due to any cause (up to approximately 58 months) | Time-to-event |
Secondary endpoints with posted formal analyses
| Endpoint | Time frame | Analysis method | Effect measure |
|---|---|---|---|
| Objective response rate in all randomized participants | Baseline up to approximately 34 months | Cochran-Mantel-Haenszel test | Difference in overall response rates |
| Objective response rate in participants with detectable PD-L1 | Baseline up to approximately 34 months | Cochran-Mantel-Haenszel test | Difference in overall response rates |
| Duration of response in all randomized participants | Baseline up to approximately 34 months | Log-rank test | Hazard ratio |
| Duration of response in participants with detectable PD-L1 | Baseline up to approximately 34 months | Log-rank test | Hazard ratio |
| Time to deterioration in global health status/health-related quality of life in all randomized participants | Baseline up to approximately 58 months | Log-rank test | Hazard ratio |
| Time to deterioration in global health status/health-related quality of life in participants with detectable PD-L1 | Baseline up to approximately 58 months | Log-rank test | Hazard ratio |
5. Statistical Methodology
Intention-to-treat analysis
For the primary all-randomized analyses, the registry defines the ITT population as all randomized patients, whether or not the assigned study treatment was received. This preserves the randomized treatment comparison: once randomization occurs, the efficacy analysis remains tied to the assigned group rather than being redefined according to subsequent treatment exposure.
The principal statistical advantage is that the treatment groups retain the allocation created by randomization. This is especially important when treatment discontinuation, protocol deviations, or other post-randomization events occur.
Log-rank test
The registry reports the log-rank test for PFS, OS, DOR, and time to deterioration. The log-rank test is designed for comparing time-to-event distributions between randomized groups while incorporating the timing of events and the presence of right-censored observations.
The resulting P-value addresses evidence against the null comparison under the specified test. It does not itself quantify how large or clinically important the treatment difference is.
Hazard ratio
The principal effect measure for the time-to-event analyses was the hazard ratio. A hazard ratio compares the estimated instantaneous event rate between treatment groups within the analysis framework.
For example, an HR of 0.80 corresponds to an estimated instantaneous event rate that is 80% of the comparator group's rate under the fitted analysis framework. Equivalently, 0.80 corresponds to a 20% lower estimated hazard; it does not mean 20% of patients avoided the event.
Stratified analysis
The registry identifies the primary time-to-event analyses as stratified analyses. The ClinicalTrials.gov record does not specify the individual stratification variables, so they are not reproduced here. The important statistical point is that the treatment comparison was not described as an unstratified log-rank analysis.
Cochran-Mantel-Haenszel test
The objective-response analyses used the Cochran-Mantel-Haenszel test. This is a categorical-data method that can compare treatment groups while accounting for stratification. The associated effect measure reported in the registry is the difference in overall response rates.
Superiority framework
All ten registry-reported formal analyses are labeled with a superiority hypothesis type. Thus, the reported hazard ratios and response-rate differences are being interpreted in a framework intended to test whether the treatment groups differ in the favorable direction specified by the trial's hypotheses. The ClinicalTrials.gov record does not provide a non-inferiority margin or equivalence margin.
6. Results: Progression-Free Survival in All Randomized Participants
The first primary endpoint was PFS according to RECIST v1.1 in all randomized participants, assessed from baseline up to approximately 34 months. The analysis population was the ITT population, defined as all randomized patients whether or not the assigned study treatment was received.
Hazard ratio for progression or death
95% CI: 0.69–0.92 · P = 0.0025
Log-rank test · Stratified analysis · Superiority hypothesis
| Primary PFS analysis | Value |
|---|---|
| Endpoint | PFS according to RECIST v1.1 in all randomized participants |
| Analysis population | ITT |
| Comparison | Placebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.80 |
| 95% CI | 0.69–0.92 |
| P-value | 0.0025 |
An HR of 0.80 means that the estimated instantaneous rate of progression or death in the atezolizumab-plus-nab-paclitaxel group was approximately 80% of that in the placebo-plus-nab-paclitaxel group under the reported time-to-event analysis. Put another way, 0.80 corresponds to a 20% lower estimated hazard.
The HR does not mean that 20% of participants avoided progression or death, that every participant experienced exactly a 20% reduction, or that the median PFS differed by a particular number of months. Those would require different reported quantities.
The 95% CI of 0.69–0.92 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient outcomes. The interval also gives more information about precision than the point estimate alone.
The P-value of 0.0025 measures the evidence against the relevant null hypothesis under the specified test; it is not a measure of effect size. A small P-value can coexist with a modest effect, while a larger effect can be estimated imprecisely in a small dataset.
Because this is a hazard ratio, interpretation also depends on the behavior of the hazards over time and on censoring. The ClinicalTrials.gov record identifies a stratified log-rank analysis but do not provide enough information here to independently assess the proportional-hazards assumption.
7. Results: Progression-Free Survival in Participants With Detectable PD-L1
The second primary endpoint evaluated PFS according to RECIST v1.1 in participants with detectable PD-L1. The registry defines this PD-L1-selected subpopulation as patients in the ITT population whose PD-L1 status was IC1/2/3 at the time of randomization.
Hazard ratio for progression or death
95% CI: 0.49–0.78 · P < 0.0001
Log-rank test · Stratified analysis · Superiority hypothesis
| Primary PFS analysis | Value |
|---|---|
| Endpoint | PFS according to RECIST v1.1 in participants with detectable PD-L1 |
| PD-L1-selected population | Patients in the ITT population whose PD-L1 status was IC1/2/3 at randomization |
| Comparison | Placebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.62 |
| 95% CI | 0.49–0.78 |
| P-value | <0.0001 |
An HR of 0.62 corresponds to an estimated instantaneous rate of progression or death approximately 62% of the comparator rate within this PD-L1-selected analysis. Expressed as a relative hazard interpretation, this is approximately a 38% lower estimated hazard.
The estimate applies to the defined PD-L1-selected population. It should not automatically be generalized to participants outside that population, and it does not establish that the treatment effect differs from the all-randomized estimate simply because the numerical HRs are different.
The 95% CI of 0.49–0.78 communicates the statistical uncertainty around 0.62. It is narrower than a statement such as "most patients had a 38% reduction"; individual patients do not have fixed hazard ratios assigned to them by the trial.
The reported P < 0.0001 indicates strong statistical evidence against the corresponding null hypothesis under the reported test. It does not mean the probability that the treatment has an effect is less than 0.0001, and it does not measure the magnitude of clinical benefit.
Because this is a prespecified primary endpoint within a superiority framework, its interpretation must be understood in the context of the trial's overall multiplicity and testing strategy. The ClinicalTrials.gov record identifies four primary endpoints but do not provide a complete alpha-allocation scheme, so no additional multiplicity adjustment is inferred here.
8. Results: Overall Survival in All Randomized Participants
The third primary endpoint was OS in all randomized participants. OS was defined as the time from randomization to death from any cause, with the registered time frame extending from baseline until death due to any cause, up to approximately 58 months.
Hazard ratio for death
95% CI: 0.75–1.02 · P = 0.0770
Log-rank test · Stratified analysis · Superiority hypothesis
| Primary OS analysis | Value |
|---|---|
| Endpoint | Overall survival in all randomized participants |
| Definition | Time from date of randomization to date of death from any cause |
| Analysis population | ITT |
| Comparison | Placebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.87 |
| 95% CI | 0.75–1.02 |
| P-value | 0.0770 |
An HR of 0.87 corresponds to an estimated instantaneous rate of death approximately 87% of the comparator rate under the reported analysis, or approximately a 13% lower estimated hazard.
The estimate is not an absolute survival difference. It does not mean that 13% of patients survived because of treatment, nor does it specify how long individual patients lived. Median OS, survival probabilities at particular time points, and absolute risk differences are separate quantities and are not reported in the ClinicalTrials.gov record.
The 95% CI of 0.75–1.02 spans 1.00. That means the reported interval includes the null hazard-ratio value. The P-value of 0.0770 likewise does not provide conventional evidence against the null at a 0.05 threshold if that threshold were used. Importantly, the P-value itself is not an effect-size measure.
The correct statistical reading is therefore more nuanced than simply describing the HR as "good" or "bad": the point estimate is below 1, but the confidence interval reflects uncertainty that includes a hazard ratio of 1. The conclusion also depends on the trial's prespecified hypothesis-testing framework and multiplicity handling, which are not fully specified in the ClinicalTrials.gov record.
As with the PFS analyses, a Cox-type hazard-ratio interpretation should not be extended beyond what the registry analysis supports. Censoring and the behavior of hazards over time remain important considerations for any time-to-event analysis.
9. Results: Overall Survival in Participants With Detectable PD-L1
The fourth primary endpoint evaluated OS in the PD-L1-selected subpopulation. This population consisted of patients in the ITT population whose PD-L1 status was IC1/2/3 at the time of randomization.
Hazard ratio for death
95% CI: 0.53–0.86 · P = 0.0016
Log-rank test · Stratified analysis · Superiority hypothesis
| Primary OS analysis | Value |
|---|---|
| Endpoint | OS in participants with detectable PD-L1 |
| Definition | Time from date of randomization to date of death from any cause |
| PD-L1-selected population | Patients in the ITT population whose PD-L1 status was IC1/2/3 at randomization |
| Comparison | Placebo plus nab-paclitaxel vs atezolizumab plus nab-paclitaxel |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.67 |
| 95% CI | 0.53–0.86 |
| P-value | 0.0016 |
An HR of 0.67 means that the estimated instantaneous rate of death was approximately 67% of the comparator rate in the PD-L1-selected analysis, corresponding to approximately a 33% lower estimated hazard.
The estimate is specific to the PD-L1-selected population defined by the registry. It does not imply that the same hazard ratio applies to every participant in the overall ITT population.
The 95% CI of 0.53–0.86 indicates uncertainty around the point estimate while remaining below 1.00. The P-value of 0.0016 indicates statistical evidence against the relevant null hypothesis under the reported analysis, but it does not quantify the magnitude of the treatment effect.
It is also important not to compare P-values as though they measure the strength of treatment benefit across endpoints. The PFS and OS analyses answer different clinical questions, and the PD-L1-selected and all-randomized populations are different analysis populations.
Finally, the ClinicalTrials.gov record does not provide enough detail to reconstruct the full multiplicity hierarchy among the four primary endpoints. The four results should therefore be interpreted within the prespecified statistical design rather than treated as four unrelated hypothesis tests.
10. Secondary Results: Objective Response Rate
Objective response was defined as the percentage of participants with an objective response of complete response (CR) or partial response (PR) according to RECIST v1.1. The registry reports formal Cochran-Mantel-Haenszel analyses in both the all-randomized response-evaluable population and the PD-L1-selected response-evaluable population.
Objective Response in All Randomized Participants
Difference in overall response rates
95% CI: 3.40–16.84 · P = 0.0021
Cochran-Mantel-Haenszel test · Stratified analysis
The reported effect measure is a difference in overall response rates, not a hazard ratio. A value of 10.12 represents a 10.12-unit difference in the response-rate scale used by the registry analysis, comparing placebo plus nab-paclitaxel with atezolizumab plus nab-paclitaxel as listed in the analysis. It should not be converted into a hazard ratio or interpreted as a 10.12-fold effect.
The 95% CI of 3.40–16.84 quantifies uncertainty around the estimated response-rate difference. The P-value of 0.0021 addresses the hypothesis test and does not describe the size of the response difference.
The response analysis uses an evaluable population defined as patients in the ITT population with measurable disease at baseline. That differs conceptually from simply assuming that every randomized participant contributes a binary response observation.
Objective Response in Participants With Detectable PD-L1
Difference in overall response rates
95% CI: 5.67–26.92 · P = 0.0016
Cochran-Mantel-Haenszel test · Stratified analysis
The reported difference in overall response rates is 16.30 in the PD-L1-selected response-evaluable analysis. Its 95% CI is 5.67–26.92, giving the statistical uncertainty around that estimate.
The P-value of 0.0016 indicates evidence against the corresponding null hypothesis under the Cochran-Mantel-Haenszel analysis. It does not mean there is a 0.16% probability that the observed difference occurred by chance, because a frequentist P-value is not the posterior probability of a hypothesis.
The difference in response rates also answers a different question from PFS or OS. Response is a categorical tumor-assessment outcome; PFS and OS incorporate the timing of events and censoring. A complete statistical interpretation should therefore retain these endpoints as complementary rather than interchangeable measures.
11. Secondary Results: Duration of Response
Duration of response (DOR) was analyzed among participants with an objective response. The registry reports log-rank tests and hazard ratios for both the all-participant and detectable-PD-L1 analyses.
| Analysis | HR | 95% CI | P-value | Method |
|---|---|---|---|---|
| DOR in all randomized participants | 0.78 | 0.63–0.98 | 0.0285 | Log-rank; unstratified analysis |
| DOR in participants with detectable PD-L1 | 0.60 | 0.43–0.86 | 0.0047 | Log-rank; unstratified analysis |
All randomized participants
An HR of 0.78 corresponds to an estimated 22% lower instantaneous rate of loss of response under the reported analysis. The 95% CI is 0.63–0.98 and the P-value is 0.0285.
Detectable PD-L1
An HR of 0.60 corresponds to an estimated 40% lower instantaneous rate of loss of response under the reported analysis. The 95% CI is 0.43–0.86 and the P-value is 0.0047.
DOR is inherently conditional on achieving an objective response. That makes its analysis population different from the ITT population used for the primary all-randomized PFS and OS analyses. A DOR hazard ratio therefore should not be interpreted as though it describes the event experience of every randomized participant.
12. Secondary Results: Time to Deterioration in Global Health Status / Quality of Life
The registry also reports time to deterioration (TTD) in global health status/health-related quality of life according to the EORTC QLQ-C30 v3.0. The analysis population was the PRO-evaluable population: patients in the ITT population with a baseline and at least one post-baseline PRO assessment.
| Analysis | HR | 95% CI | P-value | Method |
|---|---|---|---|---|
| TTD in all randomized participants | 0.98 | 0.81–1.18 | 0.8078 | Log-rank; stratified analysis |
| TTD in participants with detectable PD-L1 | 0.98 | 0.73–1.31 | 0.8879 | Log-rank; stratified analysis |
For both TTD analyses, the hazard ratio is 0.98. This is very close to the null value of 1.00, meaning that the estimated instantaneous rate of deterioration was approximately 98% of the comparator rate under each reported analysis.
For all randomized participants, the 95% CI is 0.81–1.18 and the P-value is 0.8078. For participants with detectable PD-L1, the 95% CI is 0.73–1.31 and the P-value is 0.8879. Both intervals include 1.00.
These results should not be translated into a claim that the two groups had identical quality-of-life experiences. Rather, the reported analyses do not provide evidence of a detectable difference in the time-to-deterioration hazard under the specified statistical framework. The confidence intervals also show that a range of underlying effects remains statistically compatible with the data.
13. Summary of Posted Statistical Analyses
| Endpoint | Role | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| PFS, all randomized | Primary | HR 0.80 | 0.69–0.92 | 0.0025 | Log-rank |
| PFS, detectable PD-L1 | Primary | HR 0.62 | 0.49–0.78 | <0.0001 | Log-rank |
| OS, all randomized | Primary | HR 0.87 | 0.75–1.02 | 0.0770 | Log-rank |
| OS, detectable PD-L1 | Primary | HR 0.67 | 0.53–0.86 | 0.0016 | Log-rank |
| ORR, all randomized | Secondary | Difference 10.12 | 3.40–16.84 | 0.0021 | Cochran-Mantel-Haenszel |
| ORR, detectable PD-L1 | Secondary | Difference 16.30 | 5.67–26.92 | 0.0016 | Cochran-Mantel-Haenszel |
| DOR, all randomized | Secondary | HR 0.78 | 0.63–0.98 | 0.0285 | Log-rank |
| DOR, detectable PD-L1 | Secondary | HR 0.60 | 0.43–0.86 | 0.0047 | Log-rank |
| TTD, all randomized | Secondary | HR 0.98 | 0.81–1.18 | 0.8078 | Log-rank |
| TTD, detectable PD-L1 | Secondary | HR 0.98 | 0.73–1.31 | 0.8879 | Log-rank |
The pattern of estimates illustrates why clinical-trial interpretation should not be reduced to a single P-value. The four primary analyses include two PFS hazard ratios below 1, an all-randomized OS hazard ratio of 0.87 with a confidence interval crossing 1, and a PD-L1-selected OS hazard ratio of 0.67 with a confidence interval below 1. The secondary analyses additionally cover tumor response, duration of response, and patient-reported deterioration.
14. Serious Adverse Events
The ClinicalTrials.gov record provides serious adverse event counts by treatment arm. These are presented as affected participants divided by participants at risk.
| Treatment group | Serious adverse events | Affected / at risk |
|---|---|---|
| Placebo (q2w) + Nab-Paclitaxel | Serious adverse events | 80 / 430 |
| Atezolizumab (q2w) + Nab-Paclitaxel | Serious adverse events | 110 / 460 |
Placebo arm
80 of 430 participants at risk were affected by a serious adverse event.
Atezolizumab arm
110 of 460 participants at risk were affected by a serious adverse event.
The ClinicalTrials.gov record does not provide a formal between-group statistical test for serious adverse events. Accordingly, this page reports the affected and at-risk counts without constructing an unreported P-value, risk ratio, confidence interval, or hypothesis test.
15. Statistical Methods Explained
Why was a log-rank test used for PFS and OS?
PFS and OS are time-to-event endpoints. A simple comparison of proportions would discard the timing of progression or death and would handle censoring poorly. The log-rank test instead compares the observed and expected pattern of events across the follow-up period while accommodating right-censored observations. That makes it a natural method for the registered PFS and OS analyses.
What does an HR of 0.80 mean?
An HR of 0.80 means that the estimated instantaneous event rate in the treatment group is 80% of that in the comparator group under the fitted time-to-event analysis. The corresponding relative hazard reduction is 20%. It does not mean a 20% absolute improvement, a 20-percentage-point increase in survival, or that every patient receives the same benefit.
Why is the confidence interval important?
The point estimate is only one summary of the treatment comparison. A 95% confidence interval shows the uncertainty surrounding that estimate under the statistical model and sampling framework. For example, the all-randomized OS estimate of 0.87 has a 95% CI of 0.75–1.02. That interval is more informative than the number 0.87 alone because it communicates the range of values compatible with the analysis at the stated confidence level.
Why does the P-value not measure effect size?
A P-value quantifies how compatible the observed data are with a specified null hypothesis under the test procedure. It is not a percentage benefit, probability of treatment success, or measure of clinical importance. Effect magnitude is described by measures such as the hazard ratio or response-rate difference, while the confidence interval describes uncertainty around that effect estimate.
Why are the PD-L1 analyses separate from the all-randomized analyses?
The registry explicitly defines a PD-L1-selected subpopulation: patients in the ITT population whose PD-L1 status was IC1/2/3 at randomization. Restricting an analysis population changes the question being answered. A treatment effect estimated in this subgroup cannot automatically be assumed to be identical to the effect in the full randomized population.
Why is DOR different from PFS?
DOR begins conceptually among participants who have achieved an objective response and asks how long that response persists. PFS begins at randomization and counts progression or death as the event. Because DOR conditions on response, its analysis population is different from the ITT population and its hazard ratio answers a different question.
Why use the Cochran-Mantel-Haenszel test for objective response?
Objective response is a binary categorical endpoint, unlike PFS and OS. The Cochran-Mantel-Haenszel test provides a way to compare treatment groups while accounting for stratification. The reported effect measure is the difference in overall response rates, so the test and effect estimate are aligned with the categorical nature of the endpoint.
16. Confidence Intervals and the Null Value
| Endpoint | Estimate | 95% CI | Null value | Does CI include null? |
|---|---|---|---|---|
| PFS, all randomized | HR 0.80 | 0.69–0.92 | 1.00 | No |
| PFS, detectable PD-L1 | HR 0.62 | 0.49–0.78 | 1.00 | No |
| OS, all randomized | HR 0.87 | 0.75–1.02 | 1.00 | Yes |
| OS, detectable PD-L1 | HR 0.67 | 0.53–0.86 | 1.00 | No |
| ORR, all randomized | Difference 10.12 | 3.40–16.84 | 0 | No |
| ORR, detectable PD-L1 | Difference 16.30 | 5.67–26.92 | 0 | No |
| DOR, all randomized | HR 0.78 | 0.63–0.98 | 1.00 | No |
| DOR, detectable PD-L1 | HR 0.60 | 0.43–0.86 | 1.00 | No |
| TTD, all randomized | HR 0.98 | 0.81–1.18 | 1.00 | Yes |
| TTD, detectable PD-L1 | HR 0.98 | 0.73–1.31 | 1.00 | Yes |
For hazard ratios, the null value is 1.00 because a ratio of 1 represents equal hazards between groups. For a difference in response rates, the null value is 0 because a difference of zero represents equal response rates.
17. Primary Endpoint Structure and Multiplicity
IMpassion130 has four registered primary endpoints: PFS in all randomized participants, PFS in participants with detectable PD-L1, OS in all randomized participants, and OS in participants with detectable PD-L1. This creates an important multiplicity issue because multiple primary hypotheses are being evaluated.
| Primary endpoint | Population | Effect measure | P-value |
|---|---|---|---|
| PFS | All randomized | HR 0.80 | 0.0025 |
| PFS | Detectable PD-L1 | HR 0.62 | <0.0001 |
| OS | All randomized | HR 0.87 | 0.0770 |
| OS | Detectable PD-L1 | HR 0.67 | 0.0016 |
Multiplicity is the statistical problem created when several hypotheses are tested within the same confirmatory trial. Without a prespecified strategy, repeatedly testing hypotheses can increase the familywise probability of a false-positive conclusion.
The ClinicalTrials.gov record identifies all four endpoints as primary and identify each analysis as a superiority hypothesis, but they do not provide the complete alpha-allocation or hierarchical testing procedure. Therefore, this page does not infer an unreported multiplicity adjustment or assign a new interpretation to the individual P-values beyond their reported values.
18. Comparing the Overall and PD-L1-Selected Analyses
The trial provides a useful statistical lesson in how an overall population and a biomarker-defined population can produce different estimates without necessarily demonstrating that the treatment effect itself is statistically different between populations.
The numerical differences between the all-randomized and PD-L1-selected estimates are descriptive. A formal claim that PD-L1 modifies the treatment effect would require an appropriate interaction or heterogeneity analysis. A smaller P-value in one subgroup is not, by itself, evidence that the treatment effect is statistically different from the effect in another subgroup.
19. Time-to-Event Analysis: What Is Being Compared?
For PFS, the event is the first occurrence of disease progression according to investigator tumor assessments using RECIST v1.1 or death from any cause, whichever occurs first. For OS, the event is death from any cause. These definitions mean that PFS and OS are not simply percentages measured at a single point in time.
A survival function describes the probability of remaining event-free beyond time t. In clinical-trial analysis, Kaplan-Meier estimation is commonly used to estimate this function when observations may be right-censored.
The registry-reported statistical-method field specifically reports log-rank testing rather than a Kaplan-Meier method. Kaplan-Meier estimation is nevertheless the standard descriptive framework associated with these time-to-event endpoints because it estimates the event-free distribution while retaining information from censored participants up to their censoring times.
20. P-Values in Context
The ten formal analyses span a wide range of P-values. The appropriate interpretation is endpoint-specific and must distinguish statistical evidence from effect magnitude.
| Endpoint | P-value | Effect estimate | Interpretive point |
|---|---|---|---|
| PFS, all randomized | 0.0025 | HR 0.80 | Evidence against the null under the reported test; effect magnitude is summarized by HR and CI. |
| PFS, detectable PD-L1 | <0.0001 | HR 0.62 | Strong statistical evidence under the reported test; the HR describes relative hazard. |
| OS, all randomized | 0.0770 | HR 0.87 | CI includes 1.00; the point estimate remains below 1 but is uncertain. |
| OS, detectable PD-L1 | 0.0016 | HR 0.67 | Evidence against the null under the reported test in the selected population. |
| ORR, all randomized | 0.0021 | Difference 10.12 | Categorical response difference, not a time-to-event effect. |
| ORR, detectable PD-L1 | 0.0016 | Difference 16.30 | Response-rate difference in the selected population. |
| DOR, all randomized | 0.0285 | HR 0.78 | Duration among responders, not the randomized population's PFS. |
| DOR, detectable PD-L1 | 0.0047 | HR 0.60 | Duration-of-response comparison among responders. |
| TTD, all randomized | 0.8078 | HR 0.98 | Estimate close to the null; CI includes 1.00. |
| TTD, detectable PD-L1 | 0.8879 | HR 0.98 | Estimate close to the null; CI includes 1.00. |
This illustrates why statistical significance should never substitute for reporting the effect estimate and confidence interval. A complete result contains at least the endpoint definition, analysis population, effect measure, point estimate, confidence interval, and P-value.
21. Limitations
- Incomplete statistical-design detail in the ClinicalTrials.gov record: the data identify stratified analyses and superiority hypotheses, but do not provide the individual stratification factors or complete multiplicity/alpha-allocation procedure.
- No inferred numerical results: median PFS, median OS, time-specific survival estimates, subgroup estimates beyond the defined PD-L1 analyses, and other quantities not reported are intentionally not reconstructed.
- Hazard-ratio interpretation: an HR is a relative time-to-event measure, not an absolute risk difference or a probability that an individual patient benefits.
- Proportional-hazards considerations: a single HR is easiest to interpret as a common relative hazard when the proportional-hazards assumption is reasonable. The ClinicalTrials.gov record does not provide enough information to independently assess that assumption.
- Different analysis populations: ITT, PD-L1-selected, ORR-evaluable, DOR-evaluable, and PRO-evaluable populations answer different questions and should not be conflated.
- Multiplicity: four registered primary endpoints create a multiple-testing problem. The ClinicalTrials.gov record does not provide the full prespecified alpha-control strategy.
- Secondary endpoints: response, DOR, and TTD are supportive analyses and should not automatically be treated as equivalent to the primary endpoints.
- Safety: only serious-adverse-event counts by arm are reported here. No additional safety comparison is inferred from those counts.
- No causal interpretation from descriptive subgroup differences alone: different HRs in the all-randomized and PD-L1-selected populations do not by themselves establish treatment-effect modification.
22. Why This Trial Matters Statistically
IMpassion130 is a useful teaching case because the registry record contains several layers of clinical-trial statistics in a single randomized study: time-to-event endpoints, a biomarker-defined analysis population, categorical response endpoints, duration-of-response analysis, patient-reported time-to-deterioration analysis, stratified testing, and ITT analysis.
| Concept | How it appears in IMpassion130 |
|---|---|
| Randomization | Two-arm randomized parallel-group phase 3 design. |
| Double blinding | Both treatment strategies were evaluated in a double-blind framework. |
| ITT analysis | Primary all-randomized analyses use an ITT population consisting of all randomized patients. |
| Time-to-event endpoints | PFS, OS, DOR, and TTD are analyzed using log-rank methods. |
| Hazard ratio | Primary and secondary time-to-event effects are reported as HRs. |
| Confidence intervals | All ten registry-reported formal analyses include 95% confidence intervals. |
| Log-rank test | Used for all four primary endpoints and the DOR and TTD secondary analyses. |
| Cochran-Mantel-Haenszel test | Used for the two reported objective-response analyses. |
| Biomarker-selected population | Separate PFS and OS primary analyses use the detectable-PD-L1 population. |
| Different analysis populations | ORR, DOR, and PRO analyses use endpoint-specific evaluable populations. |
| Multiplicity | Four registered primary endpoints require interpretation within a multiple-testing framework. |
| Safety analysis | Serious adverse events are reported as affected participants over participants at risk by arm. |
The statistical lesson is broader than any individual result. A randomized trial does not generate one universal "treatment effect." It generates a collection of estimates tied to specific endpoints, populations, time frames, estimands, and statistical methods. Understanding those connections is essential for interpreting the evidence correctly.
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
Apply the same statistical concepts with Clinical Biostats calculators:
25. Sources
- ClinicalTrials.gov: IMpassion130, NCT02425891.
- PubMed: PMID 35695609.
- PubMed: PMID 33579739.
- PubMed: PMID 32178964.
- PubMed: PMID 31786121.
- PubMed: PMID 30345906.
Continue with the statistical methods behind this trial
Explore the underlying biostatistical concepts, then apply them with Clinical Biostats calculators and related clinical-trial analyses.
26. Record Summary
IMpassion130 provides a compact example of how modern randomized-trial evidence is assembled from multiple statistical layers. The four registered primary endpoints include PFS and OS in both the all-randomized and detectable-PD-L1 populations. The reported primary hazard ratios were 0.80 and 0.62 for PFS and 0.87 and 0.67 for OS, with their corresponding confidence intervals and P-values. Secondary analyses extended the statistical picture to objective response, duration of response, and patient-reported time to deterioration.
The most important interpretive principle is to keep the endpoint, analysis population, effect measure, confidence interval, and hypothesis test connected. An HR of 0.62 for PD-L1-selected PFS cannot be substituted for the HR of 0.87 for all-randomized OS, just as a response-rate difference cannot be interpreted as a survival effect. Each statistic describes a particular comparison under a particular analysis framework.