This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record. Statistical explanations distinguish reported estimates from their interpretation.
1. Trial at a Glance
KEYNOTE-119 was a randomized, parallel-group, open-label phase 3 trial comparing single-agent pembrolizumab with single-agent chemotherapy in metastatic triple negative breast cancer. The registry reports three primary overall-survival analyses, defined in participants with PD-L1 CPS ≥10, participants with PD-L1 CPS ≥1, and all participants.
| Feature | KEYNOTE-119 |
|---|---|
| Phase | Phase 3 |
| Condition | Metastatic Triple Negative Breast Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 622 |
| Arms | 2 |
| Primary endpoints | 3 registered primary endpoints, all overall survival |
| Results posted | Yes |
| Statistical analyses posted | 12 |
| ClinicalTrials.gov | NCT02555657 |
2. Clinical Question
The central question was whether single-agent pembrolizumab produced a superior overall-survival outcome compared with single-agent chemotherapy in participants with metastatic triple negative breast cancer. The registered primary hypothesis type was superiority.
Population
Participants with metastatic triple negative breast cancer enrolled in the randomized phase 3 trial.
Intervention
Single-agent pembrolizumab.
Comparator
Single-agent chemotherapy. The registry lists capecitabine, eribulin, gemcitabine, and vinorelbine as chemotherapy interventions.
Primary question
Does pembrolizumab improve overall survival relative to chemotherapy under the prespecified superiority comparisons?
3. Trial Design
Pembrolizumab
- Single-agent pembrolizumab
- Biological intervention
- Compared with single-agent chemotherapy
Chemotherapy
- Single-agent chemotherapy
- Drug options listed in the registry include capecitabine, eribulin, gemcitabine, and vinorelbine
- Compared with single-agent pembrolizumab
The trial was randomized and used a parallel design, but it was not masked. This matters statistically because randomization establishes the framework for the between-group efficacy comparison, while the absence of masking can matter for outcomes that depend on assessment or treatment behavior. The primary endpoints reported here are overall-survival endpoints, which are less directly dependent on subjective outcome assessment than some other clinical outcomes.
Trial timeline
Trial start
The registry lists 13-Oct-2015 as the study start date.
Primary completion
The registry lists 11-Apr-2019 as the primary completion date.
Final-analysis database cutoff
The primary endpoint time frames specify follow-up through the final-analysis database cutoff date of 11-April-2019.
4. Endpoints
The registry defines overall survival as the time from randomization to death due to any cause. All three registered primary endpoints are time-to-event outcomes with the same final-analysis time frame.
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Overall Survival in Participants With PD-L1 CPS ≥10 | Overall survival was defined as the time from randomization to death due to any cause. Up to approximately 36 months, through the Final Analysis database cutoff date of 11-April-2019. | Cox proportional-hazards model |
| Overall Survival in Participants With PD-L1 CPS ≥1 | Overall survival was defined as the time from randomization to death due to any cause. Up to approximately 36 months, through the Final Analysis database cutoff date of 11-April-2019. | Cox proportional-hazards model |
| Overall Survival in All Participants | Overall survival was defined as the time from randomization to death due to any cause. Up to approximately 36 months, through the Final Analysis database cutoff date of 11-April-2019. | Cox proportional-hazards model |
Secondary endpoints with formal analyses
The registry also reports overall response rate, progression-free survival, and disease control rate analyses in the same three population definitions: PD-L1 CPS ≥10, PD-L1 CPS ≥1, and all participants.
| Secondary endpoint family | Population | Effect measure | Method |
|---|---|---|---|
| Overall Response Rate per RECIST 1.1 | PD-L1 CPS ≥10 | Risk difference | Score-based CI for proportions |
| Overall Response Rate per RECIST 1.1 | PD-L1 CPS ≥1 | Risk difference | Score-based CI for proportions |
| Overall Response Rate per RECIST 1.1 | All participants | Risk difference | Score-based CI for proportions |
| Progression-Free Survival per RECIST 1.1 | PD-L1 CPS ≥10 | Hazard ratio | Cox proportional-hazards model |
| Progression-Free Survival per RECIST 1.1 | PD-L1 CPS ≥1 | Hazard ratio | Cox proportional-hazards model |
| Progression-Free Survival per RECIST 1.1 | All participants | Hazard ratio | Cox proportional-hazards model |
| Disease Control Rate per RECIST 1.1 | PD-L1 CPS ≥10 | Risk difference | Score-based CI for proportions |
| Disease Control Rate per RECIST 1.1 | PD-L1 CPS ≥1 | Risk difference | Score-based CI for proportions |
| Disease Control Rate per RECIST 1.1 | All participants | Risk difference | Score-based CI for proportions |
5. Statistical Methodology
Cox proportional-hazards model
The registry reports a Cox proportional-hazards model for all six time-to-event analyses reported here: the three primary overall-survival comparisons and the three progression-free-survival comparisons. The effect measure is a hazard ratio comparing pembrolizumab with chemotherapy.
An HR below 1 corresponds to a lower estimated instantaneous event rate in the pembrolizumab group under the fitted model. An HR above 1 corresponds to a higher estimated instantaneous event rate.
A hazard ratio is a relative time-to-event measure. It does not directly give the probability of death by a particular time, the median survival, or the proportion of participants who benefit. Those quantities require additional information that is not contained in the statistical analyses posted on ClinicalTrials.gov.
Score-based confidence intervals for proportions
The registry reports the Miettinen & Nurminen method for the overall response rate and disease control rate comparisons. This is a score-based confidence-interval approach for differences in proportions, related to the Newcombe and Wilson methods.
A positive risk difference indicates a higher observed proportion in the pembrolizumab group; a negative risk difference indicates a lower observed proportion in that group.
Superiority framework
Each posted statistical analysis is labeled as a superiority hypothesis. This means the inferential question is whether the randomized treatment comparison provides evidence of a difference favoring the experimental strategy under the prespecified statistical framework. It is not a non-inferiority design, so there is no non-inferiority margin to interpret in the ClinicalTrials.gov record.
Analysis populations
The primary overall-survival analyses were performed in the following populations: all participants with PD-L1 CPS ≥10 who were included in a treatment group at randomization; all participants with PD-L1 CPS ≥1 who were included in a treatment group at randomization; and all participants who were included in a treatment group at randomization. These definitions are important because the CPS-restricted analyses are not simply smaller versions of the all-participant analysis: they answer questions in specifically defined biomarker populations.
What the registry does not specify here
The ClinicalTrials.gov record identifies the Cox model and score-based proportion method, but do not provide details about Kaplan-Meier estimation, log-rank testing, covariate adjustment beyond the reported Cox method, proportional-hazards diagnostics, missing-data imputation, interim-analysis boundaries, alpha spending, Bayesian methods, or multiplicity procedures. Those topics are therefore not assigned trial-specific procedures on this page.
6. Results: Primary Overall Survival Endpoints
ClinicalTrials.gov reports formal statistical analyses for all three registered primary endpoints. Each comparison uses the Cox proportional-hazards model, reports a two-sided 95% confidence interval, and tests a superiority hypothesis.
Overall Survival in Participants With PD-L1 CPS ≥10
Hazard ratio for death
95% CI: 0.57–1.06 · P = 0.0574
Analysis population: all participants with PD-L1 CPS ≥10 who were included in a treatment group at randomization.
The estimated hazard ratio of 0.78 corresponds to an estimated instantaneous hazard in the pembrolizumab group that was 78% of the chemotherapy-group hazard under the fitted Cox model. Expressed as a relative reduction, this corresponds to approximately a 22% lower estimated hazard under the model.
What the estimate means: The point estimate, HR 0.78, is below 1 and therefore points toward a lower estimated hazard of death with pembrolizumab than with chemotherapy in the PD-L1 CPS ≥10 analysis population.
What it does not mean: It does not mean that 22% of participants benefited, that each participant had exactly a 22% reduction in probability of death, or that survival probability was 22 percentage points higher.
Precision: The two-sided 95% CI is 0.57–1.06. It spans 1, so the estimate is compatible with a range extending from a substantially lower hazard to a modestly higher hazard under the model.
The p-value: P = 0.0574 measures the strength of evidence against the null hypothesis in the specified statistical test; it is not a measure of the size or clinical importance of the hazard ratio. A p-value should not be read as the probability that the treatment has no effect.
Caution: Interpretation of a Cox HR depends on the model and its proportional-hazards assumption. The ClinicalTrials.gov record does not report a diagnostic assessment of that assumption.
Overall Survival in Participants With PD-L1 CPS ≥1
Hazard ratio for death
95% CI: 0.69–1.06 · P = 0.0728
Analysis population: all participants with PD-L1 CPS ≥1 who were included in a treatment group at randomization.
The HR of 0.86 indicates that the estimated instantaneous hazard of death in the pembrolizumab group was 86% of the chemotherapy-group hazard under the Cox model, corresponding to approximately a 14% lower estimated hazard.
What the estimate means: The point estimate is below 1, indicating a lower estimated hazard of death for pembrolizumab relative to chemotherapy in the CPS ≥1 analysis population.
What it does not mean: HR 0.86 is not an 86% survival probability and does not mean that 14% of patients were protected from death. It is a relative time-to-event estimate.
Precision: The 95% CI of 0.69–1.06 crosses 1. The data therefore permit a range of effects that includes no hazard difference and modestly higher hazard as well as lower hazard.
The p-value: P = 0.0728 is evidence from the specified hypothesis test, not a direct probability statement about the treatment effect and not a measure of effect magnitude.
Caution: The analysis is restricted to the CPS ≥1 population. It should not be silently treated as interchangeable with the CPS ≥10 or all-participant analysis.
Overall Survival in All Participants
Hazard ratio for death
95% CI: 0.82–1.15 · P = 0.3802
Analysis population: all participants who were included in a treatment group at randomization.
For the full analysis population, the estimated hazard ratio was 0.97. Under the Cox model, this corresponds to an estimated hazard that was 97% of the chemotherapy-group hazard, or approximately a 3% lower estimated hazard.
What the estimate means: The all-participant point estimate is close to 1, indicating little estimated relative difference in the instantaneous hazard of death between the randomized groups under this model.
What it does not mean: HR 0.97 does not establish identical survival for every participant or prove that the two treatments have exactly the same effect. It is an estimate with uncertainty.
Precision: The 95% CI is 0.82–1.15. The interval includes 1 and permits both a lower and a higher hazard for pembrolizumab relative to chemotherapy.
The p-value: P = 0.3802 does not measure the probability that the treatment effect is zero, nor does it quantify clinical importance. It summarizes the result of the specified superiority hypothesis test.
Caution: No median survival estimates are reported in the ClinicalTrials.gov record, so a median-survival comparison should not be inferred from the hazard ratio alone.
7. Comparing the Three Primary Populations
| Primary endpoint population | HR | 95% CI | P-value | Interpretation of point estimate |
|---|---|---|---|---|
| PD-L1 CPS ≥10 | 0.78 | 0.57–1.06 | 0.0574 | Approximately 22% lower estimated hazard |
| PD-L1 CPS ≥1 | 0.86 | 0.69–1.06 | 0.0728 | Approximately 14% lower estimated hazard |
| All participants | 0.97 | 0.82–1.15 | 0.3802 | Approximately 3% lower estimated hazard |
These three estimates illustrate why the definition of the analysis population matters. The point estimate moves from 0.78 in the CPS ≥10 population to 0.86 in the CPS ≥1 population and 0.97 in all participants. That pattern is descriptive of the reported estimates; it does not, by itself, establish that the treatment effect statistically differs across the three populations.
8. Secondary Results: Overall Response Rate
The registry reports overall response rate per RECIST 1.1 as a secondary count/rate endpoint. The effect measure is the difference in percentages, that is, a risk difference. The reported method is the Miettinen & Nurminen method, a score-based approach for confidence intervals for proportions.
PD-L1 CPS ≥10
Difference in overall response rate
95% CI: −1.4 to 18.4 · P = 0.0457
The estimated difference in response percentages was 8.3 percentage points in favor of pembrolizumab under the registry's group comparison. The 95% confidence interval ranges from −1.4 to 18.4 percentage points.
The risk difference is an absolute rather than relative effect measure. An estimate of 8.3 means the observed response proportion in the pembrolizumab group exceeded that in the chemotherapy group by 8.3 percentage points in this analysis.
The confidence interval is relatively broad and crosses 0, ranging from −1.4 to 18.4 percentage points. Thus the interval includes a small difference in the opposite direction as well as substantially larger positive differences.
The p-value of 0.0457 summarizes the specified superiority test. It should not be interpreted as the probability that the 8.3-point effect is real, and it does not tell us how large the treatment effect is.
Because this is a secondary endpoint and the ClinicalTrials.gov record does not establish a multiplicity strategy for these analyses, the result should be interpreted in the context of the complete prespecified testing framework rather than as an isolated p-value.
PD-L1 CPS ≥1
Difference in overall response rate
95% CI: −3.3 to 9.2 · P = 0.1752
The estimated response-rate difference was 2.9 percentage points, with a 95% CI of −3.3 to 9.2 percentage points. The interval includes 0, and the reported p-value is 0.1752.
All participants
Difference in overall response rate
95% CI: −5.9 to 3.8 · P = 0.6629
In all participants, the estimated response-rate difference was −1.0 percentage point. The confidence interval extends from −5.9 to 3.8 percentage points. Thus, on this absolute response measure, the point estimate is slightly below zero, while the uncertainty interval includes both a modest reduction and a modest increase.
| Population | Risk difference | 95% CI | P-value |
|---|---|---|---|
| PD-L1 CPS ≥10 | 8.3 percentage points | −1.4 to 18.4 | 0.0457 |
| PD-L1 CPS ≥1 | 2.9 percentage points | −3.3 to 9.2 | 0.1752 |
| All participants | −1.0 percentage point | −5.9 to 3.8 | 0.6629 |
9. Secondary Results: Progression-Free Survival
Progression-free survival per RECIST 1.1 was analyzed using a Cox proportional-hazards model. The registry reports the same final-analysis time frame, through 11-April-2019.
| Population | HR | 95% CI | P-value |
|---|---|---|---|
| PD-L1 CPS ≥10 | 1.14 | 0.82–1.59 | 0.7936 |
| PD-L1 CPS ≥1 | 1.35 | 1.08–1.68 | 0.9964 |
| All participants | 1.60 | 1.33–1.92 | 1.0000 |
PD-L1 CPS ≥10
The HR of 1.14 means that the estimated instantaneous rate of progression or death was 14% higher in the pembrolizumab group under the fitted model. The 95% CI is 0.82–1.59, which includes 1.
PD-L1 CPS ≥1
The HR of 1.35 corresponds to an estimated instantaneous progression-or-death hazard 35% higher in the pembrolizumab group under the Cox model. The 95% CI of 1.08–1.68 lies above 1.
All participants
The HR of 1.60 corresponds to an estimated instantaneous progression-or-death hazard 60% higher in the pembrolizumab group under the fitted model. The 95% CI is 1.33–1.92.
The PFS estimates are not interchangeable with the primary OS estimates. PFS measures time to progression or death, whereas the primary endpoint here was overall survival. A treatment can have different effects on these two time-to-event endpoints because they represent different event definitions and follow-up processes.
The registry's PFS estimates are also a reminder that the direction of a hazard ratio matters. An HR above 1 means the estimated event hazard is higher in the pembrolizumab group for the event definition being analyzed. It does not mean that 1.14, 1.35, or 1.60 represents a corresponding percentage difference in the probability of an event by a fixed time.
10. Secondary Results: Disease Control Rate
Disease control rate per RECIST 1.1 was also analyzed using the Miettinen & Nurminen method, with the difference in percentages as the effect measure.
| Population | Risk difference | 95% CI | P-value |
|---|---|---|---|
| PD-L1 CPS ≥10 | 2.3 percentage points | −8.7 to 13.5 | 0.3388 |
| PD-L1 CPS ≥1 | −1.6 percentage points | −8.6 to 5.5 | 0.6701 |
| All participants | −6.5 percentage points | −12.2 to −0.8 | 0.9877 |
The CPS ≥10 analysis estimated a 2.3-percentage-point difference, while the CPS ≥1 analysis estimated −1.6 percentage points and the all-participant analysis estimated −6.5 percentage points. The confidence intervals quantify substantial uncertainty around these differences.
A risk difference describes an absolute difference in the proportion meeting a categorical response criterion. A hazard ratio describes a relative difference in the instantaneous rate of a time-to-event outcome. These measures operate on different statistical scales and should not be compared numerically as though they were interchangeable.
11. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment exposure group. They do not provide a complete adverse-event table or rates for every safety category, so this section is limited to the serious adverse-event counts reported.
| Safety group | Affected / at risk |
|---|---|
| Pembrolizumab First Course | 65 / 309 |
| Chemotherapy | 60 / 292 |
| Pembrolizumab Second Course | 1 / 8 |
These figures show the number affected relative to the number at risk in each reported exposure category. The ClinicalTrials.gov record does not establish that these categories correspond to the randomized treatment groups in a simple one-to-one fashion, particularly because a separate pembrolizumab second-course category is reported. For that reason, the safety figures should not be converted into an unqualified randomized-arm comparison beyond the wording provided.
12. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
Overall survival and progression-free survival are time-to-event endpoints. The Cox model is designed for this setting because it compares event hazards while incorporating different follow-up times and censoring. The registry specifically reports Cox regression for the primary OS analyses and secondary PFS analyses.
What does an HR of 0.78 mean?
An HR of 0.78 means that the estimated instantaneous event hazard in the pembrolizumab group was 78% of the chemotherapy-group hazard under the fitted Cox model. It can be expressed as approximately a 22% lower estimated hazard. It is not the same as saying that 22% fewer participants died or that survival probability increased by 22 percentage points.
Why does the confidence interval matter?
The point estimate is only one estimate from the observed trial data. The 95% confidence interval describes the uncertainty around that estimate under the statistical model and sampling framework. For example, the CPS ≥10 OS estimate is 0.78 with a 95% CI of 0.57–1.06, which shows considerably more uncertainty than the single number 0.78 conveys.
Why is a risk difference used for response rate?
Overall response rate is a categorical endpoint: participants either meet the response definition or they do not. A risk difference directly compares the response proportions between treatment groups. The registry reports the difference in percentages and the Miettinen & Nurminen method for its confidence interval.
Why can an HR be above 1?
An HR above 1 means that the estimated instantaneous event rate is higher in the pembrolizumab group for the specified event. For PFS, the all-participant estimate is 1.60, meaning the fitted model estimates a 60% higher instantaneous rate of progression or death in the pembrolizumab group relative to chemotherapy. This does not translate directly into a 60-percentage-point difference in the probability of progression or death.
Why should the p-value not be treated as the effect size?
A p-value is a property of a statistical test under a specified null hypothesis and analysis framework. It is influenced by the observed data and the amount of information available. The magnitude and clinical meaning of an effect are better described using the effect estimate and its confidence interval, alongside absolute measures where available.
Why does the analysis population matter?
The three primary OS analyses use different populations: CPS ≥10, CPS ≥1, and all participants. A treatment effect estimated in one population is not automatically the treatment effect in another. Differences among the three estimates can be described, but a formal statement that the treatment effect differs by PD-L1 population requires an appropriate statistical comparison.
13. Interpreting the Hazard Ratios as a Family
The visual makes the relative position of the three point estimates easy to see, but it should not be mistaken for a forest plot. The confidence intervals are the critical complement to the point estimates, and no graphical inference about subgroup differences should be drawn merely from the relative location of the three HRs.
| Population | Point estimate | Lower 95% CI | Upper 95% CI | What the interval tells us |
|---|---|---|---|---|
| PD-L1 CPS ≥10 | 0.78 | 0.57 | 1.06 | Includes 1; uncertainty spans lower and modestly higher hazard. |
| PD-L1 CPS ≥1 | 0.86 | 0.69 | 1.06 | Includes 1; uncertainty spans lower and modestly higher hazard. |
| All participants | 0.97 | 0.82 | 1.15 | Includes 1; uncertainty includes both lower and higher hazard. |
14. Primary Results: What the Numbers Do — and Do Not — Establish
What the OS estimates establish descriptively
The three reported point estimates are 0.78, 0.86, and 0.97 for CPS ≥10, CPS ≥1, and all participants, respectively.
What the confidence intervals add
Each 95% confidence interval includes 1, so the point estimates should not be interpreted without acknowledging the uncertainty represented by their intervals.
What the P-values add
The reported two-sided p-values are 0.0574, 0.0728, and 0.3802. They quantify evidence under the specified hypothesis tests rather than treatment magnitude.
What is not reported
The ClinicalTrials.gov record does not report median OS, Kaplan-Meier survival probabilities at specific times, event counts for the primary OS analyses, or subgroup forest plots.
One important statistical lesson is that an effect estimate and its hypothesis-test result answer different questions. The HR describes the estimated relative event hazard. The confidence interval describes uncertainty around that estimate. The p-value describes evidence against the null hypothesis under the specified test. None of the three should be substituted for the others.
15. Multiplicity and Multiple Primary Endpoints
KEYNOTE-119 has three registered primary endpoints, all based on overall survival but defined in different populations: PD-L1 CPS ≥10, PD-L1 CPS ≥1, and all participants.
| Primary comparison | Population | Effect measure | Hypothesis type |
|---|---|---|---|
| Overall survival | PD-L1 CPS ≥10 | Hazard ratio | Superiority |
| Overall survival | PD-L1 CPS ≥1 | Hazard ratio | Superiority |
| Overall survival | All participants | Hazard ratio | Superiority |
Because there are multiple primary comparisons, interpretation of individual p-values depends on the trial's prespecified multiplicity strategy. The ClinicalTrials.gov record does not provide an alpha-allocation or multiplicity procedure. Consequently, this page reports the three p-values exactly as registered but does not assign them a familywise-error interpretation that is not supported by the ClinicalTrials.gov record.
16. Interim Analysis, Missing Data, Crossover, and Bayesian Methods
The ClinicalTrials.gov record does not report an interim-analysis procedure, alpha-spending method, missing-data or imputation strategy, crossover rules, or Bayesian analysis. These design topics are therefore not presented as trial-specific methods.
This distinction is important for reproducibility. A statistical analysis page should identify what the registry actually documents rather than infer a complete statistical analysis plan from the endpoint type alone.
17. Limitations
- Registry-level detail: the ClinicalTrials.gov record identifies the principal statistical methods but do not provide the complete statistical analysis plan.
- Multiple primary comparisons: three primary OS analyses are reported, but the ClinicalTrials.gov record does not specify how multiplicity was controlled.
- Confidence-interval overlap with the null: all three primary OS 95% confidence intervals include 1, so the point estimates should be interpreted together with their uncertainty.
- No median survival values: median OS is not reported in the ClinicalTrials.gov record and therefore cannot be used to supplement the hazard-ratio interpretation.
- No time-specific survival probabilities: the ClinicalTrials.gov record does not report Kaplan-Meier survival probabilities at specific time points.
- Proportional-hazards assumption: Cox-model interpretation relies on model assumptions, but the ClinicalTrials.gov record does not provide diagnostics for proportional hazards.
- Secondary endpoints: response rate, PFS, and disease-control analyses are secondary endpoints and should be interpreted in the context of the overall testing framework.
- Open-label design: the registry specifies no masking. This can matter particularly for outcomes affected by assessment or treatment behavior, although overall survival itself is a death endpoint.
- Safety denominators: the reported serious-adverse-event categories include first-course and second-course pembrolizumab exposure, so they should not be treated as a simple two-arm randomized safety table.
- Subgroup comparisons: the ClinicalTrials.gov record does not report a formal interaction analysis establishing that treatment effects differ across PD-L1 populations.
18. Why This Trial Matters Statistically
KEYNOTE-119 is a useful teaching case because the registry combines randomized treatment comparison with several related endpoint populations and two different classes of statistical effect measures: hazard ratios for time-to-event outcomes and risk differences for categorical response outcomes.
| Statistical concept | How it appears in KEYNOTE-119 |
|---|---|
| Randomization | The trial uses randomized allocation in a parallel-group phase 3 design. |
| Time-to-event analysis | Overall survival and progression-free survival are analyzed as time-to-event outcomes. |
| Cox proportional-hazards model | Used for all registry-reported OS and PFS formal analyses. |
| Hazard ratio | Used to quantify the relative treatment effect for OS and PFS. |
| Confidence interval | 95% two-sided intervals quantify uncertainty around the hazard-ratio and risk-difference estimates. |
| Risk difference | Used for overall response rate and disease control rate comparisons. |
| Score-based CI | The Miettinen & Nurminen method is reported for categorical response and disease-control outcomes. |
| Multiple primary endpoints | Three primary OS comparisons are defined in different PD-L1 populations. |
| Open-label design | The registry specifies no masking. |
| Analysis populations | The primary OS analyses distinguish CPS ≥10, CPS ≥1, and all participants. |
Why the combination of endpoint types is instructive
It is tempting to place all trial results on a single scale, but the statistical questions are different. The OS and PFS analyses ask about time until an event, while ORR and disease control rate ask about whether a participant meets a categorical outcome definition. A hazard ratio and a risk difference therefore cannot be compared by magnitude alone.
The trial also illustrates why biomarker-defined analysis populations require careful interpretation. The CPS ≥10, CPS ≥1, and all-participant analyses use progressively different populations, and the corresponding OS estimates are not identical. A statistically rigorous interpretation describes those differences without automatically attributing them to biological treatment-effect modification.
19. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary OS hazard-ratio estimates were 0.78, 0.86, and 0.97 in the CPS ≥10, CPS ≥1, and all-participant populations, respectively. Their two-sided 95% confidence intervals were 0.57–1.06, 0.69–1.06, and 0.82–1.15.
Clinical interpretation
The registry results describe relative time-to-event effects and categorical response differences. The ClinicalTrials.gov record does not include enough information to characterize median survival, long-term survival probabilities, or the complete balance of clinical benefits and harms.
Keeping these interpretations separate is important. Statistical evidence describes the uncertainty and magnitude of the observed comparison. Clinical interpretation requires understanding the endpoint, population, treatment context, absolute effects, safety, and durability. Several of those components are not available in the ClinicalTrials.gov record.
20. A Practical Reading of the Primary Analysis
The primary outcome is overall survival: time from randomization to death due to any cause.
The registry reports three populations: CPS ≥10, CPS ≥1, and all participants.
The effect measure is the hazard ratio from a Cox proportional-hazards model.
Each estimate should be read together with its two-sided 95% confidence interval.
The p-value describes evidence under the specified superiority test; it is not an effect-size measure.
Multiple primary comparisons are present, the trial is open-label, and the ClinicalTrials.gov record does not specify multiplicity or interim-analysis procedures.
This six-step framework prevents a common analytical mistake: reading a single p-value first and only later asking what endpoint, population, effect measure, and statistical model generated it.
21. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: KEYNOTE-119, NCT02555657.
- PubMed: PMID 41038856.
- PubMed: PMID 40526219.
- PubMed: PMID 37976633.
- PubMed: PMID 33676601.
The numerical analyses on this page are restricted to the ClinicalTrials.gov record. The PubMed records are provided as linked publication records; no additional numerical results from those publications are incorporated into this analysis.
Continue with the statistical methods behind the trial
Explore the survival-analysis, confidence-interval, hypothesis-testing, and categorical-data methods used to interpret randomized clinical-trial results.
24. Record Summary
KEYNOTE-119 provides a useful example of how a randomized phase 3 trial can generate several related statistical questions from the same underlying treatment comparison. Its three primary endpoints are all overall-survival analyses, but they apply to different PD-L1-defined populations. The primary analyses use Cox proportional-hazards models and report hazard ratios with two-sided 95% confidence intervals and p-values. Secondary analyses extend the statistical framework to response rate and disease control rate using score-based confidence intervals for proportions, while progression-free survival is analyzed with Cox models.
The primary OS point estimates were 0.78 for CPS ≥10, 0.86 for CPS ≥1, and 0.97 for all participants. The corresponding 95% confidence intervals were 0.57–1.06, 0.69–1.06, and 0.82–1.15. Reading those estimates correctly requires keeping the hazard-ratio scale, confidence interval, p-value, analysis population, and multiplicity context distinct.
The secondary results reinforce the same statistical lesson. Response and disease-control outcomes are expressed as absolute percentage differences, whereas PFS is expressed as a hazard ratio. These effect measures answer different questions and should not be collapsed into a single numerical judgment.