← Clinical Trials
Metastatic Triple Negative Breast Cancer Phase 3 Completed NCT02555657

KEYNOTE-119: Complete Statistical Analysis of Pembrolizumab in Metastatic Triple Negative Breast Cancer

An independent statistical analysis of the randomized phase 3 KEYNOTE-119 trial comparing single-agent pembrolizumab with single-agent chemotherapy in participants with metastatic triple negative breast cancer, using the final-analysis results reported on ClinicalTrials.gov.

Trial start: 13-Oct-2015  ·  Primary completion: 11-Apr-2019  ·  Enrollment: 622
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record. Statistical explanations distinguish reported estimates from their interpretation.

1. Trial at a Glance

KEYNOTE-119 was a randomized, parallel-group, open-label phase 3 trial comparing single-agent pembrolizumab with single-agent chemotherapy in metastatic triple negative breast cancer. The registry reports three primary overall-survival analyses, defined in participants with PD-L1 CPS ≥10, participants with PD-L1 CPS ≥1, and all participants.

622
Enrolled
2 treatment arms
3
Primary endpoints
All overall survival
0.78
OS HR · CPS ≥10
95% CI 0.57–1.06
0.97
OS HR · All
95% CI 0.82–1.15
FeatureKEYNOTE-119
PhasePhase 3
ConditionMetastatic Triple Negative Breast Cancer
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment622
Arms2
Primary endpoints3 registered primary endpoints, all overall survival
Results postedYes
Statistical analyses posted12
ClinicalTrials.govNCT02555657

2. Clinical Question

The central question was whether single-agent pembrolizumab produced a superior overall-survival outcome compared with single-agent chemotherapy in participants with metastatic triple negative breast cancer. The registered primary hypothesis type was superiority.

Population

Participants with metastatic triple negative breast cancer enrolled in the randomized phase 3 trial.

Intervention

Single-agent pembrolizumab.

Comparator

Single-agent chemotherapy. The registry lists capecitabine, eribulin, gemcitabine, and vinorelbine as chemotherapy interventions.

Primary question

Does pembrolizumab improve overall survival relative to chemotherapy under the prespecified superiority comparisons?

3. Trial Design

01
Randomize 622 participants
02
Two arms Pembrolizumab vs chemotherapy
03
Open-label No masking
04
Follow-up Overall survival and response outcomes
05
Final analysis Database cutoff 11-Apr-2019
ARM 1

Pembrolizumab

  • Single-agent pembrolizumab
  • Biological intervention
  • Compared with single-agent chemotherapy
ARM 2

Chemotherapy

  • Single-agent chemotherapy
  • Drug options listed in the registry include capecitabine, eribulin, gemcitabine, and vinorelbine
  • Compared with single-agent pembrolizumab

The trial was randomized and used a parallel design, but it was not masked. This matters statistically because randomization establishes the framework for the between-group efficacy comparison, while the absence of masking can matter for outcomes that depend on assessment or treatment behavior. The primary endpoints reported here are overall-survival endpoints, which are less directly dependent on subjective outcome assessment than some other clinical outcomes.

Trial timeline

13-Oct-2015

Trial start

The registry lists 13-Oct-2015 as the study start date.

11-Apr-2019

Primary completion

The registry lists 11-Apr-2019 as the primary completion date.

11-Apr-2019

Final-analysis database cutoff

The primary endpoint time frames specify follow-up through the final-analysis database cutoff date of 11-April-2019.

4. Endpoints

The registry defines overall survival as the time from randomization to death due to any cause. All three registered primary endpoints are time-to-event outcomes with the same final-analysis time frame.

Endpoint Registry definition / time frame Analysis
Overall Survival in Participants With PD-L1 CPS ≥10 Overall survival was defined as the time from randomization to death due to any cause. Up to approximately 36 months, through the Final Analysis database cutoff date of 11-April-2019. Cox proportional-hazards model
Overall Survival in Participants With PD-L1 CPS ≥1 Overall survival was defined as the time from randomization to death due to any cause. Up to approximately 36 months, through the Final Analysis database cutoff date of 11-April-2019. Cox proportional-hazards model
Overall Survival in All Participants Overall survival was defined as the time from randomization to death due to any cause. Up to approximately 36 months, through the Final Analysis database cutoff date of 11-April-2019. Cox proportional-hazards model

Secondary endpoints with formal analyses

The registry also reports overall response rate, progression-free survival, and disease control rate analyses in the same three population definitions: PD-L1 CPS ≥10, PD-L1 CPS ≥1, and all participants.

Secondary endpoint familyPopulationEffect measureMethod
Overall Response Rate per RECIST 1.1PD-L1 CPS ≥10Risk differenceScore-based CI for proportions
Overall Response Rate per RECIST 1.1PD-L1 CPS ≥1Risk differenceScore-based CI for proportions
Overall Response Rate per RECIST 1.1All participantsRisk differenceScore-based CI for proportions
Progression-Free Survival per RECIST 1.1PD-L1 CPS ≥10Hazard ratioCox proportional-hazards model
Progression-Free Survival per RECIST 1.1PD-L1 CPS ≥1Hazard ratioCox proportional-hazards model
Progression-Free Survival per RECIST 1.1All participantsHazard ratioCox proportional-hazards model
Disease Control Rate per RECIST 1.1PD-L1 CPS ≥10Risk differenceScore-based CI for proportions
Disease Control Rate per RECIST 1.1PD-L1 CPS ≥1Risk differenceScore-based CI for proportions
Disease Control Rate per RECIST 1.1All participantsRisk differenceScore-based CI for proportions

5. Statistical Methodology

Cox proportional-hazards model

The registry reports a Cox proportional-hazards model for all six time-to-event analyses reported here: the three primary overall-survival comparisons and the three progression-free-survival comparisons. The effect measure is a hazard ratio comparing pembrolizumab with chemotherapy.

Hazard-ratio interpretation
HR = estimated hazard in pembrolizumab group / estimated hazard in chemotherapy group

An HR below 1 corresponds to a lower estimated instantaneous event rate in the pembrolizumab group under the fitted model. An HR above 1 corresponds to a higher estimated instantaneous event rate.

A hazard ratio is a relative time-to-event measure. It does not directly give the probability of death by a particular time, the median survival, or the proportion of participants who benefit. Those quantities require additional information that is not contained in the statistical analyses posted on ClinicalTrials.gov.

Score-based confidence intervals for proportions

The registry reports the Miettinen & Nurminen method for the overall response rate and disease control rate comparisons. This is a score-based confidence-interval approach for differences in proportions, related to the Newcombe and Wilson methods.

Risk-difference interpretation
Risk difference = response proportion in pembrolizumab group − response proportion in chemotherapy group

A positive risk difference indicates a higher observed proportion in the pembrolizumab group; a negative risk difference indicates a lower observed proportion in that group.

Superiority framework

Each posted statistical analysis is labeled as a superiority hypothesis. This means the inferential question is whether the randomized treatment comparison provides evidence of a difference favoring the experimental strategy under the prespecified statistical framework. It is not a non-inferiority design, so there is no non-inferiority margin to interpret in the ClinicalTrials.gov record.

Analysis populations

The primary overall-survival analyses were performed in the following populations: all participants with PD-L1 CPS ≥10 who were included in a treatment group at randomization; all participants with PD-L1 CPS ≥1 who were included in a treatment group at randomization; and all participants who were included in a treatment group at randomization. These definitions are important because the CPS-restricted analyses are not simply smaller versions of the all-participant analysis: they answer questions in specifically defined biomarker populations.

What the registry does not specify here

The ClinicalTrials.gov record identifies the Cox model and score-based proportion method, but do not provide details about Kaplan-Meier estimation, log-rank testing, covariate adjustment beyond the reported Cox method, proportional-hazards diagnostics, missing-data imputation, interim-analysis boundaries, alpha spending, Bayesian methods, or multiplicity procedures. Those topics are therefore not assigned trial-specific procedures on this page.

Methodological boundary: the absence of a registry-reported design detail is not evidence that the trial did not use that procedure. It means only that the provided ClinicalTrials.gov data do not establish it, so this page does not attribute an unreported method to KEYNOTE-119.

6. Results: Primary Overall Survival Endpoints

ClinicalTrials.gov reports formal statistical analyses for all three registered primary endpoints. Each comparison uses the Cox proportional-hazards model, reports a two-sided 95% confidence interval, and tests a superiority hypothesis.

Overall Survival in Participants With PD-L1 CPS ≥10

Hazard ratio for death

0.78

95% CI: 0.57–1.06   ·   P = 0.0574

Analysis population: all participants with PD-L1 CPS ≥10 who were included in a treatment group at randomization.

The estimated hazard ratio of 0.78 corresponds to an estimated instantaneous hazard in the pembrolizumab group that was 78% of the chemotherapy-group hazard under the fitted Cox model. Expressed as a relative reduction, this corresponds to approximately a 22% lower estimated hazard under the model.

Clinical Biostats interpretation

What the estimate means: The point estimate, HR 0.78, is below 1 and therefore points toward a lower estimated hazard of death with pembrolizumab than with chemotherapy in the PD-L1 CPS ≥10 analysis population.

What it does not mean: It does not mean that 22% of participants benefited, that each participant had exactly a 22% reduction in probability of death, or that survival probability was 22 percentage points higher.

Precision: The two-sided 95% CI is 0.57–1.06. It spans 1, so the estimate is compatible with a range extending from a substantially lower hazard to a modestly higher hazard under the model.

The p-value: P = 0.0574 measures the strength of evidence against the null hypothesis in the specified statistical test; it is not a measure of the size or clinical importance of the hazard ratio. A p-value should not be read as the probability that the treatment has no effect.

Caution: Interpretation of a Cox HR depends on the model and its proportional-hazards assumption. The ClinicalTrials.gov record does not report a diagnostic assessment of that assumption.

Overall Survival in Participants With PD-L1 CPS ≥1

Hazard ratio for death

0.86

95% CI: 0.69–1.06   ·   P = 0.0728

Analysis population: all participants with PD-L1 CPS ≥1 who were included in a treatment group at randomization.

The HR of 0.86 indicates that the estimated instantaneous hazard of death in the pembrolizumab group was 86% of the chemotherapy-group hazard under the Cox model, corresponding to approximately a 14% lower estimated hazard.

Clinical Biostats interpretation

What the estimate means: The point estimate is below 1, indicating a lower estimated hazard of death for pembrolizumab relative to chemotherapy in the CPS ≥1 analysis population.

What it does not mean: HR 0.86 is not an 86% survival probability and does not mean that 14% of patients were protected from death. It is a relative time-to-event estimate.

Precision: The 95% CI of 0.69–1.06 crosses 1. The data therefore permit a range of effects that includes no hazard difference and modestly higher hazard as well as lower hazard.

The p-value: P = 0.0728 is evidence from the specified hypothesis test, not a direct probability statement about the treatment effect and not a measure of effect magnitude.

Caution: The analysis is restricted to the CPS ≥1 population. It should not be silently treated as interchangeable with the CPS ≥10 or all-participant analysis.

Overall Survival in All Participants

Hazard ratio for death

0.97

95% CI: 0.82–1.15   ·   P = 0.3802

Analysis population: all participants who were included in a treatment group at randomization.

For the full analysis population, the estimated hazard ratio was 0.97. Under the Cox model, this corresponds to an estimated hazard that was 97% of the chemotherapy-group hazard, or approximately a 3% lower estimated hazard.

Clinical Biostats interpretation

What the estimate means: The all-participant point estimate is close to 1, indicating little estimated relative difference in the instantaneous hazard of death between the randomized groups under this model.

What it does not mean: HR 0.97 does not establish identical survival for every participant or prove that the two treatments have exactly the same effect. It is an estimate with uncertainty.

Precision: The 95% CI is 0.82–1.15. The interval includes 1 and permits both a lower and a higher hazard for pembrolizumab relative to chemotherapy.

The p-value: P = 0.3802 does not measure the probability that the treatment effect is zero, nor does it quantify clinical importance. It summarizes the result of the specified superiority hypothesis test.

Caution: No median survival estimates are reported in the ClinicalTrials.gov record, so a median-survival comparison should not be inferred from the hazard ratio alone.

7. Comparing the Three Primary Populations

Primary endpoint populationHR95% CIP-valueInterpretation of point estimate
PD-L1 CPS ≥100.780.57–1.060.0574Approximately 22% lower estimated hazard
PD-L1 CPS ≥10.860.69–1.060.0728Approximately 14% lower estimated hazard
All participants0.970.82–1.150.3802Approximately 3% lower estimated hazard

These three estimates illustrate why the definition of the analysis population matters. The point estimate moves from 0.78 in the CPS ≥10 population to 0.86 in the CPS ≥1 population and 0.97 in all participants. That pattern is descriptive of the reported estimates; it does not, by itself, establish that the treatment effect statistically differs across the three populations.

Do not rank the subgroups by hazard ratio alone. Comparing point estimates across nested or differently defined populations is not the same as testing treatment-effect heterogeneity. A formal claim that the effect differs between populations would require an appropriate interaction or heterogeneity analysis. Such an analysis is not reported in the ClinicalTrials.gov record.

8. Secondary Results: Overall Response Rate

The registry reports overall response rate per RECIST 1.1 as a secondary count/rate endpoint. The effect measure is the difference in percentages, that is, a risk difference. The reported method is the Miettinen & Nurminen method, a score-based approach for confidence intervals for proportions.

PD-L1 CPS ≥10

Difference in overall response rate

8.3 percentage points

95% CI: −1.4 to 18.4   ·   P = 0.0457

The estimated difference in response percentages was 8.3 percentage points in favor of pembrolizumab under the registry's group comparison. The 95% confidence interval ranges from −1.4 to 18.4 percentage points.

Clinical Biostats interpretation

The risk difference is an absolute rather than relative effect measure. An estimate of 8.3 means the observed response proportion in the pembrolizumab group exceeded that in the chemotherapy group by 8.3 percentage points in this analysis.

The confidence interval is relatively broad and crosses 0, ranging from −1.4 to 18.4 percentage points. Thus the interval includes a small difference in the opposite direction as well as substantially larger positive differences.

The p-value of 0.0457 summarizes the specified superiority test. It should not be interpreted as the probability that the 8.3-point effect is real, and it does not tell us how large the treatment effect is.

Because this is a secondary endpoint and the ClinicalTrials.gov record does not establish a multiplicity strategy for these analyses, the result should be interpreted in the context of the complete prespecified testing framework rather than as an isolated p-value.

PD-L1 CPS ≥1

Difference in overall response rate

2.9 percentage points

95% CI: −3.3 to 9.2   ·   P = 0.1752

The estimated response-rate difference was 2.9 percentage points, with a 95% CI of −3.3 to 9.2 percentage points. The interval includes 0, and the reported p-value is 0.1752.

All participants

Difference in overall response rate

−1.0 percentage point

95% CI: −5.9 to 3.8   ·   P = 0.6629

In all participants, the estimated response-rate difference was −1.0 percentage point. The confidence interval extends from −5.9 to 3.8 percentage points. Thus, on this absolute response measure, the point estimate is slightly below zero, while the uncertainty interval includes both a modest reduction and a modest increase.

PopulationRisk difference95% CIP-value
PD-L1 CPS ≥108.3 percentage points−1.4 to 18.40.0457
PD-L1 CPS ≥12.9 percentage points−3.3 to 9.20.1752
All participants−1.0 percentage point−5.9 to 3.80.6629

9. Secondary Results: Progression-Free Survival

Progression-free survival per RECIST 1.1 was analyzed using a Cox proportional-hazards model. The registry reports the same final-analysis time frame, through 11-April-2019.

PopulationHR95% CIP-value
PD-L1 CPS ≥101.140.82–1.590.7936
PD-L1 CPS ≥11.351.08–1.680.9964
All participants1.601.33–1.921.0000

PD-L1 CPS ≥10

The HR of 1.14 means that the estimated instantaneous rate of progression or death was 14% higher in the pembrolizumab group under the fitted model. The 95% CI is 0.82–1.59, which includes 1.

PD-L1 CPS ≥1

The HR of 1.35 corresponds to an estimated instantaneous progression-or-death hazard 35% higher in the pembrolizumab group under the Cox model. The 95% CI of 1.08–1.68 lies above 1.

All participants

The HR of 1.60 corresponds to an estimated instantaneous progression-or-death hazard 60% higher in the pembrolizumab group under the fitted model. The 95% CI is 1.33–1.92.

Important registry-data point: the ClinicalTrials.gov record reports P-values of 0.7936, 0.9964, and 1.0000 for these three PFS analyses. Those values are reproduced exactly. They should not be reverse-engineered or replaced with values derived from the confidence intervals.
How to read these PFS hazard ratios

The PFS estimates are not interchangeable with the primary OS estimates. PFS measures time to progression or death, whereas the primary endpoint here was overall survival. A treatment can have different effects on these two time-to-event endpoints because they represent different event definitions and follow-up processes.

The registry's PFS estimates are also a reminder that the direction of a hazard ratio matters. An HR above 1 means the estimated event hazard is higher in the pembrolizumab group for the event definition being analyzed. It does not mean that 1.14, 1.35, or 1.60 represents a corresponding percentage difference in the probability of an event by a fixed time.

10. Secondary Results: Disease Control Rate

Disease control rate per RECIST 1.1 was also analyzed using the Miettinen & Nurminen method, with the difference in percentages as the effect measure.

PopulationRisk difference95% CIP-value
PD-L1 CPS ≥102.3 percentage points−8.7 to 13.50.3388
PD-L1 CPS ≥1−1.6 percentage points−8.6 to 5.50.6701
All participants−6.5 percentage points−12.2 to −0.80.9877

The CPS ≥10 analysis estimated a 2.3-percentage-point difference, while the CPS ≥1 analysis estimated −1.6 percentage points and the all-participant analysis estimated −6.5 percentage points. The confidence intervals quantify substantial uncertainty around these differences.

Risk difference versus hazard ratio

A risk difference describes an absolute difference in the proportion meeting a categorical response criterion. A hazard ratio describes a relative difference in the instantaneous rate of a time-to-event outcome. These measures operate on different statistical scales and should not be compared numerically as though they were interchangeable.

11. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment exposure group. They do not provide a complete adverse-event table or rates for every safety category, so this section is limited to the serious adverse-event counts reported.

Safety groupAffected / at risk
Pembrolizumab First Course65 / 309
Chemotherapy60 / 292
Pembrolizumab Second Course1 / 8

These figures show the number affected relative to the number at risk in each reported exposure category. The ClinicalTrials.gov record does not establish that these categories correspond to the randomized treatment groups in a simple one-to-one fashion, particularly because a separate pembrolizumab second-course category is reported. For that reason, the safety figures should not be converted into an unqualified randomized-arm comparison beyond the wording provided.

Safety interpretation: the first-course and second-course pembrolizumab categories indicate that treatment exposure was represented separately in the registry safety reporting. A complete safety interpretation would require the underlying adverse-event definitions, denominators, timing rules, and event classifications, none of which are reported here.

12. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

Overall survival and progression-free survival are time-to-event endpoints. The Cox model is designed for this setting because it compares event hazards while incorporating different follow-up times and censoring. The registry specifically reports Cox regression for the primary OS analyses and secondary PFS analyses.

What does an HR of 0.78 mean?

An HR of 0.78 means that the estimated instantaneous event hazard in the pembrolizumab group was 78% of the chemotherapy-group hazard under the fitted Cox model. It can be expressed as approximately a 22% lower estimated hazard. It is not the same as saying that 22% fewer participants died or that survival probability increased by 22 percentage points.

Why does the confidence interval matter?

The point estimate is only one estimate from the observed trial data. The 95% confidence interval describes the uncertainty around that estimate under the statistical model and sampling framework. For example, the CPS ≥10 OS estimate is 0.78 with a 95% CI of 0.57–1.06, which shows considerably more uncertainty than the single number 0.78 conveys.

Why is a risk difference used for response rate?

Overall response rate is a categorical endpoint: participants either meet the response definition or they do not. A risk difference directly compares the response proportions between treatment groups. The registry reports the difference in percentages and the Miettinen & Nurminen method for its confidence interval.

Why can an HR be above 1?

An HR above 1 means that the estimated instantaneous event rate is higher in the pembrolizumab group for the specified event. For PFS, the all-participant estimate is 1.60, meaning the fitted model estimates a 60% higher instantaneous rate of progression or death in the pembrolizumab group relative to chemotherapy. This does not translate directly into a 60-percentage-point difference in the probability of progression or death.

Why should the p-value not be treated as the effect size?

A p-value is a property of a statistical test under a specified null hypothesis and analysis framework. It is influenced by the observed data and the amount of information available. The magnitude and clinical meaning of an effect are better described using the effect estimate and its confidence interval, alongside absolute measures where available.

Why does the analysis population matter?

The three primary OS analyses use different populations: CPS ≥10, CPS ≥1, and all participants. A treatment effect estimated in one population is not automatically the treatment effect in another. Differences among the three estimates can be described, but a formal statement that the treatment effect differs by PD-L1 population requires an appropriate statistical comparison.

13. Interpreting the Hazard Ratios as a Family

Primary overall-survival hazard ratios
CPS ≥10
0.78
CPS ≥1
0.86
All participants
0.97

The visual makes the relative position of the three point estimates easy to see, but it should not be mistaken for a forest plot. The confidence intervals are the critical complement to the point estimates, and no graphical inference about subgroup differences should be drawn merely from the relative location of the three HRs.

PopulationPoint estimateLower 95% CIUpper 95% CIWhat the interval tells us
PD-L1 CPS ≥100.780.571.06Includes 1; uncertainty spans lower and modestly higher hazard.
PD-L1 CPS ≥10.860.691.06Includes 1; uncertainty spans lower and modestly higher hazard.
All participants0.970.821.15Includes 1; uncertainty includes both lower and higher hazard.

14. Primary Results: What the Numbers Do — and Do Not — Establish

What the OS estimates establish descriptively

The three reported point estimates are 0.78, 0.86, and 0.97 for CPS ≥10, CPS ≥1, and all participants, respectively.

What the confidence intervals add

Each 95% confidence interval includes 1, so the point estimates should not be interpreted without acknowledging the uncertainty represented by their intervals.

What the P-values add

The reported two-sided p-values are 0.0574, 0.0728, and 0.3802. They quantify evidence under the specified hypothesis tests rather than treatment magnitude.

What is not reported

The ClinicalTrials.gov record does not report median OS, Kaplan-Meier survival probabilities at specific times, event counts for the primary OS analyses, or subgroup forest plots.

One important statistical lesson is that an effect estimate and its hypothesis-test result answer different questions. The HR describes the estimated relative event hazard. The confidence interval describes uncertainty around that estimate. The p-value describes evidence against the null hypothesis under the specified test. None of the three should be substituted for the others.

15. Multiplicity and Multiple Primary Endpoints

KEYNOTE-119 has three registered primary endpoints, all based on overall survival but defined in different populations: PD-L1 CPS ≥10, PD-L1 CPS ≥1, and all participants.

Primary comparisonPopulationEffect measureHypothesis type
Overall survivalPD-L1 CPS ≥10Hazard ratioSuperiority
Overall survivalPD-L1 CPS ≥1Hazard ratioSuperiority
Overall survivalAll participantsHazard ratioSuperiority

Because there are multiple primary comparisons, interpretation of individual p-values depends on the trial's prespecified multiplicity strategy. The ClinicalTrials.gov record does not provide an alpha-allocation or multiplicity procedure. Consequently, this page reports the three p-values exactly as registered but does not assign them a familywise-error interpretation that is not supported by the ClinicalTrials.gov record.

Why this matters: when several hypotheses are tested, treating every individual p-value as though it came from a single isolated confirmatory test can overstate the evidentiary strength. The correct interpretation depends on the prespecified hierarchy, alpha allocation, or other multiplicity-control method.

16. Interim Analysis, Missing Data, Crossover, and Bayesian Methods

The ClinicalTrials.gov record does not report an interim-analysis procedure, alpha-spending method, missing-data or imputation strategy, crossover rules, or Bayesian analysis. These design topics are therefore not presented as trial-specific methods.

Interim analysis
No specific interim-analysis method is reported in the ClinicalTrials.gov record.
Missing data / imputation
No specific missing-data or imputation method is reported.
Crossover
No crossover procedure is reported in the ClinicalTrials.gov record.
Bayesian methods
No Bayesian method is reported in the statistical analyses posted on ClinicalTrials.gov.

This distinction is important for reproducibility. A statistical analysis page should identify what the registry actually documents rather than infer a complete statistical analysis plan from the endpoint type alone.

17. Limitations

18. Why This Trial Matters Statistically

KEYNOTE-119 is a useful teaching case because the registry combines randomized treatment comparison with several related endpoint populations and two different classes of statistical effect measures: hazard ratios for time-to-event outcomes and risk differences for categorical response outcomes.

Statistical conceptHow it appears in KEYNOTE-119
RandomizationThe trial uses randomized allocation in a parallel-group phase 3 design.
Time-to-event analysisOverall survival and progression-free survival are analyzed as time-to-event outcomes.
Cox proportional-hazards modelUsed for all registry-reported OS and PFS formal analyses.
Hazard ratioUsed to quantify the relative treatment effect for OS and PFS.
Confidence interval95% two-sided intervals quantify uncertainty around the hazard-ratio and risk-difference estimates.
Risk differenceUsed for overall response rate and disease control rate comparisons.
Score-based CIThe Miettinen & Nurminen method is reported for categorical response and disease-control outcomes.
Multiple primary endpointsThree primary OS comparisons are defined in different PD-L1 populations.
Open-label designThe registry specifies no masking.
Analysis populationsThe primary OS analyses distinguish CPS ≥10, CPS ≥1, and all participants.

Why the combination of endpoint types is instructive

It is tempting to place all trial results on a single scale, but the statistical questions are different. The OS and PFS analyses ask about time until an event, while ORR and disease control rate ask about whether a participant meets a categorical outcome definition. A hazard ratio and a risk difference therefore cannot be compared by magnitude alone.

The trial also illustrates why biomarker-defined analysis populations require careful interpretation. The CPS ≥10, CPS ≥1, and all-participant analyses use progressively different populations, and the corresponding OS estimates are not identical. A statistically rigorous interpretation describes those differences without automatically attributing them to biological treatment-effect modification.

19. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary OS hazard-ratio estimates were 0.78, 0.86, and 0.97 in the CPS ≥10, CPS ≥1, and all-participant populations, respectively. Their two-sided 95% confidence intervals were 0.57–1.06, 0.69–1.06, and 0.82–1.15.

Clinical interpretation

The registry results describe relative time-to-event effects and categorical response differences. The ClinicalTrials.gov record does not include enough information to characterize median survival, long-term survival probabilities, or the complete balance of clinical benefits and harms.

Keeping these interpretations separate is important. Statistical evidence describes the uncertainty and magnitude of the observed comparison. Clinical interpretation requires understanding the endpoint, population, treatment context, absolute effects, safety, and durability. Several of those components are not available in the ClinicalTrials.gov record.

20. A Practical Reading of the Primary Analysis

Step 1 · Start with the estimand

The primary outcome is overall survival: time from randomization to death due to any cause.

Step 2 · Identify the analysis population

The registry reports three populations: CPS ≥10, CPS ≥1, and all participants.

Step 3 · Identify the effect measure

The effect measure is the hazard ratio from a Cox proportional-hazards model.

Step 4 · Read the uncertainty

Each estimate should be read together with its two-sided 95% confidence interval.

Step 5 · Read the hypothesis test

The p-value describes evidence under the specified superiority test; it is not an effect-size measure.

Step 6 · Check the design context

Multiple primary comparisons are present, the trial is open-label, and the ClinicalTrials.gov record does not specify multiplicity or interim-analysis procedures.

This six-step framework prevents a common analytical mistake: reading a single p-value first and only later asking what endpoint, population, effect measure, and statistical model generated it.

21. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

The numerical analyses on this page are restricted to the ClinicalTrials.gov record. The PubMed records are provided as linked publication records; no additional numerical results from those publications are incorporated into this analysis.

Continue with the statistical methods behind the trial

Explore the survival-analysis, confidence-interval, hypothesis-testing, and categorical-data methods used to interpret randomized clinical-trial results.

24. Record Summary

KEYNOTE-119 provides a useful example of how a randomized phase 3 trial can generate several related statistical questions from the same underlying treatment comparison. Its three primary endpoints are all overall-survival analyses, but they apply to different PD-L1-defined populations. The primary analyses use Cox proportional-hazards models and report hazard ratios with two-sided 95% confidence intervals and p-values. Secondary analyses extend the statistical framework to response rate and disease control rate using score-based confidence intervals for proportions, while progression-free survival is analyzed with Cox models.

The primary OS point estimates were 0.78 for CPS ≥10, 0.86 for CPS ≥1, and 0.97 for all participants. The corresponding 95% confidence intervals were 0.57–1.06, 0.69–1.06, and 0.82–1.15. Reading those estimates correctly requires keeping the hazard-ratio scale, confidence interval, p-value, analysis population, and multiplicity context distinct.

The secondary results reinforce the same statistical lesson. Response and disease-control outcomes are expressed as absolute percentage differences, whereas PFS is expressed as a hazard ratio. These effect measures answer different questions and should not be collapsed into a single numerical judgment.

Clinical Biostats methodology: A rigorous trial-results page should reconstruct the statistical structure of the reported evidence, distinguish point estimates from uncertainty, explain what each effect measure means, and avoid adding methods or numerical results that are not supported by the underlying record.