This page separates reported trial results from statistical interpretation. Numerical results are restricted to the statistical analyses posted for KEYNOTE-859 in the ClinicalTrials.gov record. Where the registry does not provide a particular result or design detail, it is not inferred from outside sources.
1. Trial at a Glance
KEYNOTE-859 was a randomized, parallel, double-masked phase 3 trial evaluating pembrolizumab plus chemotherapy versus placebo plus chemotherapy in participants with gastric or gastroesophageal junction adenocarcinoma. The registry reports 1,579 enrolled participants, three primary endpoints, and nine posted statistical analyses.
| Feature | KEYNOTE-859 |
|---|---|
| Phase | Phase 3 |
| Condition | Stomach Neoplasms |
| Trial name | KEYNOTE-859 |
| Design | Randomized, parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 1,579 |
| Primary endpoints | 3 |
| Outcome measures posted | 14 |
| Statistical analyses posted | 9 |
| Trial status | Completed |
| Start | 2018-11-08 |
| Primary completion | 2022-10-03 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03675737 |
2. Clinical Question
The central question was whether adding pembrolizumab to chemotherapy improved overall survival compared with placebo plus chemotherapy in participants with gastric or gastroesophageal junction adenocarcinoma. The registry prespecified three primary overall-survival endpoints covering the full randomized population and two PD-L1 Combined Positive Score populations.
Population
Participants with gastric or gastroesophageal junction adenocarcinoma, represented in the registry under the condition Stomach Neoplasms.
Intervention
Pembrolizumab plus chemotherapy. The registered interventions include pembrolizumab, cisplatin, 5-fluorouracil, oxaliplatin, and capecitabine.
Comparator
Placebo for pembrolizumab plus chemotherapy. The registry analyses identify the chemotherapy regimens as FP or CAPOX.
Primary question
Does pembrolizumab plus chemotherapy improve overall survival relative to placebo plus chemotherapy under the prespecified superiority framework?
3. Trial Design
Pembrolizumab combination
- Pembrolizumab
- Chemotherapy
- Registered chemotherapy interventions include cisplatin, 5-fluorouracil, oxaliplatin, and capecitabine
- Posted statistical analyses identify FP or CAPOX chemotherapy regimens
Control combination
- Placebo for pembrolizumab
- Chemotherapy
- Posted statistical analyses identify FP or CAPOX chemotherapy regimens
The ClinicalTrials.gov record identifies the allocation as randomized and the design model as parallel. The masking field is recorded as DOUBLE. The statistical analyses further show that geographic region, PD-L1 status where applicable, and chemotherapy regimen were incorporated into stratified analyses, with small strata collapsed.
4. Randomization, Stratification, and Analysis Populations
The posted primary and secondary survival analyses were not simple unadjusted comparisons. The Cox regression analyses incorporated treatment as a covariate and used stratification factors that reflect the trial's design and analysis framework.
| Analysis population | Definition reported in the registry |
|---|---|
| All-participant efficacy population | All randomized participants. |
| PD-L1 CPS ≥1 population | Randomized participants with a PD-L1 CPS of ≥1. |
| PD-L1 CPS ≥10 population | Randomized participants with a PD-L1 CPS of ≥10. |
Stratification in the survival models
For overall survival in all participants, the Cox model was stratified by geographic region, PD-L1 status (CPS <1 versus CPS ≥1), and chemotherapy regimen, with small strata collapsed. For the PD-L1 CPS ≥1 and CPS ≥10 analyses, the model was stratified by geographic region and chemotherapy regimen, again with small strata collapsed.
Stratification allows the treatment comparison to account for important categorical factors without requiring the analysis to assume that the baseline hazard is identical across those strata. It is particularly relevant here because the registry reports treatment effects within a randomized trial that used different chemotherapy regimens and evaluated PD-L1-defined populations.
The important distinction is that stratification does not mean the treatment effect is estimated separately and independently within every stratum. Instead, the Cox model estimates a common treatment hazard ratio while allowing the baseline hazard to differ across the specified strata.
5. Endpoints
The registry lists three primary endpoints, all based on overall survival. Six posted secondary analyses provide results for progression-free survival and objective response rate across the full population and PD-L1-defined populations.
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Overall Survival (OS) in All Participants | OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of the last follow-up. OS was estimated using the product-limit (Kaplan-Meier) method for censored data. Time frame: Up to 45.9 months. | Time-to-event |
| Overall Survival (OS) in Participants With PD-L1 CPS ≥1 | OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of the last follow-up. OS was estimated using the product-limit (Kaplan-Meier) method for censored data. Time frame: Up to 45.9 months. | Time-to-event |
| Overall Survival (OS) in Participants With PD-L1 CPS ≥10 | OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of the last follow-up. OS was estimated using the product-limit (Kaplan-Meier) method for censored data. Time frame: Up to 45.9 months. | Time-to-event |
| Progression Free Survival (PFS) per RECIST 1.1 assessed by BICR in all participants | Progression-free survival assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review. Time frame: Up to 49.5 months. | Time-to-event |
| PFS per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥1 | Progression-free survival assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥1. Time frame: Up to 49.5 months. | Time-to-event |
| PFS per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥10 | Progression-free survival assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥10. Time frame: Up to 49.5 months. | Time-to-event |
| Objective Response Rate (ORR) per RECIST 1.1 assessed by BICR in all participants | Objective response rate assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review. Time frame: Up to 49.5 months. | Binary |
| ORR per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥1 | Objective response rate assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥1. Time frame: Up to 49.5 months. | Binary |
| ORR per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥10 | Objective response rate assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥10. Time frame: Up to 49.5 months. | Binary |
6. Statistical Methodology
Kaplan-Meier estimation
The registry explicitly states that overall survival was estimated using the product-limit (Kaplan-Meier) method for censored data. This is appropriate for time-to-event outcomes because not every participant necessarily experiences death during the observation period.
Here, di represents events at event time ti, while ni represents participants at risk immediately before that time. Censored participants contribute information until their censoring time.
Log-rank testing
All nine posted statistical analyses used either a log-rank test for time-to-event outcomes or the Miettinen & Nurminen method for binary response outcomes. The six PFS analyses and three OS analyses used the log-rank test.
The log-rank test compares the observed and expected numbers of events between treatment groups over follow-up. It is fundamentally a test of the survival experience rather than a comparison of two isolated proportions.
Stratified Cox regression
The reported hazard ratios were based on Cox regression models with Efron's method of tie handling. Treatment was included as a covariate, while the model was stratified by prespecified factors. This gives a model-based estimate of the relative event hazard associated with treatment.
An HR below 1 indicates a lower estimated instantaneous event hazard in the pembrolizumab combination group under the fitted Cox model. It is not a percentage of patients who benefit and is not the same quantity as a risk ratio.
Miettinen & Nurminen method
The three ORR analyses used the Miettinen & Nurminen method, reported in the normalized methods field as a score-based confidence interval approach for proportions. The treatment effect was expressed as a risk difference, or difference in response percentages between groups.
A positive risk difference means that the response proportion was higher in the pembrolizumab combination group. Unlike a hazard ratio, a risk difference is expressed on an absolute percentage-point scale.
Superiority framework
All nine posted analyses identify the hypothesis type as superiority. The question is therefore whether the treatment groups differ in the prespecified direction rather than whether a new treatment is merely no worse than a control by a predefined non-inferiority margin.
7. Results: Overall Survival in All Participants
The primary overall-survival analysis included all randomized participants. The comparison was pembrolizumab plus chemotherapy versus placebo plus chemotherapy, with a two-sided 95% confidence interval and a superiority hypothesis.
Hazard ratio for overall survival
95% CI: 0.70–0.87 · P < 0.0001
Time frame: Up to 45.9 months
| Primary OS analysis | Pembrolizumab + chemotherapy vs placebo + chemotherapy |
|---|---|
| Analysis population | All randomized participants |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.78 |
| 95% CI | 0.70–0.87 |
| P-value | <0.0001 |
| Model | Cox regression with Efron's method of tie handling; treatment as a covariate; stratified by geographic region, PD-L1 status (CPS <1 versus CPS ≥1), and chemotherapy regimen with small strata collapsed |
The estimated hazard ratio of 0.78 means that, under the fitted stratified Cox model, the estimated instantaneous rate of death in the pembrolizumab plus chemotherapy group was about 78% of the estimated rate in the placebo plus chemotherapy group. Equivalently, 0.78 corresponds to a 22% lower estimated hazard under that model.
This does not mean that 22% of participants avoided death, that every participant experienced a 22% reduction in risk, or that the absolute probability of death was reduced by 22 percentage points. The hazard ratio is a relative time-to-event measure.
The 95% confidence interval of 0.70–0.87 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not an interval containing the effects that individual participants experienced.
The P-value of <0.0001 addresses evidence against the relevant null hypothesis; it does not quantify the size or clinical importance of the treatment effect. The magnitude of the estimated effect is communicated by the hazard ratio and its confidence interval.
Because the estimate comes from a Cox model, its interpretation also depends on the model's assumptions. In particular, a single hazard ratio summarizes relative event rates over time and may be less descriptive if the proportional-hazards relationship does not adequately characterize the data.
8. Results: Overall Survival in Participants With PD-L1 CPS ≥1
The second primary endpoint restricted the analysis population to randomized participants with a PD-L1 CPS of ≥1. The registry again reports a stratified log-rank analysis and a Cox model-based hazard ratio.
Hazard ratio for overall survival, PD-L1 CPS ≥1
95% CI: 0.65–0.84 · P < 0.0001
Time frame: Up to 45.9 months
| Primary OS analysis | Result |
|---|---|
| Analysis population | Randomized participants with a PD-L1 CPS of ≥1 |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.74 |
| 95% CI | 0.65–0.84 |
| P-value | <0.0001 |
| Model | Cox regression with Efron's method of tie handling; treatment as a covariate; stratified by geographic region and chemotherapy regimen with small strata collapsed |
An HR of 0.74 indicates an estimated instantaneous death hazard approximately 74% as large in the pembrolizumab plus chemotherapy group as in the placebo plus chemotherapy group, conditional on the fitted model. As a simple relative interpretation, this corresponds to a 26% lower estimated hazard.
The estimate does not mean a 26-percentage-point improvement in survival, nor does it mean that exactly 26% of patients benefit. It also does not establish that the treatment effect is identical for every individual with PD-L1 CPS ≥1.
The 95% CI of 0.65–0.84 provides the precision of this estimated relative effect under the model. The relatively narrow interval compared with the estimate itself indicates that the posted analysis provides a more constrained statistical range than would be conveyed by the point estimate alone.
The P-value of <0.0001 is evidence against the null hypothesis used in the superiority analysis. It should not be interpreted as the probability that the null hypothesis is true or as a measure of how large the treatment effect is.
This analysis is also a population-restricted comparison. Its result describes randomized participants meeting the PD-L1 CPS ≥1 criterion; it should not automatically be generalized to the full trial population without considering the distinction between the analysis populations.
9. Results: Overall Survival in Participants With PD-L1 CPS ≥10
The third primary endpoint further restricted the analysis population to randomized participants with a PD-L1 CPS of ≥10.
Hazard ratio for overall survival, PD-L1 CPS ≥10
95% CI: 0.53–0.79 · P < 0.0001
Time frame: Up to 45.9 months
| Primary OS analysis | Result |
|---|---|
| Analysis population | Randomized participants with a PD-L1 CPS of ≥10 |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.65 |
| 95% CI | 0.53–0.79 |
| P-value | <0.0001 |
| Model | Cox regression with Efron's method of tie handling; treatment as a covariate; stratified by geographic region and chemotherapy regimen with small strata collapsed |
An HR of 0.65 corresponds to an estimated death hazard about 65% of that in the comparator group under the fitted Cox model, or approximately a 35% lower estimated hazard.
Again, this is not an absolute risk reduction and does not mean that 35% of participants avoided death. It is a model-based relative comparison of event hazards over the analyzed follow-up.
The 95% CI of 0.53–0.79 describes uncertainty around the treatment-effect estimate. It does not describe variation in treatment benefit from one patient to another.
The P-value of <0.0001 measures the evidence against the null hypothesis under the specified statistical framework. It does not indicate that the probability of the observed effect being due to chance is <0.0001, and it does not measure effect magnitude.
The CPS ≥10 analysis is especially important to interpret as a prespecified population-defined endpoint rather than as a post hoc comparison of whichever subgroup happened to show the smallest hazard ratio. The ClinicalTrials.gov record identifies it explicitly as one of the three primary endpoints.
10. Primary Overall-Survival Results Together
Putting the three primary analyses side by side makes the structure of the statistical evidence clearer. All three use hazard ratios below 1, with two-sided 95% confidence intervals and P-values reported as <0.0001.
| Primary endpoint | Population | HR | 95% CI | P-value |
|---|---|---|---|---|
| OS in all participants | All randomized participants | 0.78 | 0.70–0.87 | <0.0001 |
| OS in PD-L1 CPS ≥1 | Randomized participants with CPS ≥1 | 0.74 | 0.65–0.84 | <0.0001 |
| OS in PD-L1 CPS ≥10 | Randomized participants with CPS ≥10 | 0.65 | 0.53–0.79 | <0.0001 |
The three point estimates are 0.78, 0.74, and 0.65. These values describe the estimated relative treatment effects in three different analysis populations. The CPS ≥10 estimate is numerically lower than the all-participant estimate, but a smaller point estimate in one population does not by itself establish that the treatment effect is statistically different between populations.
Formal comparison of treatment effects across subgroups generally requires an interaction or treatment-by-subgroup analysis. The ClinicalTrials.gov record does not report such an interaction test, so the three hazard ratios should be read as separate prespecified endpoint results rather than as evidence that the treatment effect necessarily changes according to PD-L1 CPS.
11. Secondary Results: Progression-Free Survival
Six secondary time-to-event analyses are posted for PFS: three in the full randomized population and three in PD-L1-defined populations. Each uses the log-rank test with a Cox regression model for the hazard ratio.
| PFS analysis | Population | HR | 95% CI | P-value |
|---|---|---|---|---|
| PFS per RECIST 1.1, BICR | All randomized participants | 0.76 | 0.67–0.85 | <0.0001 |
| PFS per RECIST 1.1, BICR | Randomized participants with PD-L1 CPS ≥1 | 0.72 | 0.63–0.82 | <0.0001 |
| PFS per RECIST 1.1, BICR | Randomized participants with PD-L1 CPS ≥10 | 0.62 | 0.51–0.76 | <0.0001 |
All three PFS analyses have the time frame up to 49.5 months. The full-population analysis was stratified by geographic region, PD-L1 status (CPS <1 versus CPS ≥1), and chemotherapy regimen with small strata collapsed. The CPS ≥1 and CPS ≥10 analyses were stratified by geographic region and chemotherapy regimen with small strata collapsed.
The PFS HR of 0.76 in all randomized participants corresponds to a 24% lower estimated hazard of progression or death under the fitted model. The CPS ≥1 and CPS ≥10 estimates of 0.72 and 0.62 correspond to 28% and 38% lower estimated hazards, respectively.
These are relative time-to-event effects. They do not tell us the absolute difference in the probability of remaining progression-free at any particular time, and they do not establish how long an individual participant will remain progression-free.
The confidence intervals are important because each point estimate has sampling uncertainty. The P-values are evidence against the corresponding null hypotheses, but they do not quantify the size of the PFS benefit.
Because PFS incorporates both progression and death, its clinical interpretation is different from OS. A PFS hazard ratio and an OS hazard ratio should not be treated as interchangeable measures of the same event.
12. Secondary Results: Objective Response Rate
Objective response rate was analyzed as a binary endpoint using the Miettinen & Nurminen method. The effect measure was a difference in percentage, normalized here as a risk difference. The registry reports two-sided 95% confidence intervals.
| ORR analysis | Population | Risk difference | 95% CI | P-value |
|---|---|---|---|---|
| ORR per RECIST 1.1, BICR | All randomized participants | 9.3 | 4.4–14.1 | 0.00009 |
| ORR per RECIST 1.1, BICR | Randomized participants with PD-L1 CPS ≥1 | 9.5 | 3.9–15.0 | 0.00041 |
| ORR per RECIST 1.1, BICR | Randomized participants with PD-L1 CPS ≥10 | 17.5 | 9.3–25.5 | 0.00002 |
Largest posted ORR difference
95% CI: 9.3–25.5 · P = 0.00002
Participants with PD-L1 CPS ≥10; risk difference in percentage points.
A risk difference of 9.3 means that the response percentage in the pembrolizumab plus chemotherapy group exceeded that in the placebo plus chemotherapy group by an estimated 9.3 percentage points in the full randomized population.
Similarly, the estimates of 9.5 and 17.5 describe absolute differences in response percentage for the CPS ≥1 and CPS ≥10 populations. These are not relative risks and should not be described as percentage reductions in the probability of nonresponse.
The 95% confidence interval gives the statistical precision of each estimated difference. For the full population, the interval is 4.4–14.1; for CPS ≥1, it is 3.9–15.0; and for CPS ≥10, it is 9.3–25.5.
The P-values indicate evidence against the relevant null hypothesis of no treatment difference in response under the specified analysis. They do not tell us whether an observed response difference is clinically important, nor do they describe the probability that the treatment has no effect.
13. Statistical Methods Explained
Why was a log-rank test used for overall survival and progression-free survival?
OS and PFS are time-to-event endpoints, meaning both the event time and the fact that some participants may be censored are part of the analysis. A log-rank test compares the event experience between treatment groups over the observed follow-up rather than reducing the outcome to a single binary proportion.
What does an HR of 0.78 mean?
Under the fitted Cox model, an HR of 0.78 means the estimated instantaneous event hazard in the pembrolizumab combination group is 78% of that in the comparator group. The complementary interpretation is a 22% lower estimated hazard. It does not mean a 22% absolute reduction in the probability of death.
Why are the CPS ≥1 and CPS ≥10 analyses different from the all-participant analysis?
They use different analysis populations. The all-participant analysis includes all randomized participants, whereas the CPS analyses restrict the population according to the registered PD-L1 CPS threshold. Because the populations differ, the resulting hazard ratios estimate treatment effects in different groups.
Why does the Cox model use stratification?
The registry reports stratification by geographic region and chemotherapy regimen for the PD-L1-defined analyses, with additional PD-L1 status stratification for the full-population OS and PFS analyses. Stratification permits the underlying baseline hazard to differ across these strata while estimating a treatment effect across them.
What does a risk difference of 17.5 mean?
The ORR risk difference of 17.5 represents an estimated 17.5-percentage-point difference in objective response between the randomized treatment groups in participants with PD-L1 CPS ≥10. It is an absolute measure, unlike the hazard ratios used for OS and PFS.
Why use the Miettinen & Nurminen method for ORR?
ORR is a binary endpoint: each participant is classified according to whether the prespecified response criterion was met. The Miettinen & Nurminen method provides a score-based approach to estimating uncertainty around the difference between two proportions, matching the registry's reported statistical method.
What does the P-value add when the confidence interval is already reported?
The P-value and confidence interval answer related but different questions. The P-value quantifies the compatibility of the observed data with the specified null hypothesis, whereas the confidence interval communicates the estimated effect and its statistical precision. Neither one replaces the effect estimate itself.
14. Understanding the Three Primary Endpoints
The three primary endpoints are all versions of the same underlying outcome, overall survival, but they answer the question in different analysis populations.
| Endpoint | Population | HR | Interpretive focus |
|---|---|---|---|
| OS in all participants | All randomized participants | 0.78 | Overall randomized treatment comparison |
| OS in CPS ≥1 | Randomized participants with CPS ≥1 | 0.74 | Treatment comparison in the CPS ≥1 population |
| OS in CPS ≥10 | Randomized participants with CPS ≥10 | 0.65 | Treatment comparison in the CPS ≥10 population |
This structure illustrates an important statistical principle: the population is part of the estimand. An effect estimate is not fully described by its numerical value alone. The endpoint definition, analysis population, treatment contrast, follow-up window, and statistical model all determine what the estimate represents.
15. Censoring and Time-to-Event Interpretation
The registry definition of OS explicitly states that participants without documented death at the time of analysis were censored at the date of the last follow-up. This is central to understanding why Kaplan-Meier and Cox methods are used rather than ordinary comparisons of means.
Event
For OS, the event is death due to any cause.
Censoring
A participant without documented death at the analysis time contributes information through the last follow-up date.
Kaplan-Meier
Estimates the survival function while accounting for censored observations.
Cox model
Estimates a relative hazard while incorporating follow-up time and censoring.
Censoring does not mean that a censored participant is treated as if the event never occurred. Rather, the analysis recognizes that the participant's event-free observation is known only up to the censoring time. The validity of standard survival analysis depends on assumptions about the censoring mechanism and the relationship between censoring and subsequent event risk.
16. Hazard Ratios: What They Do and Do Not Mean
The OS hazard ratios of 0.78, 0.74, and 0.65 all lie below 1. Under their respective Cox models, this indicates lower estimated death hazards in the pembrolizumab plus chemotherapy group than in the placebo plus chemotherapy group.
A hazard ratio is not a survival percentage, response percentage, risk difference, or median survival time. A value of 0.65 does not mean that 65% of participants survived or that survival probability was reduced by 35 percentage points.
The hazard ratio describes a treatment comparison at the population level under the fitted model. It does not imply that every participant experiences the same relative reduction in event hazard.
The confidence intervals communicate statistical precision. For example, the full-population OS estimate of 0.78 is accompanied by a 95% CI of 0.70–0.87. The interval is therefore an essential part of reporting the estimate rather than an optional add-on.
The P-values of <0.0001 for the three primary OS analyses indicate strong statistical evidence against their respective null hypotheses under the reported framework. They do not indicate the size of the treatment effect and should not be used as a substitute for the hazard ratio or its confidence interval.
17. Primary and Secondary Results: Statistical Overview
| Endpoint | Role | Type | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| OS, all participants | Primary | Time-to-event | HR | 0.78 | 0.70–0.87 | <0.0001 |
| OS, CPS ≥1 | Primary | Time-to-event | HR | 0.74 | 0.65–0.84 | <0.0001 |
| OS, CPS ≥10 | Primary | Time-to-event | HR | 0.65 | 0.53–0.79 | <0.0001 |
| PFS, all participants | Secondary | Time-to-event | HR | 0.76 | 0.67–0.85 | <0.0001 |
| PFS, CPS ≥1 | Secondary | Time-to-event | HR | 0.72 | 0.63–0.82 | <0.0001 |
| PFS, CPS ≥10 | Secondary | Time-to-event | HR | 0.62 | 0.51–0.76 | <0.0001 |
| ORR, all participants | Secondary | Binary | Risk difference | 9.3 | 4.4–14.1 | 0.00009 |
| ORR, CPS ≥1 | Secondary | Binary | Risk difference | 9.5 | 3.9–15.0 | 0.00041 |
| ORR, CPS ≥10 | Secondary | Binary | Risk difference | 17.5 | 9.3–25.5 | 0.00002 |
This table also illustrates why it is useful to report effect measures in their natural statistical scale. OS and PFS are summarized with hazard ratios because they are time-to-event outcomes, while ORR is summarized with a risk difference because it is binary.
18. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk. The reported figures are:
| Arm | Serious adverse events |
|---|---|
| Pembrolizumab + Chemotherapy (FP or CAPO) | 358/785 |
| Placebo + Chemotherapy (FP or CAPOX Regi) | 317/787 |
| Pembrolizumab Second Course | 1/12 |
The ClinicalTrials.gov record does not provide a further safety-analysis definition, confidence interval, hypothesis test, or comparative P-value for these serious-adverse-event counts. Accordingly, the safety information is reported descriptively rather than converted into an inferential treatment comparison.
Safety and efficacy answer different statistical questions. The serious-adverse-event counts describe the number of affected participants relative to the number at risk in the ClinicalTrials.gov record. They do not by themselves establish a causal treatment difference, particularly without a reported comparative analysis and its uncertainty.
It is also important not to combine the serious-adverse-event figures with the OS hazard ratio into a single numerical "benefit-risk" statistic. They arise from different outcome definitions, populations or exposure groupings, and statistical frameworks.
19. What Is and Is Not Reported in the Supplied Registry Data
The ClinicalTrials.gov record is sufficient to reconstruct the principal statistical analyses, but they do not contain every type of information that might appear in a full clinical-study report.
| Topic | What the ClinicalTrials.gov record provides |
|---|---|
| Primary efficacy | Three formal OS analyses with estimates, two-sided 95% CIs, and P-values. |
| Secondary efficacy | Six formal analyses covering PFS and ORR, with estimates, CIs, and P-values. |
| Statistical methods | Log-rank testing and Miettinen & Nurminen methodology, with Cox-model details for survival analyses. |
| Analysis populations | All randomized participants and randomized participants meeting PD-L1 CPS thresholds. |
| Baseline characteristics | Not reported in the ClinicalTrials.gov record. |
| Median OS / PFS | Not reported in the statistical analyses provided for this page. |
| Kaplan-Meier numerical time-point estimates | Not reported. |
| Subgroup hazard ratios beyond the registered CPS populations | Not reported. |
| Formal multiplicity procedure | Not reported. |
| Interim-analysis procedure | Not reported. |
| Missing-data or imputation procedure | Not reported. |
| Non-inferiority margin | Not applicable to the reported superiority analyses; no non-inferiority margin is reported. |
| Bayesian methods | Not reported in the statistical analyses posted on ClinicalTrials.gov. |
20. Limitations and Interpretation Issues
- Hazard-ratio interpretation: the OS and PFS estimates are model-based relative hazard measures, not absolute risk differences or individual-patient probabilities.
- Proportional-hazards consideration: the reported Cox hazard ratio summarizes relative event hazards over time. Its descriptive adequacy depends on the underlying hazard relationship represented by the model.
- Different analysis populations: the all-participant, CPS ≥1, and CPS ≥10 analyses estimate treatment effects in different populations and should not be treated as interchangeable.
- Subgroup comparison: numerical differences between the three CPS-specific hazard ratios do not establish treatment-effect heterogeneity without a formal interaction analysis. No such analysis is provided in the ClinicalTrials.gov record.
- Censoring: OS includes participants censored at their last follow-up when death was not documented. Survival-analysis validity therefore depends on the assumptions surrounding censoring.
- Multiplicity: the ClinicalTrials.gov record identifies three primary endpoints and multiple secondary analyses but do not provide the complete multiplicity-control procedure. The P-values should therefore be interpreted within the statistical framework actually reported rather than assuming an unreported adjustment.
- Missing-data methods: the ClinicalTrials.gov record does not report an imputation strategy. No missing-data procedure is inferred here.
- Limited safety inference: serious-adverse-event counts are provided descriptively, without a formal comparative analysis in the ClinicalTrials.gov record.
- Incomplete descriptive results: the statistical analyses posted on ClinicalTrials.gov do not provide median survival times, Kaplan-Meier time-point estimates, or a baseline characteristics table, so those quantities are not included.
- Registry versus full study report: ClinicalTrials.gov is the official registry record, but the statistical detail available in the ClinicalTrials.gov record does not necessarily reproduce every component of a complete statistical analysis plan.
21. Why This Trial Matters Statistically
KEYNOTE-859 is a useful teaching case because the registry results bring several core clinical-trial concepts together in one analysis framework: randomized treatment comparison, multiple time-to-event endpoints, PD-L1-defined analysis populations, stratified Cox regression, log-rank testing, score-based confidence intervals for binary outcomes, and different effect measures for survival and response.
| Concept | How it appears in KEYNOTE-859 |
|---|---|
| Randomization | Randomized parallel phase 3 design. |
| Blinding | Double masking. |
| Time-to-event analysis | All three primary endpoints are overall survival outcomes; PFS is also analyzed as a time-to-event endpoint. |
| Kaplan-Meier estimation | OS is estimated using the product-limit method for censored data. |
| Log-rank test | Used for all posted OS and PFS formal comparisons. |
| Hazard ratio | Used for OS and PFS treatment-effect estimates. |
| Cox regression | Used with Efron's method of tie handling and stratification. |
| Stratified analysis | Geographic region, PD-L1 status where applicable, and chemotherapy regimen are incorporated into the survival models. |
| Binary endpoint analysis | ORR is analyzed as a binary response outcome. |
| Risk difference | ORR treatment effects are reported as differences in percentage. |
| Miettinen & Nurminen | Used for the three posted ORR comparisons. |
| Confidence intervals | Two-sided 95% CIs are reported for all nine posted statistical analyses. |
| Population-specific estimands | Primary OS analyses are defined for all participants, CPS ≥1, and CPS ≥10 populations. |
22. A Deeper Statistical Reading of the Results
Relative and absolute effects are complementary
The survival analyses use hazard ratios because the timing of events matters. The ORR analyses use risk differences because response is classified as a binary outcome. These measures cannot simply be substituted for one another.
For example, an OS HR of 0.78 describes a relative hazard relationship over follow-up, whereas an ORR risk difference of 9.3 describes an absolute difference in response percentage. One does not imply a particular value of the other.
The confidence interval is part of the result
Reporting only 0.78 or 0.65 would hide the uncertainty surrounding the estimates. The corresponding confidence intervals of 0.70–0.87 and 0.53–0.79 communicate how precisely the treatment effect was estimated within the statistical framework.
The P-value and the effect estimate answer different questions
A very small P-value can arise from a relatively modest effect when the information base is large. Conversely, an important-looking point estimate can have substantial uncertainty in a smaller analysis population. KEYNOTE-859 illustrates why the HR or risk difference, its confidence interval, and the P-value should be read together rather than treating the P-value as a ranking of treatment effects.
PD-L1 thresholds define populations, not necessarily effect modification
The registry defines separate primary endpoints for CPS ≥1 and CPS ≥10. That makes these populations part of the formal endpoint structure. However, comparing their numerical HRs is not equivalent to statistically testing whether PD-L1 modifies the treatment effect. A formal interaction analysis would be required for that claim, and none is included in the statistical analyses posted on ClinicalTrials.gov.
Stratification is not the same as subgroup analysis
Geographic region and chemotherapy regimen appear as stratification factors in the Cox models. Stratification allows baseline hazards to vary across categories while estimating a treatment effect across the strata. It should not be interpreted as producing a collection of independent treatment-effect tests.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: KEYNOTE-859, NCT03675737.
- PubMed: PMID 37875143.
- PubMed: PMID 40025394.
- PubMed: PMID 41964430.
- PubMed: PMID 40750941.
- PubMed: PMID 33975465.
Continue through the Clinical Biostats statistical pathway
Clinical trial results become easier to interpret when each endpoint is connected to the statistical method used to estimate, compare, and communicate it.
26. Record Summary
KEYNOTE-859 provides a clear example of a randomized phase 3 statistical framework built around multiple time-to-event endpoints and population-defined efficacy analyses. The three primary overall-survival analyses report hazard ratios of 0.78, 0.74, and 0.65 for all randomized participants, participants with PD-L1 CPS ≥1, and participants with PD-L1 CPS ≥10, respectively. Each has a two-sided 95% confidence interval and a P-value of <0.0001.
The secondary analyses extend the same framework to PFS, with hazard ratios of 0.76, 0.72, and 0.62, and to ORR, with risk differences of 9.3, 9.5, and 17.5. The statistical methods match the outcome structures: log-rank testing and stratified Cox regression for time-to-event endpoints, and the Miettinen & Nurminen method for binary response comparisons.
The central statistical lesson is that these estimates must be interpreted together with their endpoint definitions, analysis populations, censoring rules, stratification factors, effect measures, confidence intervals, and hypothesis-testing framework. A hazard ratio is not a probability, a P-value is not an effect size, and a numerically different subgroup estimate does not by itself demonstrate treatment-effect heterogeneity.