← Clinical Trials
Gastric / GEJ Adenocarcinoma Phase 3 Completed NCT03675737

KEYNOTE-859: Complete Statistical Analysis of Pembrolizumab Plus Chemotherapy in Gastric or Gastroesophageal Junction Adenocarcinoma

An independent statistical review of the randomized phase 3 KEYNOTE-859 trial evaluating pembrolizumab plus chemotherapy versus placebo plus chemotherapy in participants with gastric or gastroesophageal junction adenocarcinoma.

Phase 3  ·  Randomized  ·  Double-masked  ·  1,579 enrolled
Scope of this analysis

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the statistical analyses posted for KEYNOTE-859 in the ClinicalTrials.gov record. Where the registry does not provide a particular result or design detail, it is not inferred from outside sources.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-859 was a randomized, parallel, double-masked phase 3 trial evaluating pembrolizumab plus chemotherapy versus placebo plus chemotherapy in participants with gastric or gastroesophageal junction adenocarcinoma. The registry reports 1,579 enrolled participants, three primary endpoints, and nine posted statistical analyses.

1,579
Enrolled
Phase 3 trial
2
Arms
Parallel design
3
Primary endpoints
All time-to-event
9
Statistical analyses
3 primary + 6 secondary
FeatureKEYNOTE-859
PhasePhase 3
ConditionStomach Neoplasms
Trial nameKEYNOTE-859
DesignRandomized, parallel
MaskingDouble
Primary purposeTreatment
Enrollment1,579
Primary endpoints3
Outcome measures posted14
Statistical analyses posted9
Trial statusCompleted
Start2018-11-08
Primary completion2022-10-03
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry
ClinicalTrials.govNCT03675737

2. Clinical Question

The central question was whether adding pembrolizumab to chemotherapy improved overall survival compared with placebo plus chemotherapy in participants with gastric or gastroesophageal junction adenocarcinoma. The registry prespecified three primary overall-survival endpoints covering the full randomized population and two PD-L1 Combined Positive Score populations.

Population

Participants with gastric or gastroesophageal junction adenocarcinoma, represented in the registry under the condition Stomach Neoplasms.

Intervention

Pembrolizumab plus chemotherapy. The registered interventions include pembrolizumab, cisplatin, 5-fluorouracil, oxaliplatin, and capecitabine.

Comparator

Placebo for pembrolizumab plus chemotherapy. The registry analyses identify the chemotherapy regimens as FP or CAPOX.

Primary question

Does pembrolizumab plus chemotherapy improve overall survival relative to placebo plus chemotherapy under the prespecified superiority framework?

3. Trial Design

01
Randomize1,579 enrolled
02
Parallel arms2 treatment groups
03
Double maskPembrolizumab or placebo
04
AssessOS, PFS, response
05
AnalyzeSurvival and binary endpoints
ARM A · PEMBROLIZUMAB + CHEMOTHERAPY

Pembrolizumab combination

  • Pembrolizumab
  • Chemotherapy
  • Registered chemotherapy interventions include cisplatin, 5-fluorouracil, oxaliplatin, and capecitabine
  • Posted statistical analyses identify FP or CAPOX chemotherapy regimens
ARM B · PLACEBO + CHEMOTHERAPY

Control combination

  • Placebo for pembrolizumab
  • Chemotherapy
  • Posted statistical analyses identify FP or CAPOX chemotherapy regimens

The ClinicalTrials.gov record identifies the allocation as randomized and the design model as parallel. The masking field is recorded as DOUBLE. The statistical analyses further show that geographic region, PD-L1 status where applicable, and chemotherapy regimen were incorporated into stratified analyses, with small strata collapsed.

4. Randomization, Stratification, and Analysis Populations

The posted primary and secondary survival analyses were not simple unadjusted comparisons. The Cox regression analyses incorporated treatment as a covariate and used stratification factors that reflect the trial's design and analysis framework.

Analysis populationDefinition reported in the registry
All-participant efficacy populationAll randomized participants.
PD-L1 CPS ≥1 populationRandomized participants with a PD-L1 CPS of ≥1.
PD-L1 CPS ≥10 populationRandomized participants with a PD-L1 CPS of ≥10.

Stratification in the survival models

For overall survival in all participants, the Cox model was stratified by geographic region, PD-L1 status (CPS <1 versus CPS ≥1), and chemotherapy regimen, with small strata collapsed. For the PD-L1 CPS ≥1 and CPS ≥10 analyses, the model was stratified by geographic region and chemotherapy regimen, again with small strata collapsed.

Clinical Biostats interpretation

Stratification allows the treatment comparison to account for important categorical factors without requiring the analysis to assume that the baseline hazard is identical across those strata. It is particularly relevant here because the registry reports treatment effects within a randomized trial that used different chemotherapy regimens and evaluated PD-L1-defined populations.

The important distinction is that stratification does not mean the treatment effect is estimated separately and independently within every stratum. Instead, the Cox model estimates a common treatment hazard ratio while allowing the baseline hazard to differ across the specified strata.

5. Endpoints

The registry lists three primary endpoints, all based on overall survival. Six posted secondary analyses provide results for progression-free survival and objective response rate across the full population and PD-L1-defined populations.

EndpointRegistry definition / time frameType
Overall Survival (OS) in All Participants OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of the last follow-up. OS was estimated using the product-limit (Kaplan-Meier) method for censored data. Time frame: Up to 45.9 months. Time-to-event
Overall Survival (OS) in Participants With PD-L1 CPS ≥1 OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of the last follow-up. OS was estimated using the product-limit (Kaplan-Meier) method for censored data. Time frame: Up to 45.9 months. Time-to-event
Overall Survival (OS) in Participants With PD-L1 CPS ≥10 OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of the last follow-up. OS was estimated using the product-limit (Kaplan-Meier) method for censored data. Time frame: Up to 45.9 months. Time-to-event
Progression Free Survival (PFS) per RECIST 1.1 assessed by BICR in all participants Progression-free survival assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review. Time frame: Up to 49.5 months. Time-to-event
PFS per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥1 Progression-free survival assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥1. Time frame: Up to 49.5 months. Time-to-event
PFS per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥10 Progression-free survival assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥10. Time frame: Up to 49.5 months. Time-to-event
Objective Response Rate (ORR) per RECIST 1.1 assessed by BICR in all participants Objective response rate assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review. Time frame: Up to 49.5 months. Binary
ORR per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥1 Objective response rate assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥1. Time frame: Up to 49.5 months. Binary
ORR per RECIST 1.1 assessed by BICR in participants with PD-L1 CPS ≥10 Objective response rate assessed per Response Evaluation Criteria in Solid Tumors Version 1.1 by blinded independent central review in randomized participants with PD-L1 CPS ≥10. Time frame: Up to 49.5 months. Binary

6. Statistical Methodology

Kaplan-Meier estimation

The registry explicitly states that overall survival was estimated using the product-limit (Kaplan-Meier) method for censored data. This is appropriate for time-to-event outcomes because not every participant necessarily experiences death during the observation period.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at event time ti, while ni represents participants at risk immediately before that time. Censored participants contribute information until their censoring time.

Log-rank testing

All nine posted statistical analyses used either a log-rank test for time-to-event outcomes or the Miettinen & Nurminen method for binary response outcomes. The six PFS analyses and three OS analyses used the log-rank test.

The log-rank test compares the observed and expected numbers of events between treatment groups over follow-up. It is fundamentally a test of the survival experience rather than a comparison of two isolated proportions.

Stratified Cox regression

The reported hazard ratios were based on Cox regression models with Efron's method of tie handling. Treatment was included as a covariate, while the model was stratified by prespecified factors. This gives a model-based estimate of the relative event hazard associated with treatment.

Hazard ratio
HR = estimated hazard in pembrolizumab + chemotherapy ÷ estimated hazard in placebo + chemotherapy

An HR below 1 indicates a lower estimated instantaneous event hazard in the pembrolizumab combination group under the fitted Cox model. It is not a percentage of patients who benefit and is not the same quantity as a risk ratio.

Miettinen & Nurminen method

The three ORR analyses used the Miettinen & Nurminen method, reported in the normalized methods field as a score-based confidence interval approach for proportions. The treatment effect was expressed as a risk difference, or difference in response percentages between groups.

Risk difference
RD = P(response | pembrolizumab + chemotherapy) − P(response | placebo + chemotherapy)

A positive risk difference means that the response proportion was higher in the pembrolizumab combination group. Unlike a hazard ratio, a risk difference is expressed on an absolute percentage-point scale.

Superiority framework

All nine posted analyses identify the hypothesis type as superiority. The question is therefore whether the treatment groups differ in the prespecified direction rather than whether a new treatment is merely no worse than a control by a predefined non-inferiority margin.

What the ClinicalTrials.gov record does not establish: the ClinicalTrials.gov record does not report a non-inferiority margin, crossover rule, factorial structure, Bayesian analysis, interim-analysis boundary, alpha-spending procedure, or missing-data/imputation method. These features are therefore not inferred or added to this page.

7. Results: Overall Survival in All Participants

The primary overall-survival analysis included all randomized participants. The comparison was pembrolizumab plus chemotherapy versus placebo plus chemotherapy, with a two-sided 95% confidence interval and a superiority hypothesis.

Hazard ratio for overall survival

0.78

95% CI: 0.70–0.87   ·   P < 0.0001

Time frame: Up to 45.9 months

Primary OS analysisPembrolizumab + chemotherapy vs placebo + chemotherapy
Analysis populationAll randomized participants
MethodLog-rank test
Effect measureHazard ratio
Estimate0.78
95% CI0.70–0.87
P-value<0.0001
ModelCox regression with Efron's method of tie handling; treatment as a covariate; stratified by geographic region, PD-L1 status (CPS <1 versus CPS ≥1), and chemotherapy regimen with small strata collapsed
Clinical Biostats interpretation

The estimated hazard ratio of 0.78 means that, under the fitted stratified Cox model, the estimated instantaneous rate of death in the pembrolizumab plus chemotherapy group was about 78% of the estimated rate in the placebo plus chemotherapy group. Equivalently, 0.78 corresponds to a 22% lower estimated hazard under that model.

This does not mean that 22% of participants avoided death, that every participant experienced a 22% reduction in risk, or that the absolute probability of death was reduced by 22 percentage points. The hazard ratio is a relative time-to-event measure.

The 95% confidence interval of 0.70–0.87 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not an interval containing the effects that individual participants experienced.

The P-value of <0.0001 addresses evidence against the relevant null hypothesis; it does not quantify the size or clinical importance of the treatment effect. The magnitude of the estimated effect is communicated by the hazard ratio and its confidence interval.

Because the estimate comes from a Cox model, its interpretation also depends on the model's assumptions. In particular, a single hazard ratio summarizes relative event rates over time and may be less descriptive if the proportional-hazards relationship does not adequately characterize the data.

8. Results: Overall Survival in Participants With PD-L1 CPS ≥1

The second primary endpoint restricted the analysis population to randomized participants with a PD-L1 CPS of ≥1. The registry again reports a stratified log-rank analysis and a Cox model-based hazard ratio.

Hazard ratio for overall survival, PD-L1 CPS ≥1

0.74

95% CI: 0.65–0.84   ·   P < 0.0001

Time frame: Up to 45.9 months

Primary OS analysisResult
Analysis populationRandomized participants with a PD-L1 CPS of ≥1
MethodLog-rank test
Effect measureHazard ratio
Estimate0.74
95% CI0.65–0.84
P-value<0.0001
ModelCox regression with Efron's method of tie handling; treatment as a covariate; stratified by geographic region and chemotherapy regimen with small strata collapsed
Clinical Biostats interpretation

An HR of 0.74 indicates an estimated instantaneous death hazard approximately 74% as large in the pembrolizumab plus chemotherapy group as in the placebo plus chemotherapy group, conditional on the fitted model. As a simple relative interpretation, this corresponds to a 26% lower estimated hazard.

The estimate does not mean a 26-percentage-point improvement in survival, nor does it mean that exactly 26% of patients benefit. It also does not establish that the treatment effect is identical for every individual with PD-L1 CPS ≥1.

The 95% CI of 0.65–0.84 provides the precision of this estimated relative effect under the model. The relatively narrow interval compared with the estimate itself indicates that the posted analysis provides a more constrained statistical range than would be conveyed by the point estimate alone.

The P-value of <0.0001 is evidence against the null hypothesis used in the superiority analysis. It should not be interpreted as the probability that the null hypothesis is true or as a measure of how large the treatment effect is.

This analysis is also a population-restricted comparison. Its result describes randomized participants meeting the PD-L1 CPS ≥1 criterion; it should not automatically be generalized to the full trial population without considering the distinction between the analysis populations.

9. Results: Overall Survival in Participants With PD-L1 CPS ≥10

The third primary endpoint further restricted the analysis population to randomized participants with a PD-L1 CPS of ≥10.

Hazard ratio for overall survival, PD-L1 CPS ≥10

0.65

95% CI: 0.53–0.79   ·   P < 0.0001

Time frame: Up to 45.9 months

Primary OS analysisResult
Analysis populationRandomized participants with a PD-L1 CPS of ≥10
MethodLog-rank test
Effect measureHazard ratio
Estimate0.65
95% CI0.53–0.79
P-value<0.0001
ModelCox regression with Efron's method of tie handling; treatment as a covariate; stratified by geographic region and chemotherapy regimen with small strata collapsed
Clinical Biostats interpretation

An HR of 0.65 corresponds to an estimated death hazard about 65% of that in the comparator group under the fitted Cox model, or approximately a 35% lower estimated hazard.

Again, this is not an absolute risk reduction and does not mean that 35% of participants avoided death. It is a model-based relative comparison of event hazards over the analyzed follow-up.

The 95% CI of 0.53–0.79 describes uncertainty around the treatment-effect estimate. It does not describe variation in treatment benefit from one patient to another.

The P-value of <0.0001 measures the evidence against the null hypothesis under the specified statistical framework. It does not indicate that the probability of the observed effect being due to chance is <0.0001, and it does not measure effect magnitude.

The CPS ≥10 analysis is especially important to interpret as a prespecified population-defined endpoint rather than as a post hoc comparison of whichever subgroup happened to show the smallest hazard ratio. The ClinicalTrials.gov record identifies it explicitly as one of the three primary endpoints.

10. Primary Overall-Survival Results Together

Putting the three primary analyses side by side makes the structure of the statistical evidence clearer. All three use hazard ratios below 1, with two-sided 95% confidence intervals and P-values reported as <0.0001.

Primary endpointPopulationHR95% CIP-value
OS in all participantsAll randomized participants0.780.70–0.87<0.0001
OS in PD-L1 CPS ≥1Randomized participants with CPS ≥10.740.65–0.84<0.0001
OS in PD-L1 CPS ≥10Randomized participants with CPS ≥100.650.53–0.79<0.0001
How to compare the three hazard ratios

The three point estimates are 0.78, 0.74, and 0.65. These values describe the estimated relative treatment effects in three different analysis populations. The CPS ≥10 estimate is numerically lower than the all-participant estimate, but a smaller point estimate in one population does not by itself establish that the treatment effect is statistically different between populations.

Formal comparison of treatment effects across subgroups generally requires an interaction or treatment-by-subgroup analysis. The ClinicalTrials.gov record does not report such an interaction test, so the three hazard ratios should be read as separate prespecified endpoint results rather than as evidence that the treatment effect necessarily changes according to PD-L1 CPS.

11. Secondary Results: Progression-Free Survival

Six secondary time-to-event analyses are posted for PFS: three in the full randomized population and three in PD-L1-defined populations. Each uses the log-rank test with a Cox regression model for the hazard ratio.

PFS analysisPopulationHR95% CIP-value
PFS per RECIST 1.1, BICR All randomized participants 0.76 0.67–0.85 <0.0001
PFS per RECIST 1.1, BICR Randomized participants with PD-L1 CPS ≥1 0.72 0.63–0.82 <0.0001
PFS per RECIST 1.1, BICR Randomized participants with PD-L1 CPS ≥10 0.62 0.51–0.76 <0.0001

All three PFS analyses have the time frame up to 49.5 months. The full-population analysis was stratified by geographic region, PD-L1 status (CPS <1 versus CPS ≥1), and chemotherapy regimen with small strata collapsed. The CPS ≥1 and CPS ≥10 analyses were stratified by geographic region and chemotherapy regimen with small strata collapsed.

Clinical Biostats interpretation

The PFS HR of 0.76 in all randomized participants corresponds to a 24% lower estimated hazard of progression or death under the fitted model. The CPS ≥1 and CPS ≥10 estimates of 0.72 and 0.62 correspond to 28% and 38% lower estimated hazards, respectively.

These are relative time-to-event effects. They do not tell us the absolute difference in the probability of remaining progression-free at any particular time, and they do not establish how long an individual participant will remain progression-free.

The confidence intervals are important because each point estimate has sampling uncertainty. The P-values are evidence against the corresponding null hypotheses, but they do not quantify the size of the PFS benefit.

Because PFS incorporates both progression and death, its clinical interpretation is different from OS. A PFS hazard ratio and an OS hazard ratio should not be treated as interchangeable measures of the same event.

12. Secondary Results: Objective Response Rate

Objective response rate was analyzed as a binary endpoint using the Miettinen & Nurminen method. The effect measure was a difference in percentage, normalized here as a risk difference. The registry reports two-sided 95% confidence intervals.

ORR analysisPopulationRisk difference95% CIP-value
ORR per RECIST 1.1, BICR All randomized participants 9.3 4.4–14.1 0.00009
ORR per RECIST 1.1, BICR Randomized participants with PD-L1 CPS ≥1 9.5 3.9–15.0 0.00041
ORR per RECIST 1.1, BICR Randomized participants with PD-L1 CPS ≥10 17.5 9.3–25.5 0.00002

Largest posted ORR difference

17.5

95% CI: 9.3–25.5   ·   P = 0.00002

Participants with PD-L1 CPS ≥10; risk difference in percentage points.

Clinical Biostats interpretation

A risk difference of 9.3 means that the response percentage in the pembrolizumab plus chemotherapy group exceeded that in the placebo plus chemotherapy group by an estimated 9.3 percentage points in the full randomized population.

Similarly, the estimates of 9.5 and 17.5 describe absolute differences in response percentage for the CPS ≥1 and CPS ≥10 populations. These are not relative risks and should not be described as percentage reductions in the probability of nonresponse.

The 95% confidence interval gives the statistical precision of each estimated difference. For the full population, the interval is 4.4–14.1; for CPS ≥1, it is 3.9–15.0; and for CPS ≥10, it is 9.3–25.5.

The P-values indicate evidence against the relevant null hypothesis of no treatment difference in response under the specified analysis. They do not tell us whether an observed response difference is clinically important, nor do they describe the probability that the treatment has no effect.

13. Statistical Methods Explained

Why was a log-rank test used for overall survival and progression-free survival?

OS and PFS are time-to-event endpoints, meaning both the event time and the fact that some participants may be censored are part of the analysis. A log-rank test compares the event experience between treatment groups over the observed follow-up rather than reducing the outcome to a single binary proportion.

What does an HR of 0.78 mean?

Under the fitted Cox model, an HR of 0.78 means the estimated instantaneous event hazard in the pembrolizumab combination group is 78% of that in the comparator group. The complementary interpretation is a 22% lower estimated hazard. It does not mean a 22% absolute reduction in the probability of death.

Why are the CPS ≥1 and CPS ≥10 analyses different from the all-participant analysis?

They use different analysis populations. The all-participant analysis includes all randomized participants, whereas the CPS analyses restrict the population according to the registered PD-L1 CPS threshold. Because the populations differ, the resulting hazard ratios estimate treatment effects in different groups.

Why does the Cox model use stratification?

The registry reports stratification by geographic region and chemotherapy regimen for the PD-L1-defined analyses, with additional PD-L1 status stratification for the full-population OS and PFS analyses. Stratification permits the underlying baseline hazard to differ across these strata while estimating a treatment effect across them.

What does a risk difference of 17.5 mean?

The ORR risk difference of 17.5 represents an estimated 17.5-percentage-point difference in objective response between the randomized treatment groups in participants with PD-L1 CPS ≥10. It is an absolute measure, unlike the hazard ratios used for OS and PFS.

Why use the Miettinen & Nurminen method for ORR?

ORR is a binary endpoint: each participant is classified according to whether the prespecified response criterion was met. The Miettinen & Nurminen method provides a score-based approach to estimating uncertainty around the difference between two proportions, matching the registry's reported statistical method.

What does the P-value add when the confidence interval is already reported?

The P-value and confidence interval answer related but different questions. The P-value quantifies the compatibility of the observed data with the specified null hypothesis, whereas the confidence interval communicates the estimated effect and its statistical precision. Neither one replaces the effect estimate itself.

14. Understanding the Three Primary Endpoints

The three primary endpoints are all versions of the same underlying outcome, overall survival, but they answer the question in different analysis populations.

EndpointPopulationHRInterpretive focus
OS in all participantsAll randomized participants0.78Overall randomized treatment comparison
OS in CPS ≥1Randomized participants with CPS ≥10.74Treatment comparison in the CPS ≥1 population
OS in CPS ≥10Randomized participants with CPS ≥100.65Treatment comparison in the CPS ≥10 population

This structure illustrates an important statistical principle: the population is part of the estimand. An effect estimate is not fully described by its numerical value alone. The endpoint definition, analysis population, treatment contrast, follow-up window, and statistical model all determine what the estimate represents.

Do not over-read the numerical ordering. The sequence 0.78 → 0.74 → 0.65 does not by itself demonstrate increasing treatment efficacy with increasing PD-L1 CPS. Demonstrating effect modification would require an appropriate formal comparison of treatment effects across the populations. No such interaction analysis is included in the statistical analyses posted on ClinicalTrials.gov.

15. Censoring and Time-to-Event Interpretation

The registry definition of OS explicitly states that participants without documented death at the time of analysis were censored at the date of the last follow-up. This is central to understanding why Kaplan-Meier and Cox methods are used rather than ordinary comparisons of means.

Event

For OS, the event is death due to any cause.

Censoring

A participant without documented death at the analysis time contributes information through the last follow-up date.

Kaplan-Meier

Estimates the survival function while accounting for censored observations.

Cox model

Estimates a relative hazard while incorporating follow-up time and censoring.

Censoring does not mean that a censored participant is treated as if the event never occurred. Rather, the analysis recognizes that the participant's event-free observation is known only up to the censoring time. The validity of standard survival analysis depends on assumptions about the censoring mechanism and the relationship between censoring and subsequent event risk.

16. Hazard Ratios: What They Do and Do Not Mean

Relative effect

The OS hazard ratios of 0.78, 0.74, and 0.65 all lie below 1. Under their respective Cox models, this indicates lower estimated death hazards in the pembrolizumab plus chemotherapy group than in the placebo plus chemotherapy group.

Not an absolute probability

A hazard ratio is not a survival percentage, response percentage, risk difference, or median survival time. A value of 0.65 does not mean that 65% of participants survived or that survival probability was reduced by 35 percentage points.

Not an individual-patient guarantee

The hazard ratio describes a treatment comparison at the population level under the fitted model. It does not imply that every participant experiences the same relative reduction in event hazard.

Confidence intervals matter

The confidence intervals communicate statistical precision. For example, the full-population OS estimate of 0.78 is accompanied by a 95% CI of 0.70–0.87. The interval is therefore an essential part of reporting the estimate rather than an optional add-on.

P-values are not effect sizes

The P-values of <0.0001 for the three primary OS analyses indicate strong statistical evidence against their respective null hypotheses under the reported framework. They do not indicate the size of the treatment effect and should not be used as a substitute for the hazard ratio or its confidence interval.

17. Primary and Secondary Results: Statistical Overview

EndpointRoleTypeEffect measureEstimate95% CIP-value
OS, all participantsPrimaryTime-to-eventHR0.780.70–0.87<0.0001
OS, CPS ≥1PrimaryTime-to-eventHR0.740.65–0.84<0.0001
OS, CPS ≥10PrimaryTime-to-eventHR0.650.53–0.79<0.0001
PFS, all participantsSecondaryTime-to-eventHR0.760.67–0.85<0.0001
PFS, CPS ≥1SecondaryTime-to-eventHR0.720.63–0.82<0.0001
PFS, CPS ≥10SecondaryTime-to-eventHR0.620.51–0.76<0.0001
ORR, all participantsSecondaryBinaryRisk difference9.34.4–14.10.00009
ORR, CPS ≥1SecondaryBinaryRisk difference9.53.9–15.00.00041
ORR, CPS ≥10SecondaryBinaryRisk difference17.59.3–25.50.00002

This table also illustrates why it is useful to report effect measures in their natural statistical scale. OS and PFS are summarized with hazard ratios because they are time-to-event outcomes, while ORR is summarized with a risk difference because it is binary.

18. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk. The reported figures are:

ArmSerious adverse events
Pembrolizumab + Chemotherapy (FP or CAPO)358/785
Placebo + Chemotherapy (FP or CAPOX Regi)317/787
Pembrolizumab Second Course1/12

The ClinicalTrials.gov record does not provide a further safety-analysis definition, confidence interval, hypothesis test, or comparative P-value for these serious-adverse-event counts. Accordingly, the safety information is reported descriptively rather than converted into an inferential treatment comparison.

Clinical Biostats interpretation

Safety and efficacy answer different statistical questions. The serious-adverse-event counts describe the number of affected participants relative to the number at risk in the ClinicalTrials.gov record. They do not by themselves establish a causal treatment difference, particularly without a reported comparative analysis and its uncertainty.

It is also important not to combine the serious-adverse-event figures with the OS hazard ratio into a single numerical "benefit-risk" statistic. They arise from different outcome definitions, populations or exposure groupings, and statistical frameworks.

19. What Is and Is Not Reported in the Supplied Registry Data

The ClinicalTrials.gov record is sufficient to reconstruct the principal statistical analyses, but they do not contain every type of information that might appear in a full clinical-study report.

TopicWhat the ClinicalTrials.gov record provides
Primary efficacyThree formal OS analyses with estimates, two-sided 95% CIs, and P-values.
Secondary efficacySix formal analyses covering PFS and ORR, with estimates, CIs, and P-values.
Statistical methodsLog-rank testing and Miettinen & Nurminen methodology, with Cox-model details for survival analyses.
Analysis populationsAll randomized participants and randomized participants meeting PD-L1 CPS thresholds.
Baseline characteristicsNot reported in the ClinicalTrials.gov record.
Median OS / PFSNot reported in the statistical analyses provided for this page.
Kaplan-Meier numerical time-point estimatesNot reported.
Subgroup hazard ratios beyond the registered CPS populationsNot reported.
Formal multiplicity procedureNot reported.
Interim-analysis procedureNot reported.
Missing-data or imputation procedureNot reported.
Non-inferiority marginNot applicable to the reported superiority analyses; no non-inferiority margin is reported.
Bayesian methodsNot reported in the statistical analyses posted on ClinicalTrials.gov.
Why this matters: a complete statistical analysis should distinguish what was actually reported from what would be conventional in a clinical trial of this type. The absence of a reported median survival, baseline table, or multiplicity procedure in the ClinicalTrials.gov record is not a reason to reconstruct one from outside information.

20. Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

KEYNOTE-859 is a useful teaching case because the registry results bring several core clinical-trial concepts together in one analysis framework: randomized treatment comparison, multiple time-to-event endpoints, PD-L1-defined analysis populations, stratified Cox regression, log-rank testing, score-based confidence intervals for binary outcomes, and different effect measures for survival and response.

ConceptHow it appears in KEYNOTE-859
RandomizationRandomized parallel phase 3 design.
BlindingDouble masking.
Time-to-event analysisAll three primary endpoints are overall survival outcomes; PFS is also analyzed as a time-to-event endpoint.
Kaplan-Meier estimationOS is estimated using the product-limit method for censored data.
Log-rank testUsed for all posted OS and PFS formal comparisons.
Hazard ratioUsed for OS and PFS treatment-effect estimates.
Cox regressionUsed with Efron's method of tie handling and stratification.
Stratified analysisGeographic region, PD-L1 status where applicable, and chemotherapy regimen are incorporated into the survival models.
Binary endpoint analysisORR is analyzed as a binary response outcome.
Risk differenceORR treatment effects are reported as differences in percentage.
Miettinen & NurminenUsed for the three posted ORR comparisons.
Confidence intervalsTwo-sided 95% CIs are reported for all nine posted statistical analyses.
Population-specific estimandsPrimary OS analyses are defined for all participants, CPS ≥1, and CPS ≥10 populations.

22. A Deeper Statistical Reading of the Results

Relative and absolute effects are complementary

The survival analyses use hazard ratios because the timing of events matters. The ORR analyses use risk differences because response is classified as a binary outcome. These measures cannot simply be substituted for one another.

For example, an OS HR of 0.78 describes a relative hazard relationship over follow-up, whereas an ORR risk difference of 9.3 describes an absolute difference in response percentage. One does not imply a particular value of the other.

The confidence interval is part of the result

Reporting only 0.78 or 0.65 would hide the uncertainty surrounding the estimates. The corresponding confidence intervals of 0.70–0.87 and 0.53–0.79 communicate how precisely the treatment effect was estimated within the statistical framework.

The P-value and the effect estimate answer different questions

A very small P-value can arise from a relatively modest effect when the information base is large. Conversely, an important-looking point estimate can have substantial uncertainty in a smaller analysis population. KEYNOTE-859 illustrates why the HR or risk difference, its confidence interval, and the P-value should be read together rather than treating the P-value as a ranking of treatment effects.

PD-L1 thresholds define populations, not necessarily effect modification

The registry defines separate primary endpoints for CPS ≥1 and CPS ≥10. That makes these populations part of the formal endpoint structure. However, comparing their numerical HRs is not equivalent to statistically testing whether PD-L1 modifies the treatment effect. A formal interaction analysis would be required for that claim, and none is included in the statistical analyses posted on ClinicalTrials.gov.

Stratification is not the same as subgroup analysis

Geographic region and chemotherapy regimen appear as stratification factors in the Cox models. Stratification allows baseline hazards to vary across categories while estimating a treatment effect across the strata. It should not be interpreted as producing a collection of independent treatment-effect tests.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Clinical trial results become easier to interpret when each endpoint is connected to the statistical method used to estimate, compare, and communicate it.

26. Record Summary

KEYNOTE-859 provides a clear example of a randomized phase 3 statistical framework built around multiple time-to-event endpoints and population-defined efficacy analyses. The three primary overall-survival analyses report hazard ratios of 0.78, 0.74, and 0.65 for all randomized participants, participants with PD-L1 CPS ≥1, and participants with PD-L1 CPS ≥10, respectively. Each has a two-sided 95% confidence interval and a P-value of <0.0001.

The secondary analyses extend the same framework to PFS, with hazard ratios of 0.76, 0.72, and 0.62, and to ORR, with risk differences of 9.3, 9.5, and 17.5. The statistical methods match the outcome structures: log-rank testing and stratified Cox regression for time-to-event endpoints, and the Miettinen & Nurminen method for binary response comparisons.

The central statistical lesson is that these estimates must be interpreted together with their endpoint definitions, analysis populations, censoring rules, stratification factors, effect measures, confidence intervals, and hypothesis-testing framework. A hazard ratio is not a probability, a P-value is not an effect size, and a numerically different subgroup estimate does not by itself demonstrate treatment-effect heterogeneity.

Clinical Biostats methodology: The purpose of this page is to reconstruct the statistical story contained in the ClinicalTrials.gov record while clearly separating reported evidence from educational interpretation. Results that were not included in the ClinicalTrials.gov record—such as median survival times, baseline characteristics, additional subgroup estimates, or unreported multiplicity and interim-analysis procedures—are intentionally not inferred.