← Clinical Trials
Cervical Cancer Phase 3 Completed NCT03635567

KEYNOTE-826: Complete Statistical Analysis of Pembrolizumab in Cervical Cancer

An independent statistical analysis of the randomized phase 3 KEYNOTE-826 trial evaluating pembrolizumab plus chemotherapy versus placebo plus chemotherapy in women with persistent, recurrent, or metastatic cervical cancer.

Trial start: 2018-10-25  ·  Primary completion: 2022-10-03  ·  Enrollment: 617
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View the ClinicalTrials.gov record.

1. Trial at a Glance

KEYNOTE-826 was a randomized, parallel, triple-masked phase 3 trial evaluating first-line pembrolizumab plus chemotherapy versus placebo plus chemotherapy in women with persistent, recurrent, or metastatic cervical cancer. The registry reports 617 enrolled participants, 2 arms, 6 registered primary endpoints, 15 posted outcome measures, and 8 posted statistical analyses.

617
Enrolled
Phase 3 trial
2
Arms
Randomized parallel design
0.58
PFS HR, CPS ≥1
95% CI 0.47–0.71
0.60
OS HR, CPS ≥1
95% CI 0.49–0.74
FeatureKEYNOTE-826
Trial nameKEYNOTE-826
NCT IDNCT03635567
PhasePhase 3
StatusCompleted
ConditionCervical Cancer
AllocationRandomized
Design modelParallel
MaskingTriple
Primary purposeTreatment
Enrollment617
Arms2
Primary endpoint typesBinary; Time-to-event
Results postedYes
Outcome measures posted15
Statistical analyses posted8
Primary-endpoint analyses6
Primary analyses with estimate + CI6
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry

2. Clinical Question

The registry describes KEYNOTE-826 as a study of first-line treatment with pembrolizumab plus chemotherapy versus placebo plus chemotherapy in women with persistent, recurrent, or metastatic cervical cancer. The statistical question is whether randomized assignment to the pembrolizumab-containing regimen is associated with different progression-free survival and overall survival, including in prespecified PD-L1 Combined Positive Score populations.

Population

Women with persistent, recurrent, or metastatic cervical cancer enrolled in the phase 3 randomized trial.

Intervention

Pembrolizumab plus chemotherapy. The registered interventions include pembrolizumab, paclitaxel, cisplatin, carboplatin, and bevacizumab.

Comparator

Placebo to pembrolizumab plus chemotherapy. The registry lists placebo to pembrolizumab as a drug intervention.

Primary question

How does pembrolizumab plus chemotherapy compare with placebo plus chemotherapy for PFS and OS in the registered analysis populations?

3. Trial Design

01
Enroll617 participants
02
Randomize2 treatment arms
03
MaskTriple-masked design
04
AssessPFS, OS, response and safety
05
AnalyzeStratified treatment comparisons
Allocation
Randomized. Participants were assigned to one of two treatment arms.
Design model
Parallel. The two randomized groups were compared over the trial follow-up rather than being assigned sequentially to each treatment.
Masking
Triple masked. The registry identifies the trial as triple-masked.
Primary purpose
Treatment. The registered primary purpose is treatment.
ARM 1

Pembrolizumab + Chemotherapy

  • Pembrolizumab
  • Paclitaxel
  • Cisplatin or carboplatin
  • Bevacizumab is also listed among the registered biological/drug interventions
ARM 2

Placebo + Chemotherapy

  • Placebo to pembrolizumab
  • Paclitaxel
  • Cisplatin or carboplatin
  • Bevacizumab is also listed among the registered biological/drug interventions

The ClinicalTrials.gov record identifies the treatment comparison as pembrolizumab + chemotherapy versus placebo + chemotherapy. It does not provide an arm-specific randomized sample size in the ClinicalTrials.gov record, so this page does not infer one from the total enrollment.

4. Endpoints

The registry lists six primary endpoints. Three concern progression-free survival and three concern overall survival. The endpoint populations distinguish all randomized participants from participants meeting PD-L1 Combined Positive Score thresholds.

Primary endpointTime frameTypeAnalysis
Progression-free Survival (PFS) Per Response Evaluation Criteria in Solid Tumors Version 1.1 (RECIST 1.1) as Assessed by Investigator in Participants With Programmed Cell Death-Ligand 1 (PD-L1) Combined Positive Score (CPS) ≥1 Up to approximately 46 months Time-to-event Stratified log-rank test; HR
PFS Per RECIST 1.1 as Assessed by Investigator in All Participants Up to approximately 46 months Binary Stratified log-rank test; HR
PFS Per RECIST 1.1 as Assessed by Investigator in Participants With PD-L1 CPS ≥10 Up to approximately 46 months Binary Stratified log-rank test; HR
Overall Survival (OS) in Participants With PD-L1 CPS ≥1 Up to approximately 46 months Time-to-event Stratified log-rank test; HR
OS in All Participants Up to approximately 46 months Binary Stratified log-rank test; HR
OS in Participants With PD-L1 CPS ≥10 Up to approximately 46 months Binary Stratified log-rank test; HR

Registry definitions of the primary endpoints

Progression-free survival: PFS was defined as the time from randomization to the first documented progressive disease (PD) or death due to any cause, whichever occurs first. Per RECIST 1.1, PD was defined as ≥ 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study. In addition to the relative increase of 20%, the sum must also demonstrate an absolute increase of ≥5 mm.

Overall survival: OS was defined as the time from randomization to death due to any cause. The OS for all randomized participants with PD-L1 CPS ≥1, all randomized participants, and all randomized participants with PD-L1 CPS ≥10 was presented according to the respective registered endpoint.

Endpoint terminology matters. The registry labels two PFS and two OS endpoints as "Binary" even though the definitions reported in the registry describe time from randomization to an event. For this page, the reported analysis method and effect measure are followed as provided in the posted statistical analyses: the formal treatment comparisons use stratified log-rank testing and hazard ratios.

5. Results Overview

ClinicalTrials.gov contains formal statistical analyses for all six registered primary endpoints and two secondary endpoints in the ClinicalTrials.gov record. All six primary analyses have an estimate and 95% confidence interval. The primary time-to-event analyses consistently use a stratified log-rank framework with treatment-effect estimation through a Cox regression model.

Endpoint familyPopulationEstimate95% CIP-value
PFSPD-L1 CPS ≥1HR 0.580.47–0.71<0.0001
PFSAll participantsHR 0.610.50–0.74<0.0001
PFSPD-L1 CPS ≥10HR 0.520.40–0.68<0.0001
OSPD-L1 CPS ≥1HR 0.600.49–0.74<0.0001
OSAll participantsHR 0.630.52–0.77<0.0001
OSPD-L1 CPS ≥10HR 0.580.44–0.78<0.0001

These estimates are not interchangeable. Each belongs to a different endpoint and analysis population. The CPS ≥1 and CPS ≥10 analyses answer questions in biomarker-defined populations, whereas the all-participant analyses describe the randomized population without that PD-L1 restriction.

6. Primary Results: Progression-Free Survival

PFS in Participants With PD-L1 CPS ≥1

Hazard ratio for progression or death

0.58

95% CI: 0.47–0.71   ·   P < 0.0001

Analysis population: all randomized participants with PD-L1 CPS ≥1, analyzed according to randomized treatment group.

The registry reports a stratified log-rank treatment comparison. The treatment comparison for the hazard ratio was based on a Cox regression model using Efron's method of tie handling, with treatment as a covariate and stratification by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.

Clinical Biostats interpretation

An HR of 0.58 means that the fitted time-to-event model estimated an instantaneous rate of progression or death approximately 42% lower in the pembrolizumab-plus-chemotherapy group than in the placebo-plus-chemotherapy group, because 1 − 0.58 = 0.42.

The HR does not mean that 42% of participants avoided progression or death, that median PFS was 42% longer, or that every individual participant experienced a 42% reduction in risk. It is a relative model-based measure of the event rate over time.

The 95% CI of 0.47–0.71 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of treatment effects among individual patients.

The p-value of <0.0001 addresses the evidence against the null comparison specified by the statistical test. It does not measure the size, clinical importance, or certainty of the treatment effect. The HR and its confidence interval are needed to describe the estimated magnitude and precision.

Because the estimate is based on a Cox regression model, interpretation also depends on the appropriateness of the model and, in particular, the proportional-hazards framework. A single HR summarizes a time-to-event comparison but does not show how the separation between treatment groups evolves at each point in follow-up.

PFS in All Participants

Hazard ratio for progression or death

0.61

95% CI: 0.50–0.74   ·   P < 0.0001

Analysis population: all randomized participants.

The all-participant analysis provides a broader randomized-population estimate than the PD-L1 CPS ≥1 and CPS ≥10 analyses. The registry again specifies a stratified log-rank comparison and a Cox regression model with Efron's method of tie handling.

Clinical Biostats interpretation

An HR of 0.61 corresponds to an estimated 39% lower instantaneous rate of progression or death under the fitted model for the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group.

The confidence interval of 0.50–0.74 communicates the precision of the estimate. It is narrower than simply reporting the point estimate because it describes a range of values compatible with the specified statistical framework and observed data.

The p-value of <0.0001 is evidence from the stated hypothesis-testing procedure; it is not a probability that the treatment effect is true, nor does it quantify the clinical magnitude of the difference.

The comparison is based on randomized treatment assignment rather than observed treatment received, which is important when interpreting an efficacy analysis in a randomized trial. The ClinicalTrials.gov record does not provide additional information about treatment discontinuation, crossover, or missing-data handling, so those issues should not be inferred from the HR alone.

PFS in Participants With PD-L1 CPS ≥10

Hazard ratio for progression or death

0.52

95% CI: 0.40–0.68   ·   P < 0.0001

Analysis population: all randomized participants with PD-L1 CPS ≥10.

Clinical Biostats interpretation

An HR of 0.52 corresponds to an estimated instantaneous rate of progression or death approximately 48% lower in the pembrolizumab-plus-chemotherapy group under the fitted model.

The 95% CI of 0.40–0.68 indicates the statistical uncertainty around the estimate. It should not be interpreted as saying that the true effect for individual patients must lie somewhere between a 32% and 60% reduction in their personal risk.

The p-value of <0.0001 indicates strong evidence under the reported testing procedure, but it is not an effect-size metric. The HR remains the primary measure of relative treatment effect for this analysis.

This is a biomarker-defined analysis population. A comparison between the CPS ≥10 HR and another subgroup's HR does not, by itself, establish that PD-L1 status modifies the treatment effect. Demonstrating effect modification requires an appropriate comparison or interaction analysis, and the registry-reported statistical analysis does not report such an interaction estimate.

7. Primary Results: Overall Survival

OS in Participants With PD-L1 CPS ≥1

Hazard ratio for death

0.60

95% CI: 0.49–0.74   ·   P < 0.0001

Analysis population: all randomized participants with PD-L1 CPS ≥1.

Clinical Biostats interpretation

An HR of 0.60 means that the fitted model estimated an instantaneous rate of death approximately 40% lower in the pembrolizumab-plus-chemotherapy group than in the placebo-plus-chemotherapy group.

This does not mean that 40% of participants survived because of treatment, that 40% of participants were cured, or that the absolute probability of death was reduced by exactly 40 percentage points. Hazard ratios are relative time-to-event measures rather than absolute risk differences.

The 95% CI of 0.49–0.74 describes uncertainty in the estimated HR. It provides more information than the point estimate alone because an estimate of 0.60 based on a very imprecise interval would carry a different statistical interpretation from the same point estimate with a narrow interval.

The p-value of <0.0001 is evidence from the stratified log-rank testing framework. It does not tell the reader how large the treatment effect is; that information comes from the HR and its confidence interval.

OS in All Participants

Hazard ratio for death

0.63

95% CI: 0.52–0.77   ·   P < 0.0001

Analysis population: all randomized participants.

Clinical Biostats interpretation

The HR of 0.63 corresponds to an estimated instantaneous rate of death approximately 37% lower in the pembrolizumab-plus-chemotherapy group under the fitted Cox model.

The 95% CI of 0.52–0.77 gives the uncertainty around that estimate. It does not give a range of individual survival probabilities and does not replace presentation of absolute survival probabilities or other clinically interpretable time-specific measures.

The p-value of <0.0001 should be read as a hypothesis-test result rather than as a measure of treatment benefit. Large datasets can produce small p-values for relatively modest effects, while small datasets can produce imprecise estimates even when the point estimate is substantial.

The analysis uses randomized treatment assignment and stratification variables specified in the registry analysis text. The ClinicalTrials.gov record does not report a separate absolute survival estimate or median OS, so this page does not infer one from the HR.

OS in Participants With PD-L1 CPS ≥10

Hazard ratio for death

0.58

95% CI: 0.44–0.78   ·   P < 0.0001

Analysis population: all randomized participants with PD-L1 CPS ≥10.

Clinical Biostats interpretation

An HR of 0.58 corresponds to an estimated instantaneous rate of death approximately 42% lower in the pembrolizumab-plus-chemotherapy group under the fitted model.

The 95% CI of 0.44–0.78 indicates uncertainty around the estimate. The interval should not be treated as a prediction interval for individual patients or as a statement that all patients experience an effect within that numerical range.

The p-value of <0.0001 indicates the result of the reported statistical comparison. It is not the probability that the null hypothesis is true and does not indicate the practical importance of the treatment effect.

Because this is a PD-L1 CPS ≥10 analysis, the estimate describes a defined subgroup rather than automatically establishing that the treatment effect is stronger or weaker than in the other registered populations.

8. Secondary Endpoint Results

Objective Response Rate

The registry reports Objective Response Rate (ORR) Per RECIST 1.1 as Assessed by Investigator as a secondary endpoint. The outcome unit is percentage of participants, and the analysis population consists of all randomized participants based on the treatment group to which they were randomized.

Risk difference in objective response rate

14.9 percentage points

95% CI: 7.4–22.3   ·   P = 0.0001

Analysis: Miettinen & Nurminen method, stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.

Clinical Biostats interpretation

The reported risk difference of 14.9 percentage points means that the observed percentage of participants achieving an objective response differed by 14.9 percentage points between the randomized treatment groups, with the direction corresponding to the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group.

This is an absolute difference, unlike a hazard ratio. It is therefore directly expressed in percentage points rather than as a relative rate over time.

The 95% CI of 7.4–22.3 describes uncertainty around the estimated difference. It is not the range of response rates for individual patients.

The p-value of 0.0001 addresses the statistical comparison under the reported Miettinen & Nurminen method. It does not measure the magnitude of the response-rate difference. The 14.9-point estimate and its confidence interval provide that information.

PFS by Blinded Independent Central Review

Hazard ratio for progression or death

0.60

95% CI: 0.49–0.74   ·   P < 0.0001

Analysis population: all randomized participants. Assessment: RECIST 1.1 by blinded independent central review.

Clinical Biostats interpretation

The BICR PFS analysis produced an HR of 0.60, corresponding to an estimated instantaneous rate of progression or death approximately 40% lower under the fitted model in the pembrolizumab-plus-chemotherapy group.

The 95% CI of 0.49–0.74 describes uncertainty around the relative treatment effect. The BICR assessment is statistically useful because it provides an assessment framework distinct from investigator assessment, but the ClinicalTrials.gov record does not provide a comparison between the two assessment approaches beyond their respective posted results.

The p-value of <0.0001 is evidence from the reported stratified log-rank comparison and should not be interpreted as an estimate of treatment magnitude.

Secondary endpointEffect measureEstimate95% CIP-value
Objective Response Rate, investigator assessed Risk difference 14.9 percentage points 7.4–22.3 0.0001
PFS, BICR assessed Hazard ratio 0.60 0.49–0.74 <0.0001

9. Statistical Methodology

Stratified log-rank test

The registry reports the stratified log-rank test as the formal method for each of the six primary time-to-event comparisons and for the secondary BICR PFS analysis. A stratified log-rank test compares the observed event pattern between randomized groups while accounting for specified strata.

For the reported treatment comparisons, the analysis text identifies three stratification dimensions: metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status. PD-L1 status was represented by the categories CPS <1, CPS 1 to <10, or CPS ≥10.

Conceptual interpretation
Compare event experience between treatment groups while respecting prespecified strata

The stratified log-rank framework is particularly suited to randomized time-to-event comparisons when important baseline factors were incorporated into the analysis structure.

Cox regression and hazard ratios

The treatment comparison for each reported hazard ratio was based on a Cox regression model with Efron's method of tie handling. Treatment was included as a covariate, and the model was stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.

Hazard-ratio interpretation
HR = estimated instantaneous event rate in treatment ÷ estimated instantaneous event rate in control

An HR below 1 indicates a lower estimated instantaneous event rate in the treatment group relative to the comparator under the fitted model. It is not an absolute risk difference and is not a direct measure of the percentage of participants who benefit.

Score-based confidence interval for proportions

The secondary ORR analysis used the Miettinen & Nurminen method, normalized in the registry data as a score-based confidence interval approach for proportions. The reported effect measure was a difference in percentage, presented here as a risk difference.

This distinction matters because a response endpoint is categorical rather than a time-to-event endpoint. A risk difference directly answers how far apart the response proportions are on an absolute percentage-point scale.

Covariate adjustment and stratification

The analysis text identifies both covariate adjustment and stratified analysis. The Cox model includes treatment as a covariate while stratifying by the listed trial factors. The response-rate analysis likewise uses a treatment comparison stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.

Randomized analysis populations

The primary efficacy analyses were conducted among all randomized participants, with subgroup analyses restricted to randomized participants satisfying the corresponding PD-L1 criterion. Participants were analyzed according to the treatment group to which they were randomized.

This preserves the central comparison created by randomization. It also means that a hazard ratio should not be interpreted as an estimate limited only to participants who completed or adhered to treatment, because the registry-reported analysis population is defined by randomized assignment.

Blinded independent central review

One secondary endpoint measured PFS according to RECIST 1.1 as assessed by blinded independent central review (BICR). This provides an assessment framework separate from investigator-assessed PFS and is relevant when radiologic progression is an endpoint.

10. Statistical Methods Explained

Why was a stratified log-rank test used?

A time-to-event endpoint such as PFS or OS contains information about both whether an event occurred and when it occurred. The stratified log-rank test compares the event experience of the randomized groups over follow-up while accounting for prespecified strata. In KEYNOTE-826, the registry-reported analysis text identifies metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status as stratification factors.

What does an HR of 0.58 mean for PFS?

An HR of 0.58 indicates that the fitted model estimates the instantaneous rate of progression or death in the pembrolizumab-plus-chemotherapy group at approximately 58% of the corresponding rate in the placebo-plus-chemotherapy group. Equivalently, 1 − 0.58 = 0.42, so the model-based relative reduction in the estimated hazard is approximately 42%. This does not mean that 42% of patients avoided progression or death.

Why is a hazard ratio different from a risk difference?

The HR is a relative time-to-event measure that incorporates the timing of events and censoring. The risk difference used for ORR is an absolute difference between proportions. A hazard ratio of 0.60 and a risk difference of 14.9 percentage points therefore describe fundamentally different aspects of a treatment comparison and should not be compared numerically.

Why use the Miettinen & Nurminen method for ORR?

ORR is a binary endpoint: a participant either meets the response definition or does not. The Miettinen & Nurminen method provides a score-based approach to constructing confidence intervals for the difference between proportions. In this trial, the reported treatment comparison was also stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.

What does the 95% confidence interval tell us?

A 95% confidence interval quantifies statistical uncertainty around the estimated treatment effect under the specified model and sampling framework. For example, the PFS HR of 0.58 has a 95% CI of 0.47–0.71. The interval communicates precision around the estimate; it is not a statement that 95% of individual patients will experience treatment effects within that interval.

Why does the p-value not measure treatment effect size?

A p-value evaluates the evidence against a specified null hypothesis under a statistical testing procedure. It depends on the magnitude of the observed difference, the variability of the data, and the amount of information available. A very small p-value therefore does not tell the reader whether an effect is large or small. The effect estimate and confidence interval must be examined to understand magnitude and precision.

Why are the PD-L1 subgroup estimates not automatically evidence of treatment-effect heterogeneity?

The CPS ≥1 and CPS ≥10 analyses are separate analysis populations. Their HRs can be numerically different without establishing that the treatment effect truly differs between populations. A formal claim that a biomarker modifies treatment effect requires an appropriate interaction or heterogeneity analysis. The statistical analyses posted on ClinicalTrials.gov do not report such an interaction estimate.

11. Understanding the Six Primary Results Together

The six primary analyses form a structured set rather than six versions of the same number. Three are PFS analyses and three are OS analyses, each evaluated in an overall or PD-L1-defined population.

QuestionPopulationHR95% CIP-value
Does treatment affect PFS?PD-L1 CPS ≥10.580.47–0.71<0.0001
Does treatment affect PFS?All participants0.610.50–0.74<0.0001
Does treatment affect PFS?PD-L1 CPS ≥100.520.40–0.68<0.0001
Does treatment affect OS?PD-L1 CPS ≥10.600.49–0.74<0.0001
Does treatment affect OS?All participants0.630.52–0.77<0.0001
Does treatment affect OS?PD-L1 CPS ≥100.580.44–0.78<0.0001

Across these six reported estimates, every HR is below 1 and every registry-reported 95% confidence interval lies below 1. The registry also reports a p-value of <0.0001 for each primary analysis. These are descriptive statements about the posted analyses; they should not be expanded into claims about unreported median survival, absolute survival probabilities, or individual patient benefit.

12. Interpreting Relative and Absolute Effects

The primary time-to-event results are expressed as hazard ratios, whereas the secondary ORR result is expressed as a risk difference. Keeping these scales separate is essential for accurate interpretation.

Relative time-to-event effect

An HR such as 0.60 summarizes the relative difference in the modeled instantaneous event rate between randomized groups.

Absolute binary effect

A risk difference such as 14.9 percentage points expresses the absolute difference in the proportion meeting the response endpoint.

Precision

The confidence interval describes uncertainty around the corresponding effect estimate. It must be interpreted on the same scale as the estimate.

Statistical evidence

The p-value describes the evidence from the specified hypothesis-testing procedure. It is not an effect-size or clinical-importance measure.

For example, the OS HR of 0.63 cannot be converted into an absolute percentage-point survival benefit without additional survival estimates. Likewise, the ORR risk difference of 14.9 percentage points cannot be converted into an HR because the response endpoint does not use time-to-event information in the reported analysis.

What the registry does not provide in the ClinicalTrials.gov record: median PFS, median OS, time-specific survival probabilities, event counts for the primary efficacy analyses, baseline characteristic tables, subgroup forest plots, or detailed missing-data/imputation specifications. Those quantities are therefore not presented or reconstructed here.

13. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. The affected and at-risk counts are presented exactly as provided.

GroupSerious adverse events affectedAt risk
Pembrolizumab + Chemotherapy157307
Placebo + Chemotherapy132309
Pembrolizumab (Second Course)412

The ClinicalTrials.gov record is specifically the number affected and the number at risk for serious adverse events. It does not provide a formal between-arm statistical analysis for serious adverse events, confidence intervals, p-values, or a definition of the individual serious adverse-event categories.

Clinical Biostats interpretation

Safety counts answer a different question from the efficacy HRs. An efficacy HR describes a modeled time-to-event comparison, whereas an affected/at-risk safety count describes how many participants experienced the reported safety outcome among those at risk.

The presence of three reported groups also matters. The Pembrolizumab (Second Course) group has 4 affected participants among 12 at risk and is not the same randomized comparison as the two primary treatment arms. It should therefore not be combined with the randomized pembrolizumab-plus-chemotherapy arm or used to construct an overall treatment effect without additional statistical specifications.

14. Stratification in the Analysis

Stratification appears repeatedly in the posted analyses. The treatment comparison for the primary and secondary time-to-event analyses was stratified by:

Stratification factorCategories reported in the analysis text
Metastatic disease at initial diagnosisFIGO (2009) stage IVB: yes or no
Bevacizumab useYes or no
PD-L1 statusCPS <1; CPS 1 to <10; CPS ≥10

Stratification is not the same as simply adjusting for a continuous covariate. In a stratified Cox model, the baseline hazard is allowed to differ across strata while the treatment effect is estimated across the stratified analysis structure. This can make the treatment comparison better aligned with the randomized design when the stratification factors are important prognostic or design variables.

The same general principle appears in the Miettinen & Nurminen response-rate analysis, where the treatment comparison is also stratified by the three registry-reported factors.

15. Analysis Populations and Randomization

For the primary analyses, the ClinicalTrials.gov record consistently define the analysis population as all randomized participants, or all randomized participants satisfying the relevant PD-L1 CPS criterion. Participants are analyzed according to the treatment group to which they were randomized.

Why randomized assignment matters
Randomized assignment → treatment groups defined before outcome information is observed

Analyzing participants according to randomized assignment preserves the treatment contrast generated by randomization. It is therefore important not to reinterpret the posted efficacy HRs as if they were comparisons only among participants who remained on therapy.

The ClinicalTrials.gov record does not specify a separate per-protocol population, a treatment-completer population, or a formal as-treated efficacy analysis. They therefore are not introduced here.

16. What the Hazard Ratios Do — and Do Not — Mean

PFS HR 0.58

The model estimates an instantaneous rate of progression or death approximately 42% lower in the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group. It does not mean that 42% of participants were protected from progression or death.

PFS HR 0.61

In the all-participant population, the estimated instantaneous rate of progression or death was approximately 39% lower under the fitted model. The result is specific to PFS and the all-randomized analysis population.

PFS HR 0.52

In participants with PD-L1 CPS ≥10, the fitted model estimated an instantaneous rate of progression or death approximately 48% lower. This does not by itself establish that the treatment effect differs from that in another PD-L1 population.

OS HR 0.60

Among randomized participants with PD-L1 CPS ≥1, the estimated instantaneous rate of death was approximately 40% lower under the fitted model. The HR is not an absolute survival probability.

OS HR 0.63

In all randomized participants, the estimated instantaneous rate of death was approximately 37% lower under the fitted model. The confidence interval of 0.52–0.77 describes the uncertainty around that estimate.

OS HR 0.58

Among participants with PD-L1 CPS ≥10, the estimated instantaneous rate of death was approximately 42% lower under the fitted model. This remains a subgroup-specific estimate rather than a formal test of biomarker-treatment interaction.

17. Confidence Intervals and Precision

All six primary analyses have posted two-sided 95% confidence intervals. Looking at the interval alongside the point estimate is essential because it reveals how precisely the treatment effect was estimated.

EndpointEstimate95% confidence intervalInterpretive scale
PFS, CPS ≥10.580.47–0.71Hazard ratio
PFS, all participants0.610.50–0.74Hazard ratio
PFS, CPS ≥100.520.40–0.68Hazard ratio
OS, CPS ≥10.600.49–0.74Hazard ratio
OS, all participants0.630.52–0.77Hazard ratio
OS, CPS ≥100.580.44–0.78Hazard ratio
ORR, all participants14.9 percentage points7.4–22.3Risk difference
BICR PFS, all participants0.600.49–0.74Hazard ratio

The confidence intervals are endpoint-specific. The interval for an HR must be interpreted on the hazard-ratio scale, while the interval for ORR must be interpreted on the percentage-point risk-difference scale.

A confidence interval should also not be interpreted as a guarantee that future studies will reproduce exactly the same range. It summarizes uncertainty associated with the particular estimate, data, model, and statistical framework used for the analysis.

18. Statistical Interpretation of the P-Values

The six primary analyses all have a reported p-value of <0.0001. The secondary BICR PFS analysis also has a reported p-value of <0.0001, while the secondary ORR analysis reports 0.0001.

P-value

Addresses evidence against the null hypothesis under the specified testing procedure.

Effect estimate

Describes the magnitude and direction of the observed treatment comparison.

Confidence interval

Describes uncertainty and precision around the corresponding effect estimate.

Clinical meaning

Requires interpretation of effect magnitude, endpoint context, precision, and the design—not the p-value alone.

For example, the difference between an HR of 0.52 and an HR of 0.63 is a difference in estimated treatment effect, not a difference that can be inferred merely from the fact that both p-values are <0.0001. Conversely, two studies could have the same HR but different p-values because their information content differs.

19. Blinding and Independent Assessment

The registry identifies KEYNOTE-826 as triple masked. The ClinicalTrials.gov record also report a secondary PFS endpoint assessed by blinded independent central review using RECIST 1.1.

Blinding is statistically relevant because knowledge of treatment assignment can influence behavior, assessment, or decisions around subjective or partially subjective outcomes. Central review provides an additional assessment framework for radiologic progression. The ClinicalTrials.gov record does not specify the identities of the individuals or roles included in the triple-masked designation, so no further interpretation of which parties were masked is added.

20. RECIST 1.1 and the PFS Event Definition

The registry defines PFS as the time from randomization to the first documented progressive disease or death from any cause, whichever occurs first. The registry-reported RECIST 1.1 definition of progression requires a ≥20% increase in the sum of diameters of target lesions, using the smallest sum on study as the reference, together with an absolute increase of ≥5 mm.

Why the endpoint is time-to-event
Randomization → first PD or death → event time; otherwise follow until censoring

The endpoint combines radiologic progression and death into a single event definition. Its statistical analysis therefore needs to account for the timing of events and participants who have not experienced the event by the end of available follow-up.

This structure explains why the primary PFS analysis uses a stratified log-rank test and hazard ratio rather than a simple comparison of percentages. A binary response endpoint such as ORR, by contrast, is analyzed using the reported Miettinen & Nurminen method.

21. Limitations

22. What This Trial Teaches About Time-to-Event Analysis

KEYNOTE-826 is a useful statistical teaching example because the registry results combine several important concepts in one randomized trial: time-to-event endpoints, stratified comparisons, Cox regression, hazard ratios, confidence intervals, PD-L1-defined analysis populations, blinded independent central review, and a binary response endpoint analyzed with a score-based method.

Statistical conceptHow it appears in KEYNOTE-826
RandomizationRandomized allocation to two parallel treatment arms
BlindingTriple-masked trial design
Time-to-event endpointPFS and OS measured from randomization to the defined event
RECIST 1.1Progression defined using the registry-reported RECIST 1.1 criteria
Stratified log-rank testFormal treatment comparison for the primary time-to-event analyses
Cox regressionUsed for the reported hazard-ratio treatment comparisons
Efron's tie handlingSpecified in the Cox regression analysis text
Hazard ratioEffect measure for PFS and OS treatment comparisons
Confidence intervalTwo-sided 95% intervals for all six primary estimates
Risk differenceEffect measure for the secondary ORR analysis
Miettinen & NurminenScore-based method used for the ORR treatment comparison
Stratified binary analysisORR comparison stratified by the registry-reported baseline factors
Blinded independent central reviewSecondary PFS endpoint assessed independently
Biomarker-defined analysisPFS and OS evaluated in PD-L1 CPS ≥1 and CPS ≥10 populations

23. Why This Trial Matters Statistically

The statistical value of KEYNOTE-826 is not contained in any single hazard ratio. The trial illustrates how a modern randomized oncology study connects design, endpoint definition, analysis population, stratification, model choice, effect measure, and uncertainty.

Endpoint definition comes first

The interpretation of a PFS HR depends on the precise definition of progression and death used to create the endpoint.

Analysis follows endpoint type

Time-to-event endpoints use survival-analysis methods, while ORR uses a binary-outcome framework.

Stratification connects design and analysis

The same registry-reported baseline factors recur in the formal treatment comparisons, linking the randomized design to the statistical model.

Precision matters

Each effect estimate should be read with its confidence interval rather than as an isolated point estimate.

The trial also demonstrates why statistical interpretation should remain separate from numerical repetition. Saying that an HR is below 1 is descriptive. Explaining that an HR of 0.58 represents an estimated 42% lower instantaneous event rate under the fitted model is statistical interpretation. Neither statement, by itself, establishes how an individual patient will respond.

24. A Statistical Reading of the Trial Results

Step 1 · Identify the estimand scale

PFS and OS are reported using hazard ratios, while ORR is reported using a risk difference. The first step in interpretation is therefore to avoid treating all estimates as if they were the same type of quantity.

Step 2 · Identify the population

The six primary estimates correspond to all randomized participants, PD-L1 CPS ≥1, or PD-L1 CPS ≥10. The population must be read before comparing the numerical estimates.

Step 3 · Read the confidence interval

The confidence interval shows how precisely the effect was estimated. For the six primary HRs, the registry-reported two-sided 95% intervals are all below 1.

Step 4 · Read the p-value separately

The p-value provides hypothesis-testing evidence. It does not replace the effect estimate or confidence interval and should not be interpreted as a probability that the treatment is effective.

Step 5 · Check the analysis method

The primary time-to-event analyses use stratified log-rank testing and Cox regression with Efron's method of tie handling. The secondary ORR analysis uses the Miettinen & Nurminen method.

Step 6 · Keep efficacy and safety separate

The serious-adverse-event counts provide safety information but are not the same statistical object as the PFS and OS hazard ratios. A complete interpretation considers both domains without collapsing them into a single numerical score.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue with the underlying statistical methods

Explore the survival-analysis, clinical-trial, confidence-interval, and binary-outcome methods that appear in KEYNOTE-826.

28. Record Summary

KEYNOTE-826 provides a detailed example of how randomized clinical-trial evidence can be analyzed across multiple endpoint definitions and populations. The registry reports six primary analyses covering PFS and OS in all randomized participants and in PD-L1 CPS-defined populations. Each primary time-to-event analysis uses a stratified log-rank framework, with treatment-effect estimation through a Cox regression model using Efron's method of tie handling. The secondary analyses add an investigator-assessed ORR comparison using the Miettinen & Nurminen method and a BICR-assessed PFS comparison.

The most informative statistical reading therefore combines the hazard ratio, the 95% confidence interval, the p-value, the analysis population, the endpoint definition, and the stratification structure. For ORR, the corresponding interpretation uses the risk difference and its confidence interval. Safety is reported separately through serious-adverse-event counts by group.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Numerical results on this page are limited to the registry-reported KEYNOTE-826 trial data, while the surrounding explanations describe what the reported statistical quantities mean and what they do not establish.