This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View the ClinicalTrials.gov record.
1. Trial at a Glance
KEYNOTE-826 was a randomized, parallel, triple-masked phase 3 trial evaluating first-line pembrolizumab plus chemotherapy versus placebo plus chemotherapy in women with persistent, recurrent, or metastatic cervical cancer. The registry reports 617 enrolled participants, 2 arms, 6 registered primary endpoints, 15 posted outcome measures, and 8 posted statistical analyses.
| Feature | KEYNOTE-826 |
|---|---|
| Trial name | KEYNOTE-826 |
| NCT ID | NCT03635567 |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Cervical Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Triple |
| Primary purpose | Treatment |
| Enrollment | 617 |
| Arms | 2 |
| Primary endpoint types | Binary; Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 15 |
| Statistical analyses posted | 8 |
| Primary-endpoint analyses | 6 |
| Primary analyses with estimate + CI | 6 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
2. Clinical Question
The registry describes KEYNOTE-826 as a study of first-line treatment with pembrolizumab plus chemotherapy versus placebo plus chemotherapy in women with persistent, recurrent, or metastatic cervical cancer. The statistical question is whether randomized assignment to the pembrolizumab-containing regimen is associated with different progression-free survival and overall survival, including in prespecified PD-L1 Combined Positive Score populations.
Population
Women with persistent, recurrent, or metastatic cervical cancer enrolled in the phase 3 randomized trial.
Intervention
Pembrolizumab plus chemotherapy. The registered interventions include pembrolizumab, paclitaxel, cisplatin, carboplatin, and bevacizumab.
Comparator
Placebo to pembrolizumab plus chemotherapy. The registry lists placebo to pembrolizumab as a drug intervention.
Primary question
How does pembrolizumab plus chemotherapy compare with placebo plus chemotherapy for PFS and OS in the registered analysis populations?
3. Trial Design
Pembrolizumab + Chemotherapy
- Pembrolizumab
- Paclitaxel
- Cisplatin or carboplatin
- Bevacizumab is also listed among the registered biological/drug interventions
Placebo + Chemotherapy
- Placebo to pembrolizumab
- Paclitaxel
- Cisplatin or carboplatin
- Bevacizumab is also listed among the registered biological/drug interventions
The ClinicalTrials.gov record identifies the treatment comparison as pembrolizumab + chemotherapy versus placebo + chemotherapy. It does not provide an arm-specific randomized sample size in the ClinicalTrials.gov record, so this page does not infer one from the total enrollment.
4. Endpoints
The registry lists six primary endpoints. Three concern progression-free survival and three concern overall survival. The endpoint populations distinguish all randomized participants from participants meeting PD-L1 Combined Positive Score thresholds.
| Primary endpoint | Time frame | Type | Analysis |
|---|---|---|---|
| Progression-free Survival (PFS) Per Response Evaluation Criteria in Solid Tumors Version 1.1 (RECIST 1.1) as Assessed by Investigator in Participants With Programmed Cell Death-Ligand 1 (PD-L1) Combined Positive Score (CPS) ≥1 | Up to approximately 46 months | Time-to-event | Stratified log-rank test; HR |
| PFS Per RECIST 1.1 as Assessed by Investigator in All Participants | Up to approximately 46 months | Binary | Stratified log-rank test; HR |
| PFS Per RECIST 1.1 as Assessed by Investigator in Participants With PD-L1 CPS ≥10 | Up to approximately 46 months | Binary | Stratified log-rank test; HR |
| Overall Survival (OS) in Participants With PD-L1 CPS ≥1 | Up to approximately 46 months | Time-to-event | Stratified log-rank test; HR |
| OS in All Participants | Up to approximately 46 months | Binary | Stratified log-rank test; HR |
| OS in Participants With PD-L1 CPS ≥10 | Up to approximately 46 months | Binary | Stratified log-rank test; HR |
Registry definitions of the primary endpoints
Progression-free survival: PFS was defined as the time from randomization to the first documented progressive disease (PD) or death due to any cause, whichever occurs first. Per RECIST 1.1, PD was defined as ≥ 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study. In addition to the relative increase of 20%, the sum must also demonstrate an absolute increase of ≥5 mm.
Overall survival: OS was defined as the time from randomization to death due to any cause. The OS for all randomized participants with PD-L1 CPS ≥1, all randomized participants, and all randomized participants with PD-L1 CPS ≥10 was presented according to the respective registered endpoint.
5. Results Overview
ClinicalTrials.gov contains formal statistical analyses for all six registered primary endpoints and two secondary endpoints in the ClinicalTrials.gov record. All six primary analyses have an estimate and 95% confidence interval. The primary time-to-event analyses consistently use a stratified log-rank framework with treatment-effect estimation through a Cox regression model.
| Endpoint family | Population | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| PFS | PD-L1 CPS ≥1 | HR 0.58 | 0.47–0.71 | <0.0001 |
| PFS | All participants | HR 0.61 | 0.50–0.74 | <0.0001 |
| PFS | PD-L1 CPS ≥10 | HR 0.52 | 0.40–0.68 | <0.0001 |
| OS | PD-L1 CPS ≥1 | HR 0.60 | 0.49–0.74 | <0.0001 |
| OS | All participants | HR 0.63 | 0.52–0.77 | <0.0001 |
| OS | PD-L1 CPS ≥10 | HR 0.58 | 0.44–0.78 | <0.0001 |
These estimates are not interchangeable. Each belongs to a different endpoint and analysis population. The CPS ≥1 and CPS ≥10 analyses answer questions in biomarker-defined populations, whereas the all-participant analyses describe the randomized population without that PD-L1 restriction.
6. Primary Results: Progression-Free Survival
PFS in Participants With PD-L1 CPS ≥1
Hazard ratio for progression or death
95% CI: 0.47–0.71 · P < 0.0001
Analysis population: all randomized participants with PD-L1 CPS ≥1, analyzed according to randomized treatment group.
The registry reports a stratified log-rank treatment comparison. The treatment comparison for the hazard ratio was based on a Cox regression model using Efron's method of tie handling, with treatment as a covariate and stratification by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.
An HR of 0.58 means that the fitted time-to-event model estimated an instantaneous rate of progression or death approximately 42% lower in the pembrolizumab-plus-chemotherapy group than in the placebo-plus-chemotherapy group, because 1 − 0.58 = 0.42.
The HR does not mean that 42% of participants avoided progression or death, that median PFS was 42% longer, or that every individual participant experienced a 42% reduction in risk. It is a relative model-based measure of the event rate over time.
The 95% CI of 0.47–0.71 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of treatment effects among individual patients.
The p-value of <0.0001 addresses the evidence against the null comparison specified by the statistical test. It does not measure the size, clinical importance, or certainty of the treatment effect. The HR and its confidence interval are needed to describe the estimated magnitude and precision.
Because the estimate is based on a Cox regression model, interpretation also depends on the appropriateness of the model and, in particular, the proportional-hazards framework. A single HR summarizes a time-to-event comparison but does not show how the separation between treatment groups evolves at each point in follow-up.
PFS in All Participants
Hazard ratio for progression or death
95% CI: 0.50–0.74 · P < 0.0001
Analysis population: all randomized participants.
The all-participant analysis provides a broader randomized-population estimate than the PD-L1 CPS ≥1 and CPS ≥10 analyses. The registry again specifies a stratified log-rank comparison and a Cox regression model with Efron's method of tie handling.
An HR of 0.61 corresponds to an estimated 39% lower instantaneous rate of progression or death under the fitted model for the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group.
The confidence interval of 0.50–0.74 communicates the precision of the estimate. It is narrower than simply reporting the point estimate because it describes a range of values compatible with the specified statistical framework and observed data.
The p-value of <0.0001 is evidence from the stated hypothesis-testing procedure; it is not a probability that the treatment effect is true, nor does it quantify the clinical magnitude of the difference.
The comparison is based on randomized treatment assignment rather than observed treatment received, which is important when interpreting an efficacy analysis in a randomized trial. The ClinicalTrials.gov record does not provide additional information about treatment discontinuation, crossover, or missing-data handling, so those issues should not be inferred from the HR alone.
PFS in Participants With PD-L1 CPS ≥10
Hazard ratio for progression or death
95% CI: 0.40–0.68 · P < 0.0001
Analysis population: all randomized participants with PD-L1 CPS ≥10.
An HR of 0.52 corresponds to an estimated instantaneous rate of progression or death approximately 48% lower in the pembrolizumab-plus-chemotherapy group under the fitted model.
The 95% CI of 0.40–0.68 indicates the statistical uncertainty around the estimate. It should not be interpreted as saying that the true effect for individual patients must lie somewhere between a 32% and 60% reduction in their personal risk.
The p-value of <0.0001 indicates strong evidence under the reported testing procedure, but it is not an effect-size metric. The HR remains the primary measure of relative treatment effect for this analysis.
This is a biomarker-defined analysis population. A comparison between the CPS ≥10 HR and another subgroup's HR does not, by itself, establish that PD-L1 status modifies the treatment effect. Demonstrating effect modification requires an appropriate comparison or interaction analysis, and the registry-reported statistical analysis does not report such an interaction estimate.
7. Primary Results: Overall Survival
OS in Participants With PD-L1 CPS ≥1
Hazard ratio for death
95% CI: 0.49–0.74 · P < 0.0001
Analysis population: all randomized participants with PD-L1 CPS ≥1.
An HR of 0.60 means that the fitted model estimated an instantaneous rate of death approximately 40% lower in the pembrolizumab-plus-chemotherapy group than in the placebo-plus-chemotherapy group.
This does not mean that 40% of participants survived because of treatment, that 40% of participants were cured, or that the absolute probability of death was reduced by exactly 40 percentage points. Hazard ratios are relative time-to-event measures rather than absolute risk differences.
The 95% CI of 0.49–0.74 describes uncertainty in the estimated HR. It provides more information than the point estimate alone because an estimate of 0.60 based on a very imprecise interval would carry a different statistical interpretation from the same point estimate with a narrow interval.
The p-value of <0.0001 is evidence from the stratified log-rank testing framework. It does not tell the reader how large the treatment effect is; that information comes from the HR and its confidence interval.
OS in All Participants
Hazard ratio for death
95% CI: 0.52–0.77 · P < 0.0001
Analysis population: all randomized participants.
The HR of 0.63 corresponds to an estimated instantaneous rate of death approximately 37% lower in the pembrolizumab-plus-chemotherapy group under the fitted Cox model.
The 95% CI of 0.52–0.77 gives the uncertainty around that estimate. It does not give a range of individual survival probabilities and does not replace presentation of absolute survival probabilities or other clinically interpretable time-specific measures.
The p-value of <0.0001 should be read as a hypothesis-test result rather than as a measure of treatment benefit. Large datasets can produce small p-values for relatively modest effects, while small datasets can produce imprecise estimates even when the point estimate is substantial.
The analysis uses randomized treatment assignment and stratification variables specified in the registry analysis text. The ClinicalTrials.gov record does not report a separate absolute survival estimate or median OS, so this page does not infer one from the HR.
OS in Participants With PD-L1 CPS ≥10
Hazard ratio for death
95% CI: 0.44–0.78 · P < 0.0001
Analysis population: all randomized participants with PD-L1 CPS ≥10.
An HR of 0.58 corresponds to an estimated instantaneous rate of death approximately 42% lower in the pembrolizumab-plus-chemotherapy group under the fitted model.
The 95% CI of 0.44–0.78 indicates uncertainty around the estimate. The interval should not be treated as a prediction interval for individual patients or as a statement that all patients experience an effect within that numerical range.
The p-value of <0.0001 indicates the result of the reported statistical comparison. It is not the probability that the null hypothesis is true and does not indicate the practical importance of the treatment effect.
Because this is a PD-L1 CPS ≥10 analysis, the estimate describes a defined subgroup rather than automatically establishing that the treatment effect is stronger or weaker than in the other registered populations.
8. Secondary Endpoint Results
Objective Response Rate
The registry reports Objective Response Rate (ORR) Per RECIST 1.1 as Assessed by Investigator as a secondary endpoint. The outcome unit is percentage of participants, and the analysis population consists of all randomized participants based on the treatment group to which they were randomized.
Risk difference in objective response rate
95% CI: 7.4–22.3 · P = 0.0001
Analysis: Miettinen & Nurminen method, stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.
The reported risk difference of 14.9 percentage points means that the observed percentage of participants achieving an objective response differed by 14.9 percentage points between the randomized treatment groups, with the direction corresponding to the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group.
This is an absolute difference, unlike a hazard ratio. It is therefore directly expressed in percentage points rather than as a relative rate over time.
The 95% CI of 7.4–22.3 describes uncertainty around the estimated difference. It is not the range of response rates for individual patients.
The p-value of 0.0001 addresses the statistical comparison under the reported Miettinen & Nurminen method. It does not measure the magnitude of the response-rate difference. The 14.9-point estimate and its confidence interval provide that information.
PFS by Blinded Independent Central Review
Hazard ratio for progression or death
95% CI: 0.49–0.74 · P < 0.0001
Analysis population: all randomized participants. Assessment: RECIST 1.1 by blinded independent central review.
The BICR PFS analysis produced an HR of 0.60, corresponding to an estimated instantaneous rate of progression or death approximately 40% lower under the fitted model in the pembrolizumab-plus-chemotherapy group.
The 95% CI of 0.49–0.74 describes uncertainty around the relative treatment effect. The BICR assessment is statistically useful because it provides an assessment framework distinct from investigator assessment, but the ClinicalTrials.gov record does not provide a comparison between the two assessment approaches beyond their respective posted results.
The p-value of <0.0001 is evidence from the reported stratified log-rank comparison and should not be interpreted as an estimate of treatment magnitude.
| Secondary endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Objective Response Rate, investigator assessed | Risk difference | 14.9 percentage points | 7.4–22.3 | 0.0001 |
| PFS, BICR assessed | Hazard ratio | 0.60 | 0.49–0.74 | <0.0001 |
9. Statistical Methodology
Stratified log-rank test
The registry reports the stratified log-rank test as the formal method for each of the six primary time-to-event comparisons and for the secondary BICR PFS analysis. A stratified log-rank test compares the observed event pattern between randomized groups while accounting for specified strata.
For the reported treatment comparisons, the analysis text identifies three stratification dimensions: metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status. PD-L1 status was represented by the categories CPS <1, CPS 1 to <10, or CPS ≥10.
The stratified log-rank framework is particularly suited to randomized time-to-event comparisons when important baseline factors were incorporated into the analysis structure.
Cox regression and hazard ratios
The treatment comparison for each reported hazard ratio was based on a Cox regression model with Efron's method of tie handling. Treatment was included as a covariate, and the model was stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.
An HR below 1 indicates a lower estimated instantaneous event rate in the treatment group relative to the comparator under the fitted model. It is not an absolute risk difference and is not a direct measure of the percentage of participants who benefit.
Score-based confidence interval for proportions
The secondary ORR analysis used the Miettinen & Nurminen method, normalized in the registry data as a score-based confidence interval approach for proportions. The reported effect measure was a difference in percentage, presented here as a risk difference.
This distinction matters because a response endpoint is categorical rather than a time-to-event endpoint. A risk difference directly answers how far apart the response proportions are on an absolute percentage-point scale.
Covariate adjustment and stratification
The analysis text identifies both covariate adjustment and stratified analysis. The Cox model includes treatment as a covariate while stratifying by the listed trial factors. The response-rate analysis likewise uses a treatment comparison stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.
Randomized analysis populations
The primary efficacy analyses were conducted among all randomized participants, with subgroup analyses restricted to randomized participants satisfying the corresponding PD-L1 criterion. Participants were analyzed according to the treatment group to which they were randomized.
This preserves the central comparison created by randomization. It also means that a hazard ratio should not be interpreted as an estimate limited only to participants who completed or adhered to treatment, because the registry-reported analysis population is defined by randomized assignment.
Blinded independent central review
One secondary endpoint measured PFS according to RECIST 1.1 as assessed by blinded independent central review (BICR). This provides an assessment framework separate from investigator-assessed PFS and is relevant when radiologic progression is an endpoint.
10. Statistical Methods Explained
Why was a stratified log-rank test used?
A time-to-event endpoint such as PFS or OS contains information about both whether an event occurred and when it occurred. The stratified log-rank test compares the event experience of the randomized groups over follow-up while accounting for prespecified strata. In KEYNOTE-826, the registry-reported analysis text identifies metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status as stratification factors.
What does an HR of 0.58 mean for PFS?
An HR of 0.58 indicates that the fitted model estimates the instantaneous rate of progression or death in the pembrolizumab-plus-chemotherapy group at approximately 58% of the corresponding rate in the placebo-plus-chemotherapy group. Equivalently, 1 − 0.58 = 0.42, so the model-based relative reduction in the estimated hazard is approximately 42%. This does not mean that 42% of patients avoided progression or death.
Why is a hazard ratio different from a risk difference?
The HR is a relative time-to-event measure that incorporates the timing of events and censoring. The risk difference used for ORR is an absolute difference between proportions. A hazard ratio of 0.60 and a risk difference of 14.9 percentage points therefore describe fundamentally different aspects of a treatment comparison and should not be compared numerically.
Why use the Miettinen & Nurminen method for ORR?
ORR is a binary endpoint: a participant either meets the response definition or does not. The Miettinen & Nurminen method provides a score-based approach to constructing confidence intervals for the difference between proportions. In this trial, the reported treatment comparison was also stratified by metastatic disease at initial diagnosis, bevacizumab use, and PD-L1 status.
What does the 95% confidence interval tell us?
A 95% confidence interval quantifies statistical uncertainty around the estimated treatment effect under the specified model and sampling framework. For example, the PFS HR of 0.58 has a 95% CI of 0.47–0.71. The interval communicates precision around the estimate; it is not a statement that 95% of individual patients will experience treatment effects within that interval.
Why does the p-value not measure treatment effect size?
A p-value evaluates the evidence against a specified null hypothesis under a statistical testing procedure. It depends on the magnitude of the observed difference, the variability of the data, and the amount of information available. A very small p-value therefore does not tell the reader whether an effect is large or small. The effect estimate and confidence interval must be examined to understand magnitude and precision.
Why are the PD-L1 subgroup estimates not automatically evidence of treatment-effect heterogeneity?
The CPS ≥1 and CPS ≥10 analyses are separate analysis populations. Their HRs can be numerically different without establishing that the treatment effect truly differs between populations. A formal claim that a biomarker modifies treatment effect requires an appropriate interaction or heterogeneity analysis. The statistical analyses posted on ClinicalTrials.gov do not report such an interaction estimate.
11. Understanding the Six Primary Results Together
The six primary analyses form a structured set rather than six versions of the same number. Three are PFS analyses and three are OS analyses, each evaluated in an overall or PD-L1-defined population.
| Question | Population | HR | 95% CI | P-value |
|---|---|---|---|---|
| Does treatment affect PFS? | PD-L1 CPS ≥1 | 0.58 | 0.47–0.71 | <0.0001 |
| Does treatment affect PFS? | All participants | 0.61 | 0.50–0.74 | <0.0001 |
| Does treatment affect PFS? | PD-L1 CPS ≥10 | 0.52 | 0.40–0.68 | <0.0001 |
| Does treatment affect OS? | PD-L1 CPS ≥1 | 0.60 | 0.49–0.74 | <0.0001 |
| Does treatment affect OS? | All participants | 0.63 | 0.52–0.77 | <0.0001 |
| Does treatment affect OS? | PD-L1 CPS ≥10 | 0.58 | 0.44–0.78 | <0.0001 |
Across these six reported estimates, every HR is below 1 and every registry-reported 95% confidence interval lies below 1. The registry also reports a p-value of <0.0001 for each primary analysis. These are descriptive statements about the posted analyses; they should not be expanded into claims about unreported median survival, absolute survival probabilities, or individual patient benefit.
12. Interpreting Relative and Absolute Effects
The primary time-to-event results are expressed as hazard ratios, whereas the secondary ORR result is expressed as a risk difference. Keeping these scales separate is essential for accurate interpretation.
Relative time-to-event effect
An HR such as 0.60 summarizes the relative difference in the modeled instantaneous event rate between randomized groups.
Absolute binary effect
A risk difference such as 14.9 percentage points expresses the absolute difference in the proportion meeting the response endpoint.
Precision
The confidence interval describes uncertainty around the corresponding effect estimate. It must be interpreted on the same scale as the estimate.
Statistical evidence
The p-value describes the evidence from the specified hypothesis-testing procedure. It is not an effect-size or clinical-importance measure.
For example, the OS HR of 0.63 cannot be converted into an absolute percentage-point survival benefit without additional survival estimates. Likewise, the ORR risk difference of 14.9 percentage points cannot be converted into an HR because the response endpoint does not use time-to-event information in the reported analysis.
13. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. The affected and at-risk counts are presented exactly as provided.
| Group | Serious adverse events affected | At risk |
|---|---|---|
| Pembrolizumab + Chemotherapy | 157 | 307 |
| Placebo + Chemotherapy | 132 | 309 |
| Pembrolizumab (Second Course) | 4 | 12 |
The ClinicalTrials.gov record is specifically the number affected and the number at risk for serious adverse events. It does not provide a formal between-arm statistical analysis for serious adverse events, confidence intervals, p-values, or a definition of the individual serious adverse-event categories.
Safety counts answer a different question from the efficacy HRs. An efficacy HR describes a modeled time-to-event comparison, whereas an affected/at-risk safety count describes how many participants experienced the reported safety outcome among those at risk.
The presence of three reported groups also matters. The Pembrolizumab (Second Course) group has 4 affected participants among 12 at risk and is not the same randomized comparison as the two primary treatment arms. It should therefore not be combined with the randomized pembrolizumab-plus-chemotherapy arm or used to construct an overall treatment effect without additional statistical specifications.
14. Stratification in the Analysis
Stratification appears repeatedly in the posted analyses. The treatment comparison for the primary and secondary time-to-event analyses was stratified by:
| Stratification factor | Categories reported in the analysis text |
|---|---|
| Metastatic disease at initial diagnosis | FIGO (2009) stage IVB: yes or no |
| Bevacizumab use | Yes or no |
| PD-L1 status | CPS <1; CPS 1 to <10; CPS ≥10 |
Stratification is not the same as simply adjusting for a continuous covariate. In a stratified Cox model, the baseline hazard is allowed to differ across strata while the treatment effect is estimated across the stratified analysis structure. This can make the treatment comparison better aligned with the randomized design when the stratification factors are important prognostic or design variables.
The same general principle appears in the Miettinen & Nurminen response-rate analysis, where the treatment comparison is also stratified by the three registry-reported factors.
15. Analysis Populations and Randomization
For the primary analyses, the ClinicalTrials.gov record consistently define the analysis population as all randomized participants, or all randomized participants satisfying the relevant PD-L1 CPS criterion. Participants are analyzed according to the treatment group to which they were randomized.
Analyzing participants according to randomized assignment preserves the treatment contrast generated by randomization. It is therefore important not to reinterpret the posted efficacy HRs as if they were comparisons only among participants who remained on therapy.
The ClinicalTrials.gov record does not specify a separate per-protocol population, a treatment-completer population, or a formal as-treated efficacy analysis. They therefore are not introduced here.
16. What the Hazard Ratios Do — and Do Not — Mean
The model estimates an instantaneous rate of progression or death approximately 42% lower in the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group. It does not mean that 42% of participants were protected from progression or death.
In the all-participant population, the estimated instantaneous rate of progression or death was approximately 39% lower under the fitted model. The result is specific to PFS and the all-randomized analysis population.
In participants with PD-L1 CPS ≥10, the fitted model estimated an instantaneous rate of progression or death approximately 48% lower. This does not by itself establish that the treatment effect differs from that in another PD-L1 population.
Among randomized participants with PD-L1 CPS ≥1, the estimated instantaneous rate of death was approximately 40% lower under the fitted model. The HR is not an absolute survival probability.
In all randomized participants, the estimated instantaneous rate of death was approximately 37% lower under the fitted model. The confidence interval of 0.52–0.77 describes the uncertainty around that estimate.
Among participants with PD-L1 CPS ≥10, the estimated instantaneous rate of death was approximately 42% lower under the fitted model. This remains a subgroup-specific estimate rather than a formal test of biomarker-treatment interaction.
17. Confidence Intervals and Precision
All six primary analyses have posted two-sided 95% confidence intervals. Looking at the interval alongside the point estimate is essential because it reveals how precisely the treatment effect was estimated.
| Endpoint | Estimate | 95% confidence interval | Interpretive scale |
|---|---|---|---|
| PFS, CPS ≥1 | 0.58 | 0.47–0.71 | Hazard ratio |
| PFS, all participants | 0.61 | 0.50–0.74 | Hazard ratio |
| PFS, CPS ≥10 | 0.52 | 0.40–0.68 | Hazard ratio |
| OS, CPS ≥1 | 0.60 | 0.49–0.74 | Hazard ratio |
| OS, all participants | 0.63 | 0.52–0.77 | Hazard ratio |
| OS, CPS ≥10 | 0.58 | 0.44–0.78 | Hazard ratio |
| ORR, all participants | 14.9 percentage points | 7.4–22.3 | Risk difference |
| BICR PFS, all participants | 0.60 | 0.49–0.74 | Hazard ratio |
The confidence intervals are endpoint-specific. The interval for an HR must be interpreted on the hazard-ratio scale, while the interval for ORR must be interpreted on the percentage-point risk-difference scale.
A confidence interval should also not be interpreted as a guarantee that future studies will reproduce exactly the same range. It summarizes uncertainty associated with the particular estimate, data, model, and statistical framework used for the analysis.
18. Statistical Interpretation of the P-Values
The six primary analyses all have a reported p-value of <0.0001. The secondary BICR PFS analysis also has a reported p-value of <0.0001, while the secondary ORR analysis reports 0.0001.
P-value
Addresses evidence against the null hypothesis under the specified testing procedure.
Effect estimate
Describes the magnitude and direction of the observed treatment comparison.
Confidence interval
Describes uncertainty and precision around the corresponding effect estimate.
Clinical meaning
Requires interpretation of effect magnitude, endpoint context, precision, and the design—not the p-value alone.
For example, the difference between an HR of 0.52 and an HR of 0.63 is a difference in estimated treatment effect, not a difference that can be inferred merely from the fact that both p-values are <0.0001. Conversely, two studies could have the same HR but different p-values because their information content differs.
19. Blinding and Independent Assessment
The registry identifies KEYNOTE-826 as triple masked. The ClinicalTrials.gov record also report a secondary PFS endpoint assessed by blinded independent central review using RECIST 1.1.
Blinding is statistically relevant because knowledge of treatment assignment can influence behavior, assessment, or decisions around subjective or partially subjective outcomes. Central review provides an additional assessment framework for radiologic progression. The ClinicalTrials.gov record does not specify the identities of the individuals or roles included in the triple-masked designation, so no further interpretation of which parties were masked is added.
20. RECIST 1.1 and the PFS Event Definition
The registry defines PFS as the time from randomization to the first documented progressive disease or death from any cause, whichever occurs first. The registry-reported RECIST 1.1 definition of progression requires a ≥20% increase in the sum of diameters of target lesions, using the smallest sum on study as the reference, together with an absolute increase of ≥5 mm.
The endpoint combines radiologic progression and death into a single event definition. Its statistical analysis therefore needs to account for the timing of events and participants who have not experienced the event by the end of available follow-up.
This structure explains why the primary PFS analysis uses a stratified log-rank test and hazard ratio rather than a simple comparison of percentages. A binary response endpoint such as ORR, by contrast, is analyzed using the reported Miettinen & Nurminen method.
21. Limitations
- Registry-level detail: this page is constrained to the ClinicalTrials.gov record. It does not add numerical results from publications or other sources.
- Missing absolute time-to-event summaries: the ClinicalTrials.gov record provides hazard ratios and confidence intervals but do not provide median PFS, median OS, or time-specific survival probabilities.
- No event counts for efficacy endpoints: the statistical analyses posted on ClinicalTrials.gov do not report the number of PFS or OS events supporting each hazard ratio.
- Subgroup interpretation: CPS ≥1 and CPS ≥10 estimates describe defined analysis populations. Differences between subgroup HRs do not establish treatment-effect heterogeneity without an appropriate interaction analysis.
- Proportional-hazards framework: the Cox model provides a compact hazard-ratio summary. The HR should not automatically be treated as a constant clinical effect at every point in time without considering the underlying event-time patterns.
- Multiplicity: six primary endpoints and additional secondary analyses create a broader inferential structure than a single hypothesis test. The ClinicalTrials.gov record identifies the endpoint and analysis structure but do not provide a complete multiplicity-adjustment specification for every posted analysis.
- Safety interpretation: serious adverse-event counts are provided, but no formal comparative statistical analysis is posted on ClinicalTrials.gov for that safety endpoint.
- Missing-data and imputation information: the ClinicalTrials.gov record does not report a detailed missing-data or imputation strategy, so none is inferred.
- Analysis-population detail: the primary efficacy populations are explicitly described as randomized participants, but the ClinicalTrials.gov record does not provide additional protocol-level populations for sensitivity analyses.
22. What This Trial Teaches About Time-to-Event Analysis
KEYNOTE-826 is a useful statistical teaching example because the registry results combine several important concepts in one randomized trial: time-to-event endpoints, stratified comparisons, Cox regression, hazard ratios, confidence intervals, PD-L1-defined analysis populations, blinded independent central review, and a binary response endpoint analyzed with a score-based method.
| Statistical concept | How it appears in KEYNOTE-826 |
|---|---|
| Randomization | Randomized allocation to two parallel treatment arms |
| Blinding | Triple-masked trial design |
| Time-to-event endpoint | PFS and OS measured from randomization to the defined event |
| RECIST 1.1 | Progression defined using the registry-reported RECIST 1.1 criteria |
| Stratified log-rank test | Formal treatment comparison for the primary time-to-event analyses |
| Cox regression | Used for the reported hazard-ratio treatment comparisons |
| Efron's tie handling | Specified in the Cox regression analysis text |
| Hazard ratio | Effect measure for PFS and OS treatment comparisons |
| Confidence interval | Two-sided 95% intervals for all six primary estimates |
| Risk difference | Effect measure for the secondary ORR analysis |
| Miettinen & Nurminen | Score-based method used for the ORR treatment comparison |
| Stratified binary analysis | ORR comparison stratified by the registry-reported baseline factors |
| Blinded independent central review | Secondary PFS endpoint assessed independently |
| Biomarker-defined analysis | PFS and OS evaluated in PD-L1 CPS ≥1 and CPS ≥10 populations |
23. Why This Trial Matters Statistically
The statistical value of KEYNOTE-826 is not contained in any single hazard ratio. The trial illustrates how a modern randomized oncology study connects design, endpoint definition, analysis population, stratification, model choice, effect measure, and uncertainty.
Endpoint definition comes first
The interpretation of a PFS HR depends on the precise definition of progression and death used to create the endpoint.
Analysis follows endpoint type
Time-to-event endpoints use survival-analysis methods, while ORR uses a binary-outcome framework.
Stratification connects design and analysis
The same registry-reported baseline factors recur in the formal treatment comparisons, linking the randomized design to the statistical model.
Precision matters
Each effect estimate should be read with its confidence interval rather than as an isolated point estimate.
The trial also demonstrates why statistical interpretation should remain separate from numerical repetition. Saying that an HR is below 1 is descriptive. Explaining that an HR of 0.58 represents an estimated 42% lower instantaneous event rate under the fitted model is statistical interpretation. Neither statement, by itself, establishes how an individual patient will respond.
24. A Statistical Reading of the Trial Results
PFS and OS are reported using hazard ratios, while ORR is reported using a risk difference. The first step in interpretation is therefore to avoid treating all estimates as if they were the same type of quantity.
The six primary estimates correspond to all randomized participants, PD-L1 CPS ≥1, or PD-L1 CPS ≥10. The population must be read before comparing the numerical estimates.
The confidence interval shows how precisely the effect was estimated. For the six primary HRs, the registry-reported two-sided 95% intervals are all below 1.
The p-value provides hypothesis-testing evidence. It does not replace the effect estimate or confidence interval and should not be interpreted as a probability that the treatment is effective.
The primary time-to-event analyses use stratified log-rank testing and Cox regression with Efron's method of tie handling. The secondary ORR analysis uses the Miettinen & Nurminen method.
The serious-adverse-event counts provide safety information but are not the same statistical object as the PFS and OS hazard ratios. A complete interpretation considers both domains without collapsing them into a single numerical score.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: NCT03635567 — KEYNOTE-826.
- PubMed: PMID 37910822.
- PubMed: PMID 42720936.
- PubMed: PMID 40590325.
- PubMed: PMID 39499492.
- PubMed: PMID 39393777.
Continue with the underlying statistical methods
Explore the survival-analysis, clinical-trial, confidence-interval, and binary-outcome methods that appear in KEYNOTE-826.
28. Record Summary
KEYNOTE-826 provides a detailed example of how randomized clinical-trial evidence can be analyzed across multiple endpoint definitions and populations. The registry reports six primary analyses covering PFS and OS in all randomized participants and in PD-L1 CPS-defined populations. Each primary time-to-event analysis uses a stratified log-rank framework, with treatment-effect estimation through a Cox regression model using Efron's method of tie handling. The secondary analyses add an investigator-assessed ORR comparison using the Miettinen & Nurminen method and a BICR-assessed PFS comparison.
The most informative statistical reading therefore combines the hazard ratio, the 95% confidence interval, the p-value, the analysis population, the endpoint definition, and the stratification structure. For ORR, the corresponding interpretation uses the risk difference and its confidence interval. Safety is reported separately through serious-adverse-event counts by group.