This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics on this page are taken from the ClinicalTrials.gov record. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-177 was a randomized, open-label, parallel phase 3 trial evaluating pembrolizumab versus standard of care in participants with MSI-H or dMMR stage IV colorectal carcinoma. The registry reports 307 enrolled participants, two arms, two registered primary endpoints, and posted formal analyses for both primary endpoints.
| Feature | KEYNOTE-177 |
|---|---|
| Trial name | KEYNOTE-177 |
| Phase | Phase 3 |
| Condition | Colorectal carcinoma |
| Population | Participants with MSI-H or dMMR stage IV colorectal carcinoma |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 307 |
| Primary endpoints | Progression-Free Survival (PFS) and Overall Survival (OS) |
| Results | Posted |
| ClinicalTrials.gov | NCT02563002 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
2. Clinical Question
The primary statistical question was whether treatment assignment to pembrolizumab versus standard of care was associated with a difference in the two registered time-to-event endpoints: progression-free survival and overall survival in participants with MSI-H or dMMR stage IV colorectal carcinoma.
Population
Participants with MSI-H or dMMR stage IV colorectal carcinoma.
Intervention
Pembrolizumab.
Comparator
Standard of care (SOC), with the registry intervention list including mFOLFOX6, FOLFIRI, bevacizumab, and cetuximab.
Primary question
How does pembrolizumab compare with standard of care for PFS and OS under the registered trial analysis?
3. Trial Design
Pembrolizumab
- Pembrolizumab is the biological intervention listed for this arm.
- The ClinicalTrials.gov record identifies this comparison group as Pembrolizumab in the primary analyses.
Standard of Care
- The primary analyses identify this comparison group as Standard of Care (SOC).
- The intervention list includes mFOLFOX6, FOLFIRI, bevacizumab, and cetuximab.
4. Endpoints
The registry lists two primary endpoints, both time-to-event outcomes, with a time frame of up to approximately 59 months. A secondary binary response endpoint is also included in the posted statistical analyses.
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Progression-Free Survival (PFS) Per RECIST1.1 As Assessed by Central Imaging Vendor | Time from randomization to the first documented disease progression (PD) per RECIST 1.1 based on blinded central imaging vendor review or death due to any cause, whichever occurs first. Per RECIST 1.1, PD was defined as ≥20% increase in the sum of diameters of target lesions. In addition to the relative increase of 20%, the sum had to demonstrate an absolute increase of ≥5 mm. The appearance of one or more new lesions was also considered PD. Hazards ratio (HR) and associated 95% confidence intervals (CIs) from a Cox proportional hazard model with Efron's method of tie handling and with a single treatment covariate was presented for the first course study treatment per protocol.. Time frame: up to approximately 59 months. | Time-to-event |
| Overall Survival (OS) | Time from randomization to death due to any cause. Participants without documented death at the time of analysis were censored at the date of last known contact. Time frame: up to approximately 59 months. | Time-to-event |
| Overall Response Rate (ORR) Per RECIST1.1 as Assessed by Central Imaging Vendor | Posted as a secondary endpoint. Time frame: up to approximately 59 months. | Binary |
The PFS definition illustrates why time-to-event analysis is different from a simple comparison of proportions. A participant may remain free of progression at the time of analysis but have incomplete follow-up. That participant contributes information until the relevant censoring point rather than being treated as though an event either definitely did or did not occur after that point.
5. Statistical Analysis Overview
| Endpoint | Population | Comparison | Method | Effect measure |
|---|---|---|---|---|
| PFS | All randomized participants | Pembrolizumab vs SOC | Log-rank test; Cox regression with Efron's method of tie handling and treatment as a covariate | Hazard ratio |
| OS | All randomized participants | Pembrolizumab vs SOC | Log-rank test; Cox regression with Efron's method of tie handling and treatment as a covariate | Hazard ratio |
| ORR | All randomized participants | Pembrolizumab vs SOC | Miettinen & Nurminen method | Risk difference |
The posted primary analyses used all randomized participants. The PFS and OS analyses used a log-rank framework, while the associated hazard-ratio estimates were based on Cox regression with Efron's method for tied event times and treatment as a covariate. The ORR comparison used the Miettinen & Nurminen method.
6. Primary Results: Progression-Free Survival
The registry reports a formal primary analysis of PFS among all randomized participants comparing pembrolizumab with standard of care. The reported effect measure is a hazard ratio.
Hazard ratio for progression or death
95% CI: 0.45–0.79 · One-sided P = 0.0001
Log-rank analysis; Cox regression with Efron's method of tie handling and treatment as a covariate.
A PFS hazard ratio of 0.59 means that, under the reported Cox model, the estimated instantaneous rate of progression or death associated with the pembrolizumab group was approximately 41% lower than the corresponding estimated rate in the SOC group.
That does not mean that 41% of participants avoided progression, that each participant experienced exactly a 41% reduction in risk, or that median PFS was reduced or increased by 41%. A hazard ratio is a relative time-to-event measure, not a percentage of patients benefiting.
The 95% confidence interval of 0.45–0.79 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of individual patient outcomes. The interval is relatively informative about the direction and magnitude of the modeled relative effect, but it still permits a range of plausible hazard ratios.
The one-sided P-value of 0.0001 addresses evidence against the relevant null hypothesis under a one-sided testing framework. It is not an effect-size measure: a very small P-value does not tell us whether an effect is clinically large, and the magnitude of the effect is better conveyed by the hazard ratio and its confidence interval.
The analysis is also subject to the usual interpretation of a Cox hazard ratio. If the proportional-hazards assumption is poor over time, one single HR may not fully describe how treatment differences evolve during follow-up. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic.
7. Primary Results: Overall Survival
The second registered primary endpoint was overall survival. OS was defined from randomization to death from any cause, with participants without documented death censored at their last known contact at the time of analysis.
Hazard ratio for death
95% CI: 0.53–1.03 · One-sided P = 0.0359
Log-rank analysis; Cox regression with Efron's method of tie handling and treatment as a covariate.
An OS hazard ratio of 0.74 corresponds to an estimated 26% lower instantaneous rate of death in the pembrolizumab group relative to SOC under the reported Cox model. This is a model-based relative comparison over the analyzed follow-up, not a statement that 26% of participants survived because of treatment.
The 95% CI of 0.53–1.03 is important because it shows appreciable uncertainty around the point estimate. The interval extends above 1.00, so the data represented by this interval are compatible with a range of relative effects, including values close to no difference under the stated model.
The registry reports a one-sided P-value of 0.0359. Because the P-value is one-sided while the reported confidence interval is two-sided at 95%, the two quantities should not be mechanically compared as though they used identical testing conventions. Interpretation of the P-value depends on the prespecified hypothesis and error-control framework.
As with PFS, the P-value does not measure the size of the treatment effect. The HR communicates relative magnitude, while the confidence interval communicates uncertainty. The OS endpoint also has censoring, and the registry definition makes clear that participants without documented death were censored at their last known contact.
8. Secondary Endpoint: Overall Response Rate
Overall Response Rate (ORR) per RECIST1.1 as assessed by the central imaging vendor was posted as a secondary binary endpoint. The analysis population was all randomized participants.
Risk difference in response rate
95% CI: 1.0–22.6 percentage points · P = 0.0159
Miettinen & Nurminen method.
A reported risk difference of 12.0 percentage points means that the estimated proportion meeting the ORR definition was 12.0 percentage points higher in the pembrolizumab group than in the SOC group in the analyzed randomized population.
This is an absolute difference, unlike the hazard ratios used for PFS and OS. A risk difference is directly expressed in percentage points and therefore answers a different question from a relative time-to-event measure.
The 95% CI of 1.0 to 22.6 percentage points expresses uncertainty around the estimated difference. It does not mean that an individual patient's response probability must lie within that interval, nor does it provide a range for individual treatment benefit.
The reported P-value of 0.0159 addresses the evidence against the relevant null comparison under the specified method. It does not quantify how clinically important a 12.0-percentage-point difference is. Effect size and statistical evidence should therefore be read together rather than treating the P-value as a magnitude measure.
9. Putting the Three Reported Effects Together
The posted analyses use three different ways of expressing treatment effects. PFS and OS are time-to-event endpoints summarized with hazard ratios, whereas ORR is binary and summarized with a risk difference. These measures should not be converted mentally into one another.
| Endpoint | Effect estimate | 95% CI | P-value | What the measure represents |
|---|---|---|---|---|
| PFS | HR 0.59 | 0.45–0.79 | 0.0001, one-sided | Relative time-to-event effect for progression or death |
| OS | HR 0.74 | 0.53–1.03 | 0.0359, one-sided | Relative time-to-event effect for death |
| ORR | Risk difference 12.0 | 1.0–22.6 percentage points | 0.0159 | Absolute difference in the proportion meeting the response definition |
There is a useful statistical distinction here. The PFS HR describes the relative event rate for progression or death over time; the OS HR describes the relative event rate for death; and the ORR risk difference describes an absolute difference in a binary response outcome. A coherent interpretation therefore considers the endpoint definition first and the effect measure second.
10. Statistical Methodology
Log-rank testing
The registry reports the log-rank test for both primary time-to-event endpoints. A log-rank comparison uses the timing of observed events and the numbers at risk over follow-up rather than reducing every participant to a simple yes/no outcome at a fixed date.
The log-rank framework compares the observed and expected pattern of events between treatment groups across follow-up. It is particularly natural for endpoints such as PFS and OS because both event timing and censoring matter.
Cox proportional-hazards regression
The analysis notes state that the hazard ratios were based on Cox regression with Efron's method of tie handling and treatment as a covariate. The registry-reported OS endpoint definition also explicitly describes the Cox proportional hazard model and Efron's method of tie handling.
Conceptually, the hazard ratio compares the modeled instantaneous event rates between groups. An HR below 1 indicates a lower modeled event rate in the treatment group, while an HR above 1 indicates a higher modeled event rate.
Efron's method for tied event times
Clinical trial data can contain multiple participants experiencing an event at the same recorded time. The registry specifies Efron's method for handling tied event times in the Cox regression. This is a technical detail of model estimation rather than a separate clinical endpoint.
Miettinen & Nurminen method
The ORR analysis used the Miettinen & Nurminen method. This is a score-based approach for comparing two binomial proportions and constructing an interval for their difference. In this trial, the resulting effect measure was the difference in percentage, reported as a risk difference.
A positive risk difference means the estimated response proportion is higher in the first group. The reported KEYNOTE-177 estimate is 12.0 percentage points.
One-sided versus two-sided inference
The primary PFS and OS analyses explicitly report one-sided P-values, while their confidence intervals are reported as two-sided 95% CIs. This distinction matters because a P-value and confidence interval encode inferential information under potentially different tail conventions. They should be interpreted according to the prespecified statistical plan rather than treated as interchangeable labels for significance.
11. Statistical Methods Explained
Why was a log-rank test used for PFS and OS?
PFS and OS are time-to-event endpoints. Participants can experience the event at different times, while others may be censored because the event has not been documented by the analysis time. The log-rank framework uses this longitudinal information rather than discarding the timing of events.
What does an HR of 0.59 mean for PFS?
It means that the Cox model estimated the instantaneous rate of progression or death in the pembrolizumab group to be 0.59 times that in the SOC group, under the model. Expressed as a relative reduction, that corresponds to approximately 41% lower estimated hazard. It does not mean that PFS was 41% longer or that 41% of patients benefited.
Why is the OS confidence interval important?
The OS estimate is HR 0.74, but its 95% CI is 0.53–1.03. The point estimate alone would provide an incomplete description because it hides uncertainty. The interval shows that the estimated relative effect is not known with arbitrary precision.
Why does the ORR analysis use a risk difference rather than a hazard ratio?
ORR is a binary endpoint: each participant either meets the response definition or does not. A risk difference compares the resulting proportions directly. A hazard ratio would instead require a time-to-event outcome and is therefore not the natural effect measure for this endpoint.
What does the Miettinen & Nurminen method add?
For two response proportions, the method provides a score-based framework for inference about their difference. It is designed around the binomial comparison itself rather than applying a normal approximation without regard to the underlying two-group proportion problem.
Why should the one-sided P-values not be read as effect sizes?
A P-value measures evidence against a specified null hypothesis under a statistical model and testing convention. It does not tell us whether an observed effect is large or small. In KEYNOTE-177, the effect sizes are conveyed by the HRs for PFS and OS and the 12.0-percentage-point risk difference for ORR; the confidence intervals add information about their precision.
12. Censoring and Time-to-Event Interpretation
The OS definition explicitly states that participants without documented death at the time of analysis were censored at the date of last known contact. This is fundamental to interpreting a survival analysis: censoring means the analysis uses the information available up to the censoring time rather than assuming that the participant subsequently experienced no event.
PFS event
The first documented RECIST 1.1 progression based on blinded central imaging vendor review or death from any cause, whichever occurs first.
OS event
Death due to any cause.
Censoring
For OS, participants without documented death were censored at their last known contact at the analysis.
Analysis horizon
Both registered primary endpoints had a time frame of up to approximately 59 months.
The distinction between event definition and censoring is important. A censored participant is not treated as an event-free participant for all future time. Instead, the participant contributes information through the point at which follow-up is available under the analysis framework.
13. Analysis Population and Randomization
Both primary analyses were conducted in all randomized participants. This is an important feature of the analysis because the treatment comparison is anchored to randomized assignment rather than restricted to participants who completed a particular treatment course.
| Endpoint | Analysis population | Groups compared |
|---|---|---|
| PFS | All randomized participants | Pembrolizumab vs Standard of Care (SOC) |
| OS | All randomized participants | Pembrolizumab vs Standard of Care (SOC) |
| ORR | All randomized participants | Pembrolizumab vs Standard of Care (SOC) |
Randomization is valuable statistically because it establishes treatment assignment before outcome information accumulates. In a properly conducted randomized comparison, baseline differences that occur by chance do not systematically arise from treatment choice. The ClinicalTrials.gov record confirms randomized allocation but does not provide the randomization ratio or stratification factors, so neither is inferred here.
14. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk. These safety figures are reported separately from the efficacy analysis populations and are presented exactly as reported in the registry.
| Safety group | Serious adverse events |
|---|---|
| Pembrolizumab First Course | 62/153 |
| Standard of Care (SOC) First Course | 76/143 |
| SOC Switched Over to Pembrolizumab | 27/57 |
| Pembrolizumab Second Course | 3/12 |
| SOC Switched Over to Pembrolizumab-Secon | 1/5 |
Safety also answers a different question from PFS, OS, and ORR. Efficacy endpoints describe outcomes associated with the randomized treatment comparison, while adverse-event summaries describe observed safety experience within defined treatment or exposure groups. A statistically favorable efficacy estimate does not itself quantify safety, and a safety count does not establish the magnitude of an efficacy effect.
15. What the Hazard Ratios Do — and Do Not — Mean
The estimated instantaneous rate of progression or death was approximately 41% lower in the pembrolizumab group than in the SOC group under the reported Cox model.
This does not mean a 41% absolute reduction in the proportion of participants progressing, a 41% increase in survival time, or a uniform 41% treatment effect for every participant.
The estimated instantaneous rate of death was approximately 26% lower in the pembrolizumab group than in the SOC group under the reported Cox model.
This does not mean that 26% of participants avoided death, that 26% of participants were cured, or that each participant experienced a 26% reduction in their individual probability of death.
The PFS 95% CI is 0.45–0.79 and the OS 95% CI is 0.53–1.03. These intervals describe uncertainty around the respective model-based estimates. They are not ranges of outcomes for individual participants.
16. Why the Confidence Intervals Matter
The point estimate is only one part of a statistical result. A hazard ratio of 0.59 is more informative when accompanied by its 95% CI of 0.45–0.79 because the interval shows the uncertainty surrounding the estimate.
| Measure | Point estimate | 95% CI | Interpretive focus |
|---|---|---|---|
| PFS hazard ratio | 0.59 | 0.45–0.79 | Magnitude and precision of the modeled relative time-to-event effect |
| OS hazard ratio | 0.74 | 0.53–1.03 | Magnitude and precision of the modeled relative effect on death |
| ORR risk difference | 12.0 percentage points | 1.0–22.6 percentage points | Magnitude and precision of the absolute difference in response proportion |
The OS interval deserves particular attention because it extends from below 1.00 to slightly above 1.00. The appropriate interpretation is not that the point estimate disappears, but that the uncertainty around that estimate is materially broader than the point estimate alone suggests.
17. Primary Endpoint Comparison
KEYNOTE-177 has two primary endpoints, both of which are time-to-event outcomes. That creates a useful teaching distinction: the same broad statistical framework can be used for both endpoints while the event definition remains different.
| Feature | PFS | OS |
|---|---|---|
| Time origin | Randomization | Randomization |
| Event | RECIST 1.1 progression or death, whichever occurs first | Death from any cause |
| Analysis population | All randomized participants | All randomized participants |
| Comparison | Pembrolizumab vs SOC | Pembrolizumab vs SOC |
| Primary method | Log-rank test | Log-rank test |
| Model-based effect | HR 0.59 | HR 0.74 |
| 95% CI | 0.45–0.79 | 0.53–1.03 |
| P-value | 0.0001, one-sided | 0.0359, one-sided |
PFS can occur because of radiographic progression or death, whereas OS requires death as the event. Consequently, the two endpoints can provide different statistical information even when they are evaluated in the same randomized population.
18. Limitations and Interpretation Issues
- Limited registry detail: the ClinicalTrials.gov record does not provide baseline characteristics, subgroup estimates, median PFS, median OS, time-specific survival probabilities, or underlying event/censoring records.
- Confidence-interval interpretation: the OS 95% CI extends to 1.03, illustrating substantial uncertainty around the reported point estimate.
- One-sided testing: the primary PFS and OS P-values are explicitly one-sided, while the reported confidence intervals are two-sided 95% intervals. These should be interpreted under their respective inferential conventions.
- Proportional-hazards assumption: Cox hazard ratios are model-based. If hazards are not reasonably proportional, a single HR may not capture changes in the treatment effect over time. The ClinicalTrials.gov record does not report a formal diagnostic.
- No reconstructed medians: median PFS and OS cannot be inferred from the HRs and confidence intervals alone.
- No subgroup extrapolation: the ClinicalTrials.gov record does not provide subgroup results, so treatment-effect consistency across patient characteristics cannot be assessed from this record.
- Safety populations differ: the serious-adverse-event data includes first-course, switched, and second-course groups. These should not be collapsed into a simple randomized comparison without the underlying safety definitions and denominators.
- Missing-data and imputation details: the ClinicalTrials.gov record does not describe a separate missing-data or imputation strategy for the posted analyses.
- Multiplicity: the ClinicalTrials.gov record identifies two primary endpoints but does not specify an alpha-allocation or multiplicity procedure. No such procedure is invented here.
- Interim analyses: the ClinicalTrials.gov record does not provide an interim-analysis schedule or alpha-spending method. None is assumed.
19. Why This Trial Matters Statistically
KEYNOTE-177 is a useful statistical teaching case because it combines randomized treatment comparison, two primary time-to-event endpoints, a binary response endpoint, Cox regression, log-rank testing, one-sided P-values, two-sided confidence intervals, and a score-based approach for comparing proportions.
| Concept | How it appears in KEYNOTE-177 |
|---|---|
| Randomization | The trial uses randomized allocation in a phase 3 parallel design. |
| Time-to-event endpoints | PFS and OS are both registered primary endpoints. |
| RECIST-based progression | PFS uses progression according to the registered RECIST 1.1 definition and central imaging vendor assessment. |
| Log-rank test | Reported for both PFS and OS. |
| Hazard ratio | Used to quantify the reported relative treatment effects for PFS and OS. |
| Cox regression | Used for HR estimation with treatment as a covariate. |
| Efron's tie handling | Specified for the Cox regression. |
| One-sided testing | Explicitly reported for the PFS and OS log-rank P-values. |
| Confidence intervals | Two-sided 95% intervals are reported for the primary HR estimates and ORR difference. |
| Risk difference | The ORR effect is reported as a difference in percentage. |
| Miettinen & Nurminen method | Used for the binary ORR comparison. |
| Analysis population | Primary and secondary posted analyses use all randomized participants. |
The central statistical lesson is that the effect measure must match the endpoint. PFS and OS require attention to event timing and censoring, making survival-analysis methods appropriate. ORR is a binary outcome, making a two-proportion method appropriate. Presenting all three results side by side is therefore more informative than trying to reduce the trial to a single number.
20. A Practical Reading Framework for KEYNOTE-177
Step 1 · Identify the endpoint
Determine whether the result concerns progression or death, death alone, or binary tumor response. The endpoint determines what the effect estimate means.
Step 2 · Read the effect measure
HRs are relative time-to-event measures. The ORR risk difference is an absolute percentage-point difference.
Step 3 · Read the CI
Use the confidence interval to understand the uncertainty surrounding the point estimate rather than relying on the point estimate alone.
Step 4 · Read the P-value
Interpret the P-value as evidence under the stated hypothesis-testing framework, not as a measure of treatment magnitude.
This framework also prevents a common error in clinical-trial interpretation: assuming that a small P-value automatically means a large treatment effect. In KEYNOTE-177, the numerical magnitude is carried by the HRs and risk difference, while the P-values address evidence against the relevant null hypotheses.
21. Registry Results Summary
| Endpoint | Role | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Progression-Free Survival | Primary | HR 0.59 | 0.45–0.79 | 0.0001, one-sided |
| Overall Survival | Primary | HR 0.74 | 0.53–1.03 | 0.0359, one-sided |
| Overall Response Rate | Secondary | Risk difference 12.0 percentage points | 1.0–22.6 percentage points | 0.0159 |
The registry therefore provides formal numerical comparisons for both primary endpoints and one secondary endpoint. The primary PFS estimate is 0.59, the primary OS estimate is 0.74, and the secondary ORR comparison reports a 12.0-percentage-point difference. Each should be interpreted using its own endpoint definition and statistical scale.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: NCT02563002 — KEYNOTE-177.
- PubMed: PMID 36369901.
- PubMed: PMID 35427471.
- PubMed: PMID 33812497.
- PubMed: PMID 33264544.
Continue through Clinical Biostats
Explore the statistical methods behind randomized clinical trials, then apply the concepts with focused tutorials and calculators.
25. Record Summary
KEYNOTE-177 provides a compact example of several important principles in clinical-trial statistics. The trial is randomized, parallel, open label, and phase 3, with 307 participants enrolled. Its two registered primary endpoints are time-to-event outcomes: PFS and OS. Both were analyzed using log-rank testing, with hazard ratios based on Cox regression using Efron's method for tied event times and treatment as a covariate. The posted PFS analysis reports an HR of 0.59 with a 95% CI of 0.45–0.79 and a one-sided P-value of 0.0001. The OS analysis reports an HR of 0.74 with a 95% CI of 0.53–1.03 and a one-sided P-value of 0.0359.
The secondary ORR analysis demonstrates a different statistical structure. Because ORR is binary, the treatment effect is reported as a risk difference of 12.0 percentage points, with a 95% CI of 1.0–22.6 percentage points and a P-value of 0.0159 using the Miettinen & Nurminen method. Reading these results correctly requires keeping the scales distinct: hazard ratios describe relative time-to-event effects, whereas risk differences describe absolute differences in proportions.