← Clinical Trials
Colorectal Carcinoma Phase 3 Completed NCT02563002

KEYNOTE-177: Complete Statistical Analysis of Pembrolizumab in Stage IV MSI-H or dMMR Colorectal Carcinoma

An independent statistical review of the randomized phase 3 KEYNOTE-177 trial comparing pembrolizumab with standard of care in participants with microsatellite instability-high (MSI-H) or mismatch repair deficient (dMMR) stage IV colorectal carcinoma.

Randomized  ·  Parallel design  ·  Enrollment 307  ·  Results posted
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics on this page are taken from the ClinicalTrials.gov record. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-177 was a randomized, open-label, parallel phase 3 trial evaluating pembrolizumab versus standard of care in participants with MSI-H or dMMR stage IV colorectal carcinoma. The registry reports 307 enrolled participants, two arms, two registered primary endpoints, and posted formal analyses for both primary endpoints.

307
Enrollment
Registry enrollment
2
Arms
Randomized parallel design
0.59
PFS HR
95% CI 0.45–0.79
0.74
OS HR
95% CI 0.53–1.03
FeatureKEYNOTE-177
Trial nameKEYNOTE-177
PhasePhase 3
ConditionColorectal carcinoma
PopulationParticipants with MSI-H or dMMR stage IV colorectal carcinoma
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
Enrollment307
Primary endpointsProgression-Free Survival (PFS) and Overall Survival (OS)
ResultsPosted
ClinicalTrials.govNCT02563002
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry

2. Clinical Question

The primary statistical question was whether treatment assignment to pembrolizumab versus standard of care was associated with a difference in the two registered time-to-event endpoints: progression-free survival and overall survival in participants with MSI-H or dMMR stage IV colorectal carcinoma.

Population

Participants with MSI-H or dMMR stage IV colorectal carcinoma.

Intervention

Pembrolizumab.

Comparator

Standard of care (SOC), with the registry intervention list including mFOLFOX6, FOLFIRI, bevacizumab, and cetuximab.

Primary question

How does pembrolizumab compare with standard of care for PFS and OS under the registered trial analysis?

3. Trial Design

01
Randomize307 enrolled
02
Parallel arms2 treatment groups
03
Open labelNo masking
04
AssessPFS / OS / ORR
05
AnalyzeSurvival and response methods
Allocation
Randomized.
Model
Parallel.
Masking
None.
Primary purpose
Treatment.
Start
2015-11-30.
Primary completion
2021-02-19.
ARM 1

Pembrolizumab

  • Pembrolizumab is the biological intervention listed for this arm.
  • The ClinicalTrials.gov record identifies this comparison group as Pembrolizumab in the primary analyses.
ARM 2

Standard of Care

  • The primary analyses identify this comparison group as Standard of Care (SOC).
  • The intervention list includes mFOLFOX6, FOLFIRI, bevacizumab, and cetuximab.
What the registry data does not establish: the ClinicalTrials.gov record does not provide a detailed dosing schedule, randomization ratio, stratification factors, crossover rules, interim-analysis schedule, multiplicity procedure, missing-data strategy, or subgroup results. Those features are therefore not reconstructed here.

4. Endpoints

The registry lists two primary endpoints, both time-to-event outcomes, with a time frame of up to approximately 59 months. A secondary binary response endpoint is also included in the posted statistical analyses.

EndpointRegistry definition / time frameEndpoint type
Progression-Free Survival (PFS) Per RECIST1.1 As Assessed by Central Imaging Vendor Time from randomization to the first documented disease progression (PD) per RECIST 1.1 based on blinded central imaging vendor review or death due to any cause, whichever occurs first. Per RECIST 1.1, PD was defined as ≥20% increase in the sum of diameters of target lesions. In addition to the relative increase of 20%, the sum had to demonstrate an absolute increase of ≥5 mm. The appearance of one or more new lesions was also considered PD. Hazards ratio (HR) and associated 95% confidence intervals (CIs) from a Cox proportional hazard model with Efron's method of tie handling and with a single treatment covariate was presented for the first course study treatment per protocol.. Time frame: up to approximately 59 months. Time-to-event
Overall Survival (OS) Time from randomization to death due to any cause. Participants without documented death at the time of analysis were censored at the date of last known contact. Time frame: up to approximately 59 months. Time-to-event
Overall Response Rate (ORR) Per RECIST1.1 as Assessed by Central Imaging Vendor Posted as a secondary endpoint. Time frame: up to approximately 59 months. Binary

The PFS definition illustrates why time-to-event analysis is different from a simple comparison of proportions. A participant may remain free of progression at the time of analysis but have incomplete follow-up. That participant contributes information until the relevant censoring point rather than being treated as though an event either definitely did or did not occur after that point.

5. Statistical Analysis Overview

EndpointPopulationComparisonMethodEffect measure
PFS All randomized participants Pembrolizumab vs SOC Log-rank test; Cox regression with Efron's method of tie handling and treatment as a covariate Hazard ratio
OS All randomized participants Pembrolizumab vs SOC Log-rank test; Cox regression with Efron's method of tie handling and treatment as a covariate Hazard ratio
ORR All randomized participants Pembrolizumab vs SOC Miettinen & Nurminen method Risk difference
Analysis conventions reported in the registry
PFS HR = 0.59   |   OS HR = 0.74   |   ORR difference = 12.0 percentage points

The posted primary analyses used all randomized participants. The PFS and OS analyses used a log-rank framework, while the associated hazard-ratio estimates were based on Cox regression with Efron's method for tied event times and treatment as a covariate. The ORR comparison used the Miettinen & Nurminen method.

6. Primary Results: Progression-Free Survival

The registry reports a formal primary analysis of PFS among all randomized participants comparing pembrolizumab with standard of care. The reported effect measure is a hazard ratio.

Hazard ratio for progression or death

0.59

95% CI: 0.45–0.79   ·   One-sided P = 0.0001

Log-rank analysis; Cox regression with Efron's method of tie handling and treatment as a covariate.

Clinical Biostats interpretation

A PFS hazard ratio of 0.59 means that, under the reported Cox model, the estimated instantaneous rate of progression or death associated with the pembrolizumab group was approximately 41% lower than the corresponding estimated rate in the SOC group.

That does not mean that 41% of participants avoided progression, that each participant experienced exactly a 41% reduction in risk, or that median PFS was reduced or increased by 41%. A hazard ratio is a relative time-to-event measure, not a percentage of patients benefiting.

The 95% confidence interval of 0.45–0.79 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of individual patient outcomes. The interval is relatively informative about the direction and magnitude of the modeled relative effect, but it still permits a range of plausible hazard ratios.

The one-sided P-value of 0.0001 addresses evidence against the relevant null hypothesis under a one-sided testing framework. It is not an effect-size measure: a very small P-value does not tell us whether an effect is clinically large, and the magnitude of the effect is better conveyed by the hazard ratio and its confidence interval.

The analysis is also subject to the usual interpretation of a Cox hazard ratio. If the proportional-hazards assumption is poor over time, one single HR may not fully describe how treatment differences evolve during follow-up. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic.

7. Primary Results: Overall Survival

The second registered primary endpoint was overall survival. OS was defined from randomization to death from any cause, with participants without documented death censored at their last known contact at the time of analysis.

Hazard ratio for death

0.74

95% CI: 0.53–1.03   ·   One-sided P = 0.0359

Log-rank analysis; Cox regression with Efron's method of tie handling and treatment as a covariate.

Clinical Biostats interpretation

An OS hazard ratio of 0.74 corresponds to an estimated 26% lower instantaneous rate of death in the pembrolizumab group relative to SOC under the reported Cox model. This is a model-based relative comparison over the analyzed follow-up, not a statement that 26% of participants survived because of treatment.

The 95% CI of 0.53–1.03 is important because it shows appreciable uncertainty around the point estimate. The interval extends above 1.00, so the data represented by this interval are compatible with a range of relative effects, including values close to no difference under the stated model.

The registry reports a one-sided P-value of 0.0359. Because the P-value is one-sided while the reported confidence interval is two-sided at 95%, the two quantities should not be mechanically compared as though they used identical testing conventions. Interpretation of the P-value depends on the prespecified hypothesis and error-control framework.

As with PFS, the P-value does not measure the size of the treatment effect. The HR communicates relative magnitude, while the confidence interval communicates uncertainty. The OS endpoint also has censoring, and the registry definition makes clear that participants without documented death were censored at their last known contact.

The ClinicalTrials.gov record does not provide median OS, median PFS, Kaplan-Meier event counts, survival probabilities at specific time points, or underlying event/censoring data. Those quantities are therefore not added or reconstructed here.

8. Secondary Endpoint: Overall Response Rate

Overall Response Rate (ORR) per RECIST1.1 as assessed by the central imaging vendor was posted as a secondary binary endpoint. The analysis population was all randomized participants.

Risk difference in response rate

12.0

95% CI: 1.0–22.6 percentage points   ·   P = 0.0159

Miettinen & Nurminen method.

Clinical Biostats interpretation

A reported risk difference of 12.0 percentage points means that the estimated proportion meeting the ORR definition was 12.0 percentage points higher in the pembrolizumab group than in the SOC group in the analyzed randomized population.

This is an absolute difference, unlike the hazard ratios used for PFS and OS. A risk difference is directly expressed in percentage points and therefore answers a different question from a relative time-to-event measure.

The 95% CI of 1.0 to 22.6 percentage points expresses uncertainty around the estimated difference. It does not mean that an individual patient's response probability must lie within that interval, nor does it provide a range for individual treatment benefit.

The reported P-value of 0.0159 addresses the evidence against the relevant null comparison under the specified method. It does not quantify how clinically important a 12.0-percentage-point difference is. Effect size and statistical evidence should therefore be read together rather than treating the P-value as a magnitude measure.

9. Putting the Three Reported Effects Together

The posted analyses use three different ways of expressing treatment effects. PFS and OS are time-to-event endpoints summarized with hazard ratios, whereas ORR is binary and summarized with a risk difference. These measures should not be converted mentally into one another.

EndpointEffect estimate95% CIP-valueWhat the measure represents
PFS HR 0.59 0.45–0.79 0.0001, one-sided Relative time-to-event effect for progression or death
OS HR 0.74 0.53–1.03 0.0359, one-sided Relative time-to-event effect for death
ORR Risk difference 12.0 1.0–22.6 percentage points 0.0159 Absolute difference in the proportion meeting the response definition

There is a useful statistical distinction here. The PFS HR describes the relative event rate for progression or death over time; the OS HR describes the relative event rate for death; and the ORR risk difference describes an absolute difference in a binary response outcome. A coherent interpretation therefore considers the endpoint definition first and the effect measure second.

10. Statistical Methodology

Log-rank testing

The registry reports the log-rank test for both primary time-to-event endpoints. A log-rank comparison uses the timing of observed events and the numbers at risk over follow-up rather than reducing every participant to a simple yes/no outcome at a fixed date.

Conceptual survival comparison
H0: survival experience is the same between randomized groups

The log-rank framework compares the observed and expected pattern of events between treatment groups across follow-up. It is particularly natural for endpoints such as PFS and OS because both event timing and censoring matter.

Cox proportional-hazards regression

The analysis notes state that the hazard ratios were based on Cox regression with Efron's method of tie handling and treatment as a covariate. The registry-reported OS endpoint definition also explicitly describes the Cox proportional hazard model and Efron's method of tie handling.

Hazard ratio
HR = htreatment(t) / hcontrol(t)

Conceptually, the hazard ratio compares the modeled instantaneous event rates between groups. An HR below 1 indicates a lower modeled event rate in the treatment group, while an HR above 1 indicates a higher modeled event rate.

Efron's method for tied event times

Clinical trial data can contain multiple participants experiencing an event at the same recorded time. The registry specifies Efron's method for handling tied event times in the Cox regression. This is a technical detail of model estimation rather than a separate clinical endpoint.

Miettinen & Nurminen method

The ORR analysis used the Miettinen & Nurminen method. This is a score-based approach for comparing two binomial proportions and constructing an interval for their difference. In this trial, the resulting effect measure was the difference in percentage, reported as a risk difference.

Risk difference
RD = ppembrolizumab − pSOC

A positive risk difference means the estimated response proportion is higher in the first group. The reported KEYNOTE-177 estimate is 12.0 percentage points.

One-sided versus two-sided inference

The primary PFS and OS analyses explicitly report one-sided P-values, while their confidence intervals are reported as two-sided 95% CIs. This distinction matters because a P-value and confidence interval encode inferential information under potentially different tail conventions. They should be interpreted according to the prespecified statistical plan rather than treated as interchangeable labels for significance.

11. Statistical Methods Explained

Why was a log-rank test used for PFS and OS?

PFS and OS are time-to-event endpoints. Participants can experience the event at different times, while others may be censored because the event has not been documented by the analysis time. The log-rank framework uses this longitudinal information rather than discarding the timing of events.

What does an HR of 0.59 mean for PFS?

It means that the Cox model estimated the instantaneous rate of progression or death in the pembrolizumab group to be 0.59 times that in the SOC group, under the model. Expressed as a relative reduction, that corresponds to approximately 41% lower estimated hazard. It does not mean that PFS was 41% longer or that 41% of patients benefited.

Why is the OS confidence interval important?

The OS estimate is HR 0.74, but its 95% CI is 0.53–1.03. The point estimate alone would provide an incomplete description because it hides uncertainty. The interval shows that the estimated relative effect is not known with arbitrary precision.

Why does the ORR analysis use a risk difference rather than a hazard ratio?

ORR is a binary endpoint: each participant either meets the response definition or does not. A risk difference compares the resulting proportions directly. A hazard ratio would instead require a time-to-event outcome and is therefore not the natural effect measure for this endpoint.

What does the Miettinen & Nurminen method add?

For two response proportions, the method provides a score-based framework for inference about their difference. It is designed around the binomial comparison itself rather than applying a normal approximation without regard to the underlying two-group proportion problem.

Why should the one-sided P-values not be read as effect sizes?

A P-value measures evidence against a specified null hypothesis under a statistical model and testing convention. It does not tell us whether an observed effect is large or small. In KEYNOTE-177, the effect sizes are conveyed by the HRs for PFS and OS and the 12.0-percentage-point risk difference for ORR; the confidence intervals add information about their precision.

12. Censoring and Time-to-Event Interpretation

The OS definition explicitly states that participants without documented death at the time of analysis were censored at the date of last known contact. This is fundamental to interpreting a survival analysis: censoring means the analysis uses the information available up to the censoring time rather than assuming that the participant subsequently experienced no event.

PFS event

The first documented RECIST 1.1 progression based on blinded central imaging vendor review or death from any cause, whichever occurs first.

OS event

Death due to any cause.

Censoring

For OS, participants without documented death were censored at their last known contact at the analysis.

Analysis horizon

Both registered primary endpoints had a time frame of up to approximately 59 months.

The distinction between event definition and censoring is important. A censored participant is not treated as an event-free participant for all future time. Instead, the participant contributes information through the point at which follow-up is available under the analysis framework.

13. Analysis Population and Randomization

Both primary analyses were conducted in all randomized participants. This is an important feature of the analysis because the treatment comparison is anchored to randomized assignment rather than restricted to participants who completed a particular treatment course.

EndpointAnalysis populationGroups compared
PFSAll randomized participantsPembrolizumab vs Standard of Care (SOC)
OSAll randomized participantsPembrolizumab vs Standard of Care (SOC)
ORRAll randomized participantsPembrolizumab vs Standard of Care (SOC)

Randomization is valuable statistically because it establishes treatment assignment before outcome information accumulates. In a properly conducted randomized comparison, baseline differences that occur by chance do not systematically arise from treatment choice. The ClinicalTrials.gov record confirms randomized allocation but does not provide the randomization ratio or stratification factors, so neither is inferred here.

14. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk. These safety figures are reported separately from the efficacy analysis populations and are presented exactly as reported in the registry.

Safety groupSerious adverse events
Pembrolizumab First Course62/153
Standard of Care (SOC) First Course76/143
SOC Switched Over to Pembrolizumab27/57
Pembrolizumab Second Course3/12
SOC Switched Over to Pembrolizumab-Secon1/5
Safety interpretation: these are reported affected/at-risk counts for the listed safety groups. They should not be combined into a single randomized-arm percentage because the ClinicalTrials.gov record distinguishes first-course, switched, and second-course groups. In particular, the switched and second-course groups do not represent the same population as the two primary randomized comparison groups.

Safety also answers a different question from PFS, OS, and ORR. Efficacy endpoints describe outcomes associated with the randomized treatment comparison, while adverse-event summaries describe observed safety experience within defined treatment or exposure groups. A statistically favorable efficacy estimate does not itself quantify safety, and a safety count does not establish the magnitude of an efficacy effect.

15. What the Hazard Ratios Do — and Do Not — Mean

PFS HR 0.59

The estimated instantaneous rate of progression or death was approximately 41% lower in the pembrolizumab group than in the SOC group under the reported Cox model.

This does not mean a 41% absolute reduction in the proportion of participants progressing, a 41% increase in survival time, or a uniform 41% treatment effect for every participant.

OS HR 0.74

The estimated instantaneous rate of death was approximately 26% lower in the pembrolizumab group than in the SOC group under the reported Cox model.

This does not mean that 26% of participants avoided death, that 26% of participants were cured, or that each participant experienced a 26% reduction in their individual probability of death.

Confidence intervals

The PFS 95% CI is 0.45–0.79 and the OS 95% CI is 0.53–1.03. These intervals describe uncertainty around the respective model-based estimates. They are not ranges of outcomes for individual participants.

16. Why the Confidence Intervals Matter

The point estimate is only one part of a statistical result. A hazard ratio of 0.59 is more informative when accompanied by its 95% CI of 0.45–0.79 because the interval shows the uncertainty surrounding the estimate.

MeasurePoint estimate95% CIInterpretive focus
PFS hazard ratio0.590.45–0.79Magnitude and precision of the modeled relative time-to-event effect
OS hazard ratio0.740.53–1.03Magnitude and precision of the modeled relative effect on death
ORR risk difference12.0 percentage points1.0–22.6 percentage pointsMagnitude and precision of the absolute difference in response proportion

The OS interval deserves particular attention because it extends from below 1.00 to slightly above 1.00. The appropriate interpretation is not that the point estimate disappears, but that the uncertainty around that estimate is materially broader than the point estimate alone suggests.

17. Primary Endpoint Comparison

KEYNOTE-177 has two primary endpoints, both of which are time-to-event outcomes. That creates a useful teaching distinction: the same broad statistical framework can be used for both endpoints while the event definition remains different.

FeaturePFSOS
Time originRandomizationRandomization
EventRECIST 1.1 progression or death, whichever occurs firstDeath from any cause
Analysis populationAll randomized participantsAll randomized participants
ComparisonPembrolizumab vs SOCPembrolizumab vs SOC
Primary methodLog-rank testLog-rank test
Model-based effectHR 0.59HR 0.74
95% CI0.45–0.790.53–1.03
P-value0.0001, one-sided0.0359, one-sided

PFS can occur because of radiographic progression or death, whereas OS requires death as the event. Consequently, the two endpoints can provide different statistical information even when they are evaluated in the same randomized population.

18. Limitations and Interpretation Issues

19. Why This Trial Matters Statistically

KEYNOTE-177 is a useful statistical teaching case because it combines randomized treatment comparison, two primary time-to-event endpoints, a binary response endpoint, Cox regression, log-rank testing, one-sided P-values, two-sided confidence intervals, and a score-based approach for comparing proportions.

ConceptHow it appears in KEYNOTE-177
RandomizationThe trial uses randomized allocation in a phase 3 parallel design.
Time-to-event endpointsPFS and OS are both registered primary endpoints.
RECIST-based progressionPFS uses progression according to the registered RECIST 1.1 definition and central imaging vendor assessment.
Log-rank testReported for both PFS and OS.
Hazard ratioUsed to quantify the reported relative treatment effects for PFS and OS.
Cox regressionUsed for HR estimation with treatment as a covariate.
Efron's tie handlingSpecified for the Cox regression.
One-sided testingExplicitly reported for the PFS and OS log-rank P-values.
Confidence intervalsTwo-sided 95% intervals are reported for the primary HR estimates and ORR difference.
Risk differenceThe ORR effect is reported as a difference in percentage.
Miettinen & Nurminen methodUsed for the binary ORR comparison.
Analysis populationPrimary and secondary posted analyses use all randomized participants.

The central statistical lesson is that the effect measure must match the endpoint. PFS and OS require attention to event timing and censoring, making survival-analysis methods appropriate. ORR is a binary outcome, making a two-proportion method appropriate. Presenting all three results side by side is therefore more informative than trying to reduce the trial to a single number.

20. A Practical Reading Framework for KEYNOTE-177

Step 1 · Identify the endpoint

Determine whether the result concerns progression or death, death alone, or binary tumor response. The endpoint determines what the effect estimate means.

Step 2 · Read the effect measure

HRs are relative time-to-event measures. The ORR risk difference is an absolute percentage-point difference.

Step 3 · Read the CI

Use the confidence interval to understand the uncertainty surrounding the point estimate rather than relying on the point estimate alone.

Step 4 · Read the P-value

Interpret the P-value as evidence under the stated hypothesis-testing framework, not as a measure of treatment magnitude.

This framework also prevents a common error in clinical-trial interpretation: assuming that a small P-value automatically means a large treatment effect. In KEYNOTE-177, the numerical magnitude is carried by the HRs and risk difference, while the P-values address evidence against the relevant null hypotheses.

21. Registry Results Summary

EndpointRoleEstimate95% CIP-value
Progression-Free Survival Primary HR 0.59 0.45–0.79 0.0001, one-sided
Overall Survival Primary HR 0.74 0.53–1.03 0.0359, one-sided
Overall Response Rate Secondary Risk difference 12.0 percentage points 1.0–22.6 percentage points 0.0159

The registry therefore provides formal numerical comparisons for both primary endpoints and one secondary endpoint. The primary PFS estimate is 0.59, the primary OS estimate is 0.74, and the secondary ORR comparison reports a 12.0-percentage-point difference. Each should be interpreted using its own endpoint definition and statistical scale.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue through Clinical Biostats

Explore the statistical methods behind randomized clinical trials, then apply the concepts with focused tutorials and calculators.

25. Record Summary

KEYNOTE-177 provides a compact example of several important principles in clinical-trial statistics. The trial is randomized, parallel, open label, and phase 3, with 307 participants enrolled. Its two registered primary endpoints are time-to-event outcomes: PFS and OS. Both were analyzed using log-rank testing, with hazard ratios based on Cox regression using Efron's method for tied event times and treatment as a covariate. The posted PFS analysis reports an HR of 0.59 with a 95% CI of 0.45–0.79 and a one-sided P-value of 0.0001. The OS analysis reports an HR of 0.74 with a 95% CI of 0.53–1.03 and a one-sided P-value of 0.0359.

The secondary ORR analysis demonstrates a different statistical structure. Because ORR is binary, the treatment effect is reported as a risk difference of 12.0 percentage points, with a 95% CI of 1.0–22.6 percentage points and a P-value of 0.0159 using the Miettinen & Nurminen method. Reading these results correctly requires keeping the scales distinct: hazard ratios describe relative time-to-event effects, whereas risk differences describe absolute differences in proportions.

Clinical Biostats methodology: A trial-results page should distinguish what the registry actually reports from what statistical theory allows us to explain. For KEYNOTE-177, that means preserving the reported estimates, confidence intervals, P-values, endpoint definitions, analysis populations, and methods while avoiding unsupported reconstruction of medians, subgroup effects, baseline characteristics, crossover, interim analyses, multiplicity procedures, or missing-data strategies.