← Clinical Trials
Renal Cell Carcinoma Phase 3 Adjuvant Treatment NCT03142334

KEYNOTE-564: Complete Statistical Analysis of Pembrolizumab in Renal Cell Carcinoma

An independent statistical analysis of the randomized phase 3 KEYNOTE-564 trial evaluating pembrolizumab monotherapy versus placebo in the adjuvant treatment of renal cell carcinoma post nephrectomy, with a focus on investigator-assessed disease-free survival.

Trial status: Completed  ·  Enrollment: 994  ·  Primary completion: 14 Dec 2020
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results are restricted to the information provided in the ClinicalTrials.gov record. Where the registry does not report a result, this page does not infer one from external publications.

Registry context: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-564 was a randomized, parallel, quadruple-masked phase 3 study evaluating pembrolizumab monotherapy versus placebo in the adjuvant treatment of renal cell carcinoma post nephrectomy.

994
Enrolled
Randomized trial
2
Arms
Pembrolizumab vs placebo
0.68
DFS HR
95% CI 0.53–0.87
0.0010
P-value
One-sided log-rank test
FeatureKEYNOTE-564
PhasePhase 3
ConditionRenal Cell Carcinoma
Brief study descriptionSafety and efficacy study of pembrolizumab as monotherapy in the adjuvant treatment of renal cell carcinoma post nephrectomy
DesignRandomized, parallel
MaskingQuadruple
AllocationRandomized
Enrollment994
Arms2
Primary endpointDisease-free Survival (DFS) as Assessed by the Investigator
Primary endpoint typeTime-to-event
ClinicalTrials.govNCT03142334
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry

2. Clinical Question

The central statistical question is whether randomized assignment to pembrolizumab rather than placebo is associated with a different time to the first event comprising local recurrence, distant kidney cancer metastasis(es), or death from any cause in the adjuvant renal cell carcinoma setting after nephrectomy.

Population

Participants enrolled in the phase 3 study of adjuvant treatment of renal cell carcinoma post nephrectomy.

Intervention

Pembrolizumab, identified in the registry as a biological intervention.

Comparator

Placebo, identified in the registry as a drug intervention.

Primary question

Does randomized assignment to pembrolizumab versus placebo change disease-free survival as assessed by the investigator?

3. Trial Design

01
Enroll994 participants
02
Randomize2 parallel arms
03
MaskQuadruple-masked design
04
AssessInvestigator-assessed DFS
05
AnalyzeLog-rank and Cox methods
ARM 1

Pembrolizumab

  • Pembrolizumab monotherapy
  • Adjuvant treatment following nephrectomy
  • Biological intervention
ARM 2

Placebo

  • Placebo
  • Comparator arm
  • Drug intervention

The design has several features that matter directly to statistical interpretation. Randomization provides the basis for comparing treatment assignments, while quadruple masking is intended to reduce the potential influence of treatment knowledge on trial conduct and outcome assessment. The parallel structure means that the two randomized groups are followed as distinct treatment assignments rather than as sequential treatment phases.

Trial timeline

09 Jun 2017

Study start

The registry lists 9 June 2017 as the study start date.

14 Dec 2020

Primary completion

The registry lists 14 December 2020 as the primary completion date and uses this date as the database cutoff for the reported primary endpoint result.

4. Endpoints

The registry identifies one primary endpoint, and it is a time-to-event endpoint. The endpoint is assessed by the investigator rather than being described in the ClinicalTrials.gov record as a patient-reported or laboratory endpoint.

EndpointRegistry definitionTime frame
Disease-free Survival (DFS) as Assessed by the Investigator DFS, as assessed by the investigator, is defined as the time from randomization to the first documented local recurrence, distant kidney cancer metastasis(es), or death due to any cause, whichever occurs first. Up to approximately 42 months (database cutoff date 14 Dec 2020)

Why DFS is a time-to-event endpoint

DFS does not simply classify participants as having or not having an event. It records when the first qualifying event occurs. A participant who has not experienced the event by the end of available follow-up contributes information up to that point and may therefore be censored in the survival analysis.

The definition also uses a composite event structure: local recurrence, distant kidney cancer metastasis(es), and death from any cause are all treated as events, with the first one determining the DFS event time. This is important because the resulting hazard ratio describes the randomized-group difference in the composite time-to-first-event endpoint, not a treatment effect on each component separately.

5. Statistical Methodology

Primary analysis: stratified log-rank test

The registry reports the log-rank test as the method for the primary endpoint comparison. The analysis notes specify that the reported one-sided p-value was based on a log-rank test stratified by metastasis status, ECOG performance status, and US participant within the M0 group by investigator.

Conceptual purpose
H0: survival experience is equivalent between randomized treatment groups

The log-rank framework compares the observed and expected pattern of events over follow-up. For a stratified analysis, the comparison is constructed within the prespecified strata and then combined rather than treating the entire population as one unstratified group.

Cox regression for the hazard ratio

The registry states that the hazard ratio and its 95% confidence interval were calculated using a Cox regression model. Treatment was included as a covariate, with the model stratified by metastasis status, ECOG performance status, and US participant within the M0 group by investigator.

The registry-reported analysis notes also specify Efron's method of tie handling. This is a technical feature of the Cox model used when multiple participants experience events at the same recorded time.

Model-based effect measure
HR = estimated instantaneous event rate in pembrolizumab relative to placebo

An HR below 1 indicates a lower estimated instantaneous rate of the defined DFS event in the pembrolizumab group under the fitted model. The HR is a relative time-to-event measure; it is not itself a percentage of participants who avoid recurrence or death.

Analysis population

The posted statistical analysis specifies all randomized participants as the analysis population. This is consistent with the central role of randomization in an efficacy comparison: treatment assignment, rather than subsequent treatment exposure or adherence, defines the groups being compared for the primary endpoint.

Stratification

The registry-reported analysis notes identify three stratification dimensions for the primary time-to-event analysis: metastasis status, ECOG performance status, and US participant within the M0 group by investigator.

Stratification can improve the alignment between the statistical analysis and the randomized design when important baseline factors are associated with prognosis. Rather than assuming that every participant contributes to one homogeneous risk set, the analysis preserves the specified strata when constructing the treatment comparison.

One-sided testing

The registry analysis reports a one-sided p-value for the log-rank test. This is distinct from the reported confidence interval, which is explicitly identified as a two-sided 95% confidence interval.

Important distinction: a one-sided hypothesis test and a two-sided confidence interval are different statistical quantities. The p-value describes evidence against the prespecified null under the reported one-sided testing framework, whereas the 95% confidence interval describes uncertainty around the estimated hazard ratio using a two-sided interval.

6. Results

Disease-free Survival (DFS) as Assessed by the Investigator

Hazard ratio for disease-free survival

0.68

95% CI: 0.53–0.87   ·   One-sided P = 0.0010

Analysis population: all randomized participants

EndpointPembrolizumab vs placeboAnalysis
Disease-free SurvivalHR 0.68Stratified log-rank test
95% confidence interval0.53–0.87Two-sided 95% CI from Cox regression
P-value0.0010One-sided log-rank test
Time frameUp to approximately 42 monthsDatabase cutoff: 14 Dec 2020
Clinical Biostats interpretation

An HR of 0.68 means that, under the fitted Cox model, the estimated instantaneous rate of experiencing the defined DFS event in the pembrolizumab group was approximately 68% of the corresponding estimated rate in the placebo group. Equivalently, the estimated hazard was approximately 32% lower under this model.

The HR does not mean that 32% of participants avoided recurrence, that 32% of participants were cured, or that each participant experienced exactly a 32% reduction in risk. It is a relative time-to-event measure derived from a statistical model.

The 95% CI of 0.53–0.87 expresses uncertainty around the estimated hazard ratio under the analysis framework. Because the interval is entirely below 1, the range of model-compatible relative hazard estimates represented by this interval is below the no-difference value of 1.

The p-value of 0.0010 is evidence against the null hypothesis under the reported one-sided log-rank testing framework. It is not a measure of the size of the treatment effect. A very small p-value can coexist with a modest effect, while a larger p-value does not by itself establish that an effect is clinically unimportant.

The result should also be interpreted in light of the endpoint definition, censoring, stratification, and the assumptions underlying the Cox model. In particular, a single hazard ratio is most straightforward to interpret when the proportional-hazards assumption is reasonably appropriate over the analyzed follow-up.

How the reported statistics fit together

Log-rank test

Provides the reported hypothesis-test comparison of the time-to-event experience between randomized groups, using the specified stratification.

Cox regression

Provides the reported hazard ratio and two-sided 95% confidence interval while incorporating the specified stratification variables.

Hazard ratio

Quantifies the relative modeled event hazard associated with pembrolizumab versus placebo.

Confidence interval

Shows statistical uncertainty around the estimated hazard ratio rather than variation in individual treatment effects.

Educational note: a Kaplan-Meier curve is not reconstructed here because the ClinicalTrials.gov record provides the hazard ratio, confidence interval, and p-value but do not provide the underlying participant-level event and censoring times needed to reproduce a valid curve.

7. Statistical Methods Explained

Why was a log-rank test used?

DFS is a time-to-event endpoint. Participants can experience events at different follow-up times, and some participants may remain event-free when their observation ends. The log-rank test is designed to compare survival distributions while using the timing of events rather than reducing the analysis to a simple event/no-event proportion.

Why was a Cox model used for the hazard ratio?

The Cox proportional-hazards model provides a way to estimate a relative hazard while incorporating covariates and stratification. In this trial, treatment was the covariate of interest, while the reported model was stratified by metastasis status, ECOG performance status, and US participant within the M0 group by investigator.

What does an HR of 0.68 mean?

An HR of 0.68 means that the estimated instantaneous event rate in the pembrolizumab group was 0.68 times that in the placebo group under the fitted model. The complementary interpretation is an approximately 32% lower estimated hazard. This is not the same as saying that 32% of participants benefited or that the probability of an event was reduced by exactly 32% at every time point.

Why does the confidence interval matter?

The point estimate is only one estimate of the underlying treatment effect. The 95% CI of 0.53–0.87 communicates the uncertainty associated with that estimate under the statistical model and sampling framework. It is therefore more informative to report the HR together with its confidence interval than to report the HR alone.

Why is the p-value not an effect size?

The p-value measures how incompatible the observed test statistic is with the null hypothesis under the specified testing framework. It does not quantify the magnitude of the treatment effect. The HR supplies the effect estimate, while the confidence interval supplies information about its precision.

Why does stratification matter?

Stratification allows the time-to-event comparison to account for the specified baseline factors without requiring the analysis to impose one common baseline hazard across those strata. In KEYNOTE-564, the registry-reported analysis notes identify metastasis status, ECOG performance status, and US participant within the M0 group by investigator as the stratification dimensions.

What assumption should be considered when interpreting a single HR?

The Cox model is commonly interpreted through the proportional-hazards framework, in which the relative hazard is assumed to be reasonably stable over time. If hazards change substantially relative to one another, a single HR may compress a more complicated time-varying treatment effect into one summary number.

8. Randomization and Masking

The study is registered as randomized with a parallel design and quadruple masking. These design features address different sources of potential bias.

Design featureStatistical relevance
RandomizationCreates the basis for a treatment-group comparison in which treatment assignment is not determined by investigators or participants.
Parallel designParticipants remain in their randomized treatment groups rather than receiving sequential randomized interventions.
Quadruple maskingReduces the potential influence of treatment knowledge on trial conduct, assessment, or other study processes.
Two intervention groupsProvides a direct randomized comparison of pembrolizumab and placebo.

Randomization and masking do not remove every source of uncertainty. They instead form part of the design framework within which the statistical analysis is interpreted. The observed HR remains an estimate, which is why its confidence interval and testing framework are essential parts of the result.

9. Covariate Adjustment and Stratified Analysis

The registry's analysis notes distinguish between the statistical role of the log-rank test and the Cox regression model. The p-value was calculated from the stratified log-rank test, whereas the HR and 95% CI were calculated from the Cox regression model.

ComponentReported approach
Hypothesis testOne-sided stratified log-rank test
Effect estimateHazard ratio
Confidence intervalTwo-sided 95% CI
Regression modelCox regression
Tie handlingEfron's method
Treatment variableTreatment included as a covariate
StratificationMetastasis status; ECOG performance status; US participant within M0 group by investigator

This separation is important. It is possible for a trial report to use one statistical procedure for the formal hypothesis test and another, closely related procedure for estimation of the effect size. Here, the ClinicalTrials.gov record explicitly identify the log-rank test as the source of the p-value and Cox regression as the source of the HR and confidence interval.

10. Primary Endpoint Analysis in Detail

Endpoint construction

The DFS clock begins at randomization. The event time is the time to the first documented occurrence of any one of three qualifying events: local recurrence, distant kidney cancer metastasis(es), or death due to any cause.

Conceptual DFS definition
DFS = time from randomization → first qualifying event

The qualifying event is whichever occurs first: local recurrence, distant kidney cancer metastasis(es), or death due to any cause.

This construction means that DFS combines several clinically distinct pathways into a single time-to-first-event outcome. A DFS hazard ratio therefore summarizes the treatment comparison for the composite endpoint as registered; it should not automatically be interpreted as a separate hazard ratio for local recurrence, metastasis, and death.

Censoring

Time-to-event analyses generally require a rule for participants who have not experienced the event by the time their available follow-up ends. Such observations can be censored, allowing the participant's available event-free follow-up to contribute information without treating the participant as if an event had occurred at the end of observation.

The ClinicalTrials.gov record does not provide the individual censoring rules or participant-level follow-up records. Accordingly, this page explains the statistical role of censoring without adding an unreported trial-specific censoring convention.

Why the analysis uses both a test and an estimate

A statistical analysis benefits from answering two separate questions. First, is the observed difference sufficiently inconsistent with the null hypothesis under the prespecified testing framework? Second, what is the estimated magnitude of the difference, and how precise is that estimate? The log-rank p-value addresses the first question, while the Cox HR and confidence interval address the second.

11. Interpretation of the Hazard Ratio

What HR 0.68 says

The estimated hazard of the registered DFS event was lower for pembrolizumab than placebo, with a modeled hazard ratio of 0.68.

What HR 0.68 does not say

It does not state the percentage of participants who experienced an event, the absolute difference in event probability, the median DFS, or the probability that an individual participant will remain disease-free for a particular duration.

What the 95% CI adds

The 0.53–0.87 confidence interval indicates the precision of the estimated relative hazard under the reported Cox model. It should not be interpreted as a range in which individual participants' treatment effects are expected to fall.

What the p-value adds

The one-sided P = 0.0010 summarizes evidence against the null hypothesis using the reported stratified log-rank test. It does not tell us that there is a 0.1% probability that the null hypothesis is true, nor does it quantify clinical magnitude.

12. Confidence Intervals and Statistical Precision

The reported HR of 0.68 is accompanied by a two-sided 95% confidence interval from 0.53 to 0.87. Reporting both is essential because a point estimate without an uncertainty interval can give a misleading impression of precision.

Reported hazard-ratio interval
Lower CI
0.53
Point estimate
0.68
Upper CI
0.87

The interval remains below the null hazard ratio of 1.00. That feature is relevant to the statistical interpretation of the reported result, but the width of the interval also matters: it shows that the precise magnitude of the relative hazard reduction is not known exactly.

Confidence-interval caution: a 95% confidence interval is not a statement that there is a 95% probability that the true hazard ratio lies between 0.53 and 0.87. In the conventional frequentist framework, the interval is a procedure designed to capture the fixed parameter at the stated long-run rate under repeated sampling.

13. One-Sided Testing

The registry-reported statistical analysis reports a one-sided p-value of 0.0010. The analysis notes specifically state that this p-value was based on a stratified log-rank test using metastasis status, ECOG performance status, and US participant within the M0 group by investigator.

One-sided test

Evaluates evidence in the prespecified direction represented by the alternative hypothesis.

Two-sided confidence interval

The reported 95% CI is two-sided and therefore communicates uncertainty on both sides of the estimated HR.

These quantities should not be conflated. A p-value is tied to a hypothesis-testing procedure, while a confidence interval is an estimation procedure. Their numerical relationship depends on the exact test and interval construction.

14. Missing Data and Imputation

The ClinicalTrials.gov record identifies the primary analysis population and the time-to-event methods, but they do not provide a specific missing-data or imputation procedure for the primary DFS analysis.

For a time-to-event endpoint, missing observations are not generally handled in the same way as a missing continuous measurement. Participants can contribute follow-up until an event or censoring time, so the analysis framework incorporates incomplete observation through survival-analysis methods rather than automatically replacing an unknown event time with a single imputed value.

Because the ClinicalTrials.gov record does not specify an additional trial-specific imputation strategy, no particular imputation method is attributed to KEYNOTE-564 on this page.

15. Multiplicity and Interim Analysis

The ClinicalTrials.gov record identifies one primary endpoint and one posted formal statistical analysis for that endpoint. They do not provide a multiplicity-adjustment strategy, alpha-spending plan, interim-analysis schedule, or a numerical non-inferiority margin.

Design topicWhat the ClinicalTrials.gov record establishes
Primary endpoints1
Primary endpoint typeTime-to-event
Posted formal primary analysisYes
Multiplicity strategyNot specified in the ClinicalTrials.gov record
Interim-analysis strategyNot specified in the ClinicalTrials.gov record
Non-inferiority marginNot applicable to the reported analysis information; no margin is reported
Factorial designNot reported; design model is parallel
CrossoverNot specified in the ClinicalTrials.gov record
Bayesian methodsNot reported

This distinction matters because absence of a detail from the registry extract is not evidence that the underlying protocol or statistical analysis plan lacked such a provision. It means only that the ClinicalTrials.gov record does not establish it, so it should not be presented as a trial-specific fact here.

16. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm using affected participants over participants at risk.

Safety measurePembrolizumabPlacebo
Serious adverse events100 / 48856 / 496
Serious adverse events: affected / at risk
Pembrolizumab
100 / 488
Placebo
56 / 496

These figures describe the number affected and the corresponding number at risk reported by the registry. They should be kept separate from the DFS efficacy analysis because safety and efficacy answer different questions and may use different analysis populations or definitions.

Interpretation caution: the ClinicalTrials.gov record provides affected/at-risk counts for serious adverse events but do not provide a formal statistical comparison, confidence interval, or p-value for this safety measure. Accordingly, this page does not infer a comparative hypothesis-test result from the counts alone.

17. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary DFS analysis reported an HR of 0.68 with a two-sided 95% CI of 0.53–0.87 and a one-sided log-rank P-value of 0.0010. The analysis used stratified log-rank testing and stratified Cox regression.

Clinical interpretation

The registry result concerns the time from randomization to the first documented local recurrence, distant kidney cancer metastasis(es), or death from any cause. The ClinicalTrials.gov record does not provide median DFS or absolute DFS probabilities, so those quantities cannot be used here to characterize the absolute clinical magnitude.

This separation is useful because a statistically clear relative effect does not automatically answer every clinical question. To understand absolute benefit, one would ordinarily also examine quantities such as event probabilities at clinically meaningful time points, median event times where estimable, and the component events making up a composite endpoint. Those quantities are not reported in the ClinicalTrials.gov recordset and therefore are not added to this page.

18. Limitations

19. Why This Trial Matters Statistically

KEYNOTE-564 is a useful teaching case because its primary analysis connects several core clinical-trial concepts in a compact statistical framework: randomized allocation, masking, a composite time-to-event endpoint, stratified log-rank testing, Cox regression, hazard-ratio estimation, confidence intervals, and one-sided hypothesis testing.

ConceptHow it appears in KEYNOTE-564
RandomizationThe registry identifies the allocation as randomized.
Parallel designThe study uses a parallel design with two arms.
BlindingThe study is quadruple-masked.
Time-to-event endpointDFS is measured from randomization to the first qualifying event.
Kaplan-Meier frameworkTime-to-event data are naturally represented through survival-function estimation, although no Kaplan-Meier estimates are posted in the registry-reported analysis.
Hazard ratioThe reported primary effect measure is HR 0.68.
Confidence intervalThe HR has a two-sided 95% CI of 0.53–0.87.
Log-rank testingThe primary p-value comes from a stratified log-rank test.
Cox regressionHR and 95% CI were calculated using a stratified Cox regression model.
Covariate adjustmentTreatment was included as a covariate in the Cox model.
Stratified analysisThe analysis was stratified by metastasis status, ECOG performance status, and US participant within the M0 group by investigator.
One-sided testingThe reported log-rank p-value is one-sided.
Tie handlingEfron's method was used in the Cox model.
Safety analysisSerious adverse events are reported by arm as affected/at risk.

20. A Statistical Reading of the Primary Result

The most complete way to read the primary result is to consider the endpoint, effect estimate, uncertainty, and hypothesis test together.

QuestionKEYNOTE-564 result
What was measured?Time from randomization to the first documented local recurrence, distant kidney cancer metastasis(es), or death due to any cause.
What was the relative effect?HR 0.68 for pembrolizumab versus placebo.
How precise was the estimate?Two-sided 95% CI 0.53–0.87.
What was the hypothesis-test result?One-sided log-rank P = 0.0010.
Which population was analyzed?All randomized participants.
How was the p-value calculated?Stratified log-rank test.
How was the HR calculated?Cox regression with treatment as a covariate and the specified stratification factors.

That sequence prevents a common statistical error: treating the p-value as though it were the primary result. The p-value tells us about evidence under the testing framework; the HR tells us the estimated relative effect; and the confidence interval communicates uncertainty around that effect.

21. What the Registry Does Not Establish

The ClinicalTrials.gov record is sufficient to characterize the primary statistical result, but they are not sufficient to answer every possible efficacy question. In particular, the dataset does not provide median DFS, time-specific DFS probabilities, subgroup estimates, component-specific recurrence estimates, or a forest plot.

It also does not provide trial-specific details for multiplicity control, interim monitoring, missing-data imputation, crossover, or Bayesian analysis. These topics therefore should not be retroactively attributed to KEYNOTE-564 based only on the fact that they are common features of modern clinical trials.

Methodological principle: absence of a number in a registry extract is different from a number being zero, not reached, or unavailable in the underlying trial. This page reports only what the ClinicalTrials.gov record establishes.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue through the Clinical Biostats statistical tutorials

Use this trial as a practical example of randomized treatment comparison, survival analysis, hazard ratios, confidence intervals, and stratified hypothesis testing.

25. Record Summary

KEYNOTE-564 provides a focused example of a randomized phase 3 time-to-event analysis. The registry identifies a randomized, parallel, quadruple-masked design with 994 enrolled participants and one primary endpoint: investigator-assessed disease-free survival. The primary analysis used a stratified log-rank test for the hypothesis test and a stratified Cox regression model with Efron's method of tie handling for the hazard ratio and 95% confidence interval.

The reported DFS result was HR 0.68, with a two-sided 95% CI of 0.53–0.87 and a one-sided log-rank P-value of 0.0010. The most appropriate interpretation is therefore a statistically supported relative difference in the time-to-first-event DFS endpoint under the specified analysis framework, with the exact magnitude represented by the HR and its confidence interval rather than by the p-value alone.

Clinical Biostats methodology: A trial-results page should not merely repeat a registry result. The goal is to explain how the endpoint was constructed, how the statistical comparison was performed, what the effect estimate means, what the confidence interval contributes, and which interpretations are not supported by the available data.