This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-426 was a randomized, parallel-group, open-label phase 3 trial in renal cell carcinoma comparing pembrolizumab plus axitinib with sunitinib monotherapy. The registry reports two primary time-to-event endpoints: progression-free survival and overall survival.
| Feature | KEYNOTE-426 |
|---|---|
| Phase | Phase 3 |
| Condition | Renal Cell Carcinoma |
| Design | Randomized, parallel-group |
| Masking | None |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 861 |
| Arms | 2 |
| Primary endpoints | Progression-free survival and overall survival |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| Trial status | COMPLETED |
| Start | 16-Sep-2016 |
| Primary completion | 18-Oct-2018 |
| Lead sponsor | Merck Sharp & Dohme LLC |
2. Clinical Question
The central statistical question was whether the pembrolizumab plus axitinib combination produced better time-to-event outcomes than sunitinib monotherapy in participants with renal cell carcinoma. The registry classifies both primary hypotheses as superiority questions.
Population
Participants with renal cell carcinoma enrolled in the phase 3 randomized trial.
Intervention
Pembrolizumab plus axitinib.
Comparator
Sunitinib monotherapy.
Primary question
Does pembrolizumab plus axitinib improve progression-free survival and overall survival relative to sunitinib?
3. Trial Design
Combination treatment
- Pembrolizumab
- Axitinib
- Combination treatment evaluated against sunitinib monotherapy
Monotherapy control
- Sunitinib
- Monotherapy comparator
The design is explicitly listed as randomized, parallel, and unmasked. These features are important when interpreting the evidence: randomization supports the causal treatment comparison, while the absence of masking can be relevant to outcomes whose assessment may be influenced by knowledge of treatment assignment. For the primary PFS endpoint, however, the registry specifies blinded independent central imaging review.
4. Randomization, Stratification, and Analysis Populations
The registry's primary endpoint analyses used all randomized participants as the analysis population. Both PFS and OS analyses were based on the randomized comparison between pembrolizumab plus axitinib and sunitinib.
| Analysis | Population / approach |
|---|---|
| Primary PFS analysis | All randomized participants |
| Primary OS analysis | All randomized participants |
| PFS comparison | Pembrolizumab + axitinib vs sunitinib |
| OS comparison | Pembrolizumab + axitinib vs sunitinib |
| Stratification | IMDC risk group and geographic region |
The reported Cox models stratified by IMDC risk group — favorable, intermediate, and poor — and by geographic region — North America, Western Europe, and Rest of World. This allows the treatment comparison to account for these prespecified grouping factors without treating them as ordinary covariates in the same way as an unstratified regression model.
5. Primary Endpoints
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Progression-Free Survival (PFS) | PFS was defined as the time from randomization to the first documented progressive disease (PD) or death due to any cause, whichever occurred first. Per RECIST 1.1, PD was defined as ≥20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study. In addition to the relative increase of 20%, the sum must also demonstrate an absolute increase of ≥5 mm. | Through Database Cutoff Date of 24-Aug-2018 (up to approximately 22 months) |
| Overall Survival (OS) | OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the interim analysis were censored at the date of the last follow up. The OS was calculated using the product-limit (Kaplan-Meier) method for censored data. | Through Database Cutoff Date of 24-Aug-2018 (up to approximately 22 months) |
Both primary endpoints are time-to-event endpoints. That distinction drives the statistical methodology: patients can contribute follow-up even if they have not yet experienced the event by the database cutoff, and those observations are handled through censoring rather than simply being treated as event-free failures.
6. Statistical Methodology
Kaplan-Meier estimation
The registry explicitly identifies the Kaplan-Meier, or product-limit, approach for OS. Kaplan-Meier estimation is designed for right-censored time-to-event data. At each observed event time, the estimated survival probability is updated according to the number of events and the number of participants still at risk.
Here di is the number of events at time ti, while ni is the number at risk immediately before that event time.
Log-rank test
The registry reports a log-rank test for both primary endpoints. The log-rank framework compares the observed and expected numbers of events between randomized groups over follow-up. It is particularly suited to randomized time-to-event comparisons because it uses the ordering of event times while accommodating censoring.
Stratified Cox regression
The hazard ratios were based on Cox regression with treatment as a covariate and stratification by IMDC risk group and geographic region. Efron's method was used for handling tied event times.
A hazard ratio is a relative time-to-event measure. It is not an absolute risk reduction, a probability that a patient will benefit, or a statement that every individual patient experiences the same proportional reduction.
Score-based confidence intervals for proportions
The registry reports the Miettinen & Nurminen method for the objective response rate analysis and identifies score-based confidence intervals for proportions more generally in the posted analyses. The treatment effect was expressed as a difference in percentages, which is a risk-difference scale.
A positive difference means that the percentage meeting the response definition was higher in the pembrolizumab-plus-axitinib group than in the sunitinib group.
7. Primary Results: Progression-Free Survival
The primary PFS analysis included all randomized participants and compared pembrolizumab plus axitinib with sunitinib through the database cutoff date of 24-Aug-2018, up to approximately 22 months.
Hazard ratio for progression or death
95% CI: 0.56–0.84 · P = 0.00012
Analysis: stratified Cox regression with a log-rank test; two-sided 95% confidence interval.
| Feature | PFS primary analysis |
|---|---|
| Endpoint | Progression-Free Survival (PFS) per RECIST 1.1, assessed by blinded independent central imaging review |
| Analysis population | All randomized participants |
| Groups compared | Pembrolizumab + axitinib vs sunitinib |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.69 |
| 95% CI | 0.56–0.84 |
| P-value | 0.00012 |
| Hypothesis | Superiority |
The estimated PFS hazard ratio of 0.69 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death in the pembrolizumab-plus-axitinib group was approximately 69% of the corresponding rate in the sunitinib group. Equivalently, the estimate corresponds to an approximately 31% lower estimated hazard for progression or death.
The hazard ratio does not mean that 31% of patients avoided progression, that every patient had exactly a 31% reduction in risk, or that median PFS was reduced or increased by 31%. It is a relative model-based measure of the event rate over follow-up.
The two-sided 95% confidence interval of 0.56–0.84 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual patient outcomes. The interval remains below 1, which is consistent with the direction of the superiority hypothesis.
The P-value of 0.00012 addresses the evidence against the null hypothesis under the specified testing framework. It is not a measure of effect size. The size of the treatment effect is better communicated by the hazard ratio and its confidence interval.
The interpretation also depends on the Cox-model framework and its assumptions. In particular, a single hazard ratio is most straightforward when the proportional-hazards assumption is a reasonable description of the treatment-specific hazards. The registry result does not provide enough information here to independently assess that assumption from the underlying event-time data.
8. Primary Results: Overall Survival
The primary OS analysis used the same randomized analysis population and the same database cutoff date. The registry defines OS as time from randomization to death from any cause, with participants without documented death censored at their last follow-up.
Hazard ratio for death
95% CI: 0.38–0.74 · P = 0.00005
Analysis: stratified Cox regression with a log-rank test; two-sided 95% confidence interval.
| Feature | OS primary analysis |
|---|---|
| Endpoint | Overall Survival |
| Analysis population | All randomized participants |
| Groups compared | Pembrolizumab + axitinib vs sunitinib |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.53 |
| 95% CI | 0.38–0.74 |
| P-value | 0.00005 |
| Hypothesis | Superiority |
The estimated OS hazard ratio of 0.53 means that, under the fitted Cox model, the estimated instantaneous rate of death in the pembrolizumab-plus-axitinib group was approximately 53% of the rate in the sunitinib group. Expressed as a relative reduction in estimated hazard, this corresponds to approximately 47% lower estimated hazard.
This does not mean that 47% of patients survived, that 47% of deaths were prevented, or that an individual patient's probability of death was reduced by exactly 47%. Those would be different statistical quantities.
The 95% confidence interval of 0.38–0.74 gives a range of values reflecting uncertainty around the model-based estimate. A confidence interval is more informative than the point estimate alone because it shows how precisely the treatment effect was estimated within the statistical framework.
The P-value of 0.00005 indicates the strength of evidence against the specified null hypothesis under the trial's analysis framework. It does not tell us whether the effect is clinically large or small; effect magnitude is conveyed by the hazard ratio, confidence interval, and, where available, absolute survival measures.
As with PFS, the interpretation requires care with censoring and the Cox-model framework. The registry does not provide enough underlying data on this page to independently evaluate the proportional-hazards assumption or reconstruct the complete survival curve.
9. Primary Endpoint Comparison
The two primary estimates point in the same direction, with hazard ratios below 1 for both progression-free survival and overall survival. The OS estimate is numerically lower than the PFS estimate, but the two hazard ratios should not be interpreted as interchangeable quantities: one concerns progression or death, while the other concerns death alone.
| Primary endpoint | HR | 95% CI | P-value | Interpretive scale |
|---|---|---|---|---|
| PFS | 0.69 | 0.56–0.84 | 0.00012 | Progression or death |
| OS | 0.53 | 0.38–0.74 | 0.00005 | Death |
Because both endpoints were designated primary endpoints, their interpretation belongs to the confirmatory portion of the trial rather than being treated as exploratory response findings. The registry classifies the hypotheses as superiority hypotheses.
10. Secondary Endpoint: Objective Response Rate
Objective Response Rate (ORR) was a secondary binary endpoint assessed according to RECIST 1.1 by blinded independent central imaging review. The registry reports the treatment effect as a difference in percentages, with the Miettinen & Nurminen method used for the confidence interval.
Difference in objective response rate
95% CI: 17.2–29.9 · P < 0.0001
Two-sided analysis; difference in percentages stratified by IMDC risk group and geographic region.
| Feature | ORR analysis |
|---|---|
| Endpoint | Objective Response Rate per RECIST 1.1 as assessed by blinded independent central imaging review |
| Time frame | Through Database Cutoff Date of 24-Aug-2018 (up to approximately 22 months) |
| Endpoint type | Binary |
| Analysis population | All randomized participants |
| Groups compared | Pembrolizumab + axitinib vs sunitinib |
| Method | Miettinen & Nurminen method |
| Effect measure | Difference in percentages |
| Estimate | 23.6 |
| 95% CI | 17.2–29.9 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The reported difference in percentages of 23.6 means that the percentage of participants meeting the objective-response definition was estimated to be 23.6 percentage points higher in the pembrolizumab-plus-axitinib group than in the sunitinib group.
This is an absolute difference in response percentages, not a 23.6% relative increase and not a statement that 23.6% of all treated participants experienced an additional response. The distinction between percentage-point difference and relative effect is important for binary endpoints.
The two-sided 95% confidence interval of 17.2–29.9 describes uncertainty around the estimated difference. It does not represent the range of individual patient responses.
The P-value of <0.0001 provides evidence against the null hypothesis that the difference in percentages is zero under the reported testing framework. It does not quantify how clinically meaningful the response difference is.
The confidence interval was calculated using the Miettinen & Nurminen method and the comparison was stratified by IMDC risk group and geographic region. This is distinct from the Cox-model approach used for the primary time-to-event endpoints.
11. Secondary Endpoint: Disease Control Rate
Disease Control Rate (DCR) was another binary endpoint. The registry defines its analysis population as all randomized participants who experienced a partial response, complete response, or stable disease for at least 6 months.
Difference in disease control rate
95% CI: 4.8–17.0
Difference in percentages based on the Miettinen & Nurminen method, stratified by IMDC risk group and geographic region.
| Feature | DCR analysis |
|---|---|
| Endpoint | Disease Control Rate per RECIST 1.1 as assessed by blinded independent central imaging review |
| Time frame | Through Database Cutoff Date of 24-Aug-2018 (up to approximately 22 months) |
| Endpoint type | Binary |
| Analysis population | All randomized participants who experienced a PR, CR or SD for ≥6 months |
| Groups compared | Pembrolizumab + axitinib vs sunitinib |
| Method | Score-based confidence interval for proportions using the Miettinen & Nurminen approach |
| Effect measure | Difference in percentages |
| Estimate | 11.0 |
| 95% CI | 4.8–17.0 |
| Hypothesis | Superiority |
The DCR result requires a different interpretation from ORR because the registry-defined analysis population is explicitly restricted to randomized participants who experienced a PR, CR, or SD for at least 6 months. It therefore should not be read as though it were simply another time-to-event estimate calculated over the same population and event definition as PFS or OS.
12. Statistical Methods Explained
Why was a log-rank test used for PFS and OS?
PFS and OS are time-to-event endpoints, so the timing of events matters as well as whether an event occurred. The log-rank test compares the treatment groups across the observed follow-up while accommodating censoring. It is therefore suited to randomized survival comparisons in which some participants have not experienced the event by the analysis cutoff.
What does a hazard ratio of 0.69 mean?
A hazard ratio of 0.69 indicates that the estimated instantaneous rate of progression or death was approximately 69% of the comparator rate under the fitted Cox model. The corresponding relative reduction in estimated hazard is approximately 31%. It does not mean that 31% of patients were protected from progression or death.
Why is the OS hazard ratio different from the PFS hazard ratio?
PFS and OS measure different events. PFS counts either documented progression or death, whichever occurs first, whereas OS counts death from any cause. Because their event definitions differ, their hazard ratios quantify different treatment effects and should not be compared as though they were measurements of the same endpoint.
Why was stratification used in the Cox model?
The reported Cox models were stratified by IMDC risk group and geographic region. Stratification allows the baseline hazard to differ across those strata while estimating the treatment effect within the overall stratified framework. This can account for important differences in underlying event rates associated with the stratification factors.
What does the 23.6 percentage-point ORR difference mean?
The risk-difference scale expresses the treatment contrast as an absolute difference between response percentages. A value of 23.6 means the estimated response percentage was 23.6 percentage points higher with pembrolizumab plus axitinib than with sunitinib. It is not a relative risk ratio and should not be described as a 23.6% relative increase.
Why use the Miettinen & Nurminen method for response?
For a binary endpoint such as objective response, the treatment effect can be expressed as a difference between proportions. The Miettinen & Nurminen method provides a score-based confidence interval for this difference and, in this trial, was used with stratification by IMDC risk group and geographic region.
Why does censoring matter?
Participants who have not experienced progression or death by the analysis cutoff are not automatically treated as if they will never experience the event. Instead, their observed follow-up contributes information until the time at which they are censored. Kaplan-Meier and Cox methods are designed to incorporate this incomplete follow-up.
13. Understanding the Confidence Intervals
| Endpoint | Estimate | 95% CI | What the interval communicates |
|---|---|---|---|
| PFS hazard ratio | 0.69 | 0.56–0.84 | Uncertainty around the model-based relative event-rate estimate |
| OS hazard ratio | 0.53 | 0.38–0.74 | Uncertainty around the model-based relative death-rate estimate |
| ORR difference | 23.6 | 17.2–29.9 | Uncertainty around the absolute difference in response percentages |
| DCR difference | 11.0 | 4.8–17.0 | Uncertainty around the absolute difference in disease-control percentages |
The confidence intervals are on different statistical scales. The PFS and OS intervals are ratios, while the ORR and DCR intervals are differences in percentages. This is why the numerical ranges should not be compared directly as though they represented the same quantity.
The 95% confidence intervals do not say that 95% of individual patients will experience an effect inside the reported range. They describe uncertainty around the estimated population-level treatment effect under the relevant statistical model and sampling framework.
14. Why the P-value Does Not Measure Effect Size
KEYNOTE-426 illustrates why a P-value and an effect estimate answer different questions. The PFS result has a hazard ratio of 0.69 and a P-value of 0.00012; the OS result has a hazard ratio of 0.53 and a P-value of 0.00005. The hazard ratios describe the estimated treatment effects, while the P-values describe evidence against the corresponding null hypotheses under the testing framework.
A very small P-value does not itself tell us whether an effect is large, clinically important, or precisely estimated. Conversely, an effect estimate cannot be interpreted without considering its uncertainty. For this reason, a statistically literate trial summary should report the estimate, confidence interval, and P-value together.
15. Safety Results
The registry reports serious adverse events by treatment arm as follows:
| Safety measure | Pembrolizumab + Axitinib | Sunitinib |
|---|---|---|
| Serious adverse events | 173/429 | 133/425 |
| Denominator | 429 | 425 |
These are reported counts of affected participants over the reported at-risk denominators. They are not time-to-event efficacy measures and should not be interpreted through the PFS or OS hazard ratios. The ClinicalTrials.gov record does not provide additional safety estimates or formal hypothesis tests for these serious adverse-event counts.
16. Design Features That Matter Statistically
Randomization
Random allocation creates the principal basis for comparing outcomes between treatment groups while reducing systematic confounding from measured and unmeasured baseline characteristics.
Open-label design
The trial was unmasked. Knowledge of treatment assignment can matter for some assessments, which makes the blinded independent central imaging review specified for PFS particularly relevant.
Superiority hypothesis
The registry classifies the primary hypotheses as superiority rather than non-inferiority. The interpretation therefore focuses on evidence that the treatment effect differs from the null in the favorable direction.
Stratified analysis
The primary Cox models accounted for IMDC risk group and geographic region through stratification, aligning the analysis with clinically relevant trial-design factors.
17. What Is Not Supported by the Registry Data Provided Here
The ClinicalTrials.gov record contains the primary and selected secondary statistical analyses, but they do not provide several quantities that would ordinarily appear in a more extensive clinical-trial report. In accordance with the registry-only data rule for this page, those quantities are not reconstructed or imported from external publications.
| Topic | Available in the ClinicalTrials.gov record? |
|---|---|
| Primary PFS estimate and CI | Yes |
| Primary OS estimate and CI | Yes |
| ORR difference and CI | Yes |
| DCR difference and CI | Yes |
| Serious adverse events by arm | Yes |
| Median PFS | Not reported in the ClinicalTrials.gov record |
| Median OS | Not reported in the ClinicalTrials.gov record |
| Time-specific survival percentages | Not reported in the ClinicalTrials.gov record |
| Subgroup-specific efficacy estimates | Not reported in the ClinicalTrials.gov record |
| Interim alpha-spending details | Not reported in the ClinicalTrials.gov record |
| Crossover analysis | Not reported in the ClinicalTrials.gov record |
| Bayesian analysis | Not reported in the ClinicalTrials.gov record |
| Missing-data / imputation strategy | Not reported in the ClinicalTrials.gov record |
The absence of these details does not mean that they were necessarily absent from the full trial protocol or statistical analysis plan. It means only that they are not part of the trial data posted on ClinicalTrials.gov for this independent registry-based analysis.
18. Statistical Interpretation of the Complete Evidence Set
The two primary time-to-event analyses report hazard ratios below 1: 0.69 for PFS and 0.53 for OS. Both two-sided 95% confidence intervals remain below 1, and the registry reports P-values of 0.00012 and 0.00005, respectively.
The secondary ORR analysis reports a 23.6 percentage-point difference, with a two-sided 95% confidence interval of 17.2–29.9 and P < 0.0001. The DCR analysis reports an 11.0 percentage-point difference with a 95% confidence interval of 4.8–17.0.
The trial therefore illustrates two distinct effect-measure families: hazard ratios for time-to-event outcomes and risk differences for binary response outcomes. Keeping these scales separate prevents common interpretation errors.
The primary PFS and OS analyses used all randomized participants. That feature is important because the randomized treatment assignment, rather than subsequent treatment exposure alone, defines the primary efficacy comparison.
19. Important Limitations and Interpretation Issues
- Open-label design: the trial was unmasked. Although the primary PFS endpoint used blinded independent central imaging review, knowledge of treatment assignment can still be relevant to other trial processes and outcomes.
- Hazard-ratio interpretation: a Cox hazard ratio is a model-based relative measure. It should not be translated directly into an absolute probability or individual patient benefit.
- Proportional-hazards assumption: the registry supplies hazard-ratio estimates but does not provide enough underlying event-time information in the ClinicalTrials.gov record to independently assess whether proportional hazards adequately describe the follow-up.
- Censoring: PFS and OS depend on appropriate handling of censored observations. The registry definition explicitly describes censoring for participants without documented death at the OS analysis.
- Different endpoint populations: the DCR analysis population is specifically defined as randomized participants who experienced PR, CR, or SD for at least 6 months, so it should not be treated as identical to the primary time-to-event analysis population without qualification.
- Secondary endpoints: ORR and DCR are secondary outcomes and should not automatically be interpreted with the same confirmatory status as the two registered primary endpoints.
- Multiplicity details: the ClinicalTrials.gov record identifies two primary endpoints but do not provide the complete multiplicity-control strategy. The P-values should therefore be interpreted within the reported trial framework rather than assuming an unreported adjustment scheme.
- Missing-data and imputation details: the ClinicalTrials.gov record does not report a detailed missing-data or imputation strategy, so none is inferred here.
- Long-term quantities: median survival and other mature follow-up estimates are not included in the ClinicalTrials.gov record and are therefore not presented.
20. Why This Trial Matters Statistically
KEYNOTE-426 is a useful statistical teaching case because it combines randomized treatment allocation with two primary time-to-event endpoints and secondary binary response measures. The resulting analysis requires several different statistical ideas to be understood together rather than treating every endpoint as if it were a simple comparison of percentages.
| Concept | How it appears in KEYNOTE-426 |
|---|---|
| Randomization | 861-participant randomized parallel-group phase 3 trial |
| Time-to-event endpoints | PFS and OS were the two registered primary endpoints |
| Kaplan-Meier estimation | OS was calculated using the product-limit Kaplan-Meier method |
| Log-rank test | Used for the primary PFS and OS treatment comparisons |
| Hazard ratio | Primary effect measure for PFS and OS |
| Cox regression | Used to estimate HRs with treatment as a covariate |
| Stratification | IMDC risk group and geographic region were used in the Cox analyses |
| Efron tie handling | Used for tied event times in the Cox model |
| RECIST 1.1 | Used for the registered PFS and response assessments |
| Blinded central review | Specified for the primary PFS and response assessments |
| Risk difference | ORR and DCR effects were expressed as differences in percentages |
| Miettinen & Nurminen | Used for the reported ORR percentage difference and confidence interval |
| Superiority testing | Primary hypotheses were classified as superiority |
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Calculators
23. Sources
- ClinicalTrials.gov: NCT02853331 — KEYNOTE-426.
- PubMed: PMID 40750932.
- PubMed: PMID 39815637.
- PubMed: PMID 37526095.
- PubMed: PMID 37146227.
- PubMed: PMID 35843776.
Continue through the Clinical Biostats statistical library
Connect the endpoints and methods in this trial to deeper statistical tutorials and practical calculation tools.
24. Record Summary
KEYNOTE-426 provides a compact example of how a modern randomized oncology trial can require multiple statistical frameworks. The two primary endpoints were time-to-event outcomes analyzed using log-rank testing and stratified Cox regression, with hazard ratios of 0.69 for PFS and 0.53 for OS. Secondary binary outcomes used a different effect scale: differences in response percentages estimated with the Miettinen & Nurminen approach. The registry also specifies blinded independent central imaging review for the primary PFS and response assessments, while the primary Cox models were stratified by IMDC risk group and geographic region.
The most useful statistical reading therefore keeps several distinctions clear: hazard ratio versus risk difference, time-to-event versus binary outcomes, point estimates versus confidence intervals, P-values versus effect sizes, and randomized efficacy populations versus endpoint-specific analysis populations. Those distinctions allow the reported results to be understood without overstating what any individual statistic can establish.