← Clinical Trials
Renal Cell Carcinoma Phase 3 Completed NCT03141177

CheckMate-9ER: Complete Statistical Analysis of Nivolumab Plus Cabozantinib in Renal Cell Carcinoma

An independent statistical analysis of the randomized phase 3 CheckMate-9ER trial comparing nivolumab combined with cabozantinib with sunitinib in previously untreated advanced or metastatic renal cell carcinoma.

Study start: July 23, 2017  ·  Primary completion: February 12, 2020  ·  Enrollment: 701
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data reported in the ClinicalTrials.gov record.

1. Trial at a Glance

CheckMate-9ER was a randomized, parallel, phase 3 treatment trial in renal cell carcinoma comparing treatment A, nivolumab combined with cabozantinib, with treatment C, sunitinib. The registry reports a total enrollment of 701 participants across 3 arms and a primary endpoint of progression-free survival.

701
Enrolled
3 treatment arms
3
Arms
Parallel design
0.51
PFS HR
95% CI 0.41–0.64
<0.0001
PFS P-value
Superiority analysis
FeatureCheckMate-9ER
PhasePhase 3
ConditionRenal Cell Carcinoma
Brief titleA Study of Nivolumab Combined With Cabozantinib Compared to Sunitinib in Previously Untreated Advanced or Metastatic Renal Cell Carcinoma
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment701
Lead sponsorBristol-Myers Squibb
Sponsor typeIndustry
StatusCompleted
ClinicalTrials.govNCT03141177

2. Clinical Question

The registered clinical question was whether nivolumab combined with cabozantinib improves the primary time-to-event endpoint of progression-free survival compared with sunitinib in previously untreated advanced or metastatic renal cell carcinoma.

Population

Participants with previously untreated advanced or metastatic renal cell carcinoma.

Intervention

Nivolumab combined with cabozantinib. The registry identifies nivolumab as a biological intervention and cabozantinib as a drug.

Comparator

Sunitinib, identified in the registry as a drug intervention.

Primary question

Does treatment A, nivolumab plus cabozantinib, improve progression-free survival relative to treatment C, sunitinib?

3. Trial Design

01
Enroll 701 participants
02
Randomize 3 parallel arms
03
Treat Randomized treatment assignment
04
Assess Time-to-event and response outcomes
05
Compare Stratified statistical tests
Allocation
Randomized.
Design model
Parallel.
Masking
None.
Primary purpose
Treatment.

The study enrolled 701 participants and used 3 arms. The statistical analyses in the ClinicalTrials.gov record compare treatment A with treatment C. The registry data identify treatment A in the objective-response analysis as the nivolumab-plus-cabozantinib group and treatment C as the sunitinib group.

TREATMENT A

Nivolumab + Cabozantinib

  • Nivolumab
  • Cabozantinib
  • Primary efficacy comparison against treatment C
TREATMENT C

Sunitinib

  • Sunitinib
  • Comparator for the reported primary analysis
  • Comparator for the reported secondary analyses
Important arm-label distinction: the registry record contains 3 treatment arms, but the formal analyses reported here compare treatment A with treatment C. Treatment B is not characterized in the ClinicalTrials.gov record, so no additional treatment comparison is inferred.

4. Trial Timeline and Registry Status

MilestoneDate / status
Study start2017-07-23
Primary completion2020-02-12
Overall statusCompleted
Results postedYes
Outcome measures posted9
Statistical analyses posted4

The ClinicalTrials.gov record contains 4 statistical analyses, including 1 analysis of the registered primary endpoint and 3 secondary analyses. This provides a compact statistical picture: one formal primary time-to-event comparison and three additional efficacy comparisons.

5. Endpoints

EndpointRegistry time frameTypeAnalysis
Progression Free Survival (PFS) From randomization date to date of first documented tumor progression or death, whichever occurs first (Up to 31 months) Time-to-event Stratified log-rank; hazard ratio
Overall Survival (OS) From randomization date to death date (Up to 31 months) Time-to-event Stratified log-rank; hazard ratio
Objective Response Rate (ORR) Up to 31 Months Binary Stratified Cochran-Mantel-Haenszel; odds ratio
Objective Response Rate (ORR) Up to 31 Months Binary Strata-adjusted difference in objective response rate

Registered primary endpoint: Progression Free Survival

PFS is defined as the time from date of randomization to the first documented tumor progression date or death due to any cause, whichever occurs first based on BICR assessment using RECIST v1.1. Participants who die without a reported progression will be considered to have progressed on the date of their death. Participants who did not progress or die will be censored on the date of their last evaluable tumor assessment on or prior to initiation of subsequent anti-cancer therapy. Progressive disease (PD); 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study

The registry classifies PFS as a time-to-event endpoint and reports the analysis population as all randomized participants of treatment A and treatment C.

6. Statistical Methodology

Stratified log-rank test

The primary PFS comparison used a stratified log-rank test. The test compares the observed pattern of event occurrence between randomized treatment groups while incorporating stratification. The registry does not identify the specific stratification factors in the ClinicalTrials.gov record, so no particular clinical variable is attributed as a stratification factor.

Time-to-event comparison
H0: no difference in the time-to-event distributions between treatment A and treatment C

The stratified log-rank procedure evaluates the evidence for a treatment difference across follow-up rather than comparing only a single fixed-time proportion.

Hazard ratio

The primary PFS effect measure is the hazard ratio (HR). The reported estimate is 0.51. In the context of the analysis, an HR below 1 indicates a lower estimated hazard of progression or death for treatment A relative to treatment C.

Conceptual interpretation
HR = hazard in treatment A / hazard in treatment C

An HR of 0.51 corresponds to an estimated hazard approximately 49% lower in treatment A than treatment C, because 1 − 0.51 = 0.49. This is a relative hazard interpretation, not an absolute probability difference.

Cochran-Mantel-Haenszel test

The reported ORR analysis used a stratified Cochran-Mantel-Haenszel test. This is appropriate for comparing a binary outcome between treatment groups while accounting for strata. Here, the outcome is objective response and the effect measure is an odds ratio.

Odds ratio

The reported ORR odds ratio is 3.52 for treatment A over treatment C. An odds ratio above 1 indicates greater odds of response in treatment A under the reported comparison. The odds ratio is not itself a risk ratio and should not be described as saying that the probability of response is 3.52 times as large.

Strata-adjusted response-rate difference

The registry also reports a difference in objective response rates of 28.6 percentage points, with the analysis note identifying this as the strata-adjusted difference in objective response rate, calculated as nivolumab plus cabozantinib minus sunitinib and based on DerSimonian and Laird.

7. Primary Result: Progression-Free Survival

The registered primary endpoint has a formal statistical analysis posted for the comparison of treatment A versus treatment C in all randomized participants in those two treatment groups.

Hazard ratio for progression or death

0.51

95% CI: 0.41–0.64   ·   P < 0.0001

Stratified log-rank test  ·  Superiority analysis

FeaturePrimary PFS analysis
EndpointProgression Free Survival (PFS)
Time frameFrom randomization date to date of first documented tumor progression or death, whichever occurs first (Up to 31 months)
PopulationAll Randomized Participants of treatment A and treatment C
ComparisonTreatment A vs Treatment C
MethodStratified log-rank test
Effect measureHazard ratio
Estimate0.51
95% CI0.41–0.64
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The HR of 0.51 means that, under the time-to-event comparison represented by the reported hazard ratio, the estimated hazard of progression or death was approximately 49% lower for treatment A than treatment C. This is a relative comparison of hazards; it does not mean that 49% of patients avoided progression or death, nor that every patient experienced exactly a 49% reduction.

The 95% CI of 0.41–0.64 describes statistical uncertainty around the estimated hazard ratio. It indicates that the observed estimate is more precise than a very wide interval would be, while still recognizing uncertainty in the treatment effect. The interval does not describe the range of effects that individual patients experienced.

The P-value of <0.0001 describes the strength of evidence against the null hypothesis under the specified statistical framework. It does not measure the magnitude of the treatment effect, the probability that the treatment works, or the clinical importance of the result.

Because this is a hazard ratio, interpretation also depends on the time-to-event modeling framework and censoring process. A single HR summarizes a relative hazard comparison over the analyzed period; it is not interchangeable with an absolute risk reduction or a difference in median PFS. The ClinicalTrials.gov record does not report a median PFS estimate, so no median-time comparison is made on this page.

8. Secondary Result: Overall Survival

Overall survival was a secondary time-to-event endpoint. The registry defines the endpoint as the time from randomization date to death date, with a time frame of up to 31 months.

Hazard ratio for death

0.60

98.89% CI: 0.40–0.89   ·   P = 0.0010

Stratified log-rank test  ·  Superiority analysis

FeatureSecondary OS analysis
EndpointOverall Survival (OS)
Time frameFrom randomization date to death date (Up to 31 months)
PopulationAll Randomized Participants of treatment A and treatment C
ComparisonTreatment A vs Treatment C
MethodStratified log-rank test
Effect measureHazard ratio
Estimate0.60
CI98.89%, two-sided: 0.40–0.89
P-value0.0010
HypothesisSuperiority
Clinical Biostats interpretation

The OS HR of 0.60 corresponds to an estimated hazard of death approximately 40% lower for treatment A relative to treatment C under the reported analysis. It does not mean that treatment A reduced the probability of death by exactly 40% for each participant, nor does it imply that 40% of participants benefited.

The 98.89% CI of 0.40–0.89 expresses the uncertainty around the estimated hazard ratio using the reported confidence level. Because the interval is entirely below 1, its endpoints are consistent with a lower estimated hazard for treatment A under the specified analysis.

The P-value of 0.0010 quantifies the statistical evidence against the relevant null hypothesis under the reported testing framework. It should not be interpreted as an effect-size measure or as the probability that the null hypothesis is true.

The OS analysis is still a time-to-event comparison. Censoring, the timing of deaths, and the relationship between hazards over time all matter. The ClinicalTrials.gov record does not provide median OS, survival probabilities at specific time points, or a Kaplan-Meier dataset, so those quantities are not reconstructed.

9. Secondary Result: Objective Response Rate

Objective response rate was evaluated as a binary endpoint up to 31 months. The registry reports two statistical representations of the same treatment comparison: a stratified Cochran-Mantel-Haenszel odds ratio and a strata-adjusted difference in response rates.

Odds-ratio analysis

Odds ratio for objective response

3.52

95% CI: 2.51–4.95   ·   P < 0.0001

Stratified Cochran-Mantel-Haenszel test  ·  Superiority analysis

FeatureORR odds-ratio analysis
EndpointObjective Response Rate (ORR)
Time frameUp to 31 Months
PopulationAll Randomized Participants
ComparisonTreatment A vs Treatment C
MethodStratified Cochran-Mantel-Haenszel
Effect measureOdds ratio
Estimate3.52
95% CI2.51–4.95
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

An OR of 3.52 means that the estimated odds of objective response were 3.52 times as high for treatment A as for treatment C under the reported stratified comparison. Odds are not probabilities, so the OR should not be read as saying that the response rate was 3.52 times higher.

The 95% CI of 2.51–4.95 gives the uncertainty range reported for the odds ratio. It does not represent a range of individual patient outcomes.

The P-value of <0.0001 indicates strong statistical evidence against the null hypothesis used for the reported test. It does not indicate the size or clinical importance of the response difference by itself.

Strata-adjusted response-rate difference

Difference in objective response rates

28.6

95% CI: 21.7–35.6

Treatment A − Treatment C  ·  Strata-adjusted difference

FeatureResponse-rate difference analysis
EndpointObjective Response Rate (ORR)
Time frameUp to 31 Months
PopulationAll Randomized Participants
ComparisonTreatment A vs Treatment C
Effect measureDifference of Objective Response Rates
Estimate28.6
95% CI21.7–35.6
Analysis noteStrata adjusted difference in objective response rate (Nivolumab+Cabozantinib - Sunitinib) based on DerSimonian and Laird.
HypothesisOther / not stated
Clinical Biostats interpretation

The reported difference of 28.6 represents the strata-adjusted difference in objective response rate calculated as treatment A minus treatment C. A positive value therefore favors treatment A for this response endpoint under the direction specified in the registry analysis.

The 95% CI of 21.7–35.6 quantifies uncertainty around that adjusted difference. Unlike an odds ratio, a response-rate difference is expressed directly on the percentage-point scale.

This estimate should not be confused with the HR of 0.51 for PFS or the OR of 3.52 for ORR. These measures answer different statistical questions and use different scales.

10. Statistical Methods Explained

Why was a stratified log-rank test used for PFS?

PFS is a time-to-event endpoint because participants can experience progression or death at different times and some participants may not experience the event during observation. The stratified log-rank test compares the treatment groups across the follow-up period while accounting for the trial's stratified analysis structure. This is more informative than simply comparing the proportion of participants who had progressed by one arbitrary date.

What does an HR of 0.51 mean?

An HR of 0.51 indicates a lower estimated hazard for treatment A relative to treatment C. Expressed as a relative difference, 1 − 0.51 = 0.49, or approximately a 49% lower estimated hazard. It is not a statement that 49% of patients were spared an event, and it is not an absolute risk reduction.

Why does the confidence interval matter?

A point estimate alone does not describe statistical uncertainty. The PFS HR is 0.51, while its 95% CI is 0.41–0.64. The interval shows how much uncertainty surrounds the estimate under the analysis framework. Confidence intervals also make it easier to judge the range of effects that remain statistically compatible with the data than a point estimate alone.

Why report both an odds ratio and a response-rate difference?

The OR and the response-rate difference describe the same broad binary outcome on different scales. The OR compares odds, whereas the difference expresses the contrast directly in percentage-point terms. Reporting both can make the result more interpretable because readers can see both a relative odds measure and an absolute-scale difference.

Why is an odds ratio of 3.52 not the same as a 252% higher response rate?

An odds ratio operates on odds rather than probabilities. When an event is common, odds and probability can differ substantially. Therefore, an OR of 3.52 cannot be converted into a percentage increase in response simply by subtracting 1 or multiplying by 100.

Why is the P-value not a measure of effect size?

The P-value reflects the compatibility of the observed data with the null hypothesis under the specified statistical model and testing procedure. It is influenced by both effect magnitude and information in the dataset. A small P-value can therefore coexist with a relatively modest effect, while a potentially meaningful effect can be estimated imprecisely in a small or information-poor dataset.

11. Confidence Intervals and the Different Effect Measures

EndpointEffect measureEstimateConfidence interval
PFSHazard ratio0.5195% CI 0.41–0.64
OSHazard ratio0.6098.89% CI 0.40–0.89
ORROdds ratio3.5295% CI 2.51–4.95
ORRDifference in objective response rates28.695% CI 21.7–35.6

These four estimates should not be placed on a single numerical ranking because they are measured on different statistical scales. The PFS and OS hazard ratios are time-to-event measures; the OR is an odds measure for a binary endpoint; and the response-rate difference is expressed directly as an adjusted difference.

Confidence-level detail: the OS analysis reports a 98.89% two-sided confidence interval rather than the 95% interval used for the other reported estimates. The confidence level should be preserved when interpreting the interval rather than silently converting it to 95%.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk. The available safety counts are reported separately from the efficacy analyses.

TreatmentSerious adverse eventsAt riskReported affected / at risk
Treatment A — Nivolumab + Cabozantinib160 affected320160/320
Treatment B28 affected5028/50
Treatment C — Sunitinib144 affected320144/320

These counts describe the serious-adverse-event data reported in the ClinicalTrials.gov record. They should not be substituted for overall adverse-event rates, grade-specific adverse-event rates, treatment discontinuations, or treatment-related mortality because those additional safety measures are not contained in the ClinicalTrials.gov record.

Safety interpretation: the serious-adverse-event counts are not an efficacy endpoint and should not be combined mathematically with the PFS, OS, or ORR estimates. Safety and efficacy address different dimensions of a randomized treatment comparison.

13. Randomization and the Analysis Population

The trial was randomized, and the primary PFS and secondary OS analyses were specified for all randomized participants of treatment A and treatment C. The ORR analyses were specified for all randomized participants.

Why randomization matters

Randomization establishes the treatment assignment mechanism before outcomes are observed. This provides the foundation for a causal comparison between the randomized groups, subject to the trial design and analysis assumptions.

Why the population is stated explicitly

An effect estimate is meaningful only when its analysis population is clear. The ClinicalTrials.gov record explicitly identifies the randomized population for each reported analysis.

The use of the randomized population is particularly important for interpreting the hazard ratios. The HR of 0.51 is not an estimate restricted to only participants who completed treatment or remained on treatment; it is reported for all randomized participants in treatment A and treatment C.

14. What the PFS Result Does — and Does Not — Mean

Statistical interpretation

The PFS hazard ratio of 0.51 indicates a lower estimated hazard of progression or death for nivolumab plus cabozantinib relative to sunitinib under the reported analysis.

It does not mean that the median PFS was reduced or increased by 49%, because no median PFS is reported in the ClinicalTrials.gov record. It also does not mean that exactly 49% of participants benefited.

Precision

The 95% CI of 0.41–0.64 communicates uncertainty around the estimated HR. It is narrower than an interval that would indicate very limited information, but it still contains a range of plausible treatment-effect estimates under the specified statistical framework.

Statistical significance

The P-value of <0.0001 indicates strong evidence against the null hypothesis used for the primary comparison. It should not be interpreted as the probability that the treatment effect is real, nor does it measure the magnitude of benefit.

Absolute versus relative effects

Relative hazard measures and absolute event probabilities answer different questions. The ClinicalTrials.gov record provides the HR and confidence interval but do not provide the underlying Kaplan-Meier estimates needed to reconstruct absolute PFS probabilities at specific time points.

15. Comparing the Primary and Secondary Evidence

EndpointStatistical questionReported result
PFS Does treatment A differ from treatment C in time to progression or death? HR 0.51; 95% CI 0.41–0.64; P < 0.0001
OS Does treatment A differ from treatment C in time to death? HR 0.60; 98.89% CI 0.40–0.89; P = 0.0010
ORR Are the odds of objective response different between treatments? OR 3.52; 95% CI 2.51–4.95; P < 0.0001
ORR difference What is the strata-adjusted difference in objective response rates? 28.6; 95% CI 21.7–35.6

The endpoints complement one another. PFS and OS are time-to-event outcomes, while ORR is binary. Consequently, the treatment effect cannot be summarized adequately by a single statistic. The HR describes relative event hazards, whereas the OR and response-rate difference describe response on different binary-outcome scales.

16. Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

CheckMate-9ER provides a compact example of several core clinical-trial statistical concepts. Its primary endpoint is time-to-event, the primary comparison uses a stratified log-rank test, and the treatment effect is expressed as a hazard ratio. The registry also demonstrates how the same randomized comparison can be examined through a binary response endpoint using both an odds ratio and an adjusted response-rate difference.

ConceptHow it appears in CheckMate-9ER
RandomizationRandomized phase 3 treatment trial
Parallel design3-arm parallel study
Time-to-event analysisPFS and OS
Stratified log-rank testPrimary PFS and secondary OS comparisons
Hazard ratioRelative treatment effect for PFS and OS
Confidence intervalUncertainty around HR, OR, and response-rate difference estimates
Cochran-Mantel-Haenszel testStratified analysis of objective response rate
Odds ratioBinary response effect measure
Absolute-scale differenceStrata-adjusted difference in objective response rates
Superiority testingReported hypothesis type for the primary PFS and selected secondary analyses
Analysis populationAll randomized participants for the reported efficacy comparisons

18. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

19. Related Calculators

The same statistical concepts can be explored quantitatively with Clinical Biostats calculators:

20. Sources

Continue learning through Clinical Biostats

Explore the statistical concepts behind randomized clinical trials, time-to-event endpoints, categorical outcomes, confidence intervals, and treatment-effect measures.

21. Record Summary

CheckMate-9ER is a useful clinical-trial statistics case because the registry provides both a primary time-to-event analysis and complementary binary-response analyses. The primary PFS comparison used a stratified log-rank test and reported a hazard ratio of 0.51 with a 95% CI of 0.41–0.64 and a P-value of <0.0001. Secondary analyses reported an OS HR of 0.60 with a 98.89% CI of 0.40–0.89 and P = 0.0010, together with an ORR odds ratio of 3.52 and a strata-adjusted response-rate difference of 28.6.

The statistical lesson is that these estimates should be interpreted according to their scales and endpoint structures. Hazard ratios describe relative event hazards, odds ratios describe relative odds for a binary endpoint, and a response-rate difference expresses an adjusted absolute-scale contrast. Confidence intervals communicate uncertainty, while P-values address evidence against a null hypothesis rather than effect magnitude. The registry data also make clear why the analysis population, stratification, endpoint definition, and time frame must be stated alongside every treatment-effect estimate.

Clinical Biostats methodology: A trial-results page should separate reported numerical evidence from statistical interpretation. Where the registry does not supply a quantity, this page does not reconstruct it from external publications or unstated assumptions.