This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data reported in the ClinicalTrials.gov record.
1. Trial at a Glance
CheckMate-9ER was a randomized, parallel, phase 3 treatment trial in renal cell carcinoma comparing treatment A, nivolumab combined with cabozantinib, with treatment C, sunitinib. The registry reports a total enrollment of 701 participants across 3 arms and a primary endpoint of progression-free survival.
| Feature | CheckMate-9ER |
|---|---|
| Phase | Phase 3 |
| Condition | Renal Cell Carcinoma |
| Brief title | A Study of Nivolumab Combined With Cabozantinib Compared to Sunitinib in Previously Untreated Advanced or Metastatic Renal Cell Carcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 701 |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
| Status | Completed |
| ClinicalTrials.gov | NCT03141177 |
2. Clinical Question
The registered clinical question was whether nivolumab combined with cabozantinib improves the primary time-to-event endpoint of progression-free survival compared with sunitinib in previously untreated advanced or metastatic renal cell carcinoma.
Population
Participants with previously untreated advanced or metastatic renal cell carcinoma.
Intervention
Nivolumab combined with cabozantinib. The registry identifies nivolumab as a biological intervention and cabozantinib as a drug.
Comparator
Sunitinib, identified in the registry as a drug intervention.
Primary question
Does treatment A, nivolumab plus cabozantinib, improve progression-free survival relative to treatment C, sunitinib?
3. Trial Design
The study enrolled 701 participants and used 3 arms. The statistical analyses in the ClinicalTrials.gov record compare treatment A with treatment C. The registry data identify treatment A in the objective-response analysis as the nivolumab-plus-cabozantinib group and treatment C as the sunitinib group.
Nivolumab + Cabozantinib
- Nivolumab
- Cabozantinib
- Primary efficacy comparison against treatment C
Sunitinib
- Sunitinib
- Comparator for the reported primary analysis
- Comparator for the reported secondary analyses
4. Trial Timeline and Registry Status
| Milestone | Date / status |
|---|---|
| Study start | 2017-07-23 |
| Primary completion | 2020-02-12 |
| Overall status | Completed |
| Results posted | Yes |
| Outcome measures posted | 9 |
| Statistical analyses posted | 4 |
The ClinicalTrials.gov record contains 4 statistical analyses, including 1 analysis of the registered primary endpoint and 3 secondary analyses. This provides a compact statistical picture: one formal primary time-to-event comparison and three additional efficacy comparisons.
5. Endpoints
| Endpoint | Registry time frame | Type | Analysis |
|---|---|---|---|
| Progression Free Survival (PFS) | From randomization date to date of first documented tumor progression or death, whichever occurs first (Up to 31 months) | Time-to-event | Stratified log-rank; hazard ratio |
| Overall Survival (OS) | From randomization date to death date (Up to 31 months) | Time-to-event | Stratified log-rank; hazard ratio |
| Objective Response Rate (ORR) | Up to 31 Months | Binary | Stratified Cochran-Mantel-Haenszel; odds ratio |
| Objective Response Rate (ORR) | Up to 31 Months | Binary | Strata-adjusted difference in objective response rate |
Registered primary endpoint: Progression Free Survival
PFS is defined as the time from date of randomization to the first documented tumor progression date or death due to any cause, whichever occurs first based on BICR assessment using RECIST v1.1. Participants who die without a reported progression will be considered to have progressed on the date of their death. Participants who did not progress or die will be censored on the date of their last evaluable tumor assessment on or prior to initiation of subsequent anti-cancer therapy. Progressive disease (PD); 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study
The registry classifies PFS as a time-to-event endpoint and reports the analysis population as all randomized participants of treatment A and treatment C.
6. Statistical Methodology
Stratified log-rank test
The primary PFS comparison used a stratified log-rank test. The test compares the observed pattern of event occurrence between randomized treatment groups while incorporating stratification. The registry does not identify the specific stratification factors in the ClinicalTrials.gov record, so no particular clinical variable is attributed as a stratification factor.
The stratified log-rank procedure evaluates the evidence for a treatment difference across follow-up rather than comparing only a single fixed-time proportion.
Hazard ratio
The primary PFS effect measure is the hazard ratio (HR). The reported estimate is 0.51. In the context of the analysis, an HR below 1 indicates a lower estimated hazard of progression or death for treatment A relative to treatment C.
An HR of 0.51 corresponds to an estimated hazard approximately 49% lower in treatment A than treatment C, because 1 − 0.51 = 0.49. This is a relative hazard interpretation, not an absolute probability difference.
Cochran-Mantel-Haenszel test
The reported ORR analysis used a stratified Cochran-Mantel-Haenszel test. This is appropriate for comparing a binary outcome between treatment groups while accounting for strata. Here, the outcome is objective response and the effect measure is an odds ratio.
Odds ratio
The reported ORR odds ratio is 3.52 for treatment A over treatment C. An odds ratio above 1 indicates greater odds of response in treatment A under the reported comparison. The odds ratio is not itself a risk ratio and should not be described as saying that the probability of response is 3.52 times as large.
Strata-adjusted response-rate difference
The registry also reports a difference in objective response rates of 28.6 percentage points, with the analysis note identifying this as the strata-adjusted difference in objective response rate, calculated as nivolumab plus cabozantinib minus sunitinib and based on DerSimonian and Laird.
7. Primary Result: Progression-Free Survival
The registered primary endpoint has a formal statistical analysis posted for the comparison of treatment A versus treatment C in all randomized participants in those two treatment groups.
Hazard ratio for progression or death
95% CI: 0.41–0.64 · P < 0.0001
Stratified log-rank test · Superiority analysis
| Feature | Primary PFS analysis |
|---|---|
| Endpoint | Progression Free Survival (PFS) |
| Time frame | From randomization date to date of first documented tumor progression or death, whichever occurs first (Up to 31 months) |
| Population | All Randomized Participants of treatment A and treatment C |
| Comparison | Treatment A vs Treatment C |
| Method | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.51 |
| 95% CI | 0.41–0.64 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The HR of 0.51 means that, under the time-to-event comparison represented by the reported hazard ratio, the estimated hazard of progression or death was approximately 49% lower for treatment A than treatment C. This is a relative comparison of hazards; it does not mean that 49% of patients avoided progression or death, nor that every patient experienced exactly a 49% reduction.
The 95% CI of 0.41–0.64 describes statistical uncertainty around the estimated hazard ratio. It indicates that the observed estimate is more precise than a very wide interval would be, while still recognizing uncertainty in the treatment effect. The interval does not describe the range of effects that individual patients experienced.
The P-value of <0.0001 describes the strength of evidence against the null hypothesis under the specified statistical framework. It does not measure the magnitude of the treatment effect, the probability that the treatment works, or the clinical importance of the result.
Because this is a hazard ratio, interpretation also depends on the time-to-event modeling framework and censoring process. A single HR summarizes a relative hazard comparison over the analyzed period; it is not interchangeable with an absolute risk reduction or a difference in median PFS. The ClinicalTrials.gov record does not report a median PFS estimate, so no median-time comparison is made on this page.
8. Secondary Result: Overall Survival
Overall survival was a secondary time-to-event endpoint. The registry defines the endpoint as the time from randomization date to death date, with a time frame of up to 31 months.
Hazard ratio for death
98.89% CI: 0.40–0.89 · P = 0.0010
Stratified log-rank test · Superiority analysis
| Feature | Secondary OS analysis |
|---|---|
| Endpoint | Overall Survival (OS) |
| Time frame | From randomization date to death date (Up to 31 months) |
| Population | All Randomized Participants of treatment A and treatment C |
| Comparison | Treatment A vs Treatment C |
| Method | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.60 |
| CI | 98.89%, two-sided: 0.40–0.89 |
| P-value | 0.0010 |
| Hypothesis | Superiority |
The OS HR of 0.60 corresponds to an estimated hazard of death approximately 40% lower for treatment A relative to treatment C under the reported analysis. It does not mean that treatment A reduced the probability of death by exactly 40% for each participant, nor does it imply that 40% of participants benefited.
The 98.89% CI of 0.40–0.89 expresses the uncertainty around the estimated hazard ratio using the reported confidence level. Because the interval is entirely below 1, its endpoints are consistent with a lower estimated hazard for treatment A under the specified analysis.
The P-value of 0.0010 quantifies the statistical evidence against the relevant null hypothesis under the reported testing framework. It should not be interpreted as an effect-size measure or as the probability that the null hypothesis is true.
The OS analysis is still a time-to-event comparison. Censoring, the timing of deaths, and the relationship between hazards over time all matter. The ClinicalTrials.gov record does not provide median OS, survival probabilities at specific time points, or a Kaplan-Meier dataset, so those quantities are not reconstructed.
9. Secondary Result: Objective Response Rate
Objective response rate was evaluated as a binary endpoint up to 31 months. The registry reports two statistical representations of the same treatment comparison: a stratified Cochran-Mantel-Haenszel odds ratio and a strata-adjusted difference in response rates.
Odds-ratio analysis
Odds ratio for objective response
95% CI: 2.51–4.95 · P < 0.0001
Stratified Cochran-Mantel-Haenszel test · Superiority analysis
| Feature | ORR odds-ratio analysis |
|---|---|
| Endpoint | Objective Response Rate (ORR) |
| Time frame | Up to 31 Months |
| Population | All Randomized Participants |
| Comparison | Treatment A vs Treatment C |
| Method | Stratified Cochran-Mantel-Haenszel |
| Effect measure | Odds ratio |
| Estimate | 3.52 |
| 95% CI | 2.51–4.95 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
An OR of 3.52 means that the estimated odds of objective response were 3.52 times as high for treatment A as for treatment C under the reported stratified comparison. Odds are not probabilities, so the OR should not be read as saying that the response rate was 3.52 times higher.
The 95% CI of 2.51–4.95 gives the uncertainty range reported for the odds ratio. It does not represent a range of individual patient outcomes.
The P-value of <0.0001 indicates strong statistical evidence against the null hypothesis used for the reported test. It does not indicate the size or clinical importance of the response difference by itself.
Strata-adjusted response-rate difference
Difference in objective response rates
95% CI: 21.7–35.6
Treatment A − Treatment C · Strata-adjusted difference
| Feature | Response-rate difference analysis |
|---|---|
| Endpoint | Objective Response Rate (ORR) |
| Time frame | Up to 31 Months |
| Population | All Randomized Participants |
| Comparison | Treatment A vs Treatment C |
| Effect measure | Difference of Objective Response Rates |
| Estimate | 28.6 |
| 95% CI | 21.7–35.6 |
| Analysis note | Strata adjusted difference in objective response rate (Nivolumab+Cabozantinib - Sunitinib) based on DerSimonian and Laird. |
| Hypothesis | Other / not stated |
The reported difference of 28.6 represents the strata-adjusted difference in objective response rate calculated as treatment A minus treatment C. A positive value therefore favors treatment A for this response endpoint under the direction specified in the registry analysis.
The 95% CI of 21.7–35.6 quantifies uncertainty around that adjusted difference. Unlike an odds ratio, a response-rate difference is expressed directly on the percentage-point scale.
This estimate should not be confused with the HR of 0.51 for PFS or the OR of 3.52 for ORR. These measures answer different statistical questions and use different scales.
10. Statistical Methods Explained
Why was a stratified log-rank test used for PFS?
PFS is a time-to-event endpoint because participants can experience progression or death at different times and some participants may not experience the event during observation. The stratified log-rank test compares the treatment groups across the follow-up period while accounting for the trial's stratified analysis structure. This is more informative than simply comparing the proportion of participants who had progressed by one arbitrary date.
What does an HR of 0.51 mean?
An HR of 0.51 indicates a lower estimated hazard for treatment A relative to treatment C. Expressed as a relative difference, 1 − 0.51 = 0.49, or approximately a 49% lower estimated hazard. It is not a statement that 49% of patients were spared an event, and it is not an absolute risk reduction.
Why does the confidence interval matter?
A point estimate alone does not describe statistical uncertainty. The PFS HR is 0.51, while its 95% CI is 0.41–0.64. The interval shows how much uncertainty surrounds the estimate under the analysis framework. Confidence intervals also make it easier to judge the range of effects that remain statistically compatible with the data than a point estimate alone.
Why report both an odds ratio and a response-rate difference?
The OR and the response-rate difference describe the same broad binary outcome on different scales. The OR compares odds, whereas the difference expresses the contrast directly in percentage-point terms. Reporting both can make the result more interpretable because readers can see both a relative odds measure and an absolute-scale difference.
Why is an odds ratio of 3.52 not the same as a 252% higher response rate?
An odds ratio operates on odds rather than probabilities. When an event is common, odds and probability can differ substantially. Therefore, an OR of 3.52 cannot be converted into a percentage increase in response simply by subtracting 1 or multiplying by 100.
Why is the P-value not a measure of effect size?
The P-value reflects the compatibility of the observed data with the null hypothesis under the specified statistical model and testing procedure. It is influenced by both effect magnitude and information in the dataset. A small P-value can therefore coexist with a relatively modest effect, while a potentially meaningful effect can be estimated imprecisely in a small or information-poor dataset.
11. Confidence Intervals and the Different Effect Measures
| Endpoint | Effect measure | Estimate | Confidence interval |
|---|---|---|---|
| PFS | Hazard ratio | 0.51 | 95% CI 0.41–0.64 |
| OS | Hazard ratio | 0.60 | 98.89% CI 0.40–0.89 |
| ORR | Odds ratio | 3.52 | 95% CI 2.51–4.95 |
| ORR | Difference in objective response rates | 28.6 | 95% CI 21.7–35.6 |
These four estimates should not be placed on a single numerical ranking because they are measured on different statistical scales. The PFS and OS hazard ratios are time-to-event measures; the OR is an odds measure for a binary endpoint; and the response-rate difference is expressed directly as an adjusted difference.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk. The available safety counts are reported separately from the efficacy analyses.
| Treatment | Serious adverse events | At risk | Reported affected / at risk |
|---|---|---|---|
| Treatment A — Nivolumab + Cabozantinib | 160 affected | 320 | 160/320 |
| Treatment B | 28 affected | 50 | 28/50 |
| Treatment C — Sunitinib | 144 affected | 320 | 144/320 |
These counts describe the serious-adverse-event data reported in the ClinicalTrials.gov record. They should not be substituted for overall adverse-event rates, grade-specific adverse-event rates, treatment discontinuations, or treatment-related mortality because those additional safety measures are not contained in the ClinicalTrials.gov record.
13. Randomization and the Analysis Population
The trial was randomized, and the primary PFS and secondary OS analyses were specified for all randomized participants of treatment A and treatment C. The ORR analyses were specified for all randomized participants.
Why randomization matters
Randomization establishes the treatment assignment mechanism before outcomes are observed. This provides the foundation for a causal comparison between the randomized groups, subject to the trial design and analysis assumptions.
Why the population is stated explicitly
An effect estimate is meaningful only when its analysis population is clear. The ClinicalTrials.gov record explicitly identifies the randomized population for each reported analysis.
The use of the randomized population is particularly important for interpreting the hazard ratios. The HR of 0.51 is not an estimate restricted to only participants who completed treatment or remained on treatment; it is reported for all randomized participants in treatment A and treatment C.
14. What the PFS Result Does — and Does Not — Mean
The PFS hazard ratio of 0.51 indicates a lower estimated hazard of progression or death for nivolumab plus cabozantinib relative to sunitinib under the reported analysis.
It does not mean that the median PFS was reduced or increased by 49%, because no median PFS is reported in the ClinicalTrials.gov record. It also does not mean that exactly 49% of participants benefited.
The 95% CI of 0.41–0.64 communicates uncertainty around the estimated HR. It is narrower than an interval that would indicate very limited information, but it still contains a range of plausible treatment-effect estimates under the specified statistical framework.
The P-value of <0.0001 indicates strong evidence against the null hypothesis used for the primary comparison. It should not be interpreted as the probability that the treatment effect is real, nor does it measure the magnitude of benefit.
Relative hazard measures and absolute event probabilities answer different questions. The ClinicalTrials.gov record provides the HR and confidence interval but do not provide the underlying Kaplan-Meier estimates needed to reconstruct absolute PFS probabilities at specific time points.
15. Comparing the Primary and Secondary Evidence
| Endpoint | Statistical question | Reported result |
|---|---|---|
| PFS | Does treatment A differ from treatment C in time to progression or death? | HR 0.51; 95% CI 0.41–0.64; P < 0.0001 |
| OS | Does treatment A differ from treatment C in time to death? | HR 0.60; 98.89% CI 0.40–0.89; P = 0.0010 |
| ORR | Are the odds of objective response different between treatments? | OR 3.52; 95% CI 2.51–4.95; P < 0.0001 |
| ORR difference | What is the strata-adjusted difference in objective response rates? | 28.6; 95% CI 21.7–35.6 |
The endpoints complement one another. PFS and OS are time-to-event outcomes, while ORR is binary. Consequently, the treatment effect cannot be summarized adequately by a single statistic. The HR describes relative event hazards, whereas the OR and response-rate difference describe response on different binary-outcome scales.
16. Limitations and Interpretation Issues
- Limited registry result detail: the ClinicalTrials.gov record contains effect estimates and confidence intervals but do not contain median PFS, median OS, Kaplan-Meier estimates, or event counts for the efficacy endpoints.
- Three-arm design: the study has 3 arms, while the registry-reported formal statistical analyses compare treatment A with treatment C. The additional arm is therefore not incorporated into the reported efficacy comparisons on this page.
- Specific stratification factors are not reported: the methods identify stratified log-rank and stratified Cochran-Mantel-Haenszel analyses, but the ClinicalTrials.gov record does not identify the individual stratification variables.
- Hazard-ratio interpretation: an HR is a relative time-to-event measure, not an absolute risk difference. Its interpretation depends on the underlying survival-analysis framework and censoring.
- Proportional-hazards consideration: a single hazard ratio is most straightforward to interpret when the relative hazard is reasonably stable over time. The ClinicalTrials.gov record does not provide diagnostics for this assumption.
- No reconstructed Kaplan-Meier curves: a valid Kaplan-Meier curve requires event and censoring information or sufficiently detailed underlying data. The registry summary does not provide those data.
- Multiplicity information is limited: the ClinicalTrials.gov record identifies a superiority hypothesis for the primary PFS analysis and for the reported OS and ORR odds-ratio analyses, but they do not provide an alpha-allocation or multiplicity-adjustment scheme. No such scheme is inferred here.
- Interim-analysis information is not reported: the ClinicalTrials.gov record do not describe an interim-analysis schedule or alpha-spending procedure, so none is assumed.
- Crossover information is not reported: the ClinicalTrials.gov record contains no crossover analysis, so no crossover effect on OS is inferred.
- Missing-data and imputation information is not reported: the registry summary does not specify an imputation strategy for the reported efficacy analyses.
- Safety scope: only serious adverse-event counts by arm are reported. Broader safety conclusions would require additional safety measures.
17. Why This Trial Matters Statistically
CheckMate-9ER provides a compact example of several core clinical-trial statistical concepts. Its primary endpoint is time-to-event, the primary comparison uses a stratified log-rank test, and the treatment effect is expressed as a hazard ratio. The registry also demonstrates how the same randomized comparison can be examined through a binary response endpoint using both an odds ratio and an adjusted response-rate difference.
| Concept | How it appears in CheckMate-9ER |
|---|---|
| Randomization | Randomized phase 3 treatment trial |
| Parallel design | 3-arm parallel study |
| Time-to-event analysis | PFS and OS |
| Stratified log-rank test | Primary PFS and secondary OS comparisons |
| Hazard ratio | Relative treatment effect for PFS and OS |
| Confidence interval | Uncertainty around HR, OR, and response-rate difference estimates |
| Cochran-Mantel-Haenszel test | Stratified analysis of objective response rate |
| Odds ratio | Binary response effect measure |
| Absolute-scale difference | Strata-adjusted difference in objective response rates |
| Superiority testing | Reported hypothesis type for the primary PFS and selected secondary analyses |
| Analysis population | All randomized participants for the reported efficacy comparisons |
18. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
19. Related Calculators
The same statistical concepts can be explored quantitatively with Clinical Biostats calculators:
20. Sources
- ClinicalTrials.gov: NCT03141177 — CheckMate-9ER.
- PubMed record: PMID 42605077.
- PubMed record: PMID 41617395.
- PubMed record: PMID 41239083.
- PubMed record: PMID 40998092.
- PubMed record: PMID 39452894.
Continue learning through Clinical Biostats
Explore the statistical concepts behind randomized clinical trials, time-to-event endpoints, categorical outcomes, confidence intervals, and treatment-effect measures.
21. Record Summary
CheckMate-9ER is a useful clinical-trial statistics case because the registry provides both a primary time-to-event analysis and complementary binary-response analyses. The primary PFS comparison used a stratified log-rank test and reported a hazard ratio of 0.51 with a 95% CI of 0.41–0.64 and a P-value of <0.0001. Secondary analyses reported an OS HR of 0.60 with a 98.89% CI of 0.40–0.89 and P = 0.0010, together with an ORR odds ratio of 3.52 and a strata-adjusted response-rate difference of 28.6.
The statistical lesson is that these estimates should be interpreted according to their scales and endpoint structures. Hazard ratios describe relative event hazards, odds ratios describe relative odds for a binary endpoint, and a response-rate difference expresses an adjusted absolute-scale contrast. Confidence intervals communicate uncertainty, while P-values address evidence against a null hypothesis rather than effect magnitude. The registry data also make clear why the analysis population, stratification, endpoint definition, and time frame must be stated alongside every treatment-effect estimate.