This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record for NCT03138512. The official registry record is the authoritative source for the trial's registration details.
1. Trial at a Glance
CheckMate-914 was a randomized, parallel-group, quadruple-masked phase 3 trial in participants with renal cell carcinoma. The registry describes a study comparing nivolumab, nivolumab in combination with ipilimumab, and placebo in participants with localized kidney cancer who underwent surgery to remove part of a kidney.
| Feature | CheckMate-914 |
|---|---|
| Phase | Phase 3 |
| Condition | Carcinoma, Renal Cell |
| Population | Participants with localized kidney cancer who underwent surgery to remove part of a kidney |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 1,641 |
| Arms | 5 |
| Primary endpoint type | Time-to-event |
| Primary endpoints registered | 1 |
| Results posted | Yes |
| Outcome measures posted | 13 |
| Statistical analyses posted | 6 |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
2. Clinical Question
The registered study evaluates treatment strategies involving nivolumab and ipilimumab against placebo-based comparison groups in participants with localized kidney cancer after surgery. Statistically, the central confirmatory question is whether the randomized treatment comparisons produce a different disease-free survival experience, where disease-free survival is defined by the registry using local recurrence, distant metastasis, or death as the event.
Population
Participants with localized kidney cancer who underwent surgery to remove part of a kidney, with the registry condition listed as carcinoma, renal cell.
Interventions
The registry lists nivolumab, ipilimumab, nivolumab placebo, and ipilimumab placebo among the interventions.
Comparator structure
The posted primary analyses compare Treatment Part A Nivo + Ipi with Treatment Part A Placebo, and Treatment Part A Placebo with Treatment Part B Nivo.
Primary question
Does either prespecified randomized comparison change disease-free survival as assessed by blinded independent central review?
3. Trial Design
Nivolumab + ipilimumab
- The registry identifies Nivo + Ipi as a Treatment Part A arm.
- Serious adverse events: 196 affected among 608 at risk in the combined Treatment Part A and B safety reporting.
Placebo
- The placebo group is the comparator for the Treatment Part A Nivo + Ipi primary analysis.
- It is also the comparator against Treatment Part B Nivo in the second primary analysis.
- Serious adverse events: 49 affected among 614 at risk in Treatment Part A and B reporting.
Nivolumab
- The Treatment Part B Nivo group is compared with the Treatment Part A placebo group for the second primary DFS analysis.
- Serious adverse events: 56 affected among 408 at risk in Treatment Part B reporting.
Five registered arms
- The registry reports 5 arms overall.
- The statistical analyses posted on ClinicalTrials.gov do not provide separate treatment descriptions for every registered arm.
- The analysis page therefore does not infer additional arm-level details.
4. Trial Timeline
Trial start
The registered study start date was July 7, 2017.
Randomized parallel design
The trial was conducted as a randomized phase 3 study with quadruple masking and treatment as its primary purpose.
Primary completion
The registered primary completion date was September 28, 2023.
Results posted
The registry status is completed, with 13 outcome measures and 6 statistical analyses posted.
5. Primary Endpoint
The registered primary endpoint is Disease-Free Survival (DFS) by BICR - Treatment Part A and B. It is a time-to-event endpoint.
| Endpoint | Registry definition | Analysis framework |
|---|---|---|
| Disease-Free Survival (DFS) by BICR - Treatment Part A and B | Disease-Free Survival (DFS) is defined as the time from randomization to development of local disease recurrence (ie, recurrence of primary tumor in situ or occurrence of a secondary renal cell carcinoma (RCC) primary cancer), distance metastasis, or death, whichever came first per Blinded Independent Central Review (BICR) based on Kaplan-Meier estimates. | Time-to-event; Kaplan-Meier estimates; posted formal analyses use the log-rank test and Cox proportional-hazard effect measure. |
6. Statistical Methodology
Kaplan-Meier estimation
The registry identifies Kaplan-Meier estimates as part of the definition of the primary DFS endpoint. Kaplan-Meier estimation is appropriate for a time-to-event endpoint because participants can have different lengths of observable follow-up. A participant who has not experienced the defined DFS event by the last usable observation can contribute information up to that point without being treated as if an event occurred.
Here, di denotes the number of events at event time ti, while ni is the number at risk immediately before that time. The resulting curve estimates the probability of remaining event-free through time t.
Log-rank test
The two posted primary analyses use the log-rank test. The log-rank test compares the observed and expected pattern of events between randomized groups over follow-up rather than comparing only a single time point.
This is particularly appropriate for DFS because the endpoint is explicitly defined as a time from randomization to the first qualifying event. The test therefore uses the ordering of event times and the risk sets formed during follow-up.
Cox proportional-hazards effect measure
The registry reports Cox Proportional Hazard as the effect measure for the primary analyses and gives the corresponding hazard ratio estimates with 95% confidence intervals.
A hazard ratio is a relative, model-based measure of event hazard. It is not a probability, not a median survival time, and not the proportion of participants who benefit. Its interpretation also depends on the proportional-hazards framework underlying the Cox model.
Superiority hypothesis
Both primary statistical analyses are classified in the registry as superiority analyses. That means the statistical question is whether the randomized groups differ in the direction represented by the model and test, rather than whether one treatment is merely no worse than another within a prespecified non-inferiority margin.
Analysis populations
The primary analyses use all randomized participants in Treatment Part A and Treatment Part B according to the registry description. The registry-reported analysis population text also specifies the prespecified exclusion of the Nivo + Ipi arm from the endpoint objective relevant to the posted comparisons.
Analysis of secondary time-to-event endpoints
The four posted secondary analyses also use hazard ratios for time-to-event outcomes. Two of those analyses have a reported log-rank method, while the Treatment Part B comparisons of Nivo + Ipi versus Nivo have the method listed as not reported. The registry therefore supports reporting the latter hazard ratios and confidence intervals, but not attributing an unreported formal test or p-value to those comparisons.
7. Primary Results: Disease-Free Survival
Nivolumab + ipilimumab vs placebo
DFS hazard ratio
95% CI: 0.75–1.20 · P = 0.6676
Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.
| Primary DFS analysis | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Treatment Part A: Nivo + Ipi vs Treatment Part A: Placebo | HR 0.95 | 0.75–1.20 | 0.6676 | Log-rank; Cox proportional hazard |
The estimated hazard ratio of 0.95 means that the fitted relative hazard of a qualifying DFS event was estimated to be 0.95 for Nivo + Ipi relative to placebo in this comparison. Expressed descriptively, that estimate is close to 1, so the point estimate represents only a modest relative difference in the estimated event hazard.
The estimate does not mean that 95% of participants remained disease-free, that 5% of participants benefited, or that the treatment changed individual risk by exactly 5%. A hazard ratio summarizes the relative event rate under the statistical model; it does not directly provide an absolute treatment effect.
The 95% confidence interval of 0.75–1.20 describes the statistical uncertainty around the estimated hazard ratio. Because the interval spans 1, the reported data are compatible with both a lower and a higher relative hazard under the model. The interval is therefore important alongside the point estimate rather than being treated as an afterthought.
The P-value of 0.6676 addresses the statistical evidence against the null comparison under the reported testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the clinical importance of the observed estimate. Effect size and uncertainty are better represented by the hazard ratio and its confidence interval.
Finally, the Cox interpretation depends on the proportional-hazards framework. A single hazard ratio is most naturally interpreted as a relative hazard measure under that model; it should not automatically be translated into a constant difference in individual risk throughout follow-up.
Placebo vs nivolumab
DFS hazard ratio
95% CI: 0.67–1.28 · P = 0.6556
Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.
| Primary DFS analysis | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Treatment Part A: Placebo vs Treatment Part B: Nivo | HR 0.93 | 0.67–1.28 | 0.6556 | Log-rank; Cox proportional hazard |
The reported hazard ratio of 0.93 is the estimated relative hazard for the comparison as recorded by the registry: Treatment Part A placebo versus Treatment Part B nivolumab. Because the registry presents the comparison in that order, the direction of the numerical ratio should be read in that same order rather than silently reversing it.
A hazard ratio of 0.93 does not mean that 93% of participants experienced an event or that nivolumab reduces individual risk by 7%. It is a model-based comparison of event hazards between the two randomized groups.
The 95% confidence interval of 0.67–1.28 is relatively broad and includes 1. This means the point estimate alone should not be treated as a precise description of the underlying treatment effect. The interval communicates that appreciable uncertainty remains around the estimated relative hazard.
The reported P-value of 0.6556 is evidence from the stated statistical test, not an effect-size measure. A large P-value does not prove that two treatments are identical, just as a small P-value would not by itself establish that an effect is clinically important.
As with the other Cox analysis, interpretation should remain tied to the proportional-hazards model and to the time-to-event nature of DFS. The registry does not provide median DFS or time-specific DFS estimates in the ClinicalTrials.gov record, so those quantities are not substituted for the reported hazard ratio.
8. Secondary Endpoint Results
Overall Survival — Treatment Part A: Nivo + Ipi vs Placebo
OS hazard ratio
95% CI: 0.85–1.85 · P = 0.2436
Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.
| Secondary endpoint | Comparison | HR | 95% CI | P-value |
|---|---|---|---|---|
| Overall Survival (OS) - Treatment Part A and B | Treatment Part A: Nivo + Ipi vs Treatment Part A: Placebo | 1.26 | 0.85–1.85 | 0.2436 |
The registry defines this OS endpoint as time from randomization to the date of death, with a time frame of up to approximately 72 months. The analysis population is all randomized participants in Treatment Part A and Treatment Part B, with the Nivo + Ipi arm prespecified to be excluded from the endpoint objective where applicable.
Overall Survival — Placebo vs Nivolumab
OS hazard ratio
95% CI: 0.61–3.07 · P = 0.4500
Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.
| Secondary endpoint | Comparison | HR | 95% CI | P-value |
|---|---|---|---|---|
| Overall Survival (OS) - Treatment Part A and B | Treatment Part A: Placebo vs Treatment Part B: Nivo | 1.36 | 0.61–3.07 | 0.4500 |
The registry identifies this as the Treatment Part B analysis and defines OS as time from randomization to the date of death, with a time frame of up to approximately 72 months. The confidence interval is wide, extending from below 1 to above 3, so the point estimate should not be interpreted without its uncertainty interval.
Disease-Free Survival — Treatment Part B Nivo + Ipi vs Nivo
DFS hazard ratio
95% CI: 0.89–1.67
Two-sided 95% confidence interval; Cox proportional-hazard effect measure.
This secondary endpoint was prespecified to be collected for Treatment Part B only. The analysis population consists of all randomized Nivo + Ipi and Nivo participants in Treatment Part B, with the placebo arm prespecified to be excluded from the endpoint objective.
Overall Survival — Treatment Part B Nivo + Ipi vs Nivo
OS hazard ratio
95% CI: 0.33–1.68
Two-sided 95% confidence interval; Cox proportional-hazard effect measure.
This secondary Treatment Part B endpoint evaluates overall survival among contemporaneously randomized Nivo + Ipi and Nivo participants. The registry again lists the statistical method as not reported, so the reported hazard ratio and confidence interval are presented without adding an unreported hypothesis test.
| Secondary endpoint | Comparison | HR | 95% CI | Formal method reported |
|---|---|---|---|---|
| DFS in contemporaneously randomized combination and monotherapy participants | Treatment Part B: Nivo + Ipi vs Treatment Part B: Nivo | 1.22 | 0.89–1.67 | Not reported |
| OS in contemporaneously randomized combination and monotherapy participants | Treatment Part B: Nivo + Ipi vs Treatment Part B: Nivo | 0.75 | 0.33–1.68 | Not reported |
9. How to Read the CheckMate-914 Hazard Ratios
The direction of a hazard ratio depends on the order in which the groups are specified. For example, the primary comparison is recorded as Nivo + Ipi vs Placebo, giving HR 0.95. The second primary comparison is recorded as Placebo vs Nivo, giving HR 0.93. These numbers should not be casually compared as if both were expressed in the same treatment-versus-control direction.
A hazard ratio compares event hazards over time. It is not the ratio of the cumulative percentages of participants who experience an event. Converting an HR directly into an absolute probability or percentage of patients affected would require additional information that is not reported in the ClinicalTrials.gov record.
The confidence interval shows how much uncertainty surrounds the point estimate under the statistical model. For the primary Nivo + Ipi versus placebo DFS comparison, the interval is 0.75–1.20. For placebo versus Nivo, it is 0.67–1.28. Both intervals include 1, so the point estimates should not be interpreted as precise evidence of a directional difference by themselves.
The primary P-values, 0.6676 and 0.6556, quantify evidence under the respective statistical testing framework. They do not tell us how large a treatment effect is. The hazard ratio and confidence interval address effect magnitude and uncertainty; the P-value addresses compatibility with the null hypothesis under the specified test.
10. Statistical Methods Explained
Why was a time-to-event method used for DFS?
DFS is explicitly defined as the time from randomization until the first of several qualifying events: local disease recurrence, distant metastasis, or death. Because both the timing of events and incomplete event observation matter, a time-to-event framework preserves more information than a simple yes/no comparison at one arbitrary time point.
Why use Kaplan-Meier estimation?
Kaplan-Meier estimation provides an estimate of the event-free survival function over time. It is designed to incorporate participants whose event status is not observed for the entire follow-up period by accounting for their available follow-up through censoring.
What does the log-rank test contribute?
The log-rank test compares the survival experience of randomized groups across observed event times. Rather than asking only whether two percentages differ at a single time point, it evaluates the accumulated pattern of events while participants are at risk.
What does an HR of 0.95 mean?
For the Treatment Part A Nivo + Ipi versus placebo primary DFS comparison, an HR of 0.95 means that the estimated event hazard for the first group relative to the second was 0.95 under the Cox model. It does not mean that 95% of patients were disease-free or that every participant had a 5% lower risk.
Why does the confidence interval matter as much as the point estimate?
A point estimate is only one representation of the observed evidence. The 95% confidence interval provides information about statistical precision. In CheckMate-914, the primary DFS intervals extend across 1, so the reported point estimates should be considered together with substantial uncertainty about the underlying relative hazard.
Why does the direction of the comparison matter?
Hazard ratios are reciprocal when the same two groups are compared in the opposite order, subject to the corresponding model specification. Therefore, “Nivo + Ipi vs Placebo” and “Placebo vs Nivo” are not interchangeable labels. The comparison direction should always be read directly from the reported analysis.
Why should an unreported P-value not be invented?
The two Treatment Part B combination-versus-monotherapy secondary analyses report Cox proportional-hazard effect measures but list the statistical method as not reported. A statistically responsible analysis preserves that distinction. A hazard ratio and confidence interval can be discussed without fabricating a test statistic or P-value that the registry does not provide.
11. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk. These are the available arm-level safety results in the ClinicalTrials.gov record.
| Safety group | Serious adverse events | Affected / at risk |
|---|---|---|
| Treatment Part A and B: Nivo + Ipi | Serious adverse events | 196 / 608 |
| Treatment Part A and B: Placebo | Serious adverse events | 49 / 614 |
| Treatment Part B: Nivo | Serious adverse events | 56 / 408 |
The denominators are important because the reported safety information is expressed as affected participants relative to those at risk. The ClinicalTrials.gov record does not provide additional categories such as grade-specific adverse events, discontinuations because of adverse events, or individual adverse-event terms, so those measures are not added here.
12. Analysis Population and Endpoint Structure
The registry's analysis descriptions reveal an important feature of CheckMate-914: the trial contains multiple treatment parts and prespecified comparisons, while the primary endpoint itself is registered once as a DFS endpoint covering Treatment Part A and B.
| Analysis | Population described in registry | Comparison | Role |
|---|---|---|---|
| Primary DFS | All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective. | Nivo + Ipi vs Placebo in Treatment Part A | Primary |
| Primary DFS | All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective. | Placebo in Treatment Part A vs Nivo in Treatment Part B | Primary |
| Secondary OS | All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective where applicable | Nivo + Ipi vs Placebo in Treatment Part A | Secondary |
| Secondary OS | All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective where applicable | Placebo in Treatment Part A vs Nivo in Treatment Part B | Secondary |
| Secondary DFS | All randomized Nivo + Ipi and Nivo participants in Treatment Part B | Nivo + Ipi vs Nivo in Treatment Part B | Secondary |
| Secondary OS | All randomized Nivo + Ipi and Nivo participants in Treatment Part B | Nivo + Ipi vs Nivo in Treatment Part B | Secondary |
This structure is statistically important because a trial can contain several randomized comparisons without every comparison being the same confirmatory question. The analysis population, treatment part, comparison direction, endpoint role, and prespecified exclusion rules should therefore be retained whenever an estimate is interpreted.
13. What Is and Is Not Reported in the Supplied Registry Data
Reported
The registry provides the primary DFS endpoint, two primary statistical analyses, four secondary statistical analyses, hazard ratios, confidence intervals, P-values for four analyses, and serious adverse-event counts by arm.
Not reported here
The ClinicalTrials.gov record does not provide median DFS, median OS, time-specific survival probabilities, baseline characteristics, subgroup estimates, or reconstructed Kaplan-Meier event counts.
Method reported
The primary analyses use a log-rank test and Cox proportional-hazard effect measure. The two Treatment Part B combination-versus-monotherapy secondary analyses list the method as not reported.
Interpretive consequence
The page emphasizes relative time-to-event effects and uncertainty rather than filling gaps with values from external publications or inferred calculations.
14. Limitations
- Limited numerical detail: the ClinicalTrials.gov record contains hazard ratios and confidence intervals but do not provide median DFS or OS or time-specific survival estimates. Those quantities cannot be recovered reliably from the reported HRs alone.
- Comparison direction: the posted analyses use different group orderings. The numerical hazard ratios should therefore be interpreted using the exact comparison labels rather than assuming every estimate has the same treatment-versus-control direction.
- Confidence-interval width: some secondary estimates have wide confidence intervals, particularly the Treatment Part B OS comparison with a 95% CI of 0.33–1.68. Point estimates should therefore not be interpreted without their precision.
- Unreported methods for two secondary analyses: the registry lists the statistical method as not reported for the Treatment Part B Nivo + Ipi versus Nivo DFS and OS analyses. Formal hypothesis testing should not be invented.
- Proportional-hazards interpretation: the Cox hazard ratio is model-based. A single HR should not automatically be interpreted as a constant difference in individual risk at every time point.
- Endpoint complexity: DFS combines local recurrence, distant metastasis, and death into a single time-to-first-event endpoint. The HR therefore summarizes the composite endpoint rather than separately quantifying each event type.
- Registry-level reporting: the ClinicalTrials.gov record contains fewer numerical details than would normally be available in a complete statistical analysis plan or full clinical study report. This page does not substitute external values for missing registry information.
15. Why This Trial Matters Statistically
CheckMate-914 is a useful teaching case because it demonstrates how a randomized oncology trial can combine a composite time-to-event endpoint, multiple treatment parts, multiple randomized comparisons, blinded central assessment, and Cox-model effect estimation. The statistical story is therefore more nuanced than simply asking whether a P-value is below a threshold.
| Concept | How it appears in CheckMate-914 |
|---|---|
| Randomization | The study uses randomized allocation in a phase 3 parallel design. |
| Blinding | The registered masking level is quadruple. |
| Time-to-event analysis | The primary endpoint is DFS, defined from randomization to the first qualifying recurrence, metastasis, or death. |
| Kaplan-Meier estimation | The registry definition identifies DFS as being based on Kaplan-Meier estimates. |
| Log-rank testing | The two primary analyses use a log-rank test. |
| Hazard ratio | The primary and secondary reported analyses use a Cox proportional-hazard effect measure. |
| Confidence intervals | Primary analyses report 95% confidence intervals of 0.75–1.20 and 0.67–1.28. |
| Superiority testing | The primary analyses are classified as superiority hypotheses. |
| Analysis population | The registry explicitly describes randomized populations and prespecified exclusions from particular endpoint objectives. |
| Method transparency | Two Treatment Part B secondary analyses report hazard ratios and confidence intervals but list the statistical method as not reported. |
16. Interpreting the Primary Evidence as a Statistical Story
The two primary DFS analyses produce estimates of 0.95 and 0.93, respectively. Both are close to 1, and both corresponding 95% confidence intervals include 1. The associated P-values are 0.6676 and 0.6556. Taken together, the ClinicalTrials.gov record does not provide a basis for describing either primary comparison as a statistically demonstrated difference in DFS under its reported testing framework.
That conclusion is deliberately narrower than saying that the treatments are proven identical. A non-significant superiority test does not establish equivalence or non-inferiority. It means that the observed data, analyzed using the reported test, do not provide the statistical evidence required to reject the null hypothesis at the conventional interpretation of that test.
The confidence intervals make the distinction clearer. For the Nivo + Ipi versus placebo comparison, the interval ranges from 0.75 to 1.20. For placebo versus Nivo, it ranges from 0.67 to 1.28. Both permit a range of underlying relative hazards, rather than identifying one exact treatment effect.
The secondary results illustrate the same principle. The OS hazard ratios are 1.26 and 1.36 for the two Treatment Part A and B comparisons, while the Treatment Part B Nivo + Ipi versus Nivo analyses give DFS HR 1.22 and OS HR 0.75. These estimates should be read with their respective confidence intervals and with attention to whether a formal statistical method was actually reported.
This is a useful distinction between estimating an effect and testing a hypothesis. A hazard ratio answers “what relative effect did the fitted model estimate?” A confidence interval answers “how precisely is that effect estimated?” A P-value answers a different question about compatibility with a null hypothesis under the specified test.
17. Primary Endpoint: Why the Composite Definition Matters
DFS in this registry record is not simply “time until recurrence.” It is defined as the time from randomization to the first of local disease recurrence, distant metastasis, or death. That construction has direct statistical consequences: whichever qualifying event occurs first determines the endpoint event time.
One time origin
The endpoint begins at randomization. This creates a common starting point for the randomized comparison.
Multiple event types
Local recurrence, distant metastasis, and death are all qualifying events for DFS.
First event governs
The definition uses whichever qualifying event occurs first, making DFS a time-to-first-event composite endpoint.
BICR assessment
The registry identifies Blinded Independent Central Review as the basis for the DFS assessment.
A composite endpoint can increase the number of events available for analysis, but its interpretation is necessarily tied to the components included in its definition. An HR for DFS should therefore not be described as an HR specifically for local recurrence, distant metastasis, or death alone.
18. Confidence Intervals and Statistical Precision
| Analysis | Estimate | 95% confidence interval | What the interval tells the reader |
|---|---|---|---|
| Primary DFS: Nivo + Ipi vs Placebo | 0.95 | 0.75–1.20 | Uncertainty extends below and above a hazard ratio of 1. |
| Primary DFS: Placebo vs Nivo | 0.93 | 0.67–1.28 | Uncertainty extends below and above a hazard ratio of 1. |
| Secondary OS: Nivo + Ipi vs Placebo | 1.26 | 0.85–1.85 | The interval spans a broad range of possible relative hazards under the model. |
| Secondary OS: Placebo vs Nivo | 1.36 | 0.61–3.07 | The wide interval indicates substantial statistical uncertainty around the point estimate. |
| Secondary DFS: Nivo + Ipi vs Nivo | 1.22 | 0.89–1.67 | The interval crosses 1 and is reported without a formal test in the ClinicalTrials.gov record. |
| Secondary OS: Nivo + Ipi vs Nivo | 0.75 | 0.33–1.68 | The interval spans both below and above 1 and is wide relative to the point estimate. |
The confidence intervals should not be interpreted as the range in which the true effect has a fixed probability of lying for an individual trial. They are an inferential interval constructed under the statistical model and sampling framework. Their practical value here is that they show how much uncertainty accompanies each reported hazard-ratio estimate.
19. Why This Trial Requires Careful Comparison Labels
One of the easiest ways to misread a multi-part randomized trial is to ignore the exact comparison label. CheckMate-914 contains analyses involving Treatment Part A, Treatment Part B, placebo, Nivo, and Nivo + Ipi. The reported hazard ratios are not all expressed using the same numerator.
Conceptually, reversing the same comparison reverses the hazard ratio. The registry's published ordering should therefore be preserved when reporting and interpreting an estimate.
This is particularly relevant when comparing the two primary DFS estimates. The first is Nivo + Ipi vs Placebo, whereas the second is Placebo vs Nivo. Their numerical values, 0.95 and 0.93, cannot be interpreted as though they represent identical treatment-versus-control contrasts.
20. Results Summary Table
| Endpoint role | Endpoint | Comparison | HR | 95% CI | P-value |
|---|---|---|---|---|---|
| Primary | DFS by BICR | Nivo + Ipi vs Placebo, Treatment Part A | 0.95 | 0.75–1.20 | 0.6676 |
| Primary | DFS by BICR | Placebo, Treatment Part A vs Nivo, Treatment Part B | 0.93 | 0.67–1.28 | 0.6556 |
| Secondary | OS | Nivo + Ipi vs Placebo, Treatment Part A | 1.26 | 0.85–1.85 | 0.2436 |
| Secondary | OS | Placebo, Treatment Part A vs Nivo, Treatment Part B | 1.36 | 0.61–3.07 | 0.4500 |
| Secondary | DFS | Nivo + Ipi vs Nivo, Treatment Part B | 1.22 | 0.89–1.67 | Not reported |
| Secondary | OS | Nivo + Ipi vs Nivo, Treatment Part B | 0.75 | 0.33–1.68 | Not reported |
21. What the Registry Does Not Permit Us to Conclude
- Median DFS: not provided in the ClinicalTrials.gov record, so no median can be reported.
- Median OS: not provided in the ClinicalTrials.gov record, so no median can be reported.
- Absolute DFS or OS rates: no time-specific survival probabilities are posted on ClinicalTrials.gov for these endpoints.
- Subgroup effects: no subgroup estimates are reported, so treatment-effect heterogeneity cannot be evaluated here.
- Baseline balance: no baseline characteristic table is included in the ClinicalTrials.gov record.
- Kaplan-Meier curve shape: the registry identifies Kaplan-Meier estimation but does not provide enough underlying event and censoring information to reconstruct a curve.
- Unreported formal tests: no P-values are added for the two Treatment Part B Nivo + Ipi versus Nivo secondary analyses because their statistical method is listed as not reported.
This restraint is important in clinical-trial analysis. A missing quantity is not evidence of a particular value. The statistical interpretation should remain proportional to what the underlying record actually reports.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: NCT03138512 — CheckMate-914.
- PubMed: PubMed record for PMID 39303200.
- PubMed: PubMed record for PMID 36774933.
- PubMed: PubMed record for PMID 33526329.
Continue through Clinical Biostats
Build from the trial's statistical concepts into deeper tutorials, survival-analysis methods, and practical statistical calculators.
25. Why This Analysis Is Educationally Useful
CheckMate-914 illustrates an important principle in clinical-trial statistics: the quality of an interpretation depends on preserving the relationship among design, endpoint definition, analysis population, comparison direction, effect measure, confidence interval, and hypothesis test.
The primary endpoint is a time-to-event outcome assessed by BICR, and the registry identifies Kaplan-Meier estimation, log-rank testing, and Cox proportional-hazard modeling. That combination provides a coherent framework for comparing DFS over time. But the resulting hazard ratio should still be interpreted as a relative model-based measure rather than as an absolute probability.
The two primary comparisons also demonstrate why the exact analysis label matters. Nivo + Ipi is compared with placebo in one Treatment Part A analysis, while placebo is compared with Nivo across the specified Treatment Part A and Treatment Part B structure in the other. The estimates therefore cannot be reduced to a single generic “treatment HR.”
The secondary analyses add another useful lesson: a registry can report an effect estimate without reporting every component of the formal statistical test. For the Treatment Part B Nivo + Ipi versus Nivo analyses, the Cox hazard ratio and confidence interval are available, but the method is listed as not reported. Responsible statistical interpretation preserves that limitation rather than filling the gap with assumptions.
Finally, the safety data demonstrate why efficacy and safety should remain analytically distinct. Serious adverse events are reported as affected participants among those at risk for each specified arm, while DFS and OS are modeled as time-to-event outcomes. Each endpoint requires its own definition and statistical interpretation.
26. Record Summary
CheckMate-914 is a completed randomized phase 3 trial with 1,641 enrolled participants, 5 registered arms, a parallel design, and quadruple masking. Its registered primary endpoint is disease-free survival by Blinded Independent Central Review, defined from randomization to local disease recurrence, distant metastasis, or death, whichever came first. The two posted primary analyses use log-rank testing and Cox proportional-hazard effect measures, producing DFS hazard ratios of 0.95 and 0.93 with 95% confidence intervals of 0.75–1.20 and 0.67–1.28, and P-values of 0.6676 and 0.6556.
The secondary analyses provide additional OS and Treatment Part B DFS/OS hazard ratios, but their interpretation depends on the exact randomized comparison and, for two analyses, the absence of a reported formal statistical method. The available safety data report serious adverse events of 196/608 for Nivo + Ipi, 49/614 for placebo, and 56/408 for Nivo in the specified reporting groups.
The central statistical lesson is that a clinical-trial result is more than a P-value. The hazard ratio describes the estimated relative event hazard, the confidence interval describes uncertainty around that estimate, the P-value addresses the corresponding hypothesis test, and the endpoint definition determines exactly what event process is being compared. Keeping those elements together produces a more accurate and reproducible statistical reading of the trial.