← Clinical Trials
Renal Cell Carcinoma Phase 3 Completed NCT03138512

CheckMate-914: Complete Statistical Analysis of Nivolumab and Ipilimumab in Renal Cell Carcinoma

An independent statistical review of the randomized phase 3 CheckMate-914 trial evaluating nivolumab, nivolumab in combination with ipilimumab, and placebo in participants with localized kidney cancer who underwent surgery to remove part of a kidney.

Enrollment 1,641  ·  Randomized  ·  Quadruple-masked  ·  Parallel design
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record for NCT03138512. The official registry record is the authoritative source for the trial's registration details.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CheckMate-914 was a randomized, parallel-group, quadruple-masked phase 3 trial in participants with renal cell carcinoma. The registry describes a study comparing nivolumab, nivolumab in combination with ipilimumab, and placebo in participants with localized kidney cancer who underwent surgery to remove part of a kidney.

1,641
Enrollment
Randomized trial
5
Arms
Registry design
0.95
DFS HR
Nivo + Ipi vs placebo
0.93
DFS HR
Placebo vs Nivo
FeatureCheckMate-914
PhasePhase 3
ConditionCarcinoma, Renal Cell
PopulationParticipants with localized kidney cancer who underwent surgery to remove part of a kidney
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment1,641
Arms5
Primary endpoint typeTime-to-event
Primary endpoints registered1
Results postedYes
Outcome measures posted13
Statistical analyses posted6
Lead sponsorBristol-Myers Squibb
Sponsor typeIndustry

2. Clinical Question

The registered study evaluates treatment strategies involving nivolumab and ipilimumab against placebo-based comparison groups in participants with localized kidney cancer after surgery. Statistically, the central confirmatory question is whether the randomized treatment comparisons produce a different disease-free survival experience, where disease-free survival is defined by the registry using local recurrence, distant metastasis, or death as the event.

Population

Participants with localized kidney cancer who underwent surgery to remove part of a kidney, with the registry condition listed as carcinoma, renal cell.

Interventions

The registry lists nivolumab, ipilimumab, nivolumab placebo, and ipilimumab placebo among the interventions.

Comparator structure

The posted primary analyses compare Treatment Part A Nivo + Ipi with Treatment Part A Placebo, and Treatment Part A Placebo with Treatment Part B Nivo.

Primary question

Does either prespecified randomized comparison change disease-free survival as assessed by blinded independent central review?

3. Trial Design

01
Randomize 1,641 enrolled
02
Parallel arms 5 registered arms
03
Quadruple masking Masked study design
04
DFS assessment BICR time-to-event endpoint
05
Survival analysis Log-rank and Cox HRs
Allocation
Randomized
Model
Parallel
Masking
Quadruple
Primary purpose
Treatment
TREATMENT PART A

Nivolumab + ipilimumab

  • The registry identifies Nivo + Ipi as a Treatment Part A arm.
  • Serious adverse events: 196 affected among 608 at risk in the combined Treatment Part A and B safety reporting.
TREATMENT PART A

Placebo

  • The placebo group is the comparator for the Treatment Part A Nivo + Ipi primary analysis.
  • It is also the comparator against Treatment Part B Nivo in the second primary analysis.
  • Serious adverse events: 49 affected among 614 at risk in Treatment Part A and B reporting.
TREATMENT PART B

Nivolumab

  • The Treatment Part B Nivo group is compared with the Treatment Part A placebo group for the second primary DFS analysis.
  • Serious adverse events: 56 affected among 408 at risk in Treatment Part B reporting.
REGISTRY STRUCTURE

Five registered arms

  • The registry reports 5 arms overall.
  • The statistical analyses posted on ClinicalTrials.gov do not provide separate treatment descriptions for every registered arm.
  • The analysis page therefore does not infer additional arm-level details.
Important design distinction: the registry's primary endpoint analysis population includes all randomized participants in Treatment Part A and Treatment Part B, but explicitly states that the Nivo + Ipi arm was prespecified to be excluded from the endpoint objective in the relevant comparison. The comparison itself, rather than the existence of an arm, determines which randomized groups contribute to each reported primary analysis.

4. Trial Timeline

2017-07-07

Trial start

The registered study start date was July 7, 2017.

Phase 3

Randomized parallel design

The trial was conducted as a randomized phase 3 study with quadruple masking and treatment as its primary purpose.

2023-09-28

Primary completion

The registered primary completion date was September 28, 2023.

Completed

Results posted

The registry status is completed, with 13 outcome measures and 6 statistical analyses posted.

5. Primary Endpoint

The registered primary endpoint is Disease-Free Survival (DFS) by BICR - Treatment Part A and B. It is a time-to-event endpoint.

EndpointRegistry definitionAnalysis framework
Disease-Free Survival (DFS) by BICR - Treatment Part A and B Disease-Free Survival (DFS) is defined as the time from randomization to development of local disease recurrence (ie, recurrence of primary tumor in situ or occurrence of a secondary renal cell carcinoma (RCC) primary cancer), distance metastasis, or death, whichever came first per Blinded Independent Central Review (BICR) based on Kaplan-Meier estimates. Time-to-event; Kaplan-Meier estimates; posted formal analyses use the log-rank test and Cox proportional-hazard effect measure.

6. Statistical Methodology

Kaplan-Meier estimation

The registry identifies Kaplan-Meier estimates as part of the definition of the primary DFS endpoint. Kaplan-Meier estimation is appropriate for a time-to-event endpoint because participants can have different lengths of observable follow-up. A participant who has not experienced the defined DFS event by the last usable observation can contribute information up to that point without being treated as if an event occurred.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di denotes the number of events at event time ti, while ni is the number at risk immediately before that time. The resulting curve estimates the probability of remaining event-free through time t.

Log-rank test

The two posted primary analyses use the log-rank test. The log-rank test compares the observed and expected pattern of events between randomized groups over follow-up rather than comparing only a single time point.

This is particularly appropriate for DFS because the endpoint is explicitly defined as a time from randomization to the first qualifying event. The test therefore uses the ordering of event times and the risk sets formed during follow-up.

Cox proportional-hazards effect measure

The registry reports Cox Proportional Hazard as the effect measure for the primary analyses and gives the corresponding hazard ratio estimates with 95% confidence intervals.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the numerator group

A hazard ratio is a relative, model-based measure of event hazard. It is not a probability, not a median survival time, and not the proportion of participants who benefit. Its interpretation also depends on the proportional-hazards framework underlying the Cox model.

Superiority hypothesis

Both primary statistical analyses are classified in the registry as superiority analyses. That means the statistical question is whether the randomized groups differ in the direction represented by the model and test, rather than whether one treatment is merely no worse than another within a prespecified non-inferiority margin.

Analysis populations

The primary analyses use all randomized participants in Treatment Part A and Treatment Part B according to the registry description. The registry-reported analysis population text also specifies the prespecified exclusion of the Nivo + Ipi arm from the endpoint objective relevant to the posted comparisons.

Analysis of secondary time-to-event endpoints

The four posted secondary analyses also use hazard ratios for time-to-event outcomes. Two of those analyses have a reported log-rank method, while the Treatment Part B comparisons of Nivo + Ipi versus Nivo have the method listed as not reported. The registry therefore supports reporting the latter hazard ratios and confidence intervals, but not attributing an unreported formal test or p-value to those comparisons.

7. Primary Results: Disease-Free Survival

Nivolumab + ipilimumab vs placebo

DFS hazard ratio

0.95

95% CI: 0.75–1.20   ·   P = 0.6676

Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.

Primary DFS analysisEstimate95% CIP-valueMethod
Treatment Part A: Nivo + Ipi vs Treatment Part A: Placebo HR 0.95 0.75–1.20 0.6676 Log-rank; Cox proportional hazard
Clinical Biostats interpretation

The estimated hazard ratio of 0.95 means that the fitted relative hazard of a qualifying DFS event was estimated to be 0.95 for Nivo + Ipi relative to placebo in this comparison. Expressed descriptively, that estimate is close to 1, so the point estimate represents only a modest relative difference in the estimated event hazard.

The estimate does not mean that 95% of participants remained disease-free, that 5% of participants benefited, or that the treatment changed individual risk by exactly 5%. A hazard ratio summarizes the relative event rate under the statistical model; it does not directly provide an absolute treatment effect.

The 95% confidence interval of 0.75–1.20 describes the statistical uncertainty around the estimated hazard ratio. Because the interval spans 1, the reported data are compatible with both a lower and a higher relative hazard under the model. The interval is therefore important alongside the point estimate rather than being treated as an afterthought.

The P-value of 0.6676 addresses the statistical evidence against the null comparison under the reported testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the clinical importance of the observed estimate. Effect size and uncertainty are better represented by the hazard ratio and its confidence interval.

Finally, the Cox interpretation depends on the proportional-hazards framework. A single hazard ratio is most naturally interpreted as a relative hazard measure under that model; it should not automatically be translated into a constant difference in individual risk throughout follow-up.

Placebo vs nivolumab

DFS hazard ratio

0.93

95% CI: 0.67–1.28   ·   P = 0.6556

Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.

Primary DFS analysisEstimate95% CIP-valueMethod
Treatment Part A: Placebo vs Treatment Part B: Nivo HR 0.93 0.67–1.28 0.6556 Log-rank; Cox proportional hazard
Clinical Biostats interpretation

The reported hazard ratio of 0.93 is the estimated relative hazard for the comparison as recorded by the registry: Treatment Part A placebo versus Treatment Part B nivolumab. Because the registry presents the comparison in that order, the direction of the numerical ratio should be read in that same order rather than silently reversing it.

A hazard ratio of 0.93 does not mean that 93% of participants experienced an event or that nivolumab reduces individual risk by 7%. It is a model-based comparison of event hazards between the two randomized groups.

The 95% confidence interval of 0.67–1.28 is relatively broad and includes 1. This means the point estimate alone should not be treated as a precise description of the underlying treatment effect. The interval communicates that appreciable uncertainty remains around the estimated relative hazard.

The reported P-value of 0.6556 is evidence from the stated statistical test, not an effect-size measure. A large P-value does not prove that two treatments are identical, just as a small P-value would not by itself establish that an effect is clinically important.

As with the other Cox analysis, interpretation should remain tied to the proportional-hazards model and to the time-to-event nature of DFS. The registry does not provide median DFS or time-specific DFS estimates in the ClinicalTrials.gov record, so those quantities are not substituted for the reported hazard ratio.

Educational note: a Kaplan-Meier curve cannot be reconstructed faithfully from the reported hazard ratios, confidence intervals, and P-values alone. The underlying event and censoring information would be required for a valid reconstruction.

8. Secondary Endpoint Results

Overall Survival — Treatment Part A: Nivo + Ipi vs Placebo

OS hazard ratio

1.26

95% CI: 0.85–1.85   ·   P = 0.2436

Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.

Secondary endpointComparisonHR95% CIP-value
Overall Survival (OS) - Treatment Part A and B Treatment Part A: Nivo + Ipi vs Treatment Part A: Placebo 1.26 0.85–1.85 0.2436

The registry defines this OS endpoint as time from randomization to the date of death, with a time frame of up to approximately 72 months. The analysis population is all randomized participants in Treatment Part A and Treatment Part B, with the Nivo + Ipi arm prespecified to be excluded from the endpoint objective where applicable.

Overall Survival — Placebo vs Nivolumab

OS hazard ratio

1.36

95% CI: 0.61–3.07   ·   P = 0.4500

Log-rank test; Cox proportional-hazard effect measure; superiority hypothesis.

Secondary endpointComparisonHR95% CIP-value
Overall Survival (OS) - Treatment Part A and B Treatment Part A: Placebo vs Treatment Part B: Nivo 1.36 0.61–3.07 0.4500

The registry identifies this as the Treatment Part B analysis and defines OS as time from randomization to the date of death, with a time frame of up to approximately 72 months. The confidence interval is wide, extending from below 1 to above 3, so the point estimate should not be interpreted without its uncertainty interval.

Disease-Free Survival — Treatment Part B Nivo + Ipi vs Nivo

DFS hazard ratio

1.22

95% CI: 0.89–1.67

Two-sided 95% confidence interval; Cox proportional-hazard effect measure.

This secondary endpoint was prespecified to be collected for Treatment Part B only. The analysis population consists of all randomized Nivo + Ipi and Nivo participants in Treatment Part B, with the placebo arm prespecified to be excluded from the endpoint objective.

No formal test is reported for this comparison. The ClinicalTrials.gov record lists the statistical method as “Not reported.” Accordingly, this page reports the hazard ratio and 95% confidence interval but does not invent a log-rank test, P-value, or other formal comparison that is absent from the registry analysis.

Overall Survival — Treatment Part B Nivo + Ipi vs Nivo

OS hazard ratio

0.75

95% CI: 0.33–1.68

Two-sided 95% confidence interval; Cox proportional-hazard effect measure.

This secondary Treatment Part B endpoint evaluates overall survival among contemporaneously randomized Nivo + Ipi and Nivo participants. The registry again lists the statistical method as not reported, so the reported hazard ratio and confidence interval are presented without adding an unreported hypothesis test.

Secondary endpointComparisonHR95% CIFormal method reported
DFS in contemporaneously randomized combination and monotherapy participants Treatment Part B: Nivo + Ipi vs Treatment Part B: Nivo 1.22 0.89–1.67 Not reported
OS in contemporaneously randomized combination and monotherapy participants Treatment Part B: Nivo + Ipi vs Treatment Part B: Nivo 0.75 0.33–1.68 Not reported

9. How to Read the CheckMate-914 Hazard Ratios

Direction matters

The direction of a hazard ratio depends on the order in which the groups are specified. For example, the primary comparison is recorded as Nivo + Ipi vs Placebo, giving HR 0.95. The second primary comparison is recorded as Placebo vs Nivo, giving HR 0.93. These numbers should not be casually compared as if both were expressed in the same treatment-versus-control direction.

A hazard ratio is not a risk ratio

A hazard ratio compares event hazards over time. It is not the ratio of the cumulative percentages of participants who experience an event. Converting an HR directly into an absolute probability or percentage of patients affected would require additional information that is not reported in the ClinicalTrials.gov record.

Confidence intervals carry information

The confidence interval shows how much uncertainty surrounds the point estimate under the statistical model. For the primary Nivo + Ipi versus placebo DFS comparison, the interval is 0.75–1.20. For placebo versus Nivo, it is 0.67–1.28. Both intervals include 1, so the point estimates should not be interpreted as precise evidence of a directional difference by themselves.

P-values do not measure effect size

The primary P-values, 0.6676 and 0.6556, quantify evidence under the respective statistical testing framework. They do not tell us how large a treatment effect is. The hazard ratio and confidence interval address effect magnitude and uncertainty; the P-value addresses compatibility with the null hypothesis under the specified test.

10. Statistical Methods Explained

Why was a time-to-event method used for DFS?

DFS is explicitly defined as the time from randomization until the first of several qualifying events: local disease recurrence, distant metastasis, or death. Because both the timing of events and incomplete event observation matter, a time-to-event framework preserves more information than a simple yes/no comparison at one arbitrary time point.

Why use Kaplan-Meier estimation?

Kaplan-Meier estimation provides an estimate of the event-free survival function over time. It is designed to incorporate participants whose event status is not observed for the entire follow-up period by accounting for their available follow-up through censoring.

What does the log-rank test contribute?

The log-rank test compares the survival experience of randomized groups across observed event times. Rather than asking only whether two percentages differ at a single time point, it evaluates the accumulated pattern of events while participants are at risk.

What does an HR of 0.95 mean?

For the Treatment Part A Nivo + Ipi versus placebo primary DFS comparison, an HR of 0.95 means that the estimated event hazard for the first group relative to the second was 0.95 under the Cox model. It does not mean that 95% of patients were disease-free or that every participant had a 5% lower risk.

Why does the confidence interval matter as much as the point estimate?

A point estimate is only one representation of the observed evidence. The 95% confidence interval provides information about statistical precision. In CheckMate-914, the primary DFS intervals extend across 1, so the reported point estimates should be considered together with substantial uncertainty about the underlying relative hazard.

Why does the direction of the comparison matter?

Hazard ratios are reciprocal when the same two groups are compared in the opposite order, subject to the corresponding model specification. Therefore, “Nivo + Ipi vs Placebo” and “Placebo vs Nivo” are not interchangeable labels. The comparison direction should always be read directly from the reported analysis.

Why should an unreported P-value not be invented?

The two Treatment Part B combination-versus-monotherapy secondary analyses report Cox proportional-hazard effect measures but list the statistical method as not reported. A statistically responsible analysis preserves that distinction. A hazard ratio and confidence interval can be discussed without fabricating a test statistic or P-value that the registry does not provide.

11. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk. These are the available arm-level safety results in the ClinicalTrials.gov record.

Safety groupSerious adverse eventsAffected / at risk
Treatment Part A and B: Nivo + Ipi Serious adverse events 196 / 608
Treatment Part A and B: Placebo Serious adverse events 49 / 614
Treatment Part B: Nivo Serious adverse events 56 / 408

The denominators are important because the reported safety information is expressed as affected participants relative to those at risk. The ClinicalTrials.gov record does not provide additional categories such as grade-specific adverse events, discontinuations because of adverse events, or individual adverse-event terms, so those measures are not added here.

Safety versus efficacy: serious adverse-event counts describe a safety outcome and should not be combined mathematically with the DFS or OS hazard ratios. Efficacy and safety address different dimensions of a randomized clinical trial and require separate interpretation.

12. Analysis Population and Endpoint Structure

The registry's analysis descriptions reveal an important feature of CheckMate-914: the trial contains multiple treatment parts and prespecified comparisons, while the primary endpoint itself is registered once as a DFS endpoint covering Treatment Part A and B.

AnalysisPopulation described in registryComparisonRole
Primary DFS All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective. Nivo + Ipi vs Placebo in Treatment Part A Primary
Primary DFS All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective. Placebo in Treatment Part A vs Nivo in Treatment Part B Primary
Secondary OS All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective where applicable Nivo + Ipi vs Placebo in Treatment Part A Secondary
Secondary OS All randomized participants in Treatment Part A and Treatment Part B; Nivo + Ipi prespecified to be excluded from the endpoint objective where applicable Placebo in Treatment Part A vs Nivo in Treatment Part B Secondary
Secondary DFS All randomized Nivo + Ipi and Nivo participants in Treatment Part B Nivo + Ipi vs Nivo in Treatment Part B Secondary
Secondary OS All randomized Nivo + Ipi and Nivo participants in Treatment Part B Nivo + Ipi vs Nivo in Treatment Part B Secondary

This structure is statistically important because a trial can contain several randomized comparisons without every comparison being the same confirmatory question. The analysis population, treatment part, comparison direction, endpoint role, and prespecified exclusion rules should therefore be retained whenever an estimate is interpreted.

13. What Is and Is Not Reported in the Supplied Registry Data

Reported

The registry provides the primary DFS endpoint, two primary statistical analyses, four secondary statistical analyses, hazard ratios, confidence intervals, P-values for four analyses, and serious adverse-event counts by arm.

Not reported here

The ClinicalTrials.gov record does not provide median DFS, median OS, time-specific survival probabilities, baseline characteristics, subgroup estimates, or reconstructed Kaplan-Meier event counts.

Method reported

The primary analyses use a log-rank test and Cox proportional-hazard effect measure. The two Treatment Part B combination-versus-monotherapy secondary analyses list the method as not reported.

Interpretive consequence

The page emphasizes relative time-to-event effects and uncertainty rather than filling gaps with values from external publications or inferred calculations.

14. Limitations

15. Why This Trial Matters Statistically

CheckMate-914 is a useful teaching case because it demonstrates how a randomized oncology trial can combine a composite time-to-event endpoint, multiple treatment parts, multiple randomized comparisons, blinded central assessment, and Cox-model effect estimation. The statistical story is therefore more nuanced than simply asking whether a P-value is below a threshold.

ConceptHow it appears in CheckMate-914
RandomizationThe study uses randomized allocation in a phase 3 parallel design.
BlindingThe registered masking level is quadruple.
Time-to-event analysisThe primary endpoint is DFS, defined from randomization to the first qualifying recurrence, metastasis, or death.
Kaplan-Meier estimationThe registry definition identifies DFS as being based on Kaplan-Meier estimates.
Log-rank testingThe two primary analyses use a log-rank test.
Hazard ratioThe primary and secondary reported analyses use a Cox proportional-hazard effect measure.
Confidence intervalsPrimary analyses report 95% confidence intervals of 0.75–1.20 and 0.67–1.28.
Superiority testingThe primary analyses are classified as superiority hypotheses.
Analysis populationThe registry explicitly describes randomized populations and prespecified exclusions from particular endpoint objectives.
Method transparencyTwo Treatment Part B secondary analyses report hazard ratios and confidence intervals but list the statistical method as not reported.

16. Interpreting the Primary Evidence as a Statistical Story

The two primary DFS analyses produce estimates of 0.95 and 0.93, respectively. Both are close to 1, and both corresponding 95% confidence intervals include 1. The associated P-values are 0.6676 and 0.6556. Taken together, the ClinicalTrials.gov record does not provide a basis for describing either primary comparison as a statistically demonstrated difference in DFS under its reported testing framework.

That conclusion is deliberately narrower than saying that the treatments are proven identical. A non-significant superiority test does not establish equivalence or non-inferiority. It means that the observed data, analyzed using the reported test, do not provide the statistical evidence required to reject the null hypothesis at the conventional interpretation of that test.

The confidence intervals make the distinction clearer. For the Nivo + Ipi versus placebo comparison, the interval ranges from 0.75 to 1.20. For placebo versus Nivo, it ranges from 0.67 to 1.28. Both permit a range of underlying relative hazards, rather than identifying one exact treatment effect.

The secondary results illustrate the same principle. The OS hazard ratios are 1.26 and 1.36 for the two Treatment Part A and B comparisons, while the Treatment Part B Nivo + Ipi versus Nivo analyses give DFS HR 1.22 and OS HR 0.75. These estimates should be read with their respective confidence intervals and with attention to whether a formal statistical method was actually reported.

This is a useful distinction between estimating an effect and testing a hypothesis. A hazard ratio answers “what relative effect did the fitted model estimate?” A confidence interval answers “how precisely is that effect estimated?” A P-value answers a different question about compatibility with a null hypothesis under the specified test.

17. Primary Endpoint: Why the Composite Definition Matters

DFS in this registry record is not simply “time until recurrence.” It is defined as the time from randomization to the first of local disease recurrence, distant metastasis, or death. That construction has direct statistical consequences: whichever qualifying event occurs first determines the endpoint event time.

One time origin

The endpoint begins at randomization. This creates a common starting point for the randomized comparison.

Multiple event types

Local recurrence, distant metastasis, and death are all qualifying events for DFS.

First event governs

The definition uses whichever qualifying event occurs first, making DFS a time-to-first-event composite endpoint.

BICR assessment

The registry identifies Blinded Independent Central Review as the basis for the DFS assessment.

A composite endpoint can increase the number of events available for analysis, but its interpretation is necessarily tied to the components included in its definition. An HR for DFS should therefore not be described as an HR specifically for local recurrence, distant metastasis, or death alone.

18. Confidence Intervals and Statistical Precision

AnalysisEstimate95% confidence intervalWhat the interval tells the reader
Primary DFS: Nivo + Ipi vs Placebo 0.95 0.75–1.20 Uncertainty extends below and above a hazard ratio of 1.
Primary DFS: Placebo vs Nivo 0.93 0.67–1.28 Uncertainty extends below and above a hazard ratio of 1.
Secondary OS: Nivo + Ipi vs Placebo 1.26 0.85–1.85 The interval spans a broad range of possible relative hazards under the model.
Secondary OS: Placebo vs Nivo 1.36 0.61–3.07 The wide interval indicates substantial statistical uncertainty around the point estimate.
Secondary DFS: Nivo + Ipi vs Nivo 1.22 0.89–1.67 The interval crosses 1 and is reported without a formal test in the ClinicalTrials.gov record.
Secondary OS: Nivo + Ipi vs Nivo 0.75 0.33–1.68 The interval spans both below and above 1 and is wide relative to the point estimate.

The confidence intervals should not be interpreted as the range in which the true effect has a fixed probability of lying for an individual trial. They are an inferential interval constructed under the statistical model and sampling framework. Their practical value here is that they show how much uncertainty accompanies each reported hazard-ratio estimate.

19. Why This Trial Requires Careful Comparison Labels

One of the easiest ways to misread a multi-part randomized trial is to ignore the exact comparison label. CheckMate-914 contains analyses involving Treatment Part A, Treatment Part B, placebo, Nivo, and Nivo + Ipi. The reported hazard ratios are not all expressed using the same numerator.

Comparison direction
HR(A vs B) = 1 / HR(B vs A)

Conceptually, reversing the same comparison reverses the hazard ratio. The registry's published ordering should therefore be preserved when reporting and interpreting an estimate.

This is particularly relevant when comparing the two primary DFS estimates. The first is Nivo + Ipi vs Placebo, whereas the second is Placebo vs Nivo. Their numerical values, 0.95 and 0.93, cannot be interpreted as though they represent identical treatment-versus-control contrasts.

20. Results Summary Table

Endpoint roleEndpointComparisonHR95% CIP-value
Primary DFS by BICR Nivo + Ipi vs Placebo, Treatment Part A 0.95 0.75–1.20 0.6676
Primary DFS by BICR Placebo, Treatment Part A vs Nivo, Treatment Part B 0.93 0.67–1.28 0.6556
Secondary OS Nivo + Ipi vs Placebo, Treatment Part A 1.26 0.85–1.85 0.2436
Secondary OS Placebo, Treatment Part A vs Nivo, Treatment Part B 1.36 0.61–3.07 0.4500
Secondary DFS Nivo + Ipi vs Nivo, Treatment Part B 1.22 0.89–1.67 Not reported
Secondary OS Nivo + Ipi vs Nivo, Treatment Part B 0.75 0.33–1.68 Not reported

21. What the Registry Does Not Permit Us to Conclude

This restraint is important in clinical-trial analysis. A missing quantity is not evidence of a particular value. The statistical interpretation should remain proportional to what the underlying record actually reports.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue through Clinical Biostats

Build from the trial's statistical concepts into deeper tutorials, survival-analysis methods, and practical statistical calculators.

25. Why This Analysis Is Educationally Useful

CheckMate-914 illustrates an important principle in clinical-trial statistics: the quality of an interpretation depends on preserving the relationship among design, endpoint definition, analysis population, comparison direction, effect measure, confidence interval, and hypothesis test.

The primary endpoint is a time-to-event outcome assessed by BICR, and the registry identifies Kaplan-Meier estimation, log-rank testing, and Cox proportional-hazard modeling. That combination provides a coherent framework for comparing DFS over time. But the resulting hazard ratio should still be interpreted as a relative model-based measure rather than as an absolute probability.

The two primary comparisons also demonstrate why the exact analysis label matters. Nivo + Ipi is compared with placebo in one Treatment Part A analysis, while placebo is compared with Nivo across the specified Treatment Part A and Treatment Part B structure in the other. The estimates therefore cannot be reduced to a single generic “treatment HR.”

The secondary analyses add another useful lesson: a registry can report an effect estimate without reporting every component of the formal statistical test. For the Treatment Part B Nivo + Ipi versus Nivo analyses, the Cox hazard ratio and confidence interval are available, but the method is listed as not reported. Responsible statistical interpretation preserves that limitation rather than filling the gap with assumptions.

Finally, the safety data demonstrate why efficacy and safety should remain analytically distinct. Serious adverse events are reported as affected participants among those at risk for each specified arm, while DFS and OS are modeled as time-to-event outcomes. Each endpoint requires its own definition and statistical interpretation.

26. Record Summary

CheckMate-914 is a completed randomized phase 3 trial with 1,641 enrolled participants, 5 registered arms, a parallel design, and quadruple masking. Its registered primary endpoint is disease-free survival by Blinded Independent Central Review, defined from randomization to local disease recurrence, distant metastasis, or death, whichever came first. The two posted primary analyses use log-rank testing and Cox proportional-hazard effect measures, producing DFS hazard ratios of 0.95 and 0.93 with 95% confidence intervals of 0.75–1.20 and 0.67–1.28, and P-values of 0.6676 and 0.6556.

The secondary analyses provide additional OS and Treatment Part B DFS/OS hazard ratios, but their interpretation depends on the exact randomized comparison and, for two analyses, the absence of a reported formal statistical method. The available safety data report serious adverse events of 196/608 for Nivo + Ipi, 49/614 for placebo, and 56/408 for Nivo in the specified reporting groups.

The central statistical lesson is that a clinical-trial result is more than a P-value. The hazard ratio describes the estimated relative event hazard, the confidence interval describes uncertainty around that estimate, the P-value addresses the corresponding hypothesis test, and the endpoint definition determines exactly what event process is being compared. Keeping those elements together produces a more accurate and reproducible statistical reading of the trial.

Clinical Biostats methodology: This analysis reports the numerical evidence posted on ClinicalTrials.gov for the registered CheckMate-914 record and distinguishes reported results from statistical interpretation. Where the ClinicalTrials.gov record does not provide a quantity or formal method, the page does not substitute an inferred value or analysis.