← Clinical Trials
Head & Neck Cancer Phase 3 Completed NCT02741570

CheckMate-651: Complete Statistical Analysis of Nivolumab + Ipilimumab in Recurrent or Metastatic Squamous Cell Carcinoma of the Head and Neck

An independent statistical analysis of the randomized phase 3 CheckMate-651 trial evaluating nivolumab in combination with ipilimumab versus the EXTREME regimen as first-line treatment in participants with recurrent or metastatic squamous cell carcinoma of the head and neck.

Randomized phase 3  ·  Enrollment 947  ·  Primary completion May 10, 2021
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses posted in the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record. View the CheckMate-651 registry record.

1. Trial at a Glance

CheckMate-651 was a randomized, parallel-group, open-label phase 3 trial evaluating nivolumab plus ipilimumab against the EXTREME regimen as first-line treatment in participants with recurrent or metastatic squamous cell carcinoma of the head and neck. The registry reports an enrollment of 947 participants and two treatment arms.

947
Enrollment
Participants
2
Treatment arms
Parallel design
0.78
Primary OS HR
CPS ≥20
0.95
Primary OS HR
All randomized
FeatureCheckMate-651
PhasePhase 3
ConditionHead and Neck Cancer
PopulationParticipants with recurrent or metastatic squamous cell carcinoma of the head and neck receiving first-line treatment
DesignRandomized, parallel-group
MaskingNone
Primary purposeTreatment
Enrollment947
Primary endpointsTwo overall survival endpoints
Primary endpoint typeTime-to-event
Primary statistical methodLog-rank test
Primary effect measureHazard ratio
Hypothesis typeSuperiority
Trial statusCompleted

2. Clinical Question

The central statistical question was whether first-line treatment with nivolumab plus ipilimumab improved overall survival compared with the EXTREME regimen in participants with recurrent or metastatic squamous cell carcinoma of the head and neck.

Population

Participants with recurrent or metastatic squamous cell carcinoma of the head and neck receiving first-line treatment.

Intervention

Nivolumab in combination with ipilimumab.

Comparator

The EXTREME regimen, consisting in the registry intervention list of cetuximab/Erbitux, cisplatin/Platinol or carboplatin/Paraplatin, and fluorouracil/Adrucil.

Primary question

Does nivolumab plus ipilimumab improve overall survival relative to the EXTREME regimen under a superiority framework?

3. Trial Design

01
Randomize 947 participants
02
Two arms Nivolumab + ipilimumab or EXTREME
03
Follow Time-to-event outcomes
04
Compare Overall survival and PFS
05
Interpret Hazard ratios and confidence intervals
Allocation
Randomized allocation to two parallel treatment arms.
Masking
None. The registry identifies the trial as unmasked.
Primary purpose
Treatment.
Statistical framework
Superiority comparison using time-to-event methods and hazard ratios.
ARM A

Nivolumab + Ipilimumab

  • Nivolumab
  • Ipilimumab
ARM B

EXTREME Regimen

  • Cetuximab / Erbitux
  • Cisplatin / Platinol or carboplatin / Paraplatin
  • Fluorouracil / Adrucil

The trial was conducted from October 5, 2016 through primary completion on May 10, 2021. The registry identifies Bristol-Myers Squibb as the lead sponsor and the sponsor type as industry.

4. Endpoints

EndpointRegistry definition / time frameStatistical role
Overall Survival (OS) in Participants With PD-L1 With a Combined Positive Score (CPS) ≥20 From randomization to date of death or date the participant was last known to be alive (Up to approximately 55 months). Overall survival is defined as the time between randomization and death. Participants without documentation of death are censored on the last date they were known to be alive; participants randomized without follow-up are censored at randomization. Primary
Overall Survival (OS) in All Randomized Participants From randomization to date of death or date the participant was last known to be alive (Up to approximately 55 months). Overall survival is defined as the time between randomization and death. Participants without documentation of death are censored on the last date they were known to be alive; participants randomized without follow-up are censored at randomization. Primary
Overall Survival (OS) in Randomized Participants With PD-L1 With a Combined Positive Score (CPS) ≥1 From randomization to date of death or date the participant was last known to be alive (Up to approximately 65 months) Secondary
Progression Free Survival (PFS) From randomization to disease progression or death (Up to approximately 65 months) Secondary

The registry therefore defines the primary outcomes as time-to-event endpoints. The same underlying survival-analysis framework is used for overall survival and progression-free survival, but the event definition differs: OS uses death, whereas PFS uses disease progression or death.

5. Primary Results

ClinicalTrials.gov reports formal statistical analyses for both primary overall-survival endpoints. Both analyses compare nivolumab plus ipilimumab with the EXTREME regimen using a stratified regular log-rank test and a hazard-ratio effect measure.

Overall Survival in Participants With PD-L1 CPS ≥20

Hazard ratio for death

0.78

97.51% two-sided CI: 0.59–1.03  ·  P = 0.0469

Analysis population: all randomized PD-L1 CPS ≥20 participants

ElementReported result
ComparisonNivolumab + ipilimumab vs EXTREME regimen
Analysis populationAll randomized PD-L1 CPS ≥20 participants
MethodLog Rank; analysis notes specify a stratified regular log-rank test
Effect measureHazard Ratio (HR)
Estimate0.78
Confidence interval97.51% two-sided CI: 0.59–1.03
P-value0.0469
HypothesisSuperiority
Clinical Biostats interpretation

An HR of 0.78 means that, under the time-to-event comparison represented by the reported hazard ratio, the estimated instantaneous rate of death in the nivolumab-plus-ipilimumab group was about 22% lower than in the EXTREME group. This is a relative hazard measure, not a statement that 22% of patients avoided death or that every individual patient's probability of death was reduced by exactly 22%.

The 97.51% two-sided confidence interval of 0.59–1.03 describes the statistical uncertainty around the estimated hazard ratio under the analysis framework. Its width indicates that the estimate is not known with arbitrary precision. The interval extends above 1, so the reported estimate should not be interpreted as establishing a uniformly lower hazard across all plausible values represented by the interval.

The P-value of 0.0469 addresses the compatibility of the observed test statistic with the null hypothesis under the specified testing framework. It does not measure the size or clinical importance of the treatment effect. The magnitude of the effect is described by the HR, while its statistical uncertainty is described by the confidence interval.

Because this is a time-to-event analysis, the HR should also not be interpreted as a simple ratio of cumulative probabilities. It summarizes relative event rates within the model-based survival-analysis framework and should be interpreted alongside the endpoint definition, censoring rules, and the analysis population.

Overall Survival in All Randomized Participants

Hazard ratio for death

0.95

97.9% two-sided CI: 0.80–1.13  ·  P = 0.4951

Analysis population: all randomized participants

ElementReported result
ComparisonNivolumab + ipilimumab vs EXTREME regimen
Analysis populationAll randomized participants
MethodLog Rank; analysis notes specify a stratified regular log-rank test
Effect measureCox Proportional Hazard
Estimate0.95
Confidence interval97.9% two-sided CI: 0.80–1.13
P-value0.4951
HypothesisSuperiority
Clinical Biostats interpretation

An HR of 0.95 corresponds to an estimated hazard that is approximately 5% lower in the nivolumab-plus-ipilimumab group than in the EXTREME group under the reported model. That relative estimate should not be translated into a 5% absolute reduction in mortality or into a statement about individual patient outcomes.

The 97.9% two-sided confidence interval of 0.80–1.13 spans 1.00. Thus, the interval includes values corresponding to a lower estimated hazard as well as values corresponding to a higher estimated hazard for nivolumab plus ipilimumab. The interval is therefore important for understanding the uncertainty around the point estimate rather than focusing on the HR alone.

The P-value of 0.4951 is not an effect-size measure. A larger P-value does not prove that the treatments are identical, just as a smaller P-value would not by itself establish that an effect is clinically important. The estimate and confidence interval provide the information about magnitude and precision.

This analysis also illustrates why the definition of the analysis population matters. The result for all randomized participants is not interchangeable with the result in the PD-L1 CPS ≥20 population. They answer different prespecified population questions.

Comparing the two primary estimates: the registry reports HR 0.78 for randomized participants with PD-L1 CPS ≥20 and HR 0.95 for all randomized participants. These are estimates in different analysis populations and should not be treated as if they were two estimates from the same population. A difference between subgroup-specific and overall estimates does not by itself establish treatment-effect heterogeneity.

6. Secondary Endpoint Results

Overall Survival in Randomized Participants With PD-L1 CPS ≥1

Hazard ratio for death

0.80

95% two-sided CI: 0.68–0.95

Analysis population: all randomized PD-L1 CPS ≥1 participants

EndpointEstimateConfidence intervalMethod reported
OS, PD-L1 CPS ≥1 HR 0.80 95% two-sided CI: 0.68–0.95 Not reported

The registry reports a Cox proportional-hazard effect measure for this secondary endpoint but does not provide a formal statistical-analysis method in the ClinicalTrials.gov record. An HR of 0.80 represents an estimated 20% lower hazard under the reported model, while the confidence interval describes uncertainty around that estimate.

Interpretation caution: the registry data posted on ClinicalTrials.gov for this endpoint do not include a P-value or a reported analysis-method field. The result should therefore be described using the posted estimate and confidence interval rather than adding an unreported formal test.

Progression Free Survival

Analysis populationEstimateConfidence intervalMethod reported
All randomized participants HR 1.40 95% two-sided CI: 1.19–1.63 Not reported
All randomized PD-L1 CPS ≥20 participants HR 1.00 95% two-sided CI: 0.77–1.30 Not reported

The registry defines PFS as the time from randomization to disease progression or death, with a time frame of up to approximately 65 months. The reported effect measure is a Cox proportional-hazard estimate.

For all randomized participants, the reported HR of 1.40 corresponds to an estimated hazard about 40% higher in the nivolumab-plus-ipilimumab group under the reported comparison. The 95% confidence interval of 1.19–1.63 is entirely above 1.00.

For randomized PD-L1 CPS ≥20 participants, the reported HR of 1.00 indicates equal estimated hazards at the point estimate. The 95% confidence interval of 0.77–1.30 indicates substantial uncertainty around that point estimate and includes both values below and above 1.00.

No P-values are reported for these secondary analyses in the ClinicalTrials.gov record. The appropriate interpretation is therefore based on the reported hazard ratios and confidence intervals. A confidence interval crossing 1.00 is not equivalent to a formal P-value, and a confidence interval excluding 1.00 should not be assigned an unreported P-value.

7. Extended Overall Survival Analyses

The registry also contains two post-hoc extended-collection overall-survival analyses. These are analytically distinct from the two registered primary endpoints and should be labeled accordingly.

Post-hoc endpointAnalysis populationEstimateConfidence interval
OS in participants with PD-L1 CPS ≥20 — Extended Collection All randomized PD-L1 CPS ≥20 participants HR 0.76 95% two-sided CI: 0.60–0.97
OS in all randomized participants — Extended Collection All randomized participants HR 0.94 95% two-sided CI: 0.82–1.08

The extended-collection HR of 0.76 in the PD-L1 CPS ≥20 population represents an estimated 24% lower hazard under the reported effect measure. Its 95% confidence interval, 0.60–0.97, does not include 1.00.

The extended-collection HR of 0.94 in all randomized participants represents an estimated 6% lower hazard at the point estimate. Its 95% confidence interval of 0.82–1.08 includes 1.00.

Post-hoc status matters: these extended-collection analyses are labeled Post_Hoc in the ClinicalTrials.gov record. They should not be presented as additional primary endpoints simply because they contain longer follow-up.

8. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The denominator is explicitly described as participants at risk for the safety summary.

Safety measureAffectedAt risk
Nivolumab + Ipilimumab 313 468
EXTREME Regimen 290 441

These figures describe the number of participants affected by serious adverse events and the corresponding number at risk in each arm. They should not be substituted for efficacy analyses: safety is summarized according to treatment exposure and adverse-event definitions, whereas the primary efficacy comparison is anchored to randomized treatment assignment and time-to-event methodology.

Clinical Biostats interpretation

The serious-adverse-event counts show that serious events occurred in both randomized groups. The affected and at-risk counts should be interpreted together; the raw counts alone are not a valid comparison when the numbers of participants at risk differ.

The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse events, severity-specific breakdown, timing, or attribution. Those quantities should not be inferred from the two affected counts.

9. Statistical Methodology

Time-to-event analysis

Overall survival and progression-free survival are time-to-event endpoints. Instead of asking only whether an event occurred, time-to-event analysis incorporates when the event occurred and can account for participants whose event time is not observed during follow-up.

Conceptual survival function
S(t) = P(T > t)

The survival function represents the probability that the event time T exceeds a specified time t. For overall survival, the event is death; for PFS, the event is disease progression or death.

Kaplan-Meier estimation

A Kaplan-Meier estimator is a standard descriptive method for estimating a survival curve in the presence of right censoring. The ClinicalTrials.gov record classifies both primary endpoints as time-to-event and explicitly define censoring for overall survival, but the registry ClinicalTrials.gov record do not identify Kaplan-Meier as the formal comparison method.

For a participant who remains alive without documentation of death, the registry specifies censoring at the last date the participant was known to be alive. Participants who were randomized but had no follow-up are censored at randomization. These rules determine how available follow-up contributes to the survival analysis.

Stratified log-rank test

The formal primary analyses use a stratified regular log-rank test. The log-rank test compares the observed and expected numbers of events between randomized treatment groups over follow-up. A stratified version performs this comparison while accounting for prespecified strata.

The ClinicalTrials.gov record identifies stratified analysis as an analysis concept, but do not provide the specific stratification variables. Accordingly, no specific stratification factors are asserted here.

Cox proportional-hazards model

The registry reports Cox proportional-hazard effect measures for several analyses. The Cox model expresses the treatment comparison through a hazard ratio. Conceptually, an HR compares the instantaneous event rates between treatment groups within the model.

Hazard-ratio interpretation
HR = hazard in nivolumab + ipilimumab group / hazard in EXTREME group

An HR below 1 corresponds to a lower estimated event hazard for the numerator group; an HR above 1 corresponds to a higher estimated event hazard.

Superiority testing

The primary analyses are identified as superiority hypotheses. In a superiority framework, the statistical question is whether the randomized intervention and comparator differ in the prespecified direction under the specified testing procedure, rather than whether the intervention is merely not unacceptably worse.

Confidence intervals

A confidence interval gives a range of parameter values compatible with the statistical model and sampling framework at the stated confidence level. For a hazard ratio, an interval containing 1.00 includes no relative-hazard difference as one of the values consistent with the interval.

10. Statistical Methods Explained

Why was a log-rank test used?

Overall survival and progression-free survival are time-to-event endpoints, so the analysis needs to use information about event timing and censoring. The log-rank test is designed specifically to compare survival distributions between randomized groups over follow-up. In CheckMate-651, the registry identifies the primary comparison as a stratified regular log-rank test.

What does an HR of 0.78 mean?

An HR of 0.78 means that the estimated instantaneous rate of death was about 22% lower for nivolumab plus ipilimumab than for the EXTREME regimen under the reported model. It does not mean that 22% of participants avoided death, nor does it imply a 22% reduction in every patient's probability of death.

Why does the confidence interval matter?

The point estimate is only one summary of the evidence. For the PD-L1 CPS ≥20 primary analysis, the HR is 0.78, but the 97.51% two-sided confidence interval ranges from 0.59 to 1.03. That interval communicates uncertainty that the point estimate alone cannot show.

Why is a P-value not an effect-size measure?

A P-value describes evidence against a specified null hypothesis under the statistical testing framework. It is affected by the amount of information in the analysis as well as the observed data. The HR describes relative effect magnitude, while the confidence interval communicates uncertainty around that magnitude. These quantities answer different statistical questions.

Why are the PD-L1 CPS ≥20 and all-randomized results not interchangeable?

The two primary endpoints use different analysis populations. The CPS ≥20 analysis has an HR of 0.78, whereas the all-randomized analysis has an HR of 0.95. A subgroup-specific estimate describes the randomized participants meeting that biomarker criterion; the all-randomized estimate describes the full randomized population. Differences between them should not automatically be interpreted as evidence that treatment effects truly differ by biomarker status.

What does an HR of 1.40 for PFS mean?

For all randomized participants, the PFS HR is 1.40. Under the reported effect measure, this corresponds to an estimated 40% higher hazard of progression or death in the nivolumab-plus-ipilimumab group relative to the EXTREME group. The interpretation concerns the hazard of the composite PFS event, not overall survival or a simple probability of progression.

Why is censoring important?

Some participants do not experience the event during observed follow-up. Instead of treating them as if the event never could occur, time-to-event methods use their observed follow-up and censor them according to prespecified rules. For OS, the registry specifically states that participants without documentation of death are censored at the last date they were known to be alive.

11. Interpreting the Primary Statistical Evidence

PD-L1 CPS ≥20

The primary OS estimate was HR 0.78 with a 97.51% two-sided CI of 0.59–1.03 and P = 0.0469. The point estimate indicates lower estimated hazard, while the confidence interval shows uncertainty around that estimate.

All randomized participants

The primary OS estimate was HR 0.95 with a 97.9% two-sided CI of 0.80–1.13 and P = 0.4951. The point estimate is close to 1.00 and the confidence interval includes 1.00.

Secondary OS

For randomized participants with PD-L1 CPS ≥1, the reported OS HR was 0.80 with a 95% two-sided CI of 0.68–0.95. No P-value is provided in the ClinicalTrials.gov record.

Secondary PFS

The PFS HR was 1.40 in all randomized participants and 1.00 in randomized PD-L1 CPS ≥20 participants. The registry supplies confidence intervals but no formal P-values for these analyses.

One of the most important statistical lessons is that these estimates should be read as a collection rather than as isolated numbers. The analysis population, endpoint definition, effect measure, confidence interval, and hypothesis-testing framework all determine what a reported result means.

12. Confidence Intervals Across the Reported Analyses

EndpointHRConfidence intervalWhat the interval indicates
Primary OS, PD-L1 CPS ≥20 0.78 97.51%: 0.59–1.03 Includes 1.00; considerable uncertainty remains around the point estimate.
Primary OS, all randomized 0.95 97.9%: 0.80–1.13 Includes 1.00; the interval permits both lower and higher estimated hazards.
Secondary OS, PD-L1 CPS ≥1 0.80 95%: 0.68–0.95 Does not include 1.00.
Secondary PFS, all randomized 1.40 95%: 1.19–1.63 Does not include 1.00 and is above 1.00 throughout the interval.
Secondary PFS, PD-L1 CPS ≥20 1.00 95%: 0.77–1.30 Centered on 1.00 and includes both lower and higher hazards.
Post-hoc OS, PD-L1 CPS ≥20 extended collection 0.76 95%: 0.60–0.97 Does not include 1.00.
Post-hoc OS, all randomized extended collection 0.94 95%: 0.82–1.08 Includes 1.00.

This table also illustrates why confidence intervals should accompany hazard ratios. Two HRs can look different numerically while their uncertainty ranges provide a more complete picture of what can reasonably be inferred from the analysis.

13. Primary Endpoint Analysis vs Later Analyses

The ClinicalTrials.gov record contains three distinct categories of efficacy analyses: registered primary endpoints, secondary endpoints, and post-hoc extended-collection analyses. Keeping those categories separate is important for statistical interpretation.

Analysis categoryExamples in CheckMate-651Interpretive role
Primary OS in PD-L1 CPS ≥20; OS in all randomized participants Prespecified primary questions under a superiority framework.
Secondary OS in PD-L1 CPS ≥1; PFS Additional efficacy questions reported separately from the primary endpoints.
Post-hoc Extended-collection OS analyses Later analyses that should not be relabeled as primary confirmatory endpoints.

The ClinicalTrials.gov record does not report an interim-analysis procedure, alpha-spending strategy, multiplicity-adjustment scheme, crossover analysis, missing-data imputation method, or Bayesian method. Those topics are therefore not incorporated into the statistical claims on this page.

14. Why the Hazard Ratio Should Not Be Read as a Survival Percentage

Relative effect versus absolute outcome

A hazard ratio is a relative time-to-event measure. For example, HR 0.78 means a lower estimated event hazard under the reported model. It does not tell the reader what percentage of participants were alive at a particular time point, how many additional months an individual participant would live, or what proportion of participants were cured.

Those questions require different summaries, such as survival probabilities at specified times or median survival, when those quantities are reported. The registry-reported CheckMate-651 registry data do not provide those additional numerical summaries, so they are not added here.

Why the PFS HR is a different quantity

The PFS event is disease progression or death. Consequently, an HR for PFS describes the relative hazard of reaching either component of that composite endpoint. It should not be interpreted as a hazard ratio for death alone.

Why analysis populations matter

An estimate from all randomized participants and an estimate from randomized participants meeting a PD-L1 CPS threshold answer different questions. A numerical difference between them is descriptive. Establishing that treatment effect genuinely differs between populations requires an appropriate formal assessment of interaction or heterogeneity, which is not reported in the ClinicalTrials.gov record.

15. Limitations

16. Why This Trial Matters Statistically

CheckMate-651 is a useful teaching example because its registry results demonstrate how the same randomized comparison can produce different statistical summaries depending on the endpoint and analysis population.

ConceptHow it appears in CheckMate-651
RandomizationRandomized allocation to two parallel treatment arms.
Time-to-event endpointsBoth primary endpoints are overall-survival measures; PFS is a secondary time-to-event endpoint.
Log-rank testingThe primary analyses use a stratified regular log-rank test.
Hazard ratioPrimary and secondary efficacy analyses use hazard ratios as effect measures.
Confidence intervalsReported around every posted efficacy hazard-ratio estimate in the analyses posted on ClinicalTrials.gov.
Analysis populationsResults are reported for all randomized participants and PD-L1 CPS-defined populations.
SuperiorityThe two primary endpoint analyses are classified as superiority hypotheses.
Post-hoc analysisExtended-collection OS analyses are separately identified as post-hoc.
CensoringThe OS definition specifies censoring at the last date a participant was known to be alive when death is undocumented.
Safety analysisSerious adverse events are reported by randomized arm using affected and at-risk counts.

The particularly useful statistical lesson is that one number is never the whole analysis. The HR must be read together with its endpoint definition, analysis population, confidence interval, P-value when reported, and the statistical role of the analysis.

17. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

18. Related Statistical Calculators

19. Sources

Continue through the Clinical Biostats statistical pathway

Use the trial's survival endpoints and reported hazard ratios as a practical starting point for learning time-to-event analysis, confidence intervals, log-rank testing, and related clinical-trial methods.

20. Record Summary

CheckMate-651 provides a useful example of how randomized oncology-trial results should be interpreted across multiple time-to-event endpoints and analysis populations. The two registered primary endpoints were overall-survival analyses in participants with PD-L1 CPS ≥20 and in all randomized participants. Their reported hazard ratios were 0.78 and 0.95, respectively, with different confidence levels and P-values. Secondary analyses added overall survival in participants with PD-L1 CPS ≥1 and progression-free survival, while the registry also reports post-hoc extended-collection overall-survival analyses.

The statistical story is therefore more nuanced than any single hazard ratio. The primary analyses use stratified log-rank testing, hazard ratios quantify relative event hazards, confidence intervals describe uncertainty, and the analysis population determines the population to which each estimate applies. The reported P-values address hypothesis testing rather than effect magnitude. The safety results provide a separate perspective through serious-adverse-event counts by treatment arm.

Clinical Biostats methodology: A trial-results page should distinguish the registered endpoint, analysis population, statistical method, effect measure, uncertainty, and hypothesis-testing result. When the registry does not report a quantity, this page does not manufacture one from other sources or from statistical reconstruction.