← Clinical Trials
Early-Stage NSCLC Phase 3 Completed NCT02998528

CheckMate-816: Complete Statistical Analysis of Nivolumab Plus Chemotherapy in Early-Stage NSCLC

An independent statistical review of the randomized phase 3 CheckMate-816 trial evaluating nivolumab plus platinum-doublet chemotherapy versus platinum-doublet chemotherapy alone in participants with early-stage non-small cell lung cancer.

Trial start: 2017-03-04  ·  Primary completion: 2021-09-08  ·  Enrollment: 505
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the trial data posted on ClinicalTrials.gov for CheckMate-816. The registry provides the official trial record and the statistical analyses reported here.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CheckMate-816 was a completed, randomized, open-label, parallel-group phase 3 trial in non-small cell lung cancer. The trial enrolled 505 participants across three arms. The posted formal primary analyses reported here compare the concurrently randomized nivolumab-plus-platinum-doublet chemotherapy arm with the platinum-doublet chemotherapy arm.

505
Enrollment
Randomized trial
3
Arms
Parallel design
0.63
Primary EFS HR
97.38% CI 0.43–0.91
21.6
pCR Difference
99% CI 13.0–30.3
FeatureCheckMate-816
Trial nameCheckMate-816
PhasePhase 3
ConditionNon Small Cell Lung Cancer
Brief titleA Neoadjuvant Study of Nivolumab Plus Ipilimumab or Nivolumab Plus Chemotherapy Versus Chemotherapy Alone in Early Stage Non-Small Cell Lung Cancer (NSCLC)
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment505
Arms3
StatusCompleted
Lead sponsorBristol-Myers Squibb
Sponsor typeIndustry
Start2017-03-04
Primary completion2021-09-08
ClinicalTrials.govNCT02998528

2. Clinical Question

The trial addresses whether adding nivolumab to platinum-doublet chemotherapy improves important pathological and time-to-event outcomes compared with platinum-doublet chemotherapy alone in participants with early-stage non-small cell lung cancer.

Population

Participants with non-small cell lung cancer enrolled in the phase 3 randomized trial of neoadjuvant nivolumab-containing strategies and chemotherapy.

Intervention of interest

Nivolumab 360 mg plus platinum-doublet chemotherapy, identified as Arm C in the posted primary analyses.

Comparator

Platinum Doublet Chemo, identified as Arm B in the posted primary analyses.

Primary question

Among concurrently randomized participants in Arms C and B, does nivolumab plus platinum-doublet chemotherapy improve event-free survival and pathologic complete response relative to chemotherapy alone?

The broader trial included a third arm, Arm A, described in the registry ClinicalTrials.gov record as nivolumab 3 mg/kg plus ipilimumab 1 mg/kg. The formal primary statistical analyses in the ClinicalTrials.gov record compare Arm C with concurrent Arm B; therefore, the numerical efficacy results below are not generalized to Arm A.

3. Trial Design

01
Randomize 505 participants
02
Three arms Parallel randomized design
03
Study treatment Nivolumab-containing or chemotherapy strategy
04
Assess endpoints EFS and pathological response
05
Follow-up Time-to-event outcomes
ARM A

Nivolumab + ipilimumab

  • Nivolumab 3 mg/kg
  • Ipilimumab 1 mg/kg
ARM B · COMPARATOR

Platinum Doublet Chemo

  • Platinum doublet chemotherapy
  • Used as the concurrent comparator in the posted primary analyses
ARM C · PRIMARY COMPARISON

Nivolumab + Platinum Doublet Chemo

  • Nivolumab 360 mg
  • Platinum doublet chemotherapy
  • Compared with concurrent Arm B in the posted analyses

The trial was unmasked, or open-label, according to the registry design profile. That feature is important statistically because treatment assignment was not concealed from participants or investigators during treatment. For the pathological complete response endpoint, however, the registry defines assessment using blinded independent pathological review; the EFS definition also refers to blinded independent central review for the post-surgery disease assessment.

Interventions represented in the ClinicalTrials.gov record

The ClinicalTrials.gov record lists nivolumab and ipilimumab as biological interventions and cisplatin, vinorelbine, gemcitabine, docetaxel, pemetrexed, carboplatin, and paclitaxel as drug interventions.

4. Randomization, Stratification, and Analysis Populations

The trial was randomized with a parallel design. The formal efficacy analyses in the ClinicalTrials.gov record use all concurrently randomized participants in Arm C and Arm B. This is a particularly important detail because the overall enrollment of 505 participants and the analysis population used for a specific endpoint are not interchangeable quantities.

Analysis featureRegistry-supported description
Overall enrollment505 participants
Overall designRandomized, parallel
Primary EFS analysis populationAll concurrently randomized participants in Arm C and Arm B
Primary pCR analysis populationAll concurrently randomized participants in Arm C and Arm B
EFS analysisStratified log-rank framework with Cox proportional-hazard effect measure
pCR analysisStratified Cochran-Mantel-Haenszel analysis

The registry analysis notes identify three stratification variables for the reported Arm C versus Arm B analyses: PD-L1 status, disease stage, and sex. Specifically, the strata were PD-L1 at ≥1% versus <1%/unevaluable/indeterminate, disease stage IB/II versus IIIA, and male versus female.

Why this matters: the reported effect estimates are not simple unadjusted comparisons. The EFS hazard ratio and the pCR treatment effects were generated within analyses that account for the prespecified or reported stratification structure. When interpreting the numbers, it is therefore more accurate to describe them as stratified treatment effects rather than as crude differences between the two arms.

5. Primary Endpoints

EndpointRegistered definition / time frameStatistical type
Event-Free Survival (EFS) From randomization to disease progression, reoccurrence, or death due to any cause. (Up to a median of 30 months) Time-to-event
Pathologic Complete Response (pCR) Rate From randomization up to a median of 30 months after randomization. Binary

Event-Free Survival

The registry defines EFS as the length of time from randomization to any of the following events: any progression of disease precluding surgery, progression or recurrence disease based on blinded independent central review (BICR) assessment per response evaluation criteria in solid tumors (RECIST) 1.1 after surgery, or death due to any cause.

The registry definition continues beyond the text available in the trial data, so this page does not add further event or censoring rules that are not explicitly provided in the source data.

Pathologic Complete Response

Pathologic complete response is defined as the number of randomized participants with absence of residual tumor in lung and lymph nodes as evaluated by blinded independent pathological review (BIPR).

Endpoint distinction: EFS and pCR are fundamentally different statistical outcomes. EFS records the time until an event and therefore incorporates both event timing and censoring. pCR is binary: participants are classified according to whether the defined absence of residual tumor was observed. A hazard ratio is therefore appropriate for EFS, whereas a difference in response rates and an odds ratio are natural effect measures for pCR.

6. Statistical Methodology

Log-rank test for EFS

The primary EFS comparison between Arm B and Arm C used a log-rank test. The registry-reported analysis identifies the method as a stratified analysis. The analysis was stratified by PD-L1 status, disease stage, and sex.

The log-rank framework compares the observed and expected pattern of events between randomized treatment groups across follow-up. It is therefore a method for comparing the entire observed time-to-event experience rather than simply comparing the proportion of participants who have experienced an event at one fixed time point.

Cox proportional-hazard model

The EFS effect measure was reported as a Cox proportional hazard estimate, normalized here as a hazard ratio. The primary estimate was 0.63 with a two-sided 97.38% confidence interval of 0.43 to 0.91.

Hazard-ratio interpretation
HR = instantaneous event rate in Arm C relative to Arm B, under the fitted Cox model

An HR below 1 indicates a lower estimated instantaneous event rate in the nivolumab-plus-chemotherapy arm than in the chemotherapy comparator under the fitted model. It does not directly report an absolute probability of remaining event-free.

Cochran-Mantel-Haenszel test for pCR

The primary pCR analysis used the Cochran-Mantel-Haenszel (CMH) test. This approach is appropriate when a binary outcome is compared between treatment groups while accounting for categorical strata. Here, the reported strata were PD-L1 status, disease stage, and sex.

The registry reports two complementary pCR effect measures. The first is a strata-adjusted percentage difference, defined as the Arm C minus concurrent Arm B difference based on CMH weighting. The second is a strata-adjusted odds ratio for Arm C over concurrent Arm B.

Confidence intervals

The confidence interval attached to an effect estimate describes the statistical uncertainty around the estimated treatment effect under the corresponding analysis framework. A confidence interval is not a prediction interval for individual participants, and it does not describe the range of biological responses that individual patients might experience.

Superiority framework

All of the registry-reported formal primary and secondary comparisons are identified as superiority hypotheses. The interpretation therefore concerns evidence that the nivolumab-containing comparison differs favorably from the chemotherapy comparator under the specified endpoint and effect measure; it is not a non-inferiority framework with a prespecified acceptable loss of efficacy.

7. Primary Results: Event-Free Survival

The primary EFS analysis compared all concurrently randomized participants in Arm C, nivolumab 360 mg plus platinum-doublet chemotherapy, with Arm B, platinum-doublet chemotherapy alone. The reported analysis used a stratified log-rank test and a Cox proportional-hazard effect measure.

Primary EFS hazard ratio

0.63

97.38% two-sided CI: 0.43–0.91   ·   P = 0.0052

Analysis: stratified log-rank test with Cox proportional-hazard effect measure.

Primary EFS analysisReported value
EndpointEvent-Free Survival (EFS)
ComparisonArm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo
Analysis populationAll concurrently randomized participants in Arm C and Arm B
MethodLog Rank
Effect measureCox Proportional Hazard / Hazard ratio
Estimate0.63
Confidence interval97.38% two-sided CI 0.43–0.91
P-value0.0052
HypothesisSuperiority
StratificationPD-L1 (≥1% vs <1%/unevaluable/indeterminate); disease stage (IB/II vs IIIA); sex (male vs female)
Clinical Biostats interpretation

The hazard ratio of 0.63 means that, under the fitted Cox model, the estimated instantaneous rate of an EFS event in Arm C was approximately 63% of the estimated rate in Arm B. Equivalently, 1 − 0.63 = 0.37, so the estimated hazard was approximately 37% lower in Arm C than in Arm B.

This does not mean that 37% of participants avoided an event, that every participant experienced a 37% reduction in risk, or that the absolute probability of remaining event-free was 37 percentage points higher. The hazard ratio is a relative, model-based time-to-event measure.

The 97.38% confidence interval from 0.43 to 0.91 communicates uncertainty around the estimated hazard ratio. The interval is compatible with a range of relative effects rather than a single exact treatment effect. It does not describe the range of individual patient outcomes.

The reported P = 0.0052 addresses evidence against the relevant null hypothesis under the trial's statistical testing framework. A p-value does not measure the magnitude of benefit, the probability that the treatment works, or the probability that the null hypothesis is true.

Because the effect measure is a Cox hazard ratio, interpretation also depends on the proportional-hazards framework underlying the model. A single hazard ratio summarizes relative event rates over the analyzed follow-up; it should not automatically be read as a constant individual-level risk reduction at every time point.

Why the stratification matters

The EFS analysis did not ignore the three reported stratification factors. PD-L1 status, disease stage, and sex were incorporated into the stratified analysis. This matters because randomization and the analysis model can be coordinated through the same clinically relevant strata, improving alignment between the treatment comparison and the trial's design.

The analysis therefore answers a more specific question than "what proportion had an EFS event in each arm?" It estimates the relative time-to-event treatment effect while respecting the stated stratification structure.

8. Primary Results: Pathologic Complete Response

The second primary endpoint was pCR rate. The primary comparison again used all concurrently randomized participants in Arms C and B. The registry reports a CMH analysis with two complementary effect measures: a strata-adjusted percentage difference and a strata-adjusted odds ratio.

Primary pCR rate difference

21.6

99% two-sided CI: 13.0–30.3

Strata-adjusted difference, Arm C minus concurrent Arm B, using CMH weighting.

Primary pCR analysisReported value
EndpointPathologic Complete Response (pCR) Rate
ComparisonArm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo
Analysis populationAll concurrently randomized participants in Arm C and Arm B
MethodCochran-Mantel-Haenszel test
Effect measure% Difference / Other difference
Estimate21.6
Confidence interval99% two-sided CI 13.0–30.3
HypothesisSuperiority
Interpretation of differenceStrata-adjusted difference, Arm C minus concurrent Arm B, based on CMH weighting
Clinical Biostats interpretation

The reported pCR difference of 21.6 is a strata-adjusted difference in percentage points between Arm C and concurrent Arm B, with the direction defined as Arm C minus Arm B. Thus, the estimate describes a 21.6-percentage-point difference in the pCR rate after the CMH stratification structure is taken into account.

This does not mean that every participant in Arm C had a 21.6% greater chance of pCR, nor does it mean that pCR translates one-for-one into longer survival. pCR is a pathological binary endpoint; EFS is a subsequent time-to-event endpoint. They answer different statistical questions.

The 99% confidence interval of 13.0 to 30.3 describes uncertainty around the strata-adjusted percentage difference. It provides a range of plausible values for the treatment effect under the stated statistical framework, rather than a range of individual responses.

No p-value is posted on ClinicalTrials.gov for this particular percentage-difference analysis in the trial data. The absence of a posted p-value does not justify calculating one independently from the confidence interval. The reported estimate and confidence interval should therefore be presented as reported.

Primary pCR odds ratio

Strata-adjusted odds ratio

13.94

99% two-sided CI: 3.49–55.75   ·   P < 0.0001

Odds ratio for Arm C over concurrent Arm B using the Mantel-Haenszel method.

pCR odds-ratio analysisReported value
MethodCochran-Mantel-Haenszel
Effect measureOdds Ratio (OR)
Estimate13.94
Confidence interval99% two-sided CI 3.49–55.75
P-value<0.0001
DirectionArm C over concurrent Arm B
HypothesisSuperiority
Clinical Biostats interpretation

An odds ratio of 13.94 means that the estimated odds of pCR in Arm C were 13.94 times the estimated odds in concurrent Arm B under the reported strata-adjusted analysis. Odds are not probabilities, so the odds ratio should not be translated directly into a 13.94-fold increase in the pCR percentage.

The 99% confidence interval of 3.49 to 55.75 is wide relative to the point estimate. This is important: the analysis indicates a large estimated difference in odds, but the precision of that estimate is limited enough that a broad range of effect magnitudes remains compatible with the data.

The reported P < 0.0001 provides evidence against the null hypothesis under the stated testing framework. It does not measure the size of the treatment effect and should not be substituted for the effect estimate or its confidence interval.

The pCR odds ratio and the pCR percentage difference are complementary rather than interchangeable. The percentage difference is easier to translate into an absolute difference in response rates, whereas the odds ratio expresses the relative odds of response and is naturally suited to a CMH analysis across strata.

9. Secondary Endpoint Results

Major Pathologic Response Rate

Major Pathologic Response (MPR) Rate was reported as a secondary binary endpoint for the same concurrently randomized Arm C versus Arm B comparison. The registry reports a strata-adjusted percentage difference using CMH weighting.

MPR rate difference

27.9

95% two-sided CI: 19.6–36.1

Strata-adjusted difference, Arm C minus concurrent Arm B.

Secondary MPR analysisReported value
EndpointMajor Pathologic Response (MPR) Rate
ComparisonArm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo
Analysis populationAll concurrently randomized participants in Arm C and Arm B
MethodCochran-Mantel-Haenszel test
Effect measure% Difference
Estimate27.9
Confidence interval95% two-sided CI 19.6–36.1
HypothesisSuperiority
StratificationPD-L1 status; disease stage; sex

Because the registry reports this as a strata-adjusted percentage difference, the estimate is most naturally read as a 27.9-percentage-point difference in MPR rate in the direction Arm C minus Arm B. The confidence interval indicates uncertainty around that adjusted difference.

MPR Rate Odds Ratio

A second posted MPR analysis reports an odds ratio, but the ClinicalTrials.gov record identifies the method as not reported for this particular analysis. The analysis is nevertheless identified as stratified by PD-L1 status, disease stage, and sex.

MeasureReported value
EndpointMajor Pathologic Response (MPR) Rate
Effect measureOdds Ratio (OR)
Estimate5.70
95% two-sided CI3.16–10.26
MethodNot reported in the ClinicalTrials.gov record
HypothesisSuperiority

The odds ratio of 5.70 indicates substantially higher estimated odds of MPR in Arm C than in Arm B under the reported analysis. It should not be interpreted as a 5.70-fold increase in the percentage of participants with MPR, because odds and probabilities are different quantities.

Time to Death or Distant Metastases

TTDM hazard ratio

0.66

95% two-sided CI: 0.49–0.90

Stratified analysis; method not reported in the ClinicalTrials.gov record.

TTDM analysisReported value
EndpointTime to Death or Distant Metastases (TTDM)
ComparisonArm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo
Analysis populationAll concurrently randomized participants in Arm C and Arm B
Effect measureHazard Ratio (HR)
Estimate0.66
95% CI0.49–0.90, two-sided
MethodNot reported in the ClinicalTrials.gov record
StratificationPD-L1 status; disease stage; sex
HypothesisSuperiority

An HR of 0.66 corresponds to an estimated instantaneous TTDM event rate approximately 34% lower in Arm C than Arm B under the fitted hazard-ratio interpretation. The confidence interval, 0.49 to 0.90, expresses the uncertainty around that estimate. No p-value is posted on ClinicalTrials.gov for this secondary analysis, so no additional significance calculation is presented here.

10. Post-Hoc Event-Free Survival Analysis

The ClinicalTrials.gov record also contain a post-hoc EFS analysis using a longer time frame, up to a median of 69 months. It compares the same concurrently randomized Arm C and Arm B populations.

Post-hoc EFS hazard ratio

0.68

95% two-sided CI: 0.51–0.91

Time frame: from randomization to disease progression, reoccurrence, or death due to any cause, up to a median of 69 months.

FeaturePost-hoc EFS analysis
EndpointEvent-Free Survival (EFS)
Analysis rolePost-hoc
ComparisonArm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo
Analysis populationAll concurrently randomized participants in Arm C and Arm B
Effect measureCox Proportional Hazard / Hazard ratio
Estimate0.68
95% CI0.51–0.91, two-sided
MethodNot reported in the registry-reported post-hoc analysis record
StratificationPD-L1 (≥1% vs <1%/unevaluable/indeterminate); disease stage (IB/II vs IIIA); sex (male vs female)
Clinical Biostats interpretation

The post-hoc HR of 0.68 indicates an estimated instantaneous EFS event rate approximately 32% lower in Arm C than Arm B under the Cox hazard-ratio interpretation. The estimate is directionally consistent with the primary EFS result of 0.63, although the two estimates correspond to different follow-up descriptions and the later analysis is explicitly classified as post-hoc.

The 95% CI of 0.51–0.91 indicates uncertainty around the longer-follow-up estimate. It should not be interpreted as evidence that the true treatment effect must remain constant throughout follow-up.

Because the analysis is labeled Post_Hoc and no formal method or p-value is reported in the ClinicalTrials.gov record, it should not be silently treated as another prespecified confirmatory primary analysis. Its most appropriate role on this page is descriptive: it shows the reported longer-term EFS effect estimate while preserving the distinction between the primary and post-hoc analyses.

11. Safety: Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants over participants at risk. These data provide an arm-specific safety summary without requiring an inference about events that are not listed in the registry extract.

ArmSerious adverse events, affected / at riskReported proportion
Arm A: Nivo 3 mg/kg + Ipi 1 mg/kg 30 / 111 30/111
Arm B: Platinum Doublet Chemo 58 / 208 58/208
Arm C: Nivo 360 mg + Platinum Doublet Ch 53 / 176 53/176

The serious-adverse-event counts should be kept separate from the efficacy analysis populations and from the primary endpoint results. Safety is exposure-related information, whereas the primary efficacy comparisons are defined by randomized treatment assignment and the endpoint-specific analysis populations.

Safety interpretation: the ClinicalTrials.gov record reports serious adverse events as affected participants divided by participants at risk. They do not provide a formal statistical comparison, confidence interval, or p-value for serious adverse events. This page therefore does not construct an unreported comparative hypothesis test.

12. Statistical Methods Explained

Why was a stratified log-rank test used for EFS?

EFS is a time-to-event endpoint, so simply comparing the number of events between treatment groups would discard information about when those events occurred and about participants who were censored before an event. The log-rank test is designed to compare event-time distributions while using the available follow-up information.

The reported analysis was stratified by PD-L1 status, disease stage, and sex. Stratification allows the comparison to account for these categorical factors rather than treating all participants as though the trial contained no such structure.

What does an EFS hazard ratio of 0.63 mean?

An HR of 0.63 means the estimated instantaneous rate of an EFS event in Arm C was approximately 63% of the estimated rate in Arm B under the fitted Cox model. The corresponding relative reduction in the estimated hazard is approximately 37%.

The important word is hazard. A hazard ratio is not the same as a risk ratio, an odds ratio, a percentage-point difference, or a probability of being event-free at a particular time.

Why is the pCR analysis different from the EFS analysis?

pCR is binary, while EFS is time-to-event. For pCR, the statistical question is whether the proportion with the defined pathological response differs between groups after accounting for the analysis strata. The CMH method is designed for this type of stratified categorical comparison.

For EFS, the analysis must account for event timing and censoring, which is why the trial uses a log-rank framework and a Cox hazard-ratio effect measure.

What does a pCR odds ratio of 13.94 mean?

The reported odds ratio is the estimated odds of pCR in Arm C divided by the corresponding odds in concurrent Arm B under the Mantel-Haenszel stratified analysis. An OR of 13.94 therefore means that the estimated odds are 13.94 times as high in Arm C.

It is important not to say that the "probability" or "pCR rate" is 13.94 times higher. The relationship between odds and probability is nonlinear, and the odds ratio cannot be converted to a probability ratio without the underlying event rates.

Why are there two pCR effect measures?

The percentage difference and odds ratio describe the same binary endpoint from different perspectives. The percentage difference is an absolute measure expressed in percentage points. The odds ratio is a relative measure on the odds scale.

Using both provides more information than using either alone. The difference communicates the magnitude of separation in response rates, while the odds ratio describes the relative odds after accounting for the trial's strata.

Why is the primary EFS confidence interval 97.38% rather than 95%?

The ClinicalTrials.gov record explicitly reports a 97.38% two-sided confidence interval for the primary EFS hazard ratio. This page preserves that value rather than replacing it with a conventional 95% interval. The different confidence level is part of the reported statistical framework and should not be silently standardized.

By contrast, the primary pCR odds-ratio and percentage-difference analyses report 99% two-sided confidence intervals. The secondary MPR and TTDM analyses report 95% two-sided intervals. Confidence-level differences therefore matter when comparing the apparent width of intervals across endpoints.

Why should the post-hoc EFS result not be treated like the primary EFS result?

The primary EFS analysis and the longer-term post-hoc EFS analysis are different statistical records. The primary analysis has a defined primary-endpoint role, a reported p-value, and a 97.38% confidence interval. The later analysis is explicitly classified as post-hoc, has no reported method in the ClinicalTrials.gov record, and provides a 95% confidence interval without a p-value.

Preserving that distinction prevents a later descriptive analysis from being mistaken for an additional prespecified confirmatory test.

13. Understanding the Confidence Intervals

EndpointEstimateConfidence intervalConfidence level
Primary EFSHR 0.630.43–0.9197.38%
Primary pCR difference21.613.0–30.399%
Primary pCR odds ratio13.943.49–55.7599%
MPR difference27.919.6–36.195%
MPR odds ratio5.703.16–10.2695%
TTDMHR 0.660.49–0.9095%
Post-hoc EFSHR 0.680.51–0.9195%

The table illustrates why confidence intervals should be interpreted together with the effect measure. A hazard-ratio interval is centered on the multiplicative hazard scale, whereas the pCR and MPR percentage-difference intervals are expressed on an absolute percentage-point scale.

A practical reading rule

First identify what is being estimated. Then identify the direction of the comparison, the confidence level, and the width of the interval. Only after those steps should a p-value be considered. This prevents a common statistical mistake: treating every number below 0.05 as though it were an effect-size measure.

14. P-Values and Effect Sizes

The CheckMate-816 statistical record demonstrates why p-values and effect estimates should be reported together.

AnalysisEffect estimateP-value reported?
Primary EFSHR 0.630.0052
Primary pCR percentage difference21.6Not reported
Primary pCR odds ratio13.94<0.0001
MPR percentage difference27.9Not reported
MPR odds ratio5.70Not reported
TTDM hazard ratio0.66Not reported
Post-hoc EFS hazard ratio0.68Not reported

A p-value answers a question about compatibility with a null hypothesis under a specified statistical model and testing framework. An effect estimate answers a different question: how large is the estimated difference or relative effect?

For example, the primary pCR result contains both an odds ratio of 13.94 and a p-value of <0.0001. The p-value supports the statistical evidence against the null, while the odds ratio and its 99% confidence interval communicate the magnitude and uncertainty of the estimated treatment effect.

15. Multiplicity, Interim Analysis, and Other Design Features

The ClinicalTrials.gov record identifies the formal hypotheses as superiority and identify seven statistical analyses in total, including three primary-endpoint analyses. However, the ClinicalTrials.gov record does not provide a detailed alpha-spending plan, interim-analysis schedule, multiplicity hierarchy, non-inferiority margin, missing-data imputation strategy, or Bayesian analysis specification.

Design topicWhat the ClinicalTrials.gov record supports
SuperiorityYes. The primary and secondary reported comparisons are identified as superiority hypotheses.
StratificationYes. PD-L1 status, disease stage, and sex are identified in the analysis notes.
Log-rank analysisYes. Used for the primary EFS analysis.
Cochran-Mantel-Haenszel analysisYes. Used for primary pCR and one secondary MPR analysis.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designNot reported; the registered design model is parallel.
Bayesian methodsNot reported.
Formal interim-analysis detailsNot reported in the ClinicalTrials.gov record.
Missing-data/imputation strategyNot reported in the ClinicalTrials.gov record.
Data-discipline note: absence of a field in the ClinicalTrials.gov record is not a reason to reconstruct it from memory or from a different publication. In particular, this page does not assign an unreported multiplicity procedure, interim boundary, imputation method, or crossover adjustment to CheckMate-816.

16. Why Stratified Analysis Matters

Three variables recur in the analysis notes: PD-L1 status, disease stage, and sex. They define strata for the reported treatment comparisons. Stratification is useful because it recognizes that the randomized comparison contains clinically meaningful categorical structure that can be incorporated into the analysis.

PD-L1 strata

The reported categories are ≥1% versus <1%/unevaluable/indeterminate.

Disease-stage strata

The reported categories are IB/II versus IIIA.

Sex strata

The reported categories are male versus female.

Why combine them?

CMH and stratified time-to-event methods can account for these categorical factors while estimating the treatment comparison across the defined strata.

Importantly, stratification does not turn the analysis into three separate trials, nor does it imply that each stratum has enough information to support an independent confirmatory conclusion. The ClinicalTrials.gov record is treatment comparisons across the stratified analysis structure.

17. Interpreting the Two Primary Endpoints Together

The two primary endpoints measure different dimensions of treatment effect. EFS asks about the time from randomization until progression, recurrence, or death according to the registered definition. pCR asks whether residual tumor is absent in the lung and lymph nodes on blinded independent pathological review.

DimensionEFSpCR
Outcome typeTime-to-eventBinary
Time componentYesNo; classified by pathological assessment
Primary effect measureHazard ratioPercentage difference and odds ratio
Primary estimate0.6321.6 percentage points; OR 13.94
Confidence level reported97.38%99%
Formal p-value reported0.0052<0.0001 for OR analysis
Statistical methodStratified log-rankStratified CMH

The important statistical point is that concordant evidence across different endpoint types can be more informative than treating either endpoint as a complete summary of the trial. The EFS analysis incorporates follow-up time and censoring. The pCR analysis gives a pathological response comparison. Neither endpoint is a substitute for the other.

18. What the Primary EFS Result Does — and Does Not — Mean

What it means

The primary EFS hazard ratio of 0.63 indicates a lower estimated instantaneous EFS event rate in Arm C than in Arm B under the reported Cox model. The estimate corresponds to approximately a 37% lower estimated hazard relative to Arm B.

What it does not mean

It does not mean that 37% of patients were protected from an EFS event, that each participant had exactly 37% less risk, or that the absolute difference in EFS probability at every time point was 37 percentage points.

Why the confidence interval matters

The 97.38% confidence interval of 0.43–0.91 shows that the point estimate is not the only statistically plausible value. The interval expresses uncertainty around the treatment effect under the analysis framework and should accompany the hazard ratio whenever the result is presented.

Why the p-value matters differently

The p-value of 0.0052 quantifies evidence against the relevant null hypothesis under the stated testing framework. It is not an effect-size measure and should not be interpreted as the probability that the treatment is effective.

19. What the Primary pCR Result Does — and Does Not — Mean

The primary pCR result is particularly useful for demonstrating the difference between an absolute effect measure and an odds ratio.

Percentage difference

The reported strata-adjusted difference is 21.6 percentage points, Arm C minus Arm B, with a 99% CI of 13.0–30.3.

Odds ratio

The reported strata-adjusted OR is 13.94, with a 99% CI of 3.49–55.75.

Statistical evidence

The p-value attached to the reported odds-ratio analysis is <0.0001.

Important caution

An odds ratio should not be read as a direct probability ratio. The absolute percentage-point difference is the more direct measure of separation in response rates.

The broad 99% confidence interval for the odds ratio is also instructive. The point estimate is 13.94, but the interval extends from 3.49 to 55.75. This illustrates why a statistically strong p-value does not imply that the effect estimate itself is known with arbitrary precision.

20. Long-Term EFS: Primary vs Post-Hoc Evidence

The trial data provide both a primary EFS analysis with a median time frame of up to 30 months and a post-hoc EFS analysis with a median time frame of up to 69 months.

FeaturePrimary EFSPost-hoc EFS
RolePrimaryPost-hoc
Time frameUp to a median of 30 monthsUp to a median of 69 months
EstimateHR 0.63HR 0.68
Confidence interval97.38% CI 0.43–0.9195% CI 0.51–0.91
P-value0.0052Not reported
Method reportedLog-rankNot reported
Effect measureCox proportional hazardCox proportional hazard

The two estimates, 0.63 and 0.68, are reasonably close in direction and magnitude, but the correct statistical description is not that one "confirms" the other. They belong to different analysis roles and confidence frameworks. The later result is best treated as a longer-term descriptive estimate because the ClinicalTrials.gov record classifies it as post-hoc and does not provide a formal p-value or method.

21. Trial Timeline

2017-03-04

Trial start

The registered trial start date is March 4, 2017.

Phase 3

Randomized parallel design

The study was registered as a randomized, parallel phase 3 treatment trial with three arms and no masking.

505 participants

Overall enrollment

the ClinicalTrials.gov record reports enrollment of 505 participants.

2021-09-08

Primary completion

The registered primary completion date is September 8, 2021.

Up to median 30 months

Primary endpoint results

The primary EFS and pCR analyses are reported for the registry time frames extending up to a median of 30 months.

Up to median 69 months

Post-hoc EFS analysis

The registry analysis record includes a later post-hoc EFS estimate using a time frame up to a median of 69 months.

22. Limitations

23. Why This Trial Matters Statistically

CheckMate-816 is a useful statistical teaching example because its the ClinicalTrials.gov record span two fundamentally different endpoint classes and several complementary effect measures. The same randomized comparison can therefore be examined through survival analysis, categorical-data methods, absolute differences, odds ratios, confidence intervals, and stratified inference.

Statistical conceptHow it appears in CheckMate-816
RandomizationThe trial is registered as randomized with a parallel design.
Multiple treatment armsThree arms were included in the ClinicalTrials.gov record.
Time-to-event analysisEFS and TTDM are time-to-event endpoints.
Log-rank testUsed for the primary EFS comparison.
Cox hazard ratioUsed as the reported EFS and post-hoc EFS effect measure.
Binary endpoint analysispCR and MPR are binary response endpoints.
Cochran-Mantel-Haenszel methodUsed for the primary pCR analysis and one MPR analysis.
Odds ratioReported for pCR and MPR.
Absolute differenceReported for pCR and MPR as percentage differences.
Stratified analysisPD-L1 status, disease stage, and sex are identified in the analysis notes.
Confidence intervalsReported at 97.38%, 99%, and 95% levels depending on the analysis.
Superiority testingThe registry-reported formal analyses are identified as superiority hypotheses.
Post-hoc analysisA later EFS analysis is explicitly classified as post-hoc.

The trial is therefore particularly useful for learning that "the statistical analysis" is not one calculation. Different endpoint types require different estimands and different methods. A time-to-event endpoint cannot simply be analyzed like a binary response, and an odds ratio cannot be interpreted as a percentage-point difference.

24. A Deeper Statistical Reading of the Primary Results

Relative versus absolute effects

The primary pCR analysis gives an absolute percentage difference of 21.6, while the same endpoint gives an odds ratio of 13.94. These numbers are not competing answers. They use different scales.

The absolute difference is generally easier to communicate when the question is "how far apart are the response rates?" The odds ratio answers a different question: "how much larger are the odds of response?" A careful statistical report preserves both scales rather than converting one into the other without the necessary underlying response rates.

Why the EFS result needs censoring concepts

Unlike pCR, EFS cannot be summarized solely by counting events. Some participants may have follow-up that ends before an EFS event occurs. Those observations contribute information until censoring rather than being treated as though an event occurred.

This is one reason Kaplan-Meier estimation and time-to-event methods are central to EFS analysis. The ClinicalTrials.gov record specifically reports a log-rank analysis and a Cox proportional-hazard effect measure, which are methods designed for this type of endpoint.

Why the endpoint definition matters

The registered EFS definition includes disease progression precluding surgery, progression or recurrence based on BICR assessment after surgery, or death due to any cause. Consequently, the endpoint is not simply "time until death" or "time until radiographic progression." Its event definition determines exactly what the hazard ratio is comparing.

Why the analysis population matters

The primary analyses use all concurrently randomized participants in Arms C and B. This matters because the trial as a whole has three arms and 505 enrolled participants. The number 505 describes the trial's overall enrollment, whereas the formal efficacy estimates belong to the specified Arm C versus Arm B analysis population.

25. Statistical Concepts in the Learning Pathway

Learn more about the methods used in this trial:

26. Related Statistical Calculators

Use these calculator pathways to explore the statistical quantities that appear in the CheckMate-816 analyses:

27. Sources

The numerical trial results on this page are restricted to the registry-reported CheckMate-816 trial data. The PubMed links provide publication navigation but are not used here to introduce additional numerical results beyond the ClinicalTrials.gov recordset.

Continue through the Clinical Biostats statistical pathway

Explore the survival-analysis, categorical-data, confidence-interval, and clinical-trial methods that underlie the CheckMate-816 statistical results.

28. Record Summary

CheckMate-816 provides a useful example of how different endpoint types require different statistical analyses. The primary EFS endpoint is a time-to-event outcome analyzed with a stratified log-rank framework and a Cox proportional-hazard effect measure, producing an HR of 0.63 with a 97.38% two-sided confidence interval of 0.43–0.91 and a reported P = 0.0052. The primary pCR endpoint is binary and was analyzed with the Cochran-Mantel-Haenszel method, producing a strata-adjusted percentage difference of 21.6 with a 99% two-sided confidence interval of 13.0–30.3, together with a strata-adjusted odds ratio of 13.94 with a 99% confidence interval of 3.49–55.75 and P < 0.0001.

The secondary analyses extend the statistical picture without changing the primary endpoint definitions. MPR was associated with a reported percentage difference of 27.9 and an odds ratio of 5.70, while TTDM had a reported hazard ratio of 0.66. A longer-term post-hoc EFS analysis reported an HR of 0.68. Each result has to be interpreted according to its own endpoint, analysis role, effect measure, confidence level, and reported statistical method.

The broader statistical lesson is that a clinical-trial result is not adequately represented by a single p-value. A rigorous analysis identifies the randomized comparison, endpoint definition, analysis population, statistical method, effect measure, confidence interval, and testing role. For CheckMate-816, that framework is especially important because the trial contains three arms while the posted primary efficacy comparisons in the ClinicalTrials.gov record concern Arms C and B, and because the two primary endpoints require fundamentally different statistical approaches.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation, preserve the endpoint definitions and analysis roles reported by the registry, and avoid filling gaps with unsupported assumptions. The purpose is to make the statistical structure of the trial understandable without changing what the reported data actually show.