This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the trial data posted on ClinicalTrials.gov for CheckMate-816. The registry provides the official trial record and the statistical analyses reported here.
1. Trial at a Glance
CheckMate-816 was a completed, randomized, open-label, parallel-group phase 3 trial in non-small cell lung cancer. The trial enrolled 505 participants across three arms. The posted formal primary analyses reported here compare the concurrently randomized nivolumab-plus-platinum-doublet chemotherapy arm with the platinum-doublet chemotherapy arm.
| Feature | CheckMate-816 |
|---|---|
| Trial name | CheckMate-816 |
| Phase | Phase 3 |
| Condition | Non Small Cell Lung Cancer |
| Brief title | A Neoadjuvant Study of Nivolumab Plus Ipilimumab or Nivolumab Plus Chemotherapy Versus Chemotherapy Alone in Early Stage Non-Small Cell Lung Cancer (NSCLC) |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 505 |
| Arms | 3 |
| Status | Completed |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
| Start | 2017-03-04 |
| Primary completion | 2021-09-08 |
| ClinicalTrials.gov | NCT02998528 |
2. Clinical Question
The trial addresses whether adding nivolumab to platinum-doublet chemotherapy improves important pathological and time-to-event outcomes compared with platinum-doublet chemotherapy alone in participants with early-stage non-small cell lung cancer.
Population
Participants with non-small cell lung cancer enrolled in the phase 3 randomized trial of neoadjuvant nivolumab-containing strategies and chemotherapy.
Intervention of interest
Nivolumab 360 mg plus platinum-doublet chemotherapy, identified as Arm C in the posted primary analyses.
Comparator
Platinum Doublet Chemo, identified as Arm B in the posted primary analyses.
Primary question
Among concurrently randomized participants in Arms C and B, does nivolumab plus platinum-doublet chemotherapy improve event-free survival and pathologic complete response relative to chemotherapy alone?
The broader trial included a third arm, Arm A, described in the registry ClinicalTrials.gov record as nivolumab 3 mg/kg plus ipilimumab 1 mg/kg. The formal primary statistical analyses in the ClinicalTrials.gov record compare Arm C with concurrent Arm B; therefore, the numerical efficacy results below are not generalized to Arm A.
3. Trial Design
Nivolumab + ipilimumab
- Nivolumab 3 mg/kg
- Ipilimumab 1 mg/kg
Platinum Doublet Chemo
- Platinum doublet chemotherapy
- Used as the concurrent comparator in the posted primary analyses
Nivolumab + Platinum Doublet Chemo
- Nivolumab 360 mg
- Platinum doublet chemotherapy
- Compared with concurrent Arm B in the posted analyses
The trial was unmasked, or open-label, according to the registry design profile. That feature is important statistically because treatment assignment was not concealed from participants or investigators during treatment. For the pathological complete response endpoint, however, the registry defines assessment using blinded independent pathological review; the EFS definition also refers to blinded independent central review for the post-surgery disease assessment.
Interventions represented in the ClinicalTrials.gov record
The ClinicalTrials.gov record lists nivolumab and ipilimumab as biological interventions and cisplatin, vinorelbine, gemcitabine, docetaxel, pemetrexed, carboplatin, and paclitaxel as drug interventions.
4. Randomization, Stratification, and Analysis Populations
The trial was randomized with a parallel design. The formal efficacy analyses in the ClinicalTrials.gov record use all concurrently randomized participants in Arm C and Arm B. This is a particularly important detail because the overall enrollment of 505 participants and the analysis population used for a specific endpoint are not interchangeable quantities.
| Analysis feature | Registry-supported description |
|---|---|
| Overall enrollment | 505 participants |
| Overall design | Randomized, parallel |
| Primary EFS analysis population | All concurrently randomized participants in Arm C and Arm B |
| Primary pCR analysis population | All concurrently randomized participants in Arm C and Arm B |
| EFS analysis | Stratified log-rank framework with Cox proportional-hazard effect measure |
| pCR analysis | Stratified Cochran-Mantel-Haenszel analysis |
The registry analysis notes identify three stratification variables for the reported Arm C versus Arm B analyses: PD-L1 status, disease stage, and sex. Specifically, the strata were PD-L1 at ≥1% versus <1%/unevaluable/indeterminate, disease stage IB/II versus IIIA, and male versus female.
5. Primary Endpoints
| Endpoint | Registered definition / time frame | Statistical type |
|---|---|---|
| Event-Free Survival (EFS) | From randomization to disease progression, reoccurrence, or death due to any cause. (Up to a median of 30 months) | Time-to-event |
| Pathologic Complete Response (pCR) Rate | From randomization up to a median of 30 months after randomization. | Binary |
Event-Free Survival
The registry defines EFS as the length of time from randomization to any of the following events: any progression of disease precluding surgery, progression or recurrence disease based on blinded independent central review (BICR) assessment per response evaluation criteria in solid tumors (RECIST) 1.1 after surgery, or death due to any cause.
The registry definition continues beyond the text available in the trial data, so this page does not add further event or censoring rules that are not explicitly provided in the source data.
Pathologic Complete Response
Pathologic complete response is defined as the number of randomized participants with absence of residual tumor in lung and lymph nodes as evaluated by blinded independent pathological review (BIPR).
6. Statistical Methodology
Log-rank test for EFS
The primary EFS comparison between Arm B and Arm C used a log-rank test. The registry-reported analysis identifies the method as a stratified analysis. The analysis was stratified by PD-L1 status, disease stage, and sex.
The log-rank framework compares the observed and expected pattern of events between randomized treatment groups across follow-up. It is therefore a method for comparing the entire observed time-to-event experience rather than simply comparing the proportion of participants who have experienced an event at one fixed time point.
Cox proportional-hazard model
The EFS effect measure was reported as a Cox proportional hazard estimate, normalized here as a hazard ratio. The primary estimate was 0.63 with a two-sided 97.38% confidence interval of 0.43 to 0.91.
An HR below 1 indicates a lower estimated instantaneous event rate in the nivolumab-plus-chemotherapy arm than in the chemotherapy comparator under the fitted model. It does not directly report an absolute probability of remaining event-free.
Cochran-Mantel-Haenszel test for pCR
The primary pCR analysis used the Cochran-Mantel-Haenszel (CMH) test. This approach is appropriate when a binary outcome is compared between treatment groups while accounting for categorical strata. Here, the reported strata were PD-L1 status, disease stage, and sex.
The registry reports two complementary pCR effect measures. The first is a strata-adjusted percentage difference, defined as the Arm C minus concurrent Arm B difference based on CMH weighting. The second is a strata-adjusted odds ratio for Arm C over concurrent Arm B.
Confidence intervals
The confidence interval attached to an effect estimate describes the statistical uncertainty around the estimated treatment effect under the corresponding analysis framework. A confidence interval is not a prediction interval for individual participants, and it does not describe the range of biological responses that individual patients might experience.
Superiority framework
All of the registry-reported formal primary and secondary comparisons are identified as superiority hypotheses. The interpretation therefore concerns evidence that the nivolumab-containing comparison differs favorably from the chemotherapy comparator under the specified endpoint and effect measure; it is not a non-inferiority framework with a prespecified acceptable loss of efficacy.
7. Primary Results: Event-Free Survival
The primary EFS analysis compared all concurrently randomized participants in Arm C, nivolumab 360 mg plus platinum-doublet chemotherapy, with Arm B, platinum-doublet chemotherapy alone. The reported analysis used a stratified log-rank test and a Cox proportional-hazard effect measure.
Primary EFS hazard ratio
97.38% two-sided CI: 0.43–0.91 · P = 0.0052
Analysis: stratified log-rank test with Cox proportional-hazard effect measure.
| Primary EFS analysis | Reported value |
|---|---|
| Endpoint | Event-Free Survival (EFS) |
| Comparison | Arm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo |
| Analysis population | All concurrently randomized participants in Arm C and Arm B |
| Method | Log Rank |
| Effect measure | Cox Proportional Hazard / Hazard ratio |
| Estimate | 0.63 |
| Confidence interval | 97.38% two-sided CI 0.43–0.91 |
| P-value | 0.0052 |
| Hypothesis | Superiority |
| Stratification | PD-L1 (≥1% vs <1%/unevaluable/indeterminate); disease stage (IB/II vs IIIA); sex (male vs female) |
The hazard ratio of 0.63 means that, under the fitted Cox model, the estimated instantaneous rate of an EFS event in Arm C was approximately 63% of the estimated rate in Arm B. Equivalently, 1 − 0.63 = 0.37, so the estimated hazard was approximately 37% lower in Arm C than in Arm B.
This does not mean that 37% of participants avoided an event, that every participant experienced a 37% reduction in risk, or that the absolute probability of remaining event-free was 37 percentage points higher. The hazard ratio is a relative, model-based time-to-event measure.
The 97.38% confidence interval from 0.43 to 0.91 communicates uncertainty around the estimated hazard ratio. The interval is compatible with a range of relative effects rather than a single exact treatment effect. It does not describe the range of individual patient outcomes.
The reported P = 0.0052 addresses evidence against the relevant null hypothesis under the trial's statistical testing framework. A p-value does not measure the magnitude of benefit, the probability that the treatment works, or the probability that the null hypothesis is true.
Because the effect measure is a Cox hazard ratio, interpretation also depends on the proportional-hazards framework underlying the model. A single hazard ratio summarizes relative event rates over the analyzed follow-up; it should not automatically be read as a constant individual-level risk reduction at every time point.
Why the stratification matters
The EFS analysis did not ignore the three reported stratification factors. PD-L1 status, disease stage, and sex were incorporated into the stratified analysis. This matters because randomization and the analysis model can be coordinated through the same clinically relevant strata, improving alignment between the treatment comparison and the trial's design.
The analysis therefore answers a more specific question than "what proportion had an EFS event in each arm?" It estimates the relative time-to-event treatment effect while respecting the stated stratification structure.
8. Primary Results: Pathologic Complete Response
The second primary endpoint was pCR rate. The primary comparison again used all concurrently randomized participants in Arms C and B. The registry reports a CMH analysis with two complementary effect measures: a strata-adjusted percentage difference and a strata-adjusted odds ratio.
Primary pCR rate difference
99% two-sided CI: 13.0–30.3
Strata-adjusted difference, Arm C minus concurrent Arm B, using CMH weighting.
| Primary pCR analysis | Reported value |
|---|---|
| Endpoint | Pathologic Complete Response (pCR) Rate |
| Comparison | Arm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo |
| Analysis population | All concurrently randomized participants in Arm C and Arm B |
| Method | Cochran-Mantel-Haenszel test |
| Effect measure | % Difference / Other difference |
| Estimate | 21.6 |
| Confidence interval | 99% two-sided CI 13.0–30.3 |
| Hypothesis | Superiority |
| Interpretation of difference | Strata-adjusted difference, Arm C minus concurrent Arm B, based on CMH weighting |
The reported pCR difference of 21.6 is a strata-adjusted difference in percentage points between Arm C and concurrent Arm B, with the direction defined as Arm C minus Arm B. Thus, the estimate describes a 21.6-percentage-point difference in the pCR rate after the CMH stratification structure is taken into account.
This does not mean that every participant in Arm C had a 21.6% greater chance of pCR, nor does it mean that pCR translates one-for-one into longer survival. pCR is a pathological binary endpoint; EFS is a subsequent time-to-event endpoint. They answer different statistical questions.
The 99% confidence interval of 13.0 to 30.3 describes uncertainty around the strata-adjusted percentage difference. It provides a range of plausible values for the treatment effect under the stated statistical framework, rather than a range of individual responses.
No p-value is posted on ClinicalTrials.gov for this particular percentage-difference analysis in the trial data. The absence of a posted p-value does not justify calculating one independently from the confidence interval. The reported estimate and confidence interval should therefore be presented as reported.
Primary pCR odds ratio
Strata-adjusted odds ratio
99% two-sided CI: 3.49–55.75 · P < 0.0001
Odds ratio for Arm C over concurrent Arm B using the Mantel-Haenszel method.
| pCR odds-ratio analysis | Reported value |
|---|---|
| Method | Cochran-Mantel-Haenszel |
| Effect measure | Odds Ratio (OR) |
| Estimate | 13.94 |
| Confidence interval | 99% two-sided CI 3.49–55.75 |
| P-value | <0.0001 |
| Direction | Arm C over concurrent Arm B |
| Hypothesis | Superiority |
An odds ratio of 13.94 means that the estimated odds of pCR in Arm C were 13.94 times the estimated odds in concurrent Arm B under the reported strata-adjusted analysis. Odds are not probabilities, so the odds ratio should not be translated directly into a 13.94-fold increase in the pCR percentage.
The 99% confidence interval of 3.49 to 55.75 is wide relative to the point estimate. This is important: the analysis indicates a large estimated difference in odds, but the precision of that estimate is limited enough that a broad range of effect magnitudes remains compatible with the data.
The reported P < 0.0001 provides evidence against the null hypothesis under the stated testing framework. It does not measure the size of the treatment effect and should not be substituted for the effect estimate or its confidence interval.
The pCR odds ratio and the pCR percentage difference are complementary rather than interchangeable. The percentage difference is easier to translate into an absolute difference in response rates, whereas the odds ratio expresses the relative odds of response and is naturally suited to a CMH analysis across strata.
9. Secondary Endpoint Results
Major Pathologic Response Rate
Major Pathologic Response (MPR) Rate was reported as a secondary binary endpoint for the same concurrently randomized Arm C versus Arm B comparison. The registry reports a strata-adjusted percentage difference using CMH weighting.
MPR rate difference
95% two-sided CI: 19.6–36.1
Strata-adjusted difference, Arm C minus concurrent Arm B.
| Secondary MPR analysis | Reported value |
|---|---|
| Endpoint | Major Pathologic Response (MPR) Rate |
| Comparison | Arm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo |
| Analysis population | All concurrently randomized participants in Arm C and Arm B |
| Method | Cochran-Mantel-Haenszel test |
| Effect measure | % Difference |
| Estimate | 27.9 |
| Confidence interval | 95% two-sided CI 19.6–36.1 |
| Hypothesis | Superiority |
| Stratification | PD-L1 status; disease stage; sex |
Because the registry reports this as a strata-adjusted percentage difference, the estimate is most naturally read as a 27.9-percentage-point difference in MPR rate in the direction Arm C minus Arm B. The confidence interval indicates uncertainty around that adjusted difference.
MPR Rate Odds Ratio
A second posted MPR analysis reports an odds ratio, but the ClinicalTrials.gov record identifies the method as not reported for this particular analysis. The analysis is nevertheless identified as stratified by PD-L1 status, disease stage, and sex.
| Measure | Reported value |
|---|---|
| Endpoint | Major Pathologic Response (MPR) Rate |
| Effect measure | Odds Ratio (OR) |
| Estimate | 5.70 |
| 95% two-sided CI | 3.16–10.26 |
| Method | Not reported in the ClinicalTrials.gov record |
| Hypothesis | Superiority |
The odds ratio of 5.70 indicates substantially higher estimated odds of MPR in Arm C than in Arm B under the reported analysis. It should not be interpreted as a 5.70-fold increase in the percentage of participants with MPR, because odds and probabilities are different quantities.
Time to Death or Distant Metastases
TTDM hazard ratio
95% two-sided CI: 0.49–0.90
Stratified analysis; method not reported in the ClinicalTrials.gov record.
| TTDM analysis | Reported value |
|---|---|
| Endpoint | Time to Death or Distant Metastases (TTDM) |
| Comparison | Arm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo |
| Analysis population | All concurrently randomized participants in Arm C and Arm B |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 0.66 |
| 95% CI | 0.49–0.90, two-sided |
| Method | Not reported in the ClinicalTrials.gov record |
| Stratification | PD-L1 status; disease stage; sex |
| Hypothesis | Superiority |
An HR of 0.66 corresponds to an estimated instantaneous TTDM event rate approximately 34% lower in Arm C than Arm B under the fitted hazard-ratio interpretation. The confidence interval, 0.49 to 0.90, expresses the uncertainty around that estimate. No p-value is posted on ClinicalTrials.gov for this secondary analysis, so no additional significance calculation is presented here.
10. Post-Hoc Event-Free Survival Analysis
The ClinicalTrials.gov record also contain a post-hoc EFS analysis using a longer time frame, up to a median of 69 months. It compares the same concurrently randomized Arm C and Arm B populations.
Post-hoc EFS hazard ratio
95% two-sided CI: 0.51–0.91
Time frame: from randomization to disease progression, reoccurrence, or death due to any cause, up to a median of 69 months.
| Feature | Post-hoc EFS analysis |
|---|---|
| Endpoint | Event-Free Survival (EFS) |
| Analysis role | Post-hoc |
| Comparison | Arm C: Nivo 360 mg + Platinum Doublet Chemo vs Arm B: Platinum Doublet Chemo |
| Analysis population | All concurrently randomized participants in Arm C and Arm B |
| Effect measure | Cox Proportional Hazard / Hazard ratio |
| Estimate | 0.68 |
| 95% CI | 0.51–0.91, two-sided |
| Method | Not reported in the registry-reported post-hoc analysis record |
| Stratification | PD-L1 (≥1% vs <1%/unevaluable/indeterminate); disease stage (IB/II vs IIIA); sex (male vs female) |
The post-hoc HR of 0.68 indicates an estimated instantaneous EFS event rate approximately 32% lower in Arm C than Arm B under the Cox hazard-ratio interpretation. The estimate is directionally consistent with the primary EFS result of 0.63, although the two estimates correspond to different follow-up descriptions and the later analysis is explicitly classified as post-hoc.
The 95% CI of 0.51–0.91 indicates uncertainty around the longer-follow-up estimate. It should not be interpreted as evidence that the true treatment effect must remain constant throughout follow-up.
Because the analysis is labeled Post_Hoc and no formal method or p-value is reported in the ClinicalTrials.gov record, it should not be silently treated as another prespecified confirmatory primary analysis. Its most appropriate role on this page is descriptive: it shows the reported longer-term EFS effect estimate while preserving the distinction between the primary and post-hoc analyses.
11. Safety: Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by arm as affected participants over participants at risk. These data provide an arm-specific safety summary without requiring an inference about events that are not listed in the registry extract.
| Arm | Serious adverse events, affected / at risk | Reported proportion |
|---|---|---|
| Arm A: Nivo 3 mg/kg + Ipi 1 mg/kg | 30 / 111 | 30/111 |
| Arm B: Platinum Doublet Chemo | 58 / 208 | 58/208 |
| Arm C: Nivo 360 mg + Platinum Doublet Ch | 53 / 176 | 53/176 |
The serious-adverse-event counts should be kept separate from the efficacy analysis populations and from the primary endpoint results. Safety is exposure-related information, whereas the primary efficacy comparisons are defined by randomized treatment assignment and the endpoint-specific analysis populations.
12. Statistical Methods Explained
Why was a stratified log-rank test used for EFS?
EFS is a time-to-event endpoint, so simply comparing the number of events between treatment groups would discard information about when those events occurred and about participants who were censored before an event. The log-rank test is designed to compare event-time distributions while using the available follow-up information.
The reported analysis was stratified by PD-L1 status, disease stage, and sex. Stratification allows the comparison to account for these categorical factors rather than treating all participants as though the trial contained no such structure.
What does an EFS hazard ratio of 0.63 mean?
An HR of 0.63 means the estimated instantaneous rate of an EFS event in Arm C was approximately 63% of the estimated rate in Arm B under the fitted Cox model. The corresponding relative reduction in the estimated hazard is approximately 37%.
The important word is hazard. A hazard ratio is not the same as a risk ratio, an odds ratio, a percentage-point difference, or a probability of being event-free at a particular time.
Why is the pCR analysis different from the EFS analysis?
pCR is binary, while EFS is time-to-event. For pCR, the statistical question is whether the proportion with the defined pathological response differs between groups after accounting for the analysis strata. The CMH method is designed for this type of stratified categorical comparison.
For EFS, the analysis must account for event timing and censoring, which is why the trial uses a log-rank framework and a Cox hazard-ratio effect measure.
What does a pCR odds ratio of 13.94 mean?
The reported odds ratio is the estimated odds of pCR in Arm C divided by the corresponding odds in concurrent Arm B under the Mantel-Haenszel stratified analysis. An OR of 13.94 therefore means that the estimated odds are 13.94 times as high in Arm C.
It is important not to say that the "probability" or "pCR rate" is 13.94 times higher. The relationship between odds and probability is nonlinear, and the odds ratio cannot be converted to a probability ratio without the underlying event rates.
Why are there two pCR effect measures?
The percentage difference and odds ratio describe the same binary endpoint from different perspectives. The percentage difference is an absolute measure expressed in percentage points. The odds ratio is a relative measure on the odds scale.
Using both provides more information than using either alone. The difference communicates the magnitude of separation in response rates, while the odds ratio describes the relative odds after accounting for the trial's strata.
Why is the primary EFS confidence interval 97.38% rather than 95%?
The ClinicalTrials.gov record explicitly reports a 97.38% two-sided confidence interval for the primary EFS hazard ratio. This page preserves that value rather than replacing it with a conventional 95% interval. The different confidence level is part of the reported statistical framework and should not be silently standardized.
By contrast, the primary pCR odds-ratio and percentage-difference analyses report 99% two-sided confidence intervals. The secondary MPR and TTDM analyses report 95% two-sided intervals. Confidence-level differences therefore matter when comparing the apparent width of intervals across endpoints.
Why should the post-hoc EFS result not be treated like the primary EFS result?
The primary EFS analysis and the longer-term post-hoc EFS analysis are different statistical records. The primary analysis has a defined primary-endpoint role, a reported p-value, and a 97.38% confidence interval. The later analysis is explicitly classified as post-hoc, has no reported method in the ClinicalTrials.gov record, and provides a 95% confidence interval without a p-value.
Preserving that distinction prevents a later descriptive analysis from being mistaken for an additional prespecified confirmatory test.
13. Understanding the Confidence Intervals
| Endpoint | Estimate | Confidence interval | Confidence level |
|---|---|---|---|
| Primary EFS | HR 0.63 | 0.43–0.91 | 97.38% |
| Primary pCR difference | 21.6 | 13.0–30.3 | 99% |
| Primary pCR odds ratio | 13.94 | 3.49–55.75 | 99% |
| MPR difference | 27.9 | 19.6–36.1 | 95% |
| MPR odds ratio | 5.70 | 3.16–10.26 | 95% |
| TTDM | HR 0.66 | 0.49–0.90 | 95% |
| Post-hoc EFS | HR 0.68 | 0.51–0.91 | 95% |
The table illustrates why confidence intervals should be interpreted together with the effect measure. A hazard-ratio interval is centered on the multiplicative hazard scale, whereas the pCR and MPR percentage-difference intervals are expressed on an absolute percentage-point scale.
First identify what is being estimated. Then identify the direction of the comparison, the confidence level, and the width of the interval. Only after those steps should a p-value be considered. This prevents a common statistical mistake: treating every number below 0.05 as though it were an effect-size measure.
14. P-Values and Effect Sizes
The CheckMate-816 statistical record demonstrates why p-values and effect estimates should be reported together.
| Analysis | Effect estimate | P-value reported? |
|---|---|---|
| Primary EFS | HR 0.63 | 0.0052 |
| Primary pCR percentage difference | 21.6 | Not reported |
| Primary pCR odds ratio | 13.94 | <0.0001 |
| MPR percentage difference | 27.9 | Not reported |
| MPR odds ratio | 5.70 | Not reported |
| TTDM hazard ratio | 0.66 | Not reported |
| Post-hoc EFS hazard ratio | 0.68 | Not reported |
A p-value answers a question about compatibility with a null hypothesis under a specified statistical model and testing framework. An effect estimate answers a different question: how large is the estimated difference or relative effect?
For example, the primary pCR result contains both an odds ratio of 13.94 and a p-value of <0.0001. The p-value supports the statistical evidence against the null, while the odds ratio and its 99% confidence interval communicate the magnitude and uncertainty of the estimated treatment effect.
15. Multiplicity, Interim Analysis, and Other Design Features
The ClinicalTrials.gov record identifies the formal hypotheses as superiority and identify seven statistical analyses in total, including three primary-endpoint analyses. However, the ClinicalTrials.gov record does not provide a detailed alpha-spending plan, interim-analysis schedule, multiplicity hierarchy, non-inferiority margin, missing-data imputation strategy, or Bayesian analysis specification.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Superiority | Yes. The primary and secondary reported comparisons are identified as superiority hypotheses. |
| Stratification | Yes. PD-L1 status, disease stage, and sex are identified in the analysis notes. |
| Log-rank analysis | Yes. Used for the primary EFS analysis. |
| Cochran-Mantel-Haenszel analysis | Yes. Used for primary pCR and one secondary MPR analysis. |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Factorial design | Not reported; the registered design model is parallel. |
| Bayesian methods | Not reported. |
| Formal interim-analysis details | Not reported in the ClinicalTrials.gov record. |
| Missing-data/imputation strategy | Not reported in the ClinicalTrials.gov record. |
16. Why Stratified Analysis Matters
Three variables recur in the analysis notes: PD-L1 status, disease stage, and sex. They define strata for the reported treatment comparisons. Stratification is useful because it recognizes that the randomized comparison contains clinically meaningful categorical structure that can be incorporated into the analysis.
PD-L1 strata
The reported categories are ≥1% versus <1%/unevaluable/indeterminate.
Disease-stage strata
The reported categories are IB/II versus IIIA.
Sex strata
The reported categories are male versus female.
Why combine them?
CMH and stratified time-to-event methods can account for these categorical factors while estimating the treatment comparison across the defined strata.
Importantly, stratification does not turn the analysis into three separate trials, nor does it imply that each stratum has enough information to support an independent confirmatory conclusion. The ClinicalTrials.gov record is treatment comparisons across the stratified analysis structure.
17. Interpreting the Two Primary Endpoints Together
The two primary endpoints measure different dimensions of treatment effect. EFS asks about the time from randomization until progression, recurrence, or death according to the registered definition. pCR asks whether residual tumor is absent in the lung and lymph nodes on blinded independent pathological review.
| Dimension | EFS | pCR |
|---|---|---|
| Outcome type | Time-to-event | Binary |
| Time component | Yes | No; classified by pathological assessment |
| Primary effect measure | Hazard ratio | Percentage difference and odds ratio |
| Primary estimate | 0.63 | 21.6 percentage points; OR 13.94 |
| Confidence level reported | 97.38% | 99% |
| Formal p-value reported | 0.0052 | <0.0001 for OR analysis |
| Statistical method | Stratified log-rank | Stratified CMH |
The important statistical point is that concordant evidence across different endpoint types can be more informative than treating either endpoint as a complete summary of the trial. The EFS analysis incorporates follow-up time and censoring. The pCR analysis gives a pathological response comparison. Neither endpoint is a substitute for the other.
18. What the Primary EFS Result Does — and Does Not — Mean
The primary EFS hazard ratio of 0.63 indicates a lower estimated instantaneous EFS event rate in Arm C than in Arm B under the reported Cox model. The estimate corresponds to approximately a 37% lower estimated hazard relative to Arm B.
It does not mean that 37% of patients were protected from an EFS event, that each participant had exactly 37% less risk, or that the absolute difference in EFS probability at every time point was 37 percentage points.
The 97.38% confidence interval of 0.43–0.91 shows that the point estimate is not the only statistically plausible value. The interval expresses uncertainty around the treatment effect under the analysis framework and should accompany the hazard ratio whenever the result is presented.
The p-value of 0.0052 quantifies evidence against the relevant null hypothesis under the stated testing framework. It is not an effect-size measure and should not be interpreted as the probability that the treatment is effective.
19. What the Primary pCR Result Does — and Does Not — Mean
The primary pCR result is particularly useful for demonstrating the difference between an absolute effect measure and an odds ratio.
Percentage difference
The reported strata-adjusted difference is 21.6 percentage points, Arm C minus Arm B, with a 99% CI of 13.0–30.3.
Odds ratio
The reported strata-adjusted OR is 13.94, with a 99% CI of 3.49–55.75.
Statistical evidence
The p-value attached to the reported odds-ratio analysis is <0.0001.
Important caution
An odds ratio should not be read as a direct probability ratio. The absolute percentage-point difference is the more direct measure of separation in response rates.
The broad 99% confidence interval for the odds ratio is also instructive. The point estimate is 13.94, but the interval extends from 3.49 to 55.75. This illustrates why a statistically strong p-value does not imply that the effect estimate itself is known with arbitrary precision.
20. Long-Term EFS: Primary vs Post-Hoc Evidence
The trial data provide both a primary EFS analysis with a median time frame of up to 30 months and a post-hoc EFS analysis with a median time frame of up to 69 months.
| Feature | Primary EFS | Post-hoc EFS |
|---|---|---|
| Role | Primary | Post-hoc |
| Time frame | Up to a median of 30 months | Up to a median of 69 months |
| Estimate | HR 0.63 | HR 0.68 |
| Confidence interval | 97.38% CI 0.43–0.91 | 95% CI 0.51–0.91 |
| P-value | 0.0052 | Not reported |
| Method reported | Log-rank | Not reported |
| Effect measure | Cox proportional hazard | Cox proportional hazard |
The two estimates, 0.63 and 0.68, are reasonably close in direction and magnitude, but the correct statistical description is not that one "confirms" the other. They belong to different analysis roles and confidence frameworks. The later result is best treated as a longer-term descriptive estimate because the ClinicalTrials.gov record classifies it as post-hoc and does not provide a formal p-value or method.
21. Trial Timeline
Trial start
The registered trial start date is March 4, 2017.
Randomized parallel design
The study was registered as a randomized, parallel phase 3 treatment trial with three arms and no masking.
Overall enrollment
the ClinicalTrials.gov record reports enrollment of 505 participants.
Primary completion
The registered primary completion date is September 8, 2021.
Primary endpoint results
The primary EFS and pCR analyses are reported for the registry time frames extending up to a median of 30 months.
Post-hoc EFS analysis
The registry analysis record includes a later post-hoc EFS estimate using a time frame up to a median of 69 months.
22. Limitations
- Analysis scope: the numerical primary efficacy ClinicalTrials.gov record compare Arm C with concurrent Arm B. They should not be generalized to Arm A without additional statistical evidence.
- Open-label design: the registered masking status is none. Lack of masking can matter for outcomes susceptible to assessment or treatment-behavior influences, although the registry specifies blinded independent review for the pathological response and relevant disease assessments.
- Endpoint differences: pCR and EFS measure different clinical phenomena. A pathological response result should not be interpreted as if it were itself a time-to-event result.
- Confidence-level differences: the primary EFS analysis reports a 97.38% confidence interval, whereas the pCR analyses report 99% intervals and several secondary analyses report 95% intervals. Comparing interval widths without accounting for the confidence level can be misleading.
- Post-hoc analysis: the longer-term EFS analysis is labeled post-hoc and does not include a reported p-value or analysis method in the ClinicalTrials.gov record.
- Unreported statistical details: the ClinicalTrials.gov record does not specify a non-inferiority margin, crossover strategy, Bayesian analysis, detailed interim-analysis plan, or missing-data/imputation procedure.
- Safety comparison: serious adverse events are reported as affected participants over participants at risk, but the ClinicalTrials.gov record does not provide a formal between-arm statistical test for these events.
- Stratified analysis is not subgroup proof: the presence of stratification factors does not mean that treatment effects within each stratum are separately powered confirmatory analyses.
- Hazard-ratio interpretation: Cox hazard ratios are model-based relative time-to-event measures and should not automatically be interpreted as constant relative risks at every point in time.
23. Why This Trial Matters Statistically
CheckMate-816 is a useful statistical teaching example because its the ClinicalTrials.gov record span two fundamentally different endpoint classes and several complementary effect measures. The same randomized comparison can therefore be examined through survival analysis, categorical-data methods, absolute differences, odds ratios, confidence intervals, and stratified inference.
| Statistical concept | How it appears in CheckMate-816 |
|---|---|
| Randomization | The trial is registered as randomized with a parallel design. |
| Multiple treatment arms | Three arms were included in the ClinicalTrials.gov record. |
| Time-to-event analysis | EFS and TTDM are time-to-event endpoints. |
| Log-rank test | Used for the primary EFS comparison. |
| Cox hazard ratio | Used as the reported EFS and post-hoc EFS effect measure. |
| Binary endpoint analysis | pCR and MPR are binary response endpoints. |
| Cochran-Mantel-Haenszel method | Used for the primary pCR analysis and one MPR analysis. |
| Odds ratio | Reported for pCR and MPR. |
| Absolute difference | Reported for pCR and MPR as percentage differences. |
| Stratified analysis | PD-L1 status, disease stage, and sex are identified in the analysis notes. |
| Confidence intervals | Reported at 97.38%, 99%, and 95% levels depending on the analysis. |
| Superiority testing | The registry-reported formal analyses are identified as superiority hypotheses. |
| Post-hoc analysis | A later EFS analysis is explicitly classified as post-hoc. |
The trial is therefore particularly useful for learning that "the statistical analysis" is not one calculation. Different endpoint types require different estimands and different methods. A time-to-event endpoint cannot simply be analyzed like a binary response, and an odds ratio cannot be interpreted as a percentage-point difference.
24. A Deeper Statistical Reading of the Primary Results
Relative versus absolute effects
The primary pCR analysis gives an absolute percentage difference of 21.6, while the same endpoint gives an odds ratio of 13.94. These numbers are not competing answers. They use different scales.
The absolute difference is generally easier to communicate when the question is "how far apart are the response rates?" The odds ratio answers a different question: "how much larger are the odds of response?" A careful statistical report preserves both scales rather than converting one into the other without the necessary underlying response rates.
Why the EFS result needs censoring concepts
Unlike pCR, EFS cannot be summarized solely by counting events. Some participants may have follow-up that ends before an EFS event occurs. Those observations contribute information until censoring rather than being treated as though an event occurred.
This is one reason Kaplan-Meier estimation and time-to-event methods are central to EFS analysis. The ClinicalTrials.gov record specifically reports a log-rank analysis and a Cox proportional-hazard effect measure, which are methods designed for this type of endpoint.
Why the endpoint definition matters
The registered EFS definition includes disease progression precluding surgery, progression or recurrence based on BICR assessment after surgery, or death due to any cause. Consequently, the endpoint is not simply "time until death" or "time until radiographic progression." Its event definition determines exactly what the hazard ratio is comparing.
Why the analysis population matters
The primary analyses use all concurrently randomized participants in Arms C and B. This matters because the trial as a whole has three arms and 505 enrolled participants. The number 505 describes the trial's overall enrollment, whereas the formal efficacy estimates belong to the specified Arm C versus Arm B analysis population.
25. Statistical Concepts in the Learning Pathway
Learn more about the methods used in this trial:
26. Related Statistical Calculators
Use these calculator pathways to explore the statistical quantities that appear in the CheckMate-816 analyses:
27. Sources
- ClinicalTrials.gov: NCT02998528 — CheckMate-816.
- PubMed: Publication indexed under PMID 40454642.
- PubMed: Publication indexed under PMID 39778121.
- PubMed: Publication indexed under PMID 37903504.
- PubMed: Publication indexed under PMID 37141544.
- PubMed: Publication indexed under PMID 36815433.
The numerical trial results on this page are restricted to the registry-reported CheckMate-816 trial data. The PubMed links provide publication navigation but are not used here to introduce additional numerical results beyond the ClinicalTrials.gov recordset.
Continue through the Clinical Biostats statistical pathway
Explore the survival-analysis, categorical-data, confidence-interval, and clinical-trial methods that underlie the CheckMate-816 statistical results.
28. Record Summary
CheckMate-816 provides a useful example of how different endpoint types require different statistical analyses. The primary EFS endpoint is a time-to-event outcome analyzed with a stratified log-rank framework and a Cox proportional-hazard effect measure, producing an HR of 0.63 with a 97.38% two-sided confidence interval of 0.43–0.91 and a reported P = 0.0052. The primary pCR endpoint is binary and was analyzed with the Cochran-Mantel-Haenszel method, producing a strata-adjusted percentage difference of 21.6 with a 99% two-sided confidence interval of 13.0–30.3, together with a strata-adjusted odds ratio of 13.94 with a 99% confidence interval of 3.49–55.75 and P < 0.0001.
The secondary analyses extend the statistical picture without changing the primary endpoint definitions. MPR was associated with a reported percentage difference of 27.9 and an odds ratio of 5.70, while TTDM had a reported hazard ratio of 0.66. A longer-term post-hoc EFS analysis reported an HR of 0.68. Each result has to be interpreted according to its own endpoint, analysis role, effect measure, confidence level, and reported statistical method.
The broader statistical lesson is that a clinical-trial result is not adequately represented by a single p-value. A rigorous analysis identifies the randomized comparison, endpoint definition, analysis population, statistical method, effect measure, confidence interval, and testing role. For CheckMate-816, that framework is especially important because the trial contains three arms while the posted primary efficacy comparisons in the ClinicalTrials.gov record concern Arms C and B, and because the two primary endpoints require fundamentally different statistical approaches.