← Clinical Trials
Non-Squamous NSCLC Phase 3 Overall Survival NCT01673867

CheckMate-057: Complete Statistical Analysis of Nivolumab in Previously Treated Metastatic Non-Squamous NSCLC

An independent statistical analysis of the randomized phase 3 CheckMate-057 trial comparing nivolumab with docetaxel in previously treated metastatic non-squamous non-small cell lung cancer, with emphasis on its overall-survival endpoint, hazard ratio, log-rank test, and interpretation of the reported uncertainty.

Trial status: Completed  ·  Enrollment: 582  ·  Primary completion: February 2015
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results on this page are restricted to the information contained in the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for CheckMate-057. The registry provides one formal statistical analysis for the registered primary endpoint.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CheckMate-057 was a randomized, parallel-group phase 3 trial comparing nivolumab with docetaxel in previously treated metastatic non-squamous non-small cell lung cancer. The registered primary endpoint was overall survival, analyzed as a time-to-event outcome in all randomized participants.

582
Enrollment
All participants
2
Arms
Nivolumab vs docetaxel
0.73
OS HR
95.92% CI 0.59–0.89
0.0015
P-value
Log-rank test
FeatureCheckMate-057
Trial nameCheckMate-057
NCT IDNCT01673867
PhasePhase 3
StatusCompleted
PopulationPreviously treated metastatic non-squamous non-small cell lung cancer
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment582
InterventionsNivolumab and docetaxel
Lead sponsorBristol-Myers Squibb
Sponsor typeIndustry
StartNovember 2, 2012
Primary completionFebruary 5, 2015

2. Clinical Question

The registered trial addresses a straightforward randomized-treatment comparison: among participants with previously treated metastatic non-squamous NSCLC, how does nivolumab compare with docetaxel with respect to overall survival?

Population

Participants with non-squamous cell non-small cell lung cancer in the trial's previously treated metastatic setting.

Intervention

Nivolumab, identified in the registry as a biological intervention.

Comparator

Docetaxel, identified in the registry as a drug intervention.

Primary question

Does nivolumab produce a different overall-survival experience from docetaxel under a prespecified superiority framework?

3. Trial Design

01
Randomize 582 participants
02
Parallel arms 2 interventions
03
Treatment Nivolumab or docetaxel
04
Follow-up Overall survival
05
Primary analysis 413 deaths
Allocation
Randomized allocation was used to assign participants to the parallel treatment groups.
Masking
The registry describes the study as having no masking.
Primary purpose
Treatment.
Hypothesis type
Superiority.
INTERVENTION

Nivolumab

  • Biological intervention
  • Compared with docetaxel
  • Serious adverse events reported for the 3 mg/kg group: 172/287
  • Serious adverse events also reported for an extension phase: 5/15 and 9/17, with an additional extension entry of 1/1
COMPARATOR

Docetaxel

  • Drug intervention
  • Compared with nivolumab
  • Serious adverse events reported: 161/268

4. Trial Timeline

November 2, 2012

Trial start

The registered trial start date was November 2, 2012.

February 5, 2015

Primary completion

The registered primary completion date was February 5, 2015.

March 2015

Primary endpoint time frame

The registered primary endpoint was followed from randomization until 413 deaths, up to March 2015, described in the registry as approximately 29 months.

5. Primary Endpoint

EndpointRegistered definition and time frameAnalysis
Overall Survival (OS) Time in Months for All Randomized Participants at Primary Endpoint Overall survival was defined as the time from randomization to the date of death. A participant who has not died will be censored at last known date alive. OS was followed continuously while participants were on the study drug and every 3 months via in-person or phone contact after participants discontinued the study drug. The registered time frame was randomization until 413 deaths, up to March 2015 (approximately 29 months). Kaplan-Meier method for median and hazard ratio; log-rank test for the formal comparison.
Endpoint structure: OS is a time-to-event endpoint. The event is death, while participants who have not died at their last known follow-up contribute censored survival information rather than being treated as if an event had occurred.

6. Statistical Methodology

Kaplan-Meier estimation

The registry states that the median and hazard ratio for overall survival were computed using the Kaplan-Meier method. Kaplan-Meier estimation is designed for time-to-event data in which some participants may be censored before experiencing the event of interest.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at a particular event time and ni represents participants at risk immediately before that time. The estimator updates survival as events occur while retaining information from censored participants up to their censoring times.

Log-rank test

The formal statistical method reported in the registry was the log-rank test. The log-rank test compares the survival experience of two groups across their observed follow-up, taking account of the timing of events and censoring.

What the log-rank test asks
H0: survival distributions are the same between randomized groups

The resulting p-value addresses evidence against the null hypothesis of no difference in the survival distributions. It is not a measure of the size of the treatment effect.

Hazard ratio

The effect measure reported for the primary analysis was the hazard ratio, with the analysis note specifying HR = Nivolumab over docetaxel. An HR below 1 therefore corresponds to a lower estimated hazard of death in the nivolumab group relative to the docetaxel group, under the time-to-event model represented by the reported effect measure.

Interpretation of the reported HR
HR = hazard under nivolumab ÷ hazard under docetaxel

An HR of 0.73 can be described as a 27% lower estimated hazard because 1 − 0.73 = 0.27. This is a relative hazard interpretation, not a statement that 27% of patients avoided death or that every individual experienced a 27% reduction in risk.

Superiority hypothesis

The registry classifies the hypothesis type as superiority. This means the statistical objective is to evaluate whether the randomized treatment groups differ in the specified direction rather than to establish that the treatments are sufficiently similar within a predefined non-inferiority margin.

7. Primary Overall Survival Result

The ClinicalTrials.gov record contains one formal statistical analysis for the registered primary endpoint. The analysis population was all randomized participants, and the comparison was nivolumab versus docetaxel.

Overall survival hazard ratio

0.73

95.92% CI: 0.59–0.89   ·   P = 0.0015

Log-rank test  ·  Superiority hypothesis  ·  All randomized participants

Primary endpointNivolumab vs docetaxelMethod
Overall survival HR 0.73 (95.92% CI 0.59–0.89), P = 0.0015 Log-rank test; hazard ratio reported as nivolumab over docetaxel
Clinical Biostats interpretation

The reported hazard ratio of 0.73 means that the estimated hazard of death for nivolumab relative to docetaxel was 0.73. Expressed as a relative hazard difference, this corresponds to a 27% lower estimated hazard because 1 − 0.73 = 0.27.

The HR does not mean that 27% of participants were protected from death, that survival time increased by 27%, or that every participant experienced the same proportional reduction in risk. It is a relative time-to-event effect measure.

The 95.92% confidence interval of 0.59–0.89 describes uncertainty around the reported effect estimate under the statistical framework used for the analysis. It does not describe the range of individual patient outcomes or imply that the true effect for every participant lies inside that interval.

The P = 0.0015 value measures the strength of evidence against the relevant null hypothesis under the specified analysis. It does not measure the magnitude or clinical importance of the treatment effect. The magnitude is described by the hazard ratio and its confidence interval.

The endpoint is subject to the usual considerations for survival analysis, including censoring and the interpretation of a single hazard ratio over follow-up. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption, so that assumption should not be treated as having been demonstrated by the reported result.

Confidence-level detail: The registry reports a 95.92% two-sided confidence interval rather than the more commonly encountered 95% interval. The reported interval should therefore be reproduced exactly rather than silently converted or relabeled.

8. Reading the Primary Result Carefully

What does HR = 0.73 mean?

Because the registry defines the effect as nivolumab over docetaxel, an HR of 0.73 indicates that the estimated hazard in the nivolumab group was 73% of the estimated hazard in the docetaxel group under the reported analysis. The complementary expression, 27% lower estimated hazard, is simply another way to describe the same relative quantity.

What does the confidence interval contribute?

The point estimate alone does not communicate how precisely the treatment effect was estimated. The interval from 0.59 to 0.89 gives a range of values compatible with the statistical uncertainty represented by the reported confidence interval. Its lower and upper limits are both below 1, which is consistent with the direction of the point estimate.

What does P = 0.0015 contribute?

The p-value addresses the statistical evidence against the null hypothesis in the reported log-rank analysis. It should be interpreted alongside the effect estimate and confidence interval. A small p-value does not tell the reader whether the effect is large, small, clinically important, or important for every individual participant.

Why is the analysis population important?

The registry identifies the analysis population as all randomized participants. This is important because randomization establishes the treatment comparison at the point of allocation. Interpreting an efficacy analysis according to randomized assignment preserves the treatment groups created by randomization even when subsequent follow-up differs between participants.

Why does censoring matter?

The registry explicitly defines a participant who has not died as censored at the last known date alive. Censoring allows information available before the participant's last known survival date to contribute to the analysis without pretending that an unobserved death occurred after that point.

Why is a hazard ratio not a risk ratio?

A risk ratio compares probabilities over a specified time period. A hazard ratio instead compares event rates in a time-to-event framework. They answer different questions and should not be substituted for one another. The reported CheckMate-057 effect measure is a hazard ratio.

Why does the superiority framework matter?

The registry identifies the hypothesis type as superiority. Unlike a non-inferiority analysis, there is no reported non-inferiority margin that must be crossed. The interpretation therefore centers on evidence for a difference between the randomized groups, with the reported effect direction and uncertainty considered together.

9. Statistical Methods Explained

Why was a log-rank test used?

Overall survival is a time-to-event endpoint with censoring. The log-rank test is designed to compare survival experience between groups while incorporating the timing of observed deaths and the available follow-up for censored participants.

What does HR = 0.73 mean?

With nivolumab as the numerator and docetaxel as the denominator, 0.73 represents a lower estimated hazard in the nivolumab group. It can also be expressed as a 27% lower estimated hazard.

Why report a confidence interval?

The confidence interval communicates statistical uncertainty around the point estimate. The width of the interval provides information about precision that the hazard ratio alone cannot provide.

Why does P = 0.0015 not measure effect size?

A p-value measures evidence against a null hypothesis under the specified statistical model. Effect size is conveyed by the hazard ratio and, for relative precision, its confidence interval.

Why does randomization matter?

Randomization creates the treatment groups by allocation rather than by observed prognosis. This is the foundation for interpreting the between-group efficacy comparison as a randomized treatment comparison.

Why is Kaplan-Meier estimation useful?

Kaplan-Meier estimation allows survival probabilities to be estimated over time despite right censoring. The registry specifically states that the median and hazard ratio were computed using the Kaplan-Meier method.

10. Analysis Population and Censoring

FeatureRegistry informationStatistical implication
Analysis population All randomized participants The primary comparison is anchored to randomized treatment assignment.
Event Date of death Death is the event defining the OS endpoint.
Censoring Last known date alive for participants who have not died Participants contribute survival information through the last known alive date.
Follow-up while on study drug Continuously Survival follow-up is not restricted to a single scheduled assessment interval while on study drug.
Follow-up after discontinuation Every 3 months via in-person or phone contact OS ascertainment continued after study-drug discontinuation according to the registered procedure.

This distinction is statistically important because overall survival is not simply a measurement made at the time treatment stops. It is defined from randomization to death, with censoring used when death has not been observed by the participant's last known date alive.

11. Results Structure in the Registry

The ClinicalTrials.gov record reports 9 outcome measures and 1 statistical analysis. That single formal statistical analysis is for the registered primary overall-survival endpoint. It includes an estimate, a two-sided 95.92% confidence interval, and a p-value.

Registry result componentReported information
Results postedYes
Outcome measures posted9
Statistical analyses posted1
Primary-endpoint analyses1
Primary analyses with estimate + CI1
Primary endpoint typeTime-to-event
Formal methodLog-rank test
Effect measureHazard ratio
Hypothesis typeSuperiority
Scope of the numerical results: The ClinicalTrials.gov record does not provide additional formal statistical analyses for the other posted outcome measures. This page therefore does not manufacture secondary efficacy estimates, subgroup estimates, median survival values, or additional p-values.

12. Safety Results

The ClinicalTrials.gov record includes serious adverse events by arm. Because the denominators are explicitly reported as affected/at risk, the values below are reproduced exactly rather than converted into percentages.

Safety groupSerious adverse events
Nivolumab 3 mg/kg 172/287
Nivolumab 480 mg 5/15
Docetaxel 161/268
Extension Phase of Docetaxel Arm: Nivolu 9/17
Extension Phase of Docetaxel Arm: Nivolu 1/1
Safety denominator caution: These serious-adverse-event entries use their own reported at-risk denominators. They should not be combined into the overall enrollment of 582 or treated as mutually exclusive randomized populations without additional registry information.

Serious adverse events are a safety outcome rather than an efficacy endpoint. They therefore answer a different question from the overall-survival analysis. The presence of safety data alongside an efficacy result does not create a single composite measure of treatment benefit or harm.

13. Why the Overall-Survival Analysis Is a Time-to-Event Problem

A conventional binary endpoint might ask whether each participant experienced an event by a fixed date. The CheckMate-057 primary endpoint instead records when the event occurred and permits censoring when the participant has not died by the last known date alive.

Timing matters

A death occurring earlier and a death occurring later are not statistically equivalent observations in a time-to-event analysis.

Censoring preserves information

A participant who remains alive through a known follow-up date contributes survival information through that date even though the event itself was not observed.

Randomization establishes the time origin

The registered OS definition starts the survival clock at randomization, making the randomized groups directly comparable from a common time origin.

The HR is relative

The hazard ratio summarizes the relative event hazard between groups; it does not replace absolute survival probabilities or individual patient outcomes.

14. Understanding the Confidence Interval

Reported uncertainty around the OS effect

0.59–0.89

95.92% two-sided confidence interval for HR = 0.73

The reported point estimate is 0.73, while the confidence interval extends from 0.59 to 0.89. The interval therefore communicates that the estimated treatment effect is not known with infinite precision. It also preserves the direction of the reported effect throughout the interval because both endpoints are below 1.

What the interval does not say

The confidence interval is not a range containing the survival experience of individual participants. It is also not a prediction interval for future patients. It describes statistical uncertainty around the estimated treatment effect under the analysis framework.

Why the confidence level should be preserved

The registry reports a 95.92% confidence interval. Replacing it with 95% would change the stated statistical quantity. For reproducibility, the registered confidence level and endpoints are retained exactly on this page.

15. Understanding the P-value

Reported p-value
P = 0.0015

The registry reports this p-value for the log-rank comparison of overall survival under a superiority hypothesis.

A p-value of 0.0015 is evidence against the null hypothesis evaluated by the reported log-rank analysis. It does not quantify the magnitude of the hazard ratio. For that purpose, the HR of 0.73 and its confidence interval are the relevant reported quantities.

The p-value also should not be interpreted as the probability that the null hypothesis is true, the probability that the treatment works, or the probability that a particular patient will benefit. Those interpretations are not what a frequentist p-value represents.

16. Superiority Versus Non-Inferiority

The registry explicitly classifies the CheckMate-057 hypothesis as superiority. That distinction matters because the statistical logic differs from a non-inferiority trial.

FeatureSuperiority framework in this trialNon-inferiority framework
Primary objective Evaluate whether treatment groups differ Evaluate whether the experimental treatment is not unacceptably worse than the comparator
Margin No non-inferiority margin is reported in the ClinicalTrials.gov record A prespecified non-inferiority margin is required
Reported hypothesis type Superiority Not the registered hypothesis type

Because no non-inferiority margin is provided in the ClinicalTrials.gov record, none should be introduced into the interpretation of the CheckMate-057 result.

17. Multiplicity, Interim Analysis, and Other Design Features

The ClinicalTrials.gov record identifies the primary endpoint, its formal analysis, the effect measure, and the superiority hypothesis. They do not provide information about an alpha-spending strategy, interim-analysis boundary, multiplicity adjustment, stratification factors, crossover, or a prespecified missing-data/imputation procedure.

Design topicWhat can be established from the ClinicalTrials.gov record
MultiplicityNo multiplicity procedure is reported in the ClinicalTrials.gov record.
Interim analysisNo interim-analysis method or boundary is reported in the ClinicalTrials.gov record.
Alpha spendingNo alpha-spending method is reported.
StratificationNo stratification factors are reported in the ClinicalTrials.gov record.
CrossoverNo crossover procedure is reported in the ClinicalTrials.gov record.
Missing-data/imputationNo missing-data or imputation procedure is reported for the primary endpoint.
Bayesian methodsNo Bayesian method is reported.
Non-inferiority marginNot applicable to the registered superiority hypothesis; no margin is reported.
Methodological restraint: The absence of a method in the ClinicalTrials.gov record does not establish that the trial protocol contained no such procedure. It means only that the ClinicalTrials.gov record does not provide enough information to describe that procedure reliably.

18. What the Reported Hazard Ratio Does — and Does Not — Mean

Relative effect

The reported OS HR of 0.73 corresponds to a 27% lower estimated hazard for nivolumab relative to docetaxel because 1 − 0.73 = 0.27. This is a relative comparison of hazards, not an absolute reduction in mortality probability.

Not an individual prediction

The HR does not imply that each participant receiving nivolumab experienced exactly a 27% reduction in their individual probability of death. Participants can have very different survival experiences within the same randomized treatment group.

Not a median-survival result

The ClinicalTrials.gov record does not report a median overall-survival estimate. The HR therefore should not be translated into an unreported difference in median survival.

Not a complete clinical-benefit summary

The primary statistical result describes overall survival. It does not by itself summarize safety, quality of life, response, progression, or every other clinical outcome measured in the trial.

19. Limitations

20. Why This Trial Matters Statistically

CheckMate-057 is a useful teaching example because the ClinicalTrials.gov record contains the core structure of a randomized time-to-event analysis: randomization, a parallel-group design, an overall-survival endpoint, censoring, Kaplan-Meier estimation, a log-rank comparison, a hazard ratio, a confidence interval, and a p-value.

ConceptHow it appears in CheckMate-057
RandomizationThe allocation is explicitly randomized.
Parallel-group designThe design model is parallel, with two interventions.
Time-to-event endpointOverall survival is the registered primary endpoint.
CensoringParticipants who have not died are censored at their last known date alive.
Kaplan-Meier estimationThe registry states that median and hazard ratio are computed using the Kaplan-Meier method.
Log-rank testThe formal statistical method reported for the primary analysis is the log-rank test.
Hazard ratioThe treatment effect is reported as HR = nivolumab over docetaxel.
Confidence intervalThe HR is accompanied by a two-sided 95.92% confidence interval.
P-valueThe primary analysis reports P = 0.0015.
SuperiorityThe registered hypothesis type is superiority.
Analysis populationThe primary analysis uses all randomized participants.
Safety analysisSerious adverse events are reported with affected/at-risk counts by arm.

21. Statistical Methods Explained: A Deeper View

Why not analyze overall survival with an ordinary two-sample mean comparison?

Overall survival is not simply a continuous measurement observed completely for every participant. Some participants may be alive at their last follow-up, producing censored observations. Kaplan-Meier and related survival-analysis methods are designed specifically for this structure.

What information does a censored participant provide?

A censored participant provides information about being alive and event-free through the censoring time. The participant is not treated as if an event occurred at the censoring time, nor is follow-up after that time assumed to be known.

Why can a hazard ratio differ from a difference in survival probabilities?

A hazard ratio is a relative comparison of event hazards over time. A survival probability at a particular time is an absolute quantity. Two studies can have similar hazard ratios but different absolute survival probabilities if their underlying event patterns and baseline risks differ.

Why is the analysis population specified as all randomized participants?

Specifying all randomized participants preserves the treatment comparison created by randomization. This avoids redefining the primary efficacy groups according to treatment exposure after randomization, which can introduce post-randomization selection.

What does the 95.92% confidence interval tell us?

It describes statistical uncertainty around the reported hazard-ratio estimate under the stated confidence level. Because the interval is 0.59–0.89, it remains below 1 throughout the reported interval.

Why should the p-value and HR be read together?

The p-value addresses evidence against the null hypothesis, while the HR describes the estimated relative effect. Reading them together gives a more complete statistical picture than either quantity alone.

Why should we avoid inventing median survival from the HR?

A hazard ratio does not uniquely determine median survival. Median survival depends on the underlying survival distributions, so an unreported median cannot be recovered from the HR alone.

22. Statistical Interpretation Versus Clinical Interpretation

Statistical interpretation

The randomized comparison produced an OS hazard ratio of 0.73, with a two-sided 95.92% confidence interval of 0.59–0.89 and a log-rank p-value of 0.0015 under the registered superiority framework.

Clinical interpretation

The ClinicalTrials.gov record establishes the reported overall-survival statistical result, but they do not provide enough additional numerical information here to characterize median survival, subgroup effects, response, or other clinical outcomes.

This distinction is important. A statistically defined treatment effect is one component of clinical evidence. Clinical interpretation also depends on the full set of efficacy and safety outcomes, follow-up, patient characteristics, and other information that is not numerically reported in the ClinicalTrials.gov record.

23. What the Registry Does Not Establish From the Supplied Data

QuestionAnswer from the ClinicalTrials.gov record
What was the median overall survival?Not reported in the ClinicalTrials.gov record.
What were subgroup hazard ratios?Not reported.
Was proportional hazards formally assessed?Not reported.
What stratification factors were used?Not reported.
Was there an interim analysis?Not reported.
Was alpha spending used?Not reported.
How was multiplicity controlled?Not reported.
Was there crossover?Not reported.
What imputation method was used?Not reported.
Were Bayesian methods used?Not reported.
What were the other formal efficacy estimates?The ClinicalTrials.gov record does not provide additional statistical analyses beyond the primary OS analysis.
Why this matters: A high-quality statistical analysis is not improved by filling gaps with plausible-looking numbers. The correct approach is to distinguish what the registry formally reports from what a statistical method would ordinarily require or permit.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts behind randomized clinical trials, time-to-event endpoints, hazard ratios, confidence intervals, and survival analysis.

27. Record Summary

CheckMate-057 provides a clear example of randomized clinical-trial survival analysis. The trial enrolled 582 participants in a randomized parallel-group phase 3 design comparing nivolumab with docetaxel. Its registered primary endpoint was overall survival, defined from randomization to death, with participants who had not died censored at their last known date alive. The formal statistical analysis used a log-rank test and reported a hazard ratio of 0.73 for nivolumab over docetaxel, with a two-sided 95.92% confidence interval of 0.59–0.89 and P = 0.0015 under a superiority hypothesis.

The statistical lesson is broader than the p-value alone. The hazard ratio describes the relative treatment effect, the confidence interval describes uncertainty around that estimate, the log-rank test supplies the formal hypothesis-test result, and the Kaplan-Meier framework accommodates the censored time-to-event structure. At the same time, the ClinicalTrials.gov record does not provide enough information to reconstruct median survival, subgroup effects, multiplicity procedures, interim monitoring, or several other design details. A rigorous trial analysis therefore distinguishes reported evidence from assumptions that cannot be verified from the available record.

Clinical Biostats methodology: A trial-results page should not merely repeat a registry record. The goal is to explain the statistical structure of the trial, interpret reported estimates precisely, and make clear where the available evidence ends.