This page separates reported trial results from statistical interpretation. Numerical trial results on this page are restricted to the information contained in the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for CheckMate-057. The registry provides one formal statistical analysis for the registered primary endpoint.
1. Trial at a Glance
CheckMate-057 was a randomized, parallel-group phase 3 trial comparing nivolumab with docetaxel in previously treated metastatic non-squamous non-small cell lung cancer. The registered primary endpoint was overall survival, analyzed as a time-to-event outcome in all randomized participants.
| Feature | CheckMate-057 |
|---|---|
| Trial name | CheckMate-057 |
| NCT ID | NCT01673867 |
| Phase | Phase 3 |
| Status | Completed |
| Population | Previously treated metastatic non-squamous non-small cell lung cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 582 |
| Interventions | Nivolumab and docetaxel |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
| Start | November 2, 2012 |
| Primary completion | February 5, 2015 |
2. Clinical Question
The registered trial addresses a straightforward randomized-treatment comparison: among participants with previously treated metastatic non-squamous NSCLC, how does nivolumab compare with docetaxel with respect to overall survival?
Population
Participants with non-squamous cell non-small cell lung cancer in the trial's previously treated metastatic setting.
Intervention
Nivolumab, identified in the registry as a biological intervention.
Comparator
Docetaxel, identified in the registry as a drug intervention.
Primary question
Does nivolumab produce a different overall-survival experience from docetaxel under a prespecified superiority framework?
3. Trial Design
Nivolumab
- Biological intervention
- Compared with docetaxel
- Serious adverse events reported for the 3 mg/kg group: 172/287
- Serious adverse events also reported for an extension phase: 5/15 and 9/17, with an additional extension entry of 1/1
Docetaxel
- Drug intervention
- Compared with nivolumab
- Serious adverse events reported: 161/268
4. Trial Timeline
Trial start
The registered trial start date was November 2, 2012.
Primary completion
The registered primary completion date was February 5, 2015.
Primary endpoint time frame
The registered primary endpoint was followed from randomization until 413 deaths, up to March 2015, described in the registry as approximately 29 months.
5. Primary Endpoint
| Endpoint | Registered definition and time frame | Analysis |
|---|---|---|
| Overall Survival (OS) Time in Months for All Randomized Participants at Primary Endpoint | Overall survival was defined as the time from randomization to the date of death. A participant who has not died will be censored at last known date alive. OS was followed continuously while participants were on the study drug and every 3 months via in-person or phone contact after participants discontinued the study drug. The registered time frame was randomization until 413 deaths, up to March 2015 (approximately 29 months). | Kaplan-Meier method for median and hazard ratio; log-rank test for the formal comparison. |
6. Statistical Methodology
Kaplan-Meier estimation
The registry states that the median and hazard ratio for overall survival were computed using the Kaplan-Meier method. Kaplan-Meier estimation is designed for time-to-event data in which some participants may be censored before experiencing the event of interest.
Here, di represents events at a particular event time and ni represents participants at risk immediately before that time. The estimator updates survival as events occur while retaining information from censored participants up to their censoring times.
Log-rank test
The formal statistical method reported in the registry was the log-rank test. The log-rank test compares the survival experience of two groups across their observed follow-up, taking account of the timing of events and censoring.
The resulting p-value addresses evidence against the null hypothesis of no difference in the survival distributions. It is not a measure of the size of the treatment effect.
Hazard ratio
The effect measure reported for the primary analysis was the hazard ratio, with the analysis note specifying HR = Nivolumab over docetaxel. An HR below 1 therefore corresponds to a lower estimated hazard of death in the nivolumab group relative to the docetaxel group, under the time-to-event model represented by the reported effect measure.
An HR of 0.73 can be described as a 27% lower estimated hazard because 1 − 0.73 = 0.27. This is a relative hazard interpretation, not a statement that 27% of patients avoided death or that every individual experienced a 27% reduction in risk.
Superiority hypothesis
The registry classifies the hypothesis type as superiority. This means the statistical objective is to evaluate whether the randomized treatment groups differ in the specified direction rather than to establish that the treatments are sufficiently similar within a predefined non-inferiority margin.
7. Primary Overall Survival Result
The ClinicalTrials.gov record contains one formal statistical analysis for the registered primary endpoint. The analysis population was all randomized participants, and the comparison was nivolumab versus docetaxel.
Overall survival hazard ratio
95.92% CI: 0.59–0.89 · P = 0.0015
Log-rank test · Superiority hypothesis · All randomized participants
| Primary endpoint | Nivolumab vs docetaxel | Method |
|---|---|---|
| Overall survival | HR 0.73 (95.92% CI 0.59–0.89), P = 0.0015 | Log-rank test; hazard ratio reported as nivolumab over docetaxel |
The reported hazard ratio of 0.73 means that the estimated hazard of death for nivolumab relative to docetaxel was 0.73. Expressed as a relative hazard difference, this corresponds to a 27% lower estimated hazard because 1 − 0.73 = 0.27.
The HR does not mean that 27% of participants were protected from death, that survival time increased by 27%, or that every participant experienced the same proportional reduction in risk. It is a relative time-to-event effect measure.
The 95.92% confidence interval of 0.59–0.89 describes uncertainty around the reported effect estimate under the statistical framework used for the analysis. It does not describe the range of individual patient outcomes or imply that the true effect for every participant lies inside that interval.
The P = 0.0015 value measures the strength of evidence against the relevant null hypothesis under the specified analysis. It does not measure the magnitude or clinical importance of the treatment effect. The magnitude is described by the hazard ratio and its confidence interval.
The endpoint is subject to the usual considerations for survival analysis, including censoring and the interpretation of a single hazard ratio over follow-up. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption, so that assumption should not be treated as having been demonstrated by the reported result.
8. Reading the Primary Result Carefully
What does HR = 0.73 mean?
Because the registry defines the effect as nivolumab over docetaxel, an HR of 0.73 indicates that the estimated hazard in the nivolumab group was 73% of the estimated hazard in the docetaxel group under the reported analysis. The complementary expression, 27% lower estimated hazard, is simply another way to describe the same relative quantity.
What does the confidence interval contribute?
The point estimate alone does not communicate how precisely the treatment effect was estimated. The interval from 0.59 to 0.89 gives a range of values compatible with the statistical uncertainty represented by the reported confidence interval. Its lower and upper limits are both below 1, which is consistent with the direction of the point estimate.
What does P = 0.0015 contribute?
The p-value addresses the statistical evidence against the null hypothesis in the reported log-rank analysis. It should be interpreted alongside the effect estimate and confidence interval. A small p-value does not tell the reader whether the effect is large, small, clinically important, or important for every individual participant.
Why is the analysis population important?
The registry identifies the analysis population as all randomized participants. This is important because randomization establishes the treatment comparison at the point of allocation. Interpreting an efficacy analysis according to randomized assignment preserves the treatment groups created by randomization even when subsequent follow-up differs between participants.
Why does censoring matter?
The registry explicitly defines a participant who has not died as censored at the last known date alive. Censoring allows information available before the participant's last known survival date to contribute to the analysis without pretending that an unobserved death occurred after that point.
Why is a hazard ratio not a risk ratio?
A risk ratio compares probabilities over a specified time period. A hazard ratio instead compares event rates in a time-to-event framework. They answer different questions and should not be substituted for one another. The reported CheckMate-057 effect measure is a hazard ratio.
Why does the superiority framework matter?
The registry identifies the hypothesis type as superiority. Unlike a non-inferiority analysis, there is no reported non-inferiority margin that must be crossed. The interpretation therefore centers on evidence for a difference between the randomized groups, with the reported effect direction and uncertainty considered together.
9. Statistical Methods Explained
Why was a log-rank test used?
Overall survival is a time-to-event endpoint with censoring. The log-rank test is designed to compare survival experience between groups while incorporating the timing of observed deaths and the available follow-up for censored participants.
What does HR = 0.73 mean?
With nivolumab as the numerator and docetaxel as the denominator, 0.73 represents a lower estimated hazard in the nivolumab group. It can also be expressed as a 27% lower estimated hazard.
Why report a confidence interval?
The confidence interval communicates statistical uncertainty around the point estimate. The width of the interval provides information about precision that the hazard ratio alone cannot provide.
Why does P = 0.0015 not measure effect size?
A p-value measures evidence against a null hypothesis under the specified statistical model. Effect size is conveyed by the hazard ratio and, for relative precision, its confidence interval.
Why does randomization matter?
Randomization creates the treatment groups by allocation rather than by observed prognosis. This is the foundation for interpreting the between-group efficacy comparison as a randomized treatment comparison.
Why is Kaplan-Meier estimation useful?
Kaplan-Meier estimation allows survival probabilities to be estimated over time despite right censoring. The registry specifically states that the median and hazard ratio were computed using the Kaplan-Meier method.
10. Analysis Population and Censoring
| Feature | Registry information | Statistical implication |
|---|---|---|
| Analysis population | All randomized participants | The primary comparison is anchored to randomized treatment assignment. |
| Event | Date of death | Death is the event defining the OS endpoint. |
| Censoring | Last known date alive for participants who have not died | Participants contribute survival information through the last known alive date. |
| Follow-up while on study drug | Continuously | Survival follow-up is not restricted to a single scheduled assessment interval while on study drug. |
| Follow-up after discontinuation | Every 3 months via in-person or phone contact | OS ascertainment continued after study-drug discontinuation according to the registered procedure. |
This distinction is statistically important because overall survival is not simply a measurement made at the time treatment stops. It is defined from randomization to death, with censoring used when death has not been observed by the participant's last known date alive.
11. Results Structure in the Registry
The ClinicalTrials.gov record reports 9 outcome measures and 1 statistical analysis. That single formal statistical analysis is for the registered primary overall-survival endpoint. It includes an estimate, a two-sided 95.92% confidence interval, and a p-value.
| Registry result component | Reported information |
|---|---|
| Results posted | Yes |
| Outcome measures posted | 9 |
| Statistical analyses posted | 1 |
| Primary-endpoint analyses | 1 |
| Primary analyses with estimate + CI | 1 |
| Primary endpoint type | Time-to-event |
| Formal method | Log-rank test |
| Effect measure | Hazard ratio |
| Hypothesis type | Superiority |
12. Safety Results
The ClinicalTrials.gov record includes serious adverse events by arm. Because the denominators are explicitly reported as affected/at risk, the values below are reproduced exactly rather than converted into percentages.
| Safety group | Serious adverse events |
|---|---|
| Nivolumab 3 mg/kg | 172/287 |
| Nivolumab 480 mg | 5/15 |
| Docetaxel | 161/268 |
| Extension Phase of Docetaxel Arm: Nivolu | 9/17 |
| Extension Phase of Docetaxel Arm: Nivolu | 1/1 |
Serious adverse events are a safety outcome rather than an efficacy endpoint. They therefore answer a different question from the overall-survival analysis. The presence of safety data alongside an efficacy result does not create a single composite measure of treatment benefit or harm.
13. Why the Overall-Survival Analysis Is a Time-to-Event Problem
A conventional binary endpoint might ask whether each participant experienced an event by a fixed date. The CheckMate-057 primary endpoint instead records when the event occurred and permits censoring when the participant has not died by the last known date alive.
Timing matters
A death occurring earlier and a death occurring later are not statistically equivalent observations in a time-to-event analysis.
Censoring preserves information
A participant who remains alive through a known follow-up date contributes survival information through that date even though the event itself was not observed.
Randomization establishes the time origin
The registered OS definition starts the survival clock at randomization, making the randomized groups directly comparable from a common time origin.
The HR is relative
The hazard ratio summarizes the relative event hazard between groups; it does not replace absolute survival probabilities or individual patient outcomes.
14. Understanding the Confidence Interval
Reported uncertainty around the OS effect
95.92% two-sided confidence interval for HR = 0.73
The reported point estimate is 0.73, while the confidence interval extends from 0.59 to 0.89. The interval therefore communicates that the estimated treatment effect is not known with infinite precision. It also preserves the direction of the reported effect throughout the interval because both endpoints are below 1.
The confidence interval is not a range containing the survival experience of individual participants. It is also not a prediction interval for future patients. It describes statistical uncertainty around the estimated treatment effect under the analysis framework.
The registry reports a 95.92% confidence interval. Replacing it with 95% would change the stated statistical quantity. For reproducibility, the registered confidence level and endpoints are retained exactly on this page.
15. Understanding the P-value
The registry reports this p-value for the log-rank comparison of overall survival under a superiority hypothesis.
A p-value of 0.0015 is evidence against the null hypothesis evaluated by the reported log-rank analysis. It does not quantify the magnitude of the hazard ratio. For that purpose, the HR of 0.73 and its confidence interval are the relevant reported quantities.
The p-value also should not be interpreted as the probability that the null hypothesis is true, the probability that the treatment works, or the probability that a particular patient will benefit. Those interpretations are not what a frequentist p-value represents.
16. Superiority Versus Non-Inferiority
The registry explicitly classifies the CheckMate-057 hypothesis as superiority. That distinction matters because the statistical logic differs from a non-inferiority trial.
| Feature | Superiority framework in this trial | Non-inferiority framework |
|---|---|---|
| Primary objective | Evaluate whether treatment groups differ | Evaluate whether the experimental treatment is not unacceptably worse than the comparator |
| Margin | No non-inferiority margin is reported in the ClinicalTrials.gov record | A prespecified non-inferiority margin is required |
| Reported hypothesis type | Superiority | Not the registered hypothesis type |
Because no non-inferiority margin is provided in the ClinicalTrials.gov record, none should be introduced into the interpretation of the CheckMate-057 result.
17. Multiplicity, Interim Analysis, and Other Design Features
The ClinicalTrials.gov record identifies the primary endpoint, its formal analysis, the effect measure, and the superiority hypothesis. They do not provide information about an alpha-spending strategy, interim-analysis boundary, multiplicity adjustment, stratification factors, crossover, or a prespecified missing-data/imputation procedure.
| Design topic | What can be established from the ClinicalTrials.gov record |
|---|---|
| Multiplicity | No multiplicity procedure is reported in the ClinicalTrials.gov record. |
| Interim analysis | No interim-analysis method or boundary is reported in the ClinicalTrials.gov record. |
| Alpha spending | No alpha-spending method is reported. |
| Stratification | No stratification factors are reported in the ClinicalTrials.gov record. |
| Crossover | No crossover procedure is reported in the ClinicalTrials.gov record. |
| Missing-data/imputation | No missing-data or imputation procedure is reported for the primary endpoint. |
| Bayesian methods | No Bayesian method is reported. |
| Non-inferiority margin | Not applicable to the registered superiority hypothesis; no margin is reported. |
18. What the Reported Hazard Ratio Does — and Does Not — Mean
The reported OS HR of 0.73 corresponds to a 27% lower estimated hazard for nivolumab relative to docetaxel because 1 − 0.73 = 0.27. This is a relative comparison of hazards, not an absolute reduction in mortality probability.
The HR does not imply that each participant receiving nivolumab experienced exactly a 27% reduction in their individual probability of death. Participants can have very different survival experiences within the same randomized treatment group.
The ClinicalTrials.gov record does not report a median overall-survival estimate. The HR therefore should not be translated into an unreported difference in median survival.
The primary statistical result describes overall survival. It does not by itself summarize safety, quality of life, response, progression, or every other clinical outcome measured in the trial.
19. Limitations
- Limited formal results in the ClinicalTrials.gov record: the record contains 9 posted outcome measures but only 1 posted statistical analysis, so the numerical efficacy interpretation on this page is necessarily centered on the registered OS analysis.
- No median OS is reported: although the registry definition states that median OS is computed using Kaplan-Meier methods, the ClinicalTrials.gov record does not provide a median estimate. None is inferred here.
- No subgroup estimates are reported: no subgroup hazard ratios or interaction analyses are reported in the ClinicalTrials.gov record, so treatment-effect heterogeneity cannot be assessed from this record.
- No proportional-hazards assessment is reported: the registry reports a hazard ratio but does not provide an assessment of whether proportional hazards held over follow-up.
- No stratification factors are reported: the registry data do not identify any stratification variables used in the analysis.
- No multiplicity information is reported: the single posted formal analysis does not provide enough information to characterize any broader family of hypothesis tests.
- Safety denominators differ from total enrollment: serious-adverse-event entries have their own at-risk denominators and should not be assumed to represent the complete randomized population without additional documentation.
- Open-label design: the registry describes the masking as none. This is a design characteristic that may matter for outcomes susceptible to knowledge of treatment assignment, although the primary endpoint is overall survival.
- Registry-level information: this analysis is constrained to the ClinicalTrials.gov record. Publication-level details not contained in that data are intentionally not added.
20. Why This Trial Matters Statistically
CheckMate-057 is a useful teaching example because the ClinicalTrials.gov record contains the core structure of a randomized time-to-event analysis: randomization, a parallel-group design, an overall-survival endpoint, censoring, Kaplan-Meier estimation, a log-rank comparison, a hazard ratio, a confidence interval, and a p-value.
| Concept | How it appears in CheckMate-057 |
|---|---|
| Randomization | The allocation is explicitly randomized. |
| Parallel-group design | The design model is parallel, with two interventions. |
| Time-to-event endpoint | Overall survival is the registered primary endpoint. |
| Censoring | Participants who have not died are censored at their last known date alive. |
| Kaplan-Meier estimation | The registry states that median and hazard ratio are computed using the Kaplan-Meier method. |
| Log-rank test | The formal statistical method reported for the primary analysis is the log-rank test. |
| Hazard ratio | The treatment effect is reported as HR = nivolumab over docetaxel. |
| Confidence interval | The HR is accompanied by a two-sided 95.92% confidence interval. |
| P-value | The primary analysis reports P = 0.0015. |
| Superiority | The registered hypothesis type is superiority. |
| Analysis population | The primary analysis uses all randomized participants. |
| Safety analysis | Serious adverse events are reported with affected/at-risk counts by arm. |
21. Statistical Methods Explained: A Deeper View
Why not analyze overall survival with an ordinary two-sample mean comparison?
Overall survival is not simply a continuous measurement observed completely for every participant. Some participants may be alive at their last follow-up, producing censored observations. Kaplan-Meier and related survival-analysis methods are designed specifically for this structure.
What information does a censored participant provide?
A censored participant provides information about being alive and event-free through the censoring time. The participant is not treated as if an event occurred at the censoring time, nor is follow-up after that time assumed to be known.
Why can a hazard ratio differ from a difference in survival probabilities?
A hazard ratio is a relative comparison of event hazards over time. A survival probability at a particular time is an absolute quantity. Two studies can have similar hazard ratios but different absolute survival probabilities if their underlying event patterns and baseline risks differ.
Why is the analysis population specified as all randomized participants?
Specifying all randomized participants preserves the treatment comparison created by randomization. This avoids redefining the primary efficacy groups according to treatment exposure after randomization, which can introduce post-randomization selection.
What does the 95.92% confidence interval tell us?
It describes statistical uncertainty around the reported hazard-ratio estimate under the stated confidence level. Because the interval is 0.59–0.89, it remains below 1 throughout the reported interval.
Why should the p-value and HR be read together?
The p-value addresses evidence against the null hypothesis, while the HR describes the estimated relative effect. Reading them together gives a more complete statistical picture than either quantity alone.
Why should we avoid inventing median survival from the HR?
A hazard ratio does not uniquely determine median survival. Median survival depends on the underlying survival distributions, so an unreported median cannot be recovered from the HR alone.
22. Statistical Interpretation Versus Clinical Interpretation
Statistical interpretation
The randomized comparison produced an OS hazard ratio of 0.73, with a two-sided 95.92% confidence interval of 0.59–0.89 and a log-rank p-value of 0.0015 under the registered superiority framework.
Clinical interpretation
The ClinicalTrials.gov record establishes the reported overall-survival statistical result, but they do not provide enough additional numerical information here to characterize median survival, subgroup effects, response, or other clinical outcomes.
This distinction is important. A statistically defined treatment effect is one component of clinical evidence. Clinical interpretation also depends on the full set of efficacy and safety outcomes, follow-up, patient characteristics, and other information that is not numerically reported in the ClinicalTrials.gov record.
23. What the Registry Does Not Establish From the Supplied Data
| Question | Answer from the ClinicalTrials.gov record |
|---|---|
| What was the median overall survival? | Not reported in the ClinicalTrials.gov record. |
| What were subgroup hazard ratios? | Not reported. |
| Was proportional hazards formally assessed? | Not reported. |
| What stratification factors were used? | Not reported. |
| Was there an interim analysis? | Not reported. |
| Was alpha spending used? | Not reported. |
| How was multiplicity controlled? | Not reported. |
| Was there crossover? | Not reported. |
| What imputation method was used? | Not reported. |
| Were Bayesian methods used? | Not reported. |
| What were the other formal efficacy estimates? | The ClinicalTrials.gov record does not provide additional statistical analyses beyond the primary OS analysis. |
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Statistical Calculators
26. Sources
- ClinicalTrials.gov: NCT01673867 — CheckMate-057.
- PubMed: PMID 37264091.
- PubMed: PMID 36897427.
- PubMed: PMID 33449799.
- PubMed: PMID 30215677.
- PubMed: PMID 30103096.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts behind randomized clinical trials, time-to-event endpoints, hazard ratios, confidence intervals, and survival analysis.
27. Record Summary
CheckMate-057 provides a clear example of randomized clinical-trial survival analysis. The trial enrolled 582 participants in a randomized parallel-group phase 3 design comparing nivolumab with docetaxel. Its registered primary endpoint was overall survival, defined from randomization to death, with participants who had not died censored at their last known date alive. The formal statistical analysis used a log-rank test and reported a hazard ratio of 0.73 for nivolumab over docetaxel, with a two-sided 95.92% confidence interval of 0.59–0.89 and P = 0.0015 under a superiority hypothesis.
The statistical lesson is broader than the p-value alone. The hazard ratio describes the relative treatment effect, the confidence interval describes uncertainty around that estimate, the log-rank test supplies the formal hypothesis-test result, and the Kaplan-Meier framework accommodates the censored time-to-event structure. At the same time, the ClinicalTrials.gov record does not provide enough information to reconstruct median survival, subgroup effects, multiplicity procedures, interim monitoring, or several other design details. A rigorous trial analysis therefore distinguishes reported evidence from assumptions that cannot be verified from the available record.