This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Trial-specific numerical results and methods on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
CheckMate-078 was a randomized, parallel-group, open-label phase 3 treatment trial comparing nivolumab with docetaxel in subjects previously treated with advanced or metastatic non-small cell lung cancer. The registry reports 504 enrolled participants, two treatment arms, two registered primary endpoints, and formal statistical analyses based on stratified Cox proportional-hazards and log-rank methods.
| Feature | CheckMate-078 |
|---|---|
| Trial name | CheckMate-078 |
| ClinicalTrials.gov identifier | NCT02613507 |
| Phase | Phase 3 |
| Condition | Non-Small Cell Lung Cancer |
| Population | Subjects previously treated with advanced or metastatic non-small cell lung cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 504 |
| Interventions | Nivolumab and Docetaxel |
| Primary endpoint types | Time-to-event |
| Hypothesis type | Superiority |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
| Trial status | Completed |
2. Clinical Question
The statistical question is whether nivolumab and docetaxel differ with respect to overall survival in subjects previously treated with advanced or metastatic non-small cell lung cancer. The registered primary endpoints are both time-to-event measures: Median Overall Survival and Overall Survival Rate.
Population
Subjects previously treated with advanced or metastatic non-small cell lung cancer.
Intervention
Nivolumab.
Comparator
Docetaxel.
Primary question
Does nivolumab produce a superior overall-survival outcome compared with docetaxel?
The registry classifies the primary hypothesis as superiority. This is important because the analysis is asking whether the randomized treatment groups differ in the direction specified by the superiority framework, rather than whether nivolumab is merely no worse than docetaxel by a prespecified non-inferiority margin.
3. Trial Design
Nivolumab
- Study intervention: nivolumab
- Included in the randomized treatment comparison
- Primary efficacy comparison uses all randomized participants
Docetaxel
- Study comparator: docetaxel
- Included in the randomized treatment comparison
- Primary efficacy comparison uses all randomized participants
The ClinicalTrials.gov record does not provide a treatment allocation ratio, individual arm enrollment counts for the primary randomized population, the randomization strata, or a detailed treatment schedule. Those details are therefore not inferred here.
4. Endpoints
| Registered primary endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Median Overall Survival | From randomization to the date of death or date participant was last known to be alive (assessed from December 2015 to Oct 2017 approximately 22 months)). OS was defined as the time between the date of randomization and the date of death. For subjects without documentation of death, OS was censored on the last date the subject was known to be alive. | Time-to-event |
| Overall Survival Rate | From first dose to the date of death or date participant was last known to be alive (assessed from December 2015 to Oct). OS was defined as the time between the date of randomization and the date of death. For subjects without documentation of death, OS was censored on the last date the subject was known to be alive. Rates provided are Kaplan-Meier estimates. | Time-to-event |
Why both endpoints are time-to-event measures
Median overall survival and an overall survival rate are two ways of describing the same underlying time-to-event process. The median is a location summary of the estimated survival distribution: it is the time at which the estimated survival probability reaches 0.50 when that point is estimable. A survival rate instead reports the estimated probability of remaining alive at a specified time point.
The statistical distinction matters. A median is not itself a treatment-effect measure. The hazard ratio from a Cox model summarizes the relative event rate under the fitted model, while a Kaplan-Meier survival rate gives an absolute survival probability at a specified time point.
5. Analysis Populations and Statistical Comparison
The registry identifies the analysis population for the posted primary endpoint analyses as all randomized participants. That population is important because the treatment comparison is anchored to the original randomized assignment.
| Feature | Registry-supported description |
|---|---|
| Analysis population | All randomized participants |
| Groups compared | Nivolumab vs Docetaxel |
| Endpoint | Median Overall Survival |
| Endpoint type | Time-to-event |
| Primary hypothesis | Superiority |
| Primary effect measure | Hazard ratio |
| Reported methods | Stratified Cox Proportional Hazard Model; Stratified weighted Log-Rank; Log Rank |
Three statistical analyses are posted for the primary endpoint. One supplies a hazard-ratio estimate and confidence interval, while two supply P-values from log-rank procedures. The registry therefore presents the treatment effect through both a model-based effect estimate and hypothesis-testing comparisons of the survival distributions.
6. Results: Median Overall Survival
The primary endpoint analysis compares nivolumab with docetaxel among all randomized participants. The registry reports a stratified Cox proportional-hazards model estimate together with a two-sided 97.7% confidence interval.
Stratified Cox proportional-hazards model
97.7% two-sided CI: 0.52–0.90
Groups compared: Nivolumab vs Docetaxel
| Primary analysis | Method | Effect / result |
|---|---|---|
| Median Overall Survival | Stratified Cox Proportional Hazard Model | HR 0.68; 97.7% two-sided CI 0.52–0.90 |
| Median Overall Survival | Stratified weighted Log-Rank | P = 0.0006 |
| Median Overall Survival | Log Rank | P = 0.0017 |
The reported HR of 0.68 means that, under the stratified Cox model, the estimated instantaneous rate of death for the nivolumab group was 0.68 times that of the docetaxel group over the analyzed time-to-event data. Expressed as a relative model-based comparison, this corresponds to an estimated 32% lower hazard of death for nivolumab relative to docetaxel, because 1 − 0.68 = 0.32.
The HR does not mean that 32% of patients avoided death, that survival was extended by 32%, or that every patient experienced the same reduction in risk. It is a relative hazard measure derived from a time-to-event model.
The 97.7% two-sided confidence interval of 0.52–0.90 describes statistical uncertainty around the estimated hazard ratio under the model and sampling framework. It is an interval for the estimated treatment effect, not a range containing the individual effects experienced by patients.
The confidence interval is also informative because it remains below 1.00. That means the range of model-compatible hazard-ratio values represented by the reported interval is on the lower-hazard side of the null value of 1.00.
The P-values of 0.0006 and 0.0017 address evidence against the corresponding null hypothesis under their respective log-rank procedures. A P-value does not measure the size of the treatment effect. The HR and its confidence interval provide the effect estimate and its precision; the P-value answers a different inferential question.
Because the analysis is a Cox proportional-hazards analysis, interpretation of a single HR also depends on the proportional-hazards framework. The ClinicalTrials.gov record does not report a formal assessment of that assumption, so the HR should not be treated as an absolute risk ratio or as a statement that the hazard difference is identical at every follow-up time.
Two log-rank analyses, two reported P-values
The registry reports both a stratified weighted log-rank analysis and a regular stratified log-rank analysis. The weighted analysis is described as Fleming and Harrington with \( \rho=0,\gamma=1 \). The regular analysis is explicitly described as the regular stratified log-rank test P-value.
Both are reported as two-sided superiority analyses for the Median Overall Survival endpoint. The ClinicalTrials.gov record does not state that either P-value should replace the Cox-model hazard ratio as the primary effect estimate.
What the weighted test changes
A standard log-rank test compares the observed and expected numbers of events across the groups throughout follow-up. A weighted log-rank test modifies the contribution of events at different portions of follow-up. In the registry-reported analysis text, the weighting is identified as Fleming and Harrington with \( \rho=0,\gamma=1 \).
This distinction is statistically important: the choice of weighting can affect sensitivity to different patterns of separation between survival curves. A weighted log-rank P-value should therefore be interpreted as the result of a specified weighted comparison, not as a second independent estimate of the hazard ratio.
7. Overall Survival Rate
The registry lists Overall Survival Rate as a second registered primary endpoint. Its definition states that OS rates are Kaplan-Meier estimates and uses the same underlying OS definition: time from randomization to death, with censoring at the last date the participant was known to be alive when death was not documented.
How this endpoint would normally be analysed
A Kaplan-Meier estimator constructs an estimated survival curve from observed event and censoring times. At a prespecified time \(t\), the curve gives an estimated probability of remaining event-free through that time. Confidence intervals can be attached to the estimated survival probability to quantify uncertainty.
Here, \(d_i\) represents the number of events at an event time and \(n_i\) the number at risk immediately before that time. The estimator naturally accommodates right-censored observations.
The registry's statement that survival rates are Kaplan-Meier estimates is therefore methodologically consistent with a time-to-event endpoint. But without the numerical rate or its specified assessment time in the ClinicalTrials.gov record, a quantitative survival-rate comparison cannot be reconstructed without adding information outside the ClinicalTrials.gov record.
8. Statistical Methodology
Kaplan-Meier estimation
The registry identifies the Overall Survival Rate as a Kaplan-Meier estimate. Kaplan-Meier analysis is designed for situations in which participants can have different follow-up times and some participants have not experienced the event by the end of observation.
The key feature is that a participant who is censored still contributes information up to the time of censoring. The method does not require assigning an unobserved death date to someone who was last known to be alive.
Stratified Cox proportional-hazards model
The primary effect estimate was obtained using a Stratified Cox Proportional Hazard Model. The Cox model relates the hazard of the event to treatment and, when applicable, other covariates or strata. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the numerator treatment group relative to the comparator.
For the reported HR of 0.68, the model-based estimated hazard under nivolumab is 68% of the corresponding hazard under docetaxel, subject to the model and stratification framework.
The word stratified is important. A stratified Cox model allows the baseline hazard to differ across strata while estimating a common treatment hazard ratio across those strata. The ClinicalTrials.gov record identifies the method as stratified but do not provide the individual stratification variables.
Log-rank testing
The log-rank test compares the survival experience of two groups using the sequence of observed events and numbers at risk over follow-up. The registry reports both a regular stratified log-rank test and a stratified weighted log-rank test.
The regular stratified log-rank analysis produced P = 0.0017. The weighted Fleming-Harrington analysis produced P = 0.0006. The two tests are related but are not identical statistical procedures.
Hazard ratio versus median survival
The endpoint is named Median Overall Survival, but the posted formal effect measure is a hazard ratio. These quantities should not be conflated.
| Measure | What it describes |
|---|---|
| Median overall survival | A location summary of the estimated survival-time distribution. |
| Hazard ratio | A model-based relative comparison of instantaneous event rates. |
| Kaplan-Meier survival rate | An estimated probability of remaining alive at a specified time. |
| Log-rank P-value | Evidence against a null hypothesis under a survival-distribution comparison. |
Consequently, a hazard ratio of 0.68 should not be described as though it were a median-survival ratio. A hazard ratio and a median are different statistical quantities and can convey different aspects of the treatment comparison.
Analysis population and randomization
The registry specifies all randomized participants for the primary endpoint analyses. This is consistent with the central role of randomization in estimating the treatment effect: participants are compared according to their randomized group rather than being reclassified according to later observations.
The ClinicalTrials.gov record does not specify the exact randomization mechanism, block structure, stratification factors, or allocation ratio. Those details are not inferred.
9. Statistical Methods Explained
Why use a Cox proportional-hazards model for overall survival?
Overall survival is a time-to-event outcome, and participants can have different follow-up durations. Some may die during observation while others may remain alive when their follow-up ends. A Cox model uses the timing of events and censoring while estimating a relative hazard between treatment groups.
Its main output here is the hazard ratio. That gives a compact relative measure of the difference between nivolumab and docetaxel without requiring every participant to have the same observation time.
What does an HR of 0.68 mean?
An HR of 0.68 means the fitted model estimates the hazard in the nivolumab group at 0.68 times the hazard in the docetaxel group, under the model's assumptions. Equivalently, the estimated relative reduction in hazard is 32%.
It does not mean that 32% of patients survive, that survival time increases by 32%, or that the absolute probability of death is 32% lower at every time point.
Why report a confidence interval with the hazard ratio?
A point estimate alone does not show how precisely the treatment effect has been estimated. The reported 97.7% two-sided interval of 0.52–0.90 provides an uncertainty range around the HR of 0.68 under the specified inferential framework.
The width of the interval also matters. A narrower interval indicates greater precision than a wider interval, all else equal. The confidence interval is therefore complementary to the point estimate rather than a decorative addition to it.
Why are there both a weighted and regular log-rank test?
The standard log-rank test gives broadly distributed weight to events over follow-up, while a weighted log-rank procedure changes how events contribute to the test statistic. The registry specifies a Fleming-Harrington weighted test with \( \rho=0,\gamma=1 \) in addition to a regular stratified log-rank test.
These tests can be useful for examining the survival comparison under different weighting schemes. Their P-values should not be interpreted as two separate estimates of treatment effect.
Why is the analysis stratified?
Stratification allows the survival comparison to account for specified strata without requiring the same baseline hazard in every stratum. In a stratified Cox model, the baseline hazard can vary between strata while the treatment effect is estimated across the strata.
The ClinicalTrials.gov record confirms that the Cox and log-rank analyses were stratified, but they do not identify the variables used to define the strata.
Why is censoring important?
For OS, the registry states that participants without documentation of death were censored on the last date they were known to be alive. Censoring is therefore part of the endpoint definition and directly affects the Kaplan-Meier and Cox calculations.
Censoring is not equivalent to survival forever. It means that the analysis uses the information available through the censoring time without observing the participant's subsequent event status within the relevant observation period.
10. Primary Analysis Framework
| Component | CheckMate-078 |
|---|---|
| Primary endpoint | Median Overall Survival |
| Endpoint type | Time-to-event |
| Analysis population | All randomized participants |
| Treatment comparison | Nivolumab vs Docetaxel |
| Primary hypothesis | Superiority |
| Model-based method | Stratified Cox Proportional Hazard Model |
| Model-based effect measure | Hazard ratio |
| Model estimate | 0.68 |
| Confidence interval | 97.7% two-sided CI 0.52–0.90 |
| Weighted comparison | Stratified weighted Log-Rank |
| Weighted P-value | 0.0006 |
| Regular comparison | Log Rank |
| Regular P-value | 0.0017 |
The combination of a model-based effect estimate and log-rank testing gives two complementary views of the treatment comparison. The Cox model provides the numerical hazard-ratio estimate and its confidence interval, while the log-rank procedures test whether the survival experience differs under their respective weighting frameworks.
11. Interpreting the P-values
P = 0.0006
This is the two-sided P-value reported for the stratified weighted Fleming-Harrington log-rank analysis.
P = 0.0017
This is the two-sided P-value reported for the regular stratified log-rank analysis.
Both P-values are associated with superiority analyses of the Median Overall Survival endpoint. They provide evidence against the relevant null hypothesis under the corresponding testing procedure.
A P-value is not the probability that the treatment has no effect, nor is it the probability that the reported hazard ratio is correct. It also does not quantify the magnitude of benefit.
For magnitude, the HR of 0.68 is the relevant reported effect estimate. For uncertainty, the 97.7% two-sided CI of 0.52–0.90 is the relevant interval. The P-values address a different question: how compatible the observed test statistic is with the null hypothesis under the specified test.
12. Crossover and Safety Information
The ClinicalTrials.gov record includes serious adverse-event counts by arm. They also identify a separate crossover nivolumab group. Because these data are provided as affected participants divided by participants at risk, they are reproduced without converting them into percentages or adding an interpretation that would require additional information.
| Group | Serious adverse events |
|---|---|
| Nivolumab | 183/337 |
| Docetaxel | 73/156 |
| Cross-over Nivolumab | 4/9 |
The denominator is important. The registry data identify these as affected/at-risk counts, so the numbers should not be treated as though all groups had the same analysis population or exposure structure.
13. Trial Timeline
Trial start
The registry profile gives December 11, 2015 as the study start date.
Primary completion
The registry profile gives September 15, 2017 as the primary completion date.
Current registry status in the ClinicalTrials.gov record
the ClinicalTrials.gov record identifies the study as completed and reports that results have been posted.
14. What the Hazard Ratio Does — and Does Not — Mean
The reported HR of 0.68 indicates a lower estimated instantaneous rate of death in the nivolumab group relative to the docetaxel group under the stratified Cox model. The simple transformation \(1-0.68\) gives a 32% estimated relative reduction in hazard.
This is a model-based relative comparison. It does not mean that every patient has a 32% lower probability of death, nor does it imply a 32% increase in expected survival time.
The 97.7% two-sided CI of 0.52–0.90 quantifies uncertainty around the reported HR under the statistical framework used for the analysis. It gives information about precision that cannot be obtained from the point estimate alone.
The interval remains below 1.00, the conventional null value for a hazard ratio. That is consistent with the superiority direction represented by the reported effect estimate.
The reported P-values of 0.0006 for the stratified weighted log-rank test and 0.0017 for the regular stratified log-rank test quantify evidence under two specified testing procedures. Neither P-value is an effect-size measure.
The Overall Survival Rate endpoint is defined using Kaplan-Meier estimates, which would provide absolute survival probabilities at specified time points. The ClinicalTrials.gov record does not include numerical survival-rate estimates, so an absolute survival comparison is not added here.
15. Multiplicity, Interim Analysis, and Other Design Topics
The ClinicalTrials.gov record provides information about the primary endpoints, the superiority hypothesis, the posted statistical analyses, and the two registered time-to-event endpoints. They do not provide a prespecified multiplicity strategy, alpha-spending plan, interim-analysis schedule, sample-size calculation, non-inferiority margin, factorial structure, Bayesian analysis, or missing-data/imputation strategy.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record; the registered hypothesis is superiority. |
| Factorial design | Not reported; the design model is parallel. |
| Bayesian methods | Not reported. |
| Interim analysis | Not reported in the ClinicalTrials.gov record. |
| Multiplicity adjustment | Not reported in the ClinicalTrials.gov record. |
| Missing-data / imputation method | Not reported in the ClinicalTrials.gov record. |
| Stratified analysis | Supported: the Cox model and log-rank analyses are reported as stratified. |
| Crossover | Supported by the separate “Cross-over Nivolumab” safety group. |
This distinction is important for statistical interpretation. Absence of a method in the ClinicalTrials.gov record is not evidence that the method was absent from the underlying protocol or statistical analysis plan. It simply means that the method cannot be reconstructed from the ClinicalTrials.gov record.
16. Limitations
- Limited numerical result reporting: the statistical analyses posted on ClinicalTrials.gov provide a hazard ratio and confidence interval plus two P-values, but do not provide a numerical median overall survival estimate.
- Overall survival rate details: the registry identifies Kaplan-Meier estimates but the ClinicalTrials.gov record does not provide the numerical rates or their confidence intervals.
- Incomplete stratification information: the analyses are explicitly stratified, but the ClinicalTrials.gov record does not identify the stratification factors.
- Proportional-hazards assumption: the Cox model provides a single hazard ratio, but the ClinicalTrials.gov record does not report an assessment of proportional hazards.
- Log-rank interpretation: the weighted and regular log-rank tests use different procedures. Their P-values should not be interpreted as two independent estimates of treatment effect.
- Multiplicity: the ClinicalTrials.gov record does not describe how multiple primary endpoints or multiple analyses were handled in the type I error framework.
- Interim monitoring: no interim-analysis schedule or alpha-spending method is provided in the ClinicalTrials.gov record.
- Crossover information: a cross-over nivolumab safety group is reported, but the ClinicalTrials.gov record does not provide enough detail to quantify its impact on efficacy interpretation.
- Safety population: serious adverse events are reported as affected/at-risk counts, but the ClinicalTrials.gov record does not fully define the safety-analysis rules or exposure windows.
- Generalizability: the trial population is specifically described as subjects previously treated with advanced or metastatic non-small cell lung cancer; the ClinicalTrials.gov record does not provide a detailed baseline-characteristics table for assessing applicability to other populations.
17. Why This Trial Matters Statistically
CheckMate-078 is a useful teaching example because its primary evidence is built around the classic time-to-event framework: randomized treatment assignment, censoring, Kaplan-Meier estimation, stratified log-rank testing, and a stratified Cox proportional-hazards model.
| Statistical concept | How it appears in CheckMate-078 |
|---|---|
| Randomization | The allocation is randomized and the design is parallel. |
| Time-to-event endpoint | Median Overall Survival and Overall Survival Rate are registered primary endpoints. |
| Kaplan-Meier estimation | The registry states that Overall Survival Rates are Kaplan-Meier estimates. |
| Hazard ratio | The primary formal effect measure is a hazard ratio. |
| Cox model | A stratified Cox proportional-hazards model supplies HR 0.68. |
| Confidence interval | A 97.7% two-sided CI of 0.52–0.90 accompanies the HR. |
| Log-rank testing | Both weighted and regular stratified log-rank analyses are reported. |
| Weighted survival comparison | The weighted test uses a Fleming-Harrington specification with \( \rho=0,\gamma=1 \). |
| Superiority testing | The registry identifies the hypothesis type as superiority. |
| Censoring | Participants without documented death are censored at the last date known to be alive. |
| Analysis population | Primary analyses use all randomized participants. |
| Crossover | A separate cross-over nivolumab group is included in the reported serious-AE data. |
18. A Deeper Statistical Reading of the Primary Result
The treatment effect is relative, not absolute
The HR of 0.68 is a relative measure. Suppose two patients had different baseline risks because of differences in disease characteristics. A relative hazard comparison does not imply that their absolute changes in probability of death would be identical.
This is why time-specific Kaplan-Meier estimates and the hazard ratio answer complementary questions. The hazard ratio summarizes relative event-rate differences, whereas the Kaplan-Meier curve and survival rate describe the absolute survival experience over time.
The confidence interval carries information that the P-value does not
The P-values of 0.0006 and 0.0017 indicate strong evidence against the corresponding null hypotheses under the specified log-rank procedures. But those values alone do not tell the reader whether the estimated effect is modest or large.
The HR of 0.68 supplies that effect estimate, and the 97.7% confidence interval of 0.52–0.90 describes its uncertainty. A statistically persuasive result can still have a confidence interval that is important to inspect for clinical and statistical precision.
The median is a different quantity from the hazard ratio
The endpoint name is “Median Overall Survival,” but the posted Cox analysis does not give a median value in the ClinicalTrials.gov record. It instead gives a hazard ratio. It would therefore be incorrect to substitute the HR for the median or to infer a median from the HR alone.
In general, two survival distributions can have similar or different medians while producing different hazard-ratio behavior over time. The underlying survival curves contain information that is compressed when represented by a single summary measure.
Weighted versus unweighted evidence
The weighted Fleming-Harrington analysis gives P = 0.0006, while the regular stratified log-rank analysis gives P = 0.0017. The fact that the registry reports both is useful statistically because it shows that the analysis included more than one way of comparing survival distributions.
However, the existence of two P-values does not create two independent treatment effects. The effect estimate remains the reported HR of 0.68 from the stratified Cox model.
19. Related Tutorials
Learn more about the methods used in this trial:
20. Related Statistical Calculators
21. Sources
- ClinicalTrials.gov: NCT02613507 — CheckMate-078.
- Linked PubMed publication: PubMed record for PMID 36897427.
- Linked PubMed publication: PubMed record for PMID 30659987.
Continue through the Clinical Biostats statistical pathway
Use the trial's time-to-event methods as a starting point for deeper study of Kaplan-Meier estimation, hazard ratios, Cox models, log-rank testing, confidence intervals, and related clinical-trial methods.
22. Record Summary
CheckMate-078 is a randomized phase 3, parallel-group trial comparing nivolumab with docetaxel in previously treated subjects with advanced or metastatic non-small cell lung cancer. The ClinicalTrials.gov record defines two primary time-to-event endpoints and report three formal primary-endpoint analyses. The principal model-based result is a hazard ratio of 0.68 for nivolumab versus docetaxel, with a 97.7% two-sided confidence interval of 0.52–0.90. The registry also reports a stratified weighted log-rank P-value of 0.0006 and a regular stratified log-rank P-value of 0.0017.
The statistical interpretation is strongest when these quantities are kept conceptually separate. The hazard ratio describes the relative event-rate comparison under the Cox model; the confidence interval describes uncertainty around that estimate; and the P-values describe evidence under the specified log-rank tests. The registered Overall Survival Rate endpoint adds an absolute Kaplan-Meier perspective, but the ClinicalTrials.gov record does not contain numerical survival-rate estimates.
The trial also illustrates several broader principles of clinical-trial statistics: randomized analysis populations preserve the treatment comparison established at randomization; censoring must be incorporated into time-to-event analysis; stratification affects the analysis framework; weighted and unweighted log-rank tests answer related but distinct questions; and a hazard ratio should never be mistaken for an absolute risk, median survival, or probability of benefit.