This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
CheckMate-459 was a randomized, parallel, phase 3 study comparing nivolumab with sorafenib as a first treatment in patients with advanced hepatocellular carcinoma. The trial enrolled 743 participants and had two treatment arms.
| Feature | CheckMate-459 |
|---|---|
| Trial name | CheckMate-459 |
| ClinicalTrials.gov identifier | NCT02576509 |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Hepatocellular Carcinoma |
| Population | Patients with advanced hepatocellular carcinoma receiving a first treatment |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | Industry |
| Enrollment | 743 |
| Interventions | Nivolumab; Sorafenib |
2. Clinical Question
The central clinical question was whether nivolumab differed from sorafenib when used as a first treatment in patients with advanced hepatocellular carcinoma, with overall survival serving as the registered primary endpoint.
Population
Patients with advanced hepatocellular carcinoma receiving a first treatment.
Intervention
Nivolumab 240 mg.
Comparator
Sorafenib 400 mg.
Primary question
What is the difference between nivolumab and sorafenib in overall survival among all randomized participants?
3. Trial Design
Nivolumab
- Nivolumab 240 mg.
Sorafenib
- Sorafenib 400 mg.
The ClinicalTrials.gov record does not specify a randomization ratio, stratification factors, crossover provisions, factorial structure, or an interim-analysis design. Those features are therefore not assumed here.
4. Endpoints
The registered primary endpoint was overall survival. ClinicalTrials.gov also posted secondary analyses covering objective response rate, progression-free survival, and efficacy according to PD-L1 expression.
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Overall Survival (OS) | OS is defined as the time from the date of randomization to the date of death due to any cause in all randomized participants. Participants who are alive will be censored at the last known alive dates. Based on Kaplan-Meier Estimates. Time frame: time from the date of randomization to the date of death due to any cause, assessed up to June 2019 (approximately 41 months) | Time-to-event |
| Objective Response Rate (ORR) Per BICR RECIST 1.1 | Time frame: the date of randomization and the date of first objectively documented progression or the date of subsequent anti-cancer therapy, whichever occurs first, assessed up to May 2019 (approximately 40 months) | Binary |
| Progression-Free Survival (PFS) | Time frame: time from the date of randomization to the date of the first objectively documented tumor progression or death, assessed up to May 2019 (approximately 40 months) | Time-to-event |
| Efficacy Based on PD-L1 Expression - OS and PFS | Time frame: the date of randomization and the date of first objectively documented progression or the date of subsequent anti-cancer therapy, whichever occurs first, assessed up to May 2019 (approximately 40 months) | Time-to-event |
| Efficacy Based on PD-L1 Expression - ORR | Time frame: the date of randomization and the date of first objectively documented progression or the date of subsequent anti-cancer therapy, whichever occurs first, assessed up to May 2019 (approximately 40 months) | Binary |
5. Statistical Methodology
Overall survival and time-to-event analysis
The primary OS analysis used a log-rank test, with the analysis text specifying that the test was stratified by the stratification factors entered into the IVRS. A stratified Cox proportional-hazards model was used for the hazard-ratio estimate and confidence interval.
The reported comparison was nivolumab 240 mg versus sorafenib 400 mg in all randomized participants.
Cochran-Mantel-Haenszel analysis
The ORR comparison was analyzed using a Cochran-Mantel-Haenszel (CMH) method. The analysis text states that the estimate of the difference in ORRs was based on CMH weighting and was stratified by the stratification factors.
Odds-ratio analysis
The registry also reports an odds ratio for ORR. An odds ratio compares the odds of response between the two treatment groups; it is not the same quantity as a difference in response percentages.
Stratified analysis
Stratification allows a comparison to account for prespecified grouping factors rather than treating all participants as though they belonged to one homogeneous stratum. For the primary OS analysis, the registry specifically reports a log-rank test stratified by the stratification factors as entered into the IVRS.
Multiplicity
The primary OS analysis notes a confidence interval adjusted for multiplicity: 95.81% CI (0.72 to 1.02). The secondary ORR and PFS analyses explicitly state that no test was performed because the OS p-value was above the prior threshold. This is an important part of interpreting the reported endpoint hierarchy.
6. Primary Result: Overall Survival
The primary endpoint was overall survival among all randomized participants, comparing nivolumab 240 mg with sorafenib 400 mg. The analysis used a stratified log-rank test and a stratified Cox proportional-hazards model.
Hazard ratio for overall survival
95% CI: 0.71–1.02 · P = 0.0752
Two-sided confidence interval; superiority hypothesis.
| Primary OS analysis | Reported result |
|---|---|
| Analysis population | All randomized participants |
| Comparison | Nivolumab 240 mg vs Sorafenib 400 mg |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.85 |
| 95% CI | 0.71–1.02 |
| P-value | 0.0752 |
| Hypothesis type | Superiority |
| Multiplicity-adjusted CI noted in analysis text | 95.81% CI: 0.72 to 1.02 |
The reported HR of 0.85 means that, under the fitted proportional-hazards model, the estimated instantaneous rate of death in the nivolumab group was 0.85 times that in the sorafenib group over the analyzed follow-up. Expressed as a simple relative interpretation, this corresponds to a 15% lower estimated hazard, not a 15% reduction in the probability that an individual participant will die.
The 95% CI of 0.71–1.02 describes statistical uncertainty around the estimated hazard ratio. Because the interval extends above 1, the ClinicalTrials.gov record does not establish that the true hazard ratio is below 1 under a conventional two-sided interpretation. The separately reported multiplicity-adjusted 95.81% CI of 0.72 to 1.02 likewise extends above 1.
The p-value of 0.0752 is a measure of compatibility with the statistical testing framework under the null hypothesis; it is not a measure of how large or clinically important the treatment effect is. A p-value should therefore be interpreted together with the hazard ratio, confidence interval, analysis population, and trial design.
Finally, a Cox hazard ratio relies on a proportional-hazards model. If the relative hazards change materially over time, one summary HR can conceal important features of the survival curves. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption, so no such assessment is inferred here.
Why the primary endpoint drives the interpretation of the secondary analyses
The registry's secondary analysis records repeatedly state that no test was performed due to the OS p-value result above the prior threshold. This means that the secondary estimates can still be described as estimates with confidence intervals, but they should not be converted into claims of statistically confirmed superiority based solely on whether an interval happens to exclude a particular value.
7. Secondary Results: Objective Response Rate
Objective response rate per BICR RECIST 1.1 was analyzed among all randomized participants. The registry-reported analysis used the Cochran-Mantel-Haenszel method and reports both a difference in ORRs and an odds ratio.
Difference in Objective Response Rates
Difference in ORRs
95% CI: 3.9–12.7
Estimate defined as Nivolumab minus Sorafenib; two-sided confidence interval.
| ORR analysis | Reported result |
|---|---|
| Analysis population | All randomized participants |
| Comparison | Nivolumab 240 mg vs Sorafenib 400 mg |
| Method | Cochran-Mantel-Haenszel test / weighting method |
| Effect measure | Difference of ORRs |
| Estimate | 8.3 |
| 95% CI | 3.9–12.7 |
| Formal test | No test was performed due to OS p-value result above the prior threshold |
Odds Ratio for Objective Response
Odds ratio for response
95% CI: 1.48–3.92
Odds ratio: Nivolumab over Sorafenib.
The odds ratio of 2.41 compares the odds of response, not the response probability itself. An odds ratio of 2.41 therefore should not be described as meaning that the response rate was 2.41 times as high.
8. Secondary Result: Progression-Free Survival
Progression-free survival was analyzed among all randomized participants. The registry reports a stratified Cox proportional-hazards model comparing nivolumab 240 mg with sorafenib 400 mg.
Hazard ratio for progression-free survival
95% CI: 0.79–1.10
Two-sided confidence interval; no test was performed due to the OS p-value result above the prior threshold.
| PFS analysis | Reported result |
|---|---|
| Analysis population | All randomized participants |
| Comparison | Nivolumab 240 mg vs Sorafenib 400 mg |
| Method | Stratified Cox proportional-hazards model |
| Effect measure | Hazard ratio |
| Estimate | 0.93 |
| 95% CI | 0.79–1.10 |
| Formal test | No test was performed due to OS p-value result above the prior threshold |
An HR of 0.93 corresponds to an estimated hazard that is 0.93 times the comparator hazard under the fitted model. The confidence interval of 0.79–1.10 indicates uncertainty that includes values below and above 1. Because the registry does not report a formal p-value for this secondary comparison, the result should be interpreted as an estimated relative treatment effect rather than a separately tested superiority claim.
9. Secondary Results: Efficacy by PD-L1 Expression
The registry contains six hazard-ratio analyses under the outcome measure Efficacy Based on PD-L1 Expression - OS and PFS. The registry-reported population description includes participants with >=1% PD-L1 expression, participants with <1% PD-L1 expression, and participants without PD-L1 quantifiable. The registry does not report a statistical method for these analyses.
| PD-L1 subgroup / endpoint | HR | 95% CI | Formal test |
|---|---|---|---|
| PD-L1 >=1%, OS | 0.80 | 0.54–1.19 | No test was performed |
| PD-L1 >=1%, PFS | 0.71 | 0.48–1.03 | No test was performed |
| PD-L1 <1%, OS | 0.84 | 0.69–1.02 | No test was performed |
| PD-L1 <1%, PFS | 0.98 | 0.81–1.17 | No test was performed |
| Without PD-L1 quantifiable, OS | 1.26 | 0.34–4.74 | No test was performed |
| Without PD-L1 quantifiable, PFS | 0.98 | 0.27–3.52 | No test was performed |
These estimates illustrate why subgroup analysis should be read through the confidence interval rather than by comparing point estimates alone. For example, the subgroup without quantifiable PD-L1 has much wider intervals, including values well below and above 1. That width indicates substantially greater statistical uncertainty in those estimates.
10. Secondary Results: PD-L1 and Objective Response
The registry also reports odds ratios for ORR within two PD-L1 categories. The analysis text describes these as unstratified odds ratios with associated unstratified 95% exact confidence intervals.
| PD-L1 subgroup | OR | 95% CI | Formal test |
|---|---|---|---|
| PD-L1 >=1%, ORR | 3.79 | 1.41–10.17 | No test was performed |
| PD-L1 <1%, ORR | 1.95 | 1.10–3.45 | No test was performed |
The odds ratios are treatment-group comparisons of the odds of response within the specified PD-L1 categories. They should not be interpreted as response-rate ratios. Their confidence intervals describe uncertainty around the corresponding odds-ratio estimates, while the absence of a formal test means that these results should not be treated as independent confirmatory findings.
11. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm using affected participants divided by participants at risk.
| Safety measure | Nivolumab 240 mg | Sorafenib 400 mg |
|---|---|---|
| Serious adverse events, affected / at risk | 216 / 367 | 216 / 363 |
The two arms have the same reported number of affected participants, 216, while the denominators differ. The ClinicalTrials.gov record therefore support reporting the affected and at-risk counts directly. No additional safety categories, exposure-adjusted rates, or statistical safety comparisons are reported in the ClinicalTrials.gov record.
12. Statistical Methods Explained
Why was a log-rank test used for overall survival?
Overall survival is a time-to-event endpoint because both the event time and censoring matter. The log-rank test compares the survival experience of two randomized groups across follow-up while accounting for the timing of observed deaths and censored observations. In this trial, the registry reports a stratified log-rank analysis.
What does an OS hazard ratio of 0.85 mean?
A hazard ratio of 0.85 means that the estimated instantaneous death rate in the nivolumab group was 0.85 times the corresponding rate in the sorafenib group under the fitted Cox model. It does not mean that 15% of participants benefited, that survival increased by 15%, or that an individual's probability of death was reduced by exactly 15%.
Why is the confidence interval important?
The point estimate is only one summary of the data. The 95% confidence interval of 0.71–1.02 shows the uncertainty around the OS hazard-ratio estimate. Its upper endpoint crosses 1, so the interval does not exclude no difference in hazard under the usual two-sided reference point.
Why does the p-value not measure effect size?
A p-value summarizes how compatible the observed data are with a specified null hypothesis under the statistical testing framework. It is affected by the amount of information and does not directly quantify the magnitude of a treatment effect. The HR and its confidence interval are therefore necessary to understand the estimated effect itself.
Why use a Cochran-Mantel-Haenszel method for ORR?
ORR is a binary endpoint, so each participant contributes a response or nonresponse classification rather than an event time. The CMH method provides a way to compare treatment groups while weighting across strata. In this trial, the registry states that the difference in ORRs was based on CMH weighting and stratified by the stratification factors.
Why should the ORR odds ratio not be read as a response-rate ratio?
An odds ratio compares odds, where odds are the probability of an event divided by the probability of no event. A response odds ratio of 2.41 therefore means the estimated odds of response were 2.41 times those in the comparator group; it does not mean that the percentage responding was 2.41 times as large.
Why do the secondary analyses need a multiplicity caution?
The ClinicalTrials.gov record explicitly state that no test was performed for the secondary ORR, PFS, PD-L1 OS/PFS, and PD-L1 ORR analyses because the OS p-value was above the prior threshold. This creates an important distinction between reporting an effect estimate and declaring an independently tested statistical finding.
13. Understanding the Primary Hazard Ratio in Context
The OS estimate of 0.85 is a relative time-to-event measure. In simple terms, 0.85 indicates an estimated hazard 15% lower for nivolumab than sorafenib under the fitted model.
The 0.71–1.02 confidence interval is relatively more informative than the point estimate alone because it shows the uncertainty surrounding 0.85. The interval includes 1, so the data do not rule out no difference in the hazard ratio under the conventional reference value.
The reported P = 0.0752 should be interpreted within the trial's superiority framework and the stated multiplicity procedure. It should not be converted into an estimate of the probability that one treatment is better than the other.
The ClinicalTrials.gov record does not report median OS, median PFS, time-specific survival rates, event counts for OS or PFS, Kaplan-Meier coordinates, or formal proportional-hazards diagnostics. Those quantities are therefore not inferred or reconstructed on this page.
14. Confidence Intervals Across the Secondary Analyses
The collection of secondary estimates demonstrates how the width of a confidence interval can vary substantially across analyses. The overall PFS estimate has a 95% CI of 0.79–1.10, whereas the estimate for OS among participants without quantifiable PD-L1 has a much wider 95% CI of 0.34–4.74.
| Analysis | Estimate | 95% CI | What the interval communicates |
|---|---|---|---|
| OS, primary | HR 0.85 | 0.71–1.02 | Uncertainty around the primary hazard-ratio estimate |
| PFS | HR 0.93 | 0.79–1.10 | Uncertainty around the secondary hazard-ratio estimate |
| PD-L1 >=1%, OS | HR 0.80 | 0.54–1.19 | Subgroup uncertainty includes values below and above 1 |
| PD-L1 <1%, OS | HR 0.84 | 0.69–1.02 | Narrower subgroup interval that still extends above 1 |
| Without quantifiable PD-L1, OS | HR 1.26 | 0.34–4.74 | Wide uncertainty around the subgroup estimate |
| PD-L1 >=1%, ORR | OR 3.79 | 1.41–10.17 | Uncertainty around the odds of response |
The appropriate lesson is not that a narrower interval is automatically more important. Rather, confidence-interval width should be considered when judging how much information an estimate contains. Subgroup analyses often have fewer participants than the overall randomized population, which can produce wider intervals.
15. Multiplicity and Endpoint Hierarchy
The registry identifies overall survival as the single primary endpoint. The statistical analyses include one primary endpoint analysis and eleven secondary analyses. The secondary analysis records repeatedly note that no test was performed because the OS p-value was above the prior threshold.
| Analysis group | Role in the ClinicalTrials.gov record | Interpretation |
|---|---|---|
| Overall survival | Primary endpoint | Formal superiority analysis with log-rank testing and Cox HR estimation |
| Objective response rate | Secondary endpoint | Effect estimates reported; formal test not performed |
| Progression-free survival | Secondary endpoint | HR and CI reported; formal test not performed |
| PD-L1 OS/PFS | Secondary endpoint | Subgroup HR estimates and CIs reported; no tests performed |
| PD-L1 ORR | Secondary endpoint | OR estimates and CIs reported; no tests performed |
This hierarchy prevents a common interpretive error: treating every confidence interval in a registry as if it represents an independently tested confirmatory hypothesis. Here, the registry explicitly supplies a sequential testing context in which the OS result governs whether the listed secondary tests were performed.
16. Analysis Population and Randomization
The primary OS analysis population was all randomized participants. This is important because randomized treatment assignment is the foundation of the principal comparative analysis.
Why randomization matters
Randomization creates the basis for comparing outcomes between treatment groups without relying solely on observed baseline characteristics to explain differences.
Why ITT-style analysis matters
Analyzing participants according to randomized assignment preserves the treatment comparison established by randomization. The registry-reported primary analysis is explicitly based on all randomized participants.
Why censoring matters
For OS, participants who remain alive are censored at their last known alive dates. Censoring allows their available follow-up information to contribute without assigning an unobserved death date.
Why analysis population should be stated
A treatment effect cannot be interpreted correctly without knowing which participants contributed to the analysis. The primary OS analysis explicitly identifies all randomized participants.
17. Time-to-Event Endpoints: What the Model Is Estimating
Both OS and PFS are time-to-event endpoints. Unlike a simple binary endpoint, they use information about when the event occurred and whether a participant was censored before experiencing the event.
An HR below 1 corresponds to a lower estimated instantaneous event rate in the nivolumab group; an HR above 1 corresponds to a higher estimated instantaneous event rate.
The Cox model compresses a potentially complex time-varying pattern into a single relative effect estimate. That summary is useful, but its interpretation depends on the proportional-hazards framework. The ClinicalTrials.gov record does not provide a formal diagnostic of that assumption.
18. Objective Response: Difference Versus Odds Ratio
The registry provides two distinct effect measures for the overall ORR analysis: a difference of ORRs and an odds ratio. These quantities answer related but different questions.
| Measure | Reported estimate | 95% CI | Interpretive question |
|---|---|---|---|
| Difference of ORRs | 8.3 | 3.9–12.7 | How far apart are the response rates on the reported difference scale? |
| Odds ratio | 2.41 | 1.48–3.92 | How do the odds of response compare between the treatment groups? |
The difference is expressed on an additive scale, whereas the odds ratio is a multiplicative measure on the odds scale. They should not be substituted for one another. The registry's CMH analysis also indicates that stratification was incorporated into the ORR difference analysis.
19. PD-L1 Subgroup Interpretation
The PD-L1 analyses are useful for demonstrating the distinction between a subgroup estimate and a formal test of treatment-effect heterogeneity.
PD-L1 >=1%
The reported OS HR was 0.80 with a 95% CI of 0.54–1.19, while the PFS HR was 0.71 with a 95% CI of 0.48–1.03.
PD-L1 <1%
The reported OS HR was 0.84 with a 95% CI of 0.69–1.02, while the PFS HR was 0.98 with a 95% CI of 0.81–1.17.
PD-L1 not quantifiable
The OS HR was 1.26 with a 95% CI of 0.34–4.74, and the PFS HR was 0.98 with a 95% CI of 0.27–3.52.
Formal subgroup comparison
The ClinicalTrials.gov record does not report a treatment-by-PD-L1 interaction test. The subgroup point estimates should therefore not be used alone to establish effect modification.
20. Serious Adverse Events and Denominator Awareness
The safety data provide an instructive example of why denominators should accompany adverse-event counts.
| Arm | Affected | At risk | Reported format |
|---|---|---|---|
| Nivolumab 240 mg | 216 | 367 | 216/367 |
| Sorafenib 400 mg | 216 | 363 | 216/363 |
Both arms have 216 affected participants, but the numbers at risk differ. Consequently, simply comparing the numerators would not provide the complete context for the reported safety measure. The ClinicalTrials.gov record does not include a formal statistical comparison of serious adverse-event rates, so none is inferred here.
21. What the Registry Does Not Report in the Supplied Data
The ClinicalTrials.gov record is sufficient for a detailed statistical analysis of the reported estimates, but several commonly reported clinical-trial quantities are not included. Clinical Biostats therefore does not reconstruct them.
- Median overall survival: not reported in the ClinicalTrials.gov record.
- Median progression-free survival: not reported in the ClinicalTrials.gov record.
- Kaplan-Meier survival probabilities: not reported in the ClinicalTrials.gov record.
- OS and PFS event counts: not reported in the ClinicalTrials.gov record.
- Baseline characteristics: not reported in the ClinicalTrials.gov record.
- Formal proportional-hazards diagnostics: not reported in the ClinicalTrials.gov record.
- Detailed randomization stratification factors: not reported in the ClinicalTrials.gov record.
- Crossover information: not reported in the ClinicalTrials.gov record.
- Interim-analysis rules: not reported in the ClinicalTrials.gov record.
- Missing-data or imputation procedures: not reported in the ClinicalTrials.gov record.
- Formal subgroup interaction tests: not reported in the ClinicalTrials.gov record.
- Bayesian methods: not reported in the ClinicalTrials.gov record.
22. Limitations
- Registry-level detail: the ClinicalTrials.gov record does not contain the full protocol or statistical analysis plan, so some design details cannot be reconstructed.
- Single primary endpoint: OS is the registered primary endpoint, while the other posted analyses are secondary.
- Multiplicity: the secondary analyses state that no test was performed because the OS p-value exceeded the prior threshold. Estimates should therefore be distinguished from confirmatory testing.
- Subgroup uncertainty: PD-L1 subgroup estimates can be substantially less precise than the overall analysis, particularly where the confidence intervals are wide.
- No interaction analysis reported: the ClinicalTrials.gov record does not establish whether treatment effects differ statistically between PD-L1 categories.
- Proportional-hazards assumption: the Cox model provides a single HR summary, but the ClinicalTrials.gov record does not report a formal assessment of whether proportional hazards holds.
- Censoring: OS analysis includes censoring of participants who were alive at their last known alive dates, so interpretation depends on the underlying censoring framework.
- Safety denominator: serious adverse-event data are reported as affected participants divided by participants at risk, and no formal comparative safety analysis is reported.
23. Why This Trial Matters Statistically
CheckMate-459 is a useful statistical teaching case because the ClinicalTrials.gov record brings together randomized treatment comparison, a primary time-to-event endpoint, stratified log-rank testing, Cox proportional-hazards modeling, multiplicity adjustment, CMH analysis of a binary endpoint, odds ratios, subgroup analyses, and an explicit distinction between estimation and formal hypothesis testing.
| Concept | How it appears in CheckMate-459 |
|---|---|
| Randomization | The trial is recorded as randomized with two treatment arms. |
| Parallel-group design | The design model is recorded as parallel. |
| Time-to-event endpoint | Overall survival is the registered primary endpoint. |
| Kaplan-Meier estimation | The OS definition states that estimates are based on Kaplan-Meier estimates. |
| Log-rank test | The primary OS comparison uses a stratified log-rank test. |
| Cox proportional-hazards model | The primary analysis reports a stratified Cox model for the HR. |
| Hazard ratio | OS and PFS are summarized using HRs. |
| Confidence intervals | Primary and secondary effect estimates are accompanied by two-sided confidence intervals. |
| Cochran-Mantel-Haenszel test | The ORR difference uses CMH weighting stratified by stratification factors. |
| Odds ratio | ORR is also reported using odds ratios. |
| Multiplicity adjustment | The primary OS analysis includes a multiplicity-adjusted 95.81% CI. |
| Endpoint hierarchy | Secondary tests were not performed after the OS p-value exceeded the prior threshold. |
| Subgroup analysis | OS, PFS, and ORR are reported according to PD-L1 expression categories. |
| Safety denominators | Serious adverse events are reported as affected participants over participants at risk. |
24. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
25. Related Statistical Calculators
26. Sources
- ClinicalTrials.gov: NCT02576509 — CheckMate-459.
- PubMed: PMID 42762856.
- PubMed: PMID 39560477.
- PubMed: PMID 36852452.
- PubMed: PMID 36111952.
- PubMed: PMID 35101942.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts that underlie randomized clinical trials, time-to-event endpoints, categorical comparisons, confidence intervals, and multiplicity.
27. Record Summary
CheckMate-459 provides a compact example of how a randomized phase 3 trial can combine a primary time-to-event endpoint with several secondary efficacy measures. The primary overall-survival analysis used a stratified log-rank test and a stratified Cox proportional-hazards model, producing an HR of 0.85 with a two-sided 95% CI of 0.71–1.02 and a p-value of 0.0752. Secondary analyses registry-reported estimates for ORR, PFS, and PD-L1-defined efficacy, while explicitly stating that no formal tests were performed for those secondary comparisons because the OS p-value was above the prior threshold.
The statistical lesson is therefore broader than any single effect estimate. Proper interpretation requires keeping the analysis population, endpoint hierarchy, effect measure, confidence interval, p-value, multiplicity framework, and subgroup status together. The ClinicalTrials.gov record also illustrate why a hazard ratio, odds ratio, and difference in response rates cannot be treated as interchangeable measures.