← Clinical Trials
Hepatocellular Carcinoma Phase 3 Completed NCT02576509

CheckMate-459: Complete Statistical Analysis of Nivolumab in Advanced Hepatocellular Carcinoma

An independent statistical analysis of the randomized phase 3 CheckMate-459 trial comparing nivolumab with sorafenib as a first treatment in patients with advanced hepatocellular carcinoma.

Trial start: December 7, 2015  ·  Primary completion: May 30, 2019  ·  Enrollment: 743
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

CheckMate-459 was a randomized, parallel, phase 3 study comparing nivolumab with sorafenib as a first treatment in patients with advanced hepatocellular carcinoma. The trial enrolled 743 participants and had two treatment arms.

743
Enrolled
Randomized trial
2
Treatment Arms
Nivolumab vs sorafenib
0.85
OS HR
95% CI 0.71–1.02
0.0752
OS P-value
Two-sided log-rank test
FeatureCheckMate-459
Trial nameCheckMate-459
ClinicalTrials.gov identifierNCT02576509
PhasePhase 3
StatusCompleted
ConditionHepatocellular Carcinoma
PopulationPatients with advanced hepatocellular carcinoma receiving a first treatment
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Lead sponsorBristol-Myers Squibb
Sponsor typeIndustry
Enrollment743
InterventionsNivolumab; Sorafenib

2. Clinical Question

The central clinical question was whether nivolumab differed from sorafenib when used as a first treatment in patients with advanced hepatocellular carcinoma, with overall survival serving as the registered primary endpoint.

Population

Patients with advanced hepatocellular carcinoma receiving a first treatment.

Intervention

Nivolumab 240 mg.

Comparator

Sorafenib 400 mg.

Primary question

What is the difference between nivolumab and sorafenib in overall survival among all randomized participants?

3. Trial Design

01
Randomize743 participants
02
Two armsNivolumab or sorafenib
03
TreatmentFirst treatment
04
Follow-upTime-to-event and response outcomes
05
AnalysisOS, ORR, PFS and PD-L1 analyses
Allocation
Randomized allocation was used to compare the two treatment groups.
Design model
Parallel-group design with two treatment arms.
Masking
The registry records the study as having no masking.
Primary purpose
Treatment.
ARM A

Nivolumab

  • Nivolumab 240 mg.
ARM B

Sorafenib

  • Sorafenib 400 mg.

The ClinicalTrials.gov record does not specify a randomization ratio, stratification factors, crossover provisions, factorial structure, or an interim-analysis design. Those features are therefore not assumed here.

4. Endpoints

The registered primary endpoint was overall survival. ClinicalTrials.gov also posted secondary analyses covering objective response rate, progression-free survival, and efficacy according to PD-L1 expression.

EndpointRegistry definition / time frameType
Overall Survival (OS) OS is defined as the time from the date of randomization to the date of death due to any cause in all randomized participants. Participants who are alive will be censored at the last known alive dates. Based on Kaplan-Meier Estimates. Time frame: time from the date of randomization to the date of death due to any cause, assessed up to June 2019 (approximately 41 months) Time-to-event
Objective Response Rate (ORR) Per BICR RECIST 1.1 Time frame: the date of randomization and the date of first objectively documented progression or the date of subsequent anti-cancer therapy, whichever occurs first, assessed up to May 2019 (approximately 40 months) Binary
Progression-Free Survival (PFS) Time frame: time from the date of randomization to the date of the first objectively documented tumor progression or death, assessed up to May 2019 (approximately 40 months) Time-to-event
Efficacy Based on PD-L1 Expression - OS and PFS Time frame: the date of randomization and the date of first objectively documented progression or the date of subsequent anti-cancer therapy, whichever occurs first, assessed up to May 2019 (approximately 40 months) Time-to-event
Efficacy Based on PD-L1 Expression - ORR Time frame: the date of randomization and the date of first objectively documented progression or the date of subsequent anti-cancer therapy, whichever occurs first, assessed up to May 2019 (approximately 40 months) Binary
Endpoint hierarchy matters. The registry identifies OS as the single primary endpoint. The analyses posted on ClinicalTrials.gov for secondary endpoints state that no test was performed because the OS p-value was above the prior threshold. Thus, the secondary estimates should not be presented as though they were independent confirmatory hypothesis tests.

5. Statistical Methodology

Overall survival and time-to-event analysis

The primary OS analysis used a log-rank test, with the analysis text specifying that the test was stratified by the stratification factors entered into the IVRS. A stratified Cox proportional-hazards model was used for the hazard-ratio estimate and confidence interval.

Primary OS analysis
Log-rank test → stratified Cox proportional-hazards model → hazard ratio with confidence interval

The reported comparison was nivolumab 240 mg versus sorafenib 400 mg in all randomized participants.

Cochran-Mantel-Haenszel analysis

The ORR comparison was analyzed using a Cochran-Mantel-Haenszel (CMH) method. The analysis text states that the estimate of the difference in ORRs was based on CMH weighting and was stratified by the stratification factors.

Odds-ratio analysis

The registry also reports an odds ratio for ORR. An odds ratio compares the odds of response between the two treatment groups; it is not the same quantity as a difference in response percentages.

Stratified analysis

Stratification allows a comparison to account for prespecified grouping factors rather than treating all participants as though they belonged to one homogeneous stratum. For the primary OS analysis, the registry specifically reports a log-rank test stratified by the stratification factors as entered into the IVRS.

Multiplicity

The primary OS analysis notes a confidence interval adjusted for multiplicity: 95.81% CI (0.72 to 1.02). The secondary ORR and PFS analyses explicitly state that no test was performed because the OS p-value was above the prior threshold. This is an important part of interpreting the reported endpoint hierarchy.

Do not treat every posted estimate as a confirmatory test. The registry contains estimates and confidence intervals for multiple secondary analyses, but the ClinicalTrials.gov record explicitly state that no test was performed for these secondary endpoints because the OS p-value exceeded the prior threshold.

6. Primary Result: Overall Survival

The primary endpoint was overall survival among all randomized participants, comparing nivolumab 240 mg with sorafenib 400 mg. The analysis used a stratified log-rank test and a stratified Cox proportional-hazards model.

Hazard ratio for overall survival

0.85

95% CI: 0.71–1.02   ·   P = 0.0752

Two-sided confidence interval; superiority hypothesis.

Primary OS analysisReported result
Analysis populationAll randomized participants
ComparisonNivolumab 240 mg vs Sorafenib 400 mg
MethodLog-rank test
Effect measureHazard ratio
Estimate0.85
95% CI0.71–1.02
P-value0.0752
Hypothesis typeSuperiority
Multiplicity-adjusted CI noted in analysis text95.81% CI: 0.72 to 1.02
Clinical Biostats interpretation

The reported HR of 0.85 means that, under the fitted proportional-hazards model, the estimated instantaneous rate of death in the nivolumab group was 0.85 times that in the sorafenib group over the analyzed follow-up. Expressed as a simple relative interpretation, this corresponds to a 15% lower estimated hazard, not a 15% reduction in the probability that an individual participant will die.

The 95% CI of 0.71–1.02 describes statistical uncertainty around the estimated hazard ratio. Because the interval extends above 1, the ClinicalTrials.gov record does not establish that the true hazard ratio is below 1 under a conventional two-sided interpretation. The separately reported multiplicity-adjusted 95.81% CI of 0.72 to 1.02 likewise extends above 1.

The p-value of 0.0752 is a measure of compatibility with the statistical testing framework under the null hypothesis; it is not a measure of how large or clinically important the treatment effect is. A p-value should therefore be interpreted together with the hazard ratio, confidence interval, analysis population, and trial design.

Finally, a Cox hazard ratio relies on a proportional-hazards model. If the relative hazards change materially over time, one summary HR can conceal important features of the survival curves. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption, so no such assessment is inferred here.

Why the primary endpoint drives the interpretation of the secondary analyses

The registry's secondary analysis records repeatedly state that no test was performed due to the OS p-value result above the prior threshold. This means that the secondary estimates can still be described as estimates with confidence intervals, but they should not be converted into claims of statistically confirmed superiority based solely on whether an interval happens to exclude a particular value.

7. Secondary Results: Objective Response Rate

Objective response rate per BICR RECIST 1.1 was analyzed among all randomized participants. The registry-reported analysis used the Cochran-Mantel-Haenszel method and reports both a difference in ORRs and an odds ratio.

Difference in Objective Response Rates

Difference in ORRs

8.3

95% CI: 3.9–12.7

Estimate defined as Nivolumab minus Sorafenib; two-sided confidence interval.

ORR analysisReported result
Analysis populationAll randomized participants
ComparisonNivolumab 240 mg vs Sorafenib 400 mg
MethodCochran-Mantel-Haenszel test / weighting method
Effect measureDifference of ORRs
Estimate8.3
95% CI3.9–12.7
Formal testNo test was performed due to OS p-value result above the prior threshold

Odds Ratio for Objective Response

Odds ratio for response

2.41

95% CI: 1.48–3.92

Odds ratio: Nivolumab over Sorafenib.

The odds ratio of 2.41 compares the odds of response, not the response probability itself. An odds ratio of 2.41 therefore should not be described as meaning that the response rate was 2.41 times as high.

Confirmatory-status caution: although the reported ORR odds-ratio confidence interval is entirely above 1, the ClinicalTrials.gov record explicitly states that no test was performed because the OS p-value was above the prior threshold. The ORR result is therefore best presented as an estimated treatment difference with its uncertainty interval rather than as a separate confirmatory hypothesis-test result.

8. Secondary Result: Progression-Free Survival

Progression-free survival was analyzed among all randomized participants. The registry reports a stratified Cox proportional-hazards model comparing nivolumab 240 mg with sorafenib 400 mg.

Hazard ratio for progression-free survival

0.93

95% CI: 0.79–1.10

Two-sided confidence interval; no test was performed due to the OS p-value result above the prior threshold.

PFS analysisReported result
Analysis populationAll randomized participants
ComparisonNivolumab 240 mg vs Sorafenib 400 mg
MethodStratified Cox proportional-hazards model
Effect measureHazard ratio
Estimate0.93
95% CI0.79–1.10
Formal testNo test was performed due to OS p-value result above the prior threshold

An HR of 0.93 corresponds to an estimated hazard that is 0.93 times the comparator hazard under the fitted model. The confidence interval of 0.79–1.10 indicates uncertainty that includes values below and above 1. Because the registry does not report a formal p-value for this secondary comparison, the result should be interpreted as an estimated relative treatment effect rather than a separately tested superiority claim.

9. Secondary Results: Efficacy by PD-L1 Expression

The registry contains six hazard-ratio analyses under the outcome measure Efficacy Based on PD-L1 Expression - OS and PFS. The registry-reported population description includes participants with >=1% PD-L1 expression, participants with <1% PD-L1 expression, and participants without PD-L1 quantifiable. The registry does not report a statistical method for these analyses.

PD-L1 subgroup / endpointHR95% CIFormal test
PD-L1 >=1%, OS0.800.54–1.19No test was performed
PD-L1 >=1%, PFS0.710.48–1.03No test was performed
PD-L1 <1%, OS0.840.69–1.02No test was performed
PD-L1 <1%, PFS0.980.81–1.17No test was performed
Without PD-L1 quantifiable, OS1.260.34–4.74No test was performed
Without PD-L1 quantifiable, PFS0.980.27–3.52No test was performed

These estimates illustrate why subgroup analysis should be read through the confidence interval rather than by comparing point estimates alone. For example, the subgroup without quantifiable PD-L1 has much wider intervals, including values well below and above 1. That width indicates substantially greater statistical uncertainty in those estimates.

Subgroup caution: a hazard ratio below 1 in one subgroup and a hazard ratio closer to 1 in another does not by itself establish that the treatment effect differs between subgroups. A formal interaction or treatment-by-subgroup analysis would be needed to support a claim of effect modification. The ClinicalTrials.gov record does not report such an interaction test.

10. Secondary Results: PD-L1 and Objective Response

The registry also reports odds ratios for ORR within two PD-L1 categories. The analysis text describes these as unstratified odds ratios with associated unstratified 95% exact confidence intervals.

PD-L1 subgroupOR95% CIFormal test
PD-L1 >=1%, ORR3.791.41–10.17No test was performed
PD-L1 <1%, ORR1.951.10–3.45No test was performed

The odds ratios are treatment-group comparisons of the odds of response within the specified PD-L1 categories. They should not be interpreted as response-rate ratios. Their confidence intervals describe uncertainty around the corresponding odds-ratio estimates, while the absence of a formal test means that these results should not be treated as independent confirmatory findings.

11. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm using affected participants divided by participants at risk.

Safety measureNivolumab 240 mgSorafenib 400 mg
Serious adverse events, affected / at risk216 / 367216 / 363
Serious adverse events: affected participants / participants at risk
Nivolumab 240 mg
216 / 367
Sorafenib 400 mg
216 / 363

The two arms have the same reported number of affected participants, 216, while the denominators differ. The ClinicalTrials.gov record therefore support reporting the affected and at-risk counts directly. No additional safety categories, exposure-adjusted rates, or statistical safety comparisons are reported in the ClinicalTrials.gov record.

12. Statistical Methods Explained

Why was a log-rank test used for overall survival?

Overall survival is a time-to-event endpoint because both the event time and censoring matter. The log-rank test compares the survival experience of two randomized groups across follow-up while accounting for the timing of observed deaths and censored observations. In this trial, the registry reports a stratified log-rank analysis.

What does an OS hazard ratio of 0.85 mean?

A hazard ratio of 0.85 means that the estimated instantaneous death rate in the nivolumab group was 0.85 times the corresponding rate in the sorafenib group under the fitted Cox model. It does not mean that 15% of participants benefited, that survival increased by 15%, or that an individual's probability of death was reduced by exactly 15%.

Why is the confidence interval important?

The point estimate is only one summary of the data. The 95% confidence interval of 0.71–1.02 shows the uncertainty around the OS hazard-ratio estimate. Its upper endpoint crosses 1, so the interval does not exclude no difference in hazard under the usual two-sided reference point.

Why does the p-value not measure effect size?

A p-value summarizes how compatible the observed data are with a specified null hypothesis under the statistical testing framework. It is affected by the amount of information and does not directly quantify the magnitude of a treatment effect. The HR and its confidence interval are therefore necessary to understand the estimated effect itself.

Why use a Cochran-Mantel-Haenszel method for ORR?

ORR is a binary endpoint, so each participant contributes a response or nonresponse classification rather than an event time. The CMH method provides a way to compare treatment groups while weighting across strata. In this trial, the registry states that the difference in ORRs was based on CMH weighting and stratified by the stratification factors.

Why should the ORR odds ratio not be read as a response-rate ratio?

An odds ratio compares odds, where odds are the probability of an event divided by the probability of no event. A response odds ratio of 2.41 therefore means the estimated odds of response were 2.41 times those in the comparator group; it does not mean that the percentage responding was 2.41 times as large.

Why do the secondary analyses need a multiplicity caution?

The ClinicalTrials.gov record explicitly state that no test was performed for the secondary ORR, PFS, PD-L1 OS/PFS, and PD-L1 ORR analyses because the OS p-value was above the prior threshold. This creates an important distinction between reporting an effect estimate and declaring an independently tested statistical finding.

13. Understanding the Primary Hazard Ratio in Context

Relative effect

The OS estimate of 0.85 is a relative time-to-event measure. In simple terms, 0.85 indicates an estimated hazard 15% lower for nivolumab than sorafenib under the fitted model.

Precision

The 0.71–1.02 confidence interval is relatively more informative than the point estimate alone because it shows the uncertainty surrounding 0.85. The interval includes 1, so the data do not rule out no difference in the hazard ratio under the conventional reference value.

Statistical evidence

The reported P = 0.0752 should be interpreted within the trial's superiority framework and the stated multiplicity procedure. It should not be converted into an estimate of the probability that one treatment is better than the other.

What is not reported

The ClinicalTrials.gov record does not report median OS, median PFS, time-specific survival rates, event counts for OS or PFS, Kaplan-Meier coordinates, or formal proportional-hazards diagnostics. Those quantities are therefore not inferred or reconstructed on this page.

14. Confidence Intervals Across the Secondary Analyses

The collection of secondary estimates demonstrates how the width of a confidence interval can vary substantially across analyses. The overall PFS estimate has a 95% CI of 0.79–1.10, whereas the estimate for OS among participants without quantifiable PD-L1 has a much wider 95% CI of 0.34–4.74.

AnalysisEstimate95% CIWhat the interval communicates
OS, primaryHR 0.850.71–1.02Uncertainty around the primary hazard-ratio estimate
PFSHR 0.930.79–1.10Uncertainty around the secondary hazard-ratio estimate
PD-L1 >=1%, OSHR 0.800.54–1.19Subgroup uncertainty includes values below and above 1
PD-L1 <1%, OSHR 0.840.69–1.02Narrower subgroup interval that still extends above 1
Without quantifiable PD-L1, OSHR 1.260.34–4.74Wide uncertainty around the subgroup estimate
PD-L1 >=1%, ORROR 3.791.41–10.17Uncertainty around the odds of response

The appropriate lesson is not that a narrower interval is automatically more important. Rather, confidence-interval width should be considered when judging how much information an estimate contains. Subgroup analyses often have fewer participants than the overall randomized population, which can produce wider intervals.

15. Multiplicity and Endpoint Hierarchy

The registry identifies overall survival as the single primary endpoint. The statistical analyses include one primary endpoint analysis and eleven secondary analyses. The secondary analysis records repeatedly note that no test was performed because the OS p-value was above the prior threshold.

Analysis groupRole in the ClinicalTrials.gov recordInterpretation
Overall survivalPrimary endpointFormal superiority analysis with log-rank testing and Cox HR estimation
Objective response rateSecondary endpointEffect estimates reported; formal test not performed
Progression-free survivalSecondary endpointHR and CI reported; formal test not performed
PD-L1 OS/PFSSecondary endpointSubgroup HR estimates and CIs reported; no tests performed
PD-L1 ORRSecondary endpointOR estimates and CIs reported; no tests performed

This hierarchy prevents a common interpretive error: treating every confidence interval in a registry as if it represents an independently tested confirmatory hypothesis. Here, the registry explicitly supplies a sequential testing context in which the OS result governs whether the listed secondary tests were performed.

16. Analysis Population and Randomization

The primary OS analysis population was all randomized participants. This is important because randomized treatment assignment is the foundation of the principal comparative analysis.

Why randomization matters

Randomization creates the basis for comparing outcomes between treatment groups without relying solely on observed baseline characteristics to explain differences.

Why ITT-style analysis matters

Analyzing participants according to randomized assignment preserves the treatment comparison established by randomization. The registry-reported primary analysis is explicitly based on all randomized participants.

Why censoring matters

For OS, participants who remain alive are censored at their last known alive dates. Censoring allows their available follow-up information to contribute without assigning an unobserved death date.

Why analysis population should be stated

A treatment effect cannot be interpreted correctly without knowing which participants contributed to the analysis. The primary OS analysis explicitly identifies all randomized participants.

17. Time-to-Event Endpoints: What the Model Is Estimating

Both OS and PFS are time-to-event endpoints. Unlike a simple binary endpoint, they use information about when the event occurred and whether a participant was censored before experiencing the event.

Conceptual hazard-ratio interpretation
HR = hazard in nivolumab group / hazard in sorafenib group

An HR below 1 corresponds to a lower estimated instantaneous event rate in the nivolumab group; an HR above 1 corresponds to a higher estimated instantaneous event rate.

The Cox model compresses a potentially complex time-varying pattern into a single relative effect estimate. That summary is useful, but its interpretation depends on the proportional-hazards framework. The ClinicalTrials.gov record does not provide a formal diagnostic of that assumption.

18. Objective Response: Difference Versus Odds Ratio

The registry provides two distinct effect measures for the overall ORR analysis: a difference of ORRs and an odds ratio. These quantities answer related but different questions.

MeasureReported estimate95% CIInterpretive question
Difference of ORRs8.33.9–12.7How far apart are the response rates on the reported difference scale?
Odds ratio2.411.48–3.92How do the odds of response compare between the treatment groups?

The difference is expressed on an additive scale, whereas the odds ratio is a multiplicative measure on the odds scale. They should not be substituted for one another. The registry's CMH analysis also indicates that stratification was incorporated into the ORR difference analysis.

19. PD-L1 Subgroup Interpretation

The PD-L1 analyses are useful for demonstrating the distinction between a subgroup estimate and a formal test of treatment-effect heterogeneity.

PD-L1 >=1%

The reported OS HR was 0.80 with a 95% CI of 0.54–1.19, while the PFS HR was 0.71 with a 95% CI of 0.48–1.03.

PD-L1 <1%

The reported OS HR was 0.84 with a 95% CI of 0.69–1.02, while the PFS HR was 0.98 with a 95% CI of 0.81–1.17.

PD-L1 not quantifiable

The OS HR was 1.26 with a 95% CI of 0.34–4.74, and the PFS HR was 0.98 with a 95% CI of 0.27–3.52.

Formal subgroup comparison

The ClinicalTrials.gov record does not report a treatment-by-PD-L1 interaction test. The subgroup point estimates should therefore not be used alone to establish effect modification.

20. Serious Adverse Events and Denominator Awareness

The safety data provide an instructive example of why denominators should accompany adverse-event counts.

ArmAffectedAt riskReported format
Nivolumab 240 mg216367216/367
Sorafenib 400 mg216363216/363

Both arms have 216 affected participants, but the numbers at risk differ. Consequently, simply comparing the numerators would not provide the complete context for the reported safety measure. The ClinicalTrials.gov record does not include a formal statistical comparison of serious adverse-event rates, so none is inferred here.

21. What the Registry Does Not Report in the Supplied Data

The ClinicalTrials.gov record is sufficient for a detailed statistical analysis of the reported estimates, but several commonly reported clinical-trial quantities are not included. Clinical Biostats therefore does not reconstruct them.

Methodological discipline: absence from the ClinicalTrials.gov record is not treated as evidence that a particular design or statistical procedure was absent from the full protocol. It simply means that the feature cannot be described reliably from the ClinicalTrials.gov record.

22. Limitations

23. Why This Trial Matters Statistically

CheckMate-459 is a useful statistical teaching case because the ClinicalTrials.gov record brings together randomized treatment comparison, a primary time-to-event endpoint, stratified log-rank testing, Cox proportional-hazards modeling, multiplicity adjustment, CMH analysis of a binary endpoint, odds ratios, subgroup analyses, and an explicit distinction between estimation and formal hypothesis testing.

ConceptHow it appears in CheckMate-459
RandomizationThe trial is recorded as randomized with two treatment arms.
Parallel-group designThe design model is recorded as parallel.
Time-to-event endpointOverall survival is the registered primary endpoint.
Kaplan-Meier estimationThe OS definition states that estimates are based on Kaplan-Meier estimates.
Log-rank testThe primary OS comparison uses a stratified log-rank test.
Cox proportional-hazards modelThe primary analysis reports a stratified Cox model for the HR.
Hazard ratioOS and PFS are summarized using HRs.
Confidence intervalsPrimary and secondary effect estimates are accompanied by two-sided confidence intervals.
Cochran-Mantel-Haenszel testThe ORR difference uses CMH weighting stratified by stratification factors.
Odds ratioORR is also reported using odds ratios.
Multiplicity adjustmentThe primary OS analysis includes a multiplicity-adjusted 95.81% CI.
Endpoint hierarchySecondary tests were not performed after the OS p-value exceeded the prior threshold.
Subgroup analysisOS, PFS, and ORR are reported according to PD-L1 expression categories.
Safety denominatorsSerious adverse events are reported as affected participants over participants at risk.

24. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts that underlie randomized clinical trials, time-to-event endpoints, categorical comparisons, confidence intervals, and multiplicity.

27. Record Summary

CheckMate-459 provides a compact example of how a randomized phase 3 trial can combine a primary time-to-event endpoint with several secondary efficacy measures. The primary overall-survival analysis used a stratified log-rank test and a stratified Cox proportional-hazards model, producing an HR of 0.85 with a two-sided 95% CI of 0.71–1.02 and a p-value of 0.0752. Secondary analyses registry-reported estimates for ORR, PFS, and PD-L1-defined efficacy, while explicitly stating that no formal tests were performed for those secondary comparisons because the OS p-value was above the prior threshold.

The statistical lesson is therefore broader than any single effect estimate. Proper interpretation requires keeping the analysis population, endpoint hierarchy, effect measure, confidence interval, p-value, multiplicity framework, and subgroup status together. The ClinicalTrials.gov record also illustrate why a hazard ratio, odds ratio, and difference in response rates cannot be treated as interchangeable measures.

Clinical Biostats methodology: This analysis separates registry-reported numerical results from statistical interpretation. Where the ClinicalTrials.gov record does not report a quantity or methodological detail, the page does not reconstruct it from outside knowledge.