← Clinical Trials
Hepatocellular Carcinoma Phase 3 Overall Survival NCT01908426

CELESTIAL: Complete Statistical Analysis of Cabozantinib in Hepatocellular Carcinoma

An independent statistical analysis of the randomized phase 3 CELESTIAL trial comparing cabozantinib tablets with placebo in subjects with hepatocellular carcinoma who had received prior sorafenib, with emphasis on overall survival, progression-free survival, objective response, and the statistical methods used for these endpoints.

Trial status: Completed  ·  Enrollment: 707  ·  Primary completion: 16 October 2017
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the registry-reported CELESTIAL registry data.

1. Trial at a Glance

CELESTIAL was a randomized, parallel, quadruple-masked phase 3 trial evaluating cabozantinib tablets versus placebo in subjects with hepatocellular carcinoma who had received prior sorafenib. The trial enrolled 707 randomized subjects and posted statistical analyses for overall survival, progression-free survival, and objective response rate.

707
Randomized
470 vs 237
2
Arms
Cabozantinib / placebo
0.76
OS HR
95% CI 0.63–0.92
0.44
PFS HR
95% CI 0.36–0.52
FeatureCELESTIAL
PhasePhase 3
ConditionHepatocellular carcinoma
Brief titleStudy of Cabozantinib (XL184) vs Placebo in Subjects With Hepatocellular Carcinoma Who Have Received Prior Sorafenib
DesignRandomized, parallel, quadruple-masked
AllocationRandomized
Primary purposeTreatment
Enrollment707
Primary endpointOverall Survival (OS)
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Trial statusCompleted
Lead sponsorExelixis
Sponsor typeIndustry
ClinicalTrials.govNCT01908426

2. Clinical Question

The central statistical question was whether randomized treatment with cabozantinib tablets differed from placebo with respect to the primary time-to-event endpoint, overall survival, in subjects with hepatocellular carcinoma who had received prior sorafenib.

Population

Subjects with hepatocellular carcinoma who had received prior sorafenib.

Intervention

Cabozantinib tablets (XL184).

Comparator

Placebo tablets.

Primary question

Does cabozantinib produce a different overall-survival experience from placebo under the prespecified superiority analysis?

This framing is important because the primary endpoint is a time-to-event outcome rather than a simple binary outcome. The analysis therefore uses methods that account for both the timing of deaths and right-censoring rather than reducing every patient to a single yes/no response indicator.

3. Trial Design

01
Randomize707 subjects
02
AssignCabozantinib or placebo
03
FollowTime-to-event outcomes
04
AssessOS / PFS / ORR
05
AnalyzeStratified statistical methods
Allocation
Randomized allocation with a parallel design.
Masking
Quadruple masking.
Arms
Two intervention groups: cabozantinib tablets and placebo tablets.
Primary purpose
Treatment.
ARM A · n = 470

Cabozantinib

  • Cabozantinib tablets (XL184)
  • Randomized treatment arm
ARM B · n = 237

Placebo

  • Placebo tablets
  • Randomized comparator arm

The statistical analysis population for the second interim OS analysis was explicitly the intention-to-treat population, consisting of all 707 randomized subjects: 470 assigned to cabozantinib and 237 assigned to placebo.

4. Trial Timeline

26 September 2013

Trial start

The registry lists 26 September 2013 as the study start date.

1 June 2017

OS interim-analysis data cutoff

The primary OS analysis was based on a second planned interim analysis with a data cutoff of 01 June 2017.

16 October 2017

Primary completion

The registry lists 16 October 2017 as the primary completion date.

5. Endpoints

EndpointRegistry time frameTypeAnalysis method
Overall Survival (OS) Up to 45 months Time-to-event Stratified log-rank test; hazard ratio
Progression-Free Survival (PFS) Up to 45 months Time-to-event Stratified log-rank test; hazard ratio
Objective Response Rate (ORR) ORR is measured by radiologic assessment every 8 weeks after randomization until disease progression or discontinuation of study treatment (up to 45 months) Binary Cochran-Mantel-Haenszel test

Overall Survival

The primary analysis of OS is defined as the time from randomization to death from any cause. The analysis was based on a second planned interim analysis prespecified to be performed at approximately the 75% information fraction, at approximately 466 deaths. The data cutoff date for this event-driven analysis in the ITT population was 01 June 2017.

Progression-Free Survival

PFS was a secondary endpoint with a registry time frame of up to 45 months. The prespecified primary analysis of PFS was based on the first 707 randomized subjects: 470 cabozantinib and 237 placebo.

Objective Response Rate

ORR was measured by radiologic assessment every 8 weeks after randomization until disease progression or discontinuation. The analysis was performed in the ITT population, with response determined by Investigator per RECIST 1.1.

6. Statistical Methodology

Intention-to-treat analysis

The ITT population included all 707 randomized subjects for the second interim OS analysis. This approach preserves the treatment comparison created by randomization: subjects remain associated with the group to which they were randomized rather than being reassigned according to subsequent treatment exposure or outcome.

For this trial, the ITT principle is especially important because the OS comparison concerns the effect of the randomized treatment strategy. An ITT estimate therefore answers a different question from an analysis restricted to patients who remained on treatment.

Stratified log-rank testing

The OS log-rank test was stratified by etiology of disease, geographic region, presence of extrahepatic spread of disease and/or macrovascular invasion. The PFS log-rank test used the same stated stratification factors.

Conceptual role of the log-rank test
H0: the time-to-event distributions are equivalent across randomized treatment groups

The log-rank test compares the observed pattern of events between treatment groups across follow-up. Stratification allows the comparison to account for prespecified factors rather than treating all randomized subjects as if they came from a single homogeneous stratum.

Hazard ratio

The reported effect measure for OS and PFS was the hazard ratio. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the cabozantinib group relative to placebo under the fitted time-to-event comparison.

Conceptual interpretation
HR = hazard in cabozantinib group / hazard in placebo group

The hazard ratio is a relative time-to-event measure. It is not itself a probability, an absolute risk difference, or a statement that every individual subject experiences the same proportional change in risk.

Cochran-Mantel-Haenszel analysis

ORR was analyzed with a Cochran-Mantel-Haenszel test in the ITT population. This is a categorical-data method that can evaluate treatment-group differences while accounting for stratification when the analysis is specified in that form.

Superiority testing

The trial's stated hypothesis type was superiority. The inferential objective was therefore to determine whether the randomized treatment groups differed, rather than to demonstrate that cabozantinib was sufficiently close to placebo within a prespecified non-inferiority margin.

7. Primary Result: Overall Survival

The primary endpoint analysis compared overall survival between cabozantinib and placebo in the ITT population. The analysis used the stratified log-rank test and reported a hazard ratio with a two-sided 95% confidence interval.

Hazard ratio for overall survival

0.76

95% CI: 0.63–0.92   ·   P = 0.0049

Superiority analysis; ITT population; second planned interim analysis.

Primary endpointCabozantinibPlaceboEffect estimateP-value
Overall Survival (OS) 470 randomized 237 randomized HR 0.76
95% CI 0.63–0.92
0.0049
Clinical Biostats interpretation

An OS hazard ratio of 0.76 means that, under the time-to-event model used for this analysis, the estimated instantaneous hazard of death in the cabozantinib group was 76% of that in the placebo group. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 24% lower hazard of death.

The HR does not mean that 24% of patients were saved, that each individual patient's risk fell by exactly 24%, or that the absolute probability of death was reduced by 24 percentage points. A hazard ratio is a relative measure of event rates over time.

The two-sided 95% confidence interval of 0.63–0.92 describes the statistical uncertainty around the estimated HR under the analysis framework. Because the entire interval is below 1, the reported interval is consistent with a lower estimated hazard in the cabozantinib group relative to placebo.

The P-value of 0.0049 addresses the strength of evidence against the specified null hypothesis under the statistical testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the clinical importance of the observed HR.

The result also comes from a planned interim analysis, not simply an unplanned look at accumulating data. The registry states that this was the second planned interim analysis at approximately the 75% information fraction and approximately 466 deaths. That design context matters when interpreting the inferential result.

8. Secondary Result: Progression-Free Survival

PFS was analyzed in the prespecified primary analysis based on the first 707 randomized subjects, with 470 assigned to cabozantinib and 237 to placebo. The registry reports a stratified log-rank analysis and a hazard ratio as the effect measure.

Hazard ratio for progression-free survival

0.44

95% CI: 0.36–0.52   ·   P < 0.0001

Superiority analysis; stratified log-rank test.

Secondary endpointCabozantinibPlaceboEffect estimateP-value
Progression-Free Survival (PFS) 470 randomized 237 randomized HR 0.44
95% CI 0.36–0.52
< 0.0001
Clinical Biostats interpretation

A PFS hazard ratio of 0.44 means that the estimated instantaneous hazard of the PFS event in the cabozantinib group was 44% of the corresponding hazard in the placebo group. In relative terms, this is an estimated 56% lower hazard of the PFS event.

This does not mean that 56% of patients avoided progression or death. PFS is a time-to-event endpoint, and the HR summarizes a relative comparison of event hazards rather than an absolute event probability.

The two-sided 95% CI of 0.36–0.52 gives the uncertainty interval for the estimated HR. The interval is relatively far below 1, indicating that the reported estimate is not close to the null value within this confidence interval.

The reported P-value of < 0.0001 indicates strong statistical evidence against the null hypothesis under the specified analysis. It should not be interpreted as a measure of how large the PFS effect is; the HR and its confidence interval provide the effect-size information.

As with OS, interpretation of the HR assumes that a single hazard ratio provides an appropriate summary of the treatment comparison over follow-up. The ClinicalTrials.gov record does not provide enough information to evaluate the proportional-hazards assumption directly.

9. Secondary Result: Objective Response Rate

ORR was measured by radiologic assessment every 8 weeks after randomization until disease progression or discontinuation. The analysis was performed in the ITT population, with response determined by Investigator per RECIST 1.1.

Cochran-Mantel-Haenszel test

P = 0.0086

Cabozantinib vs placebo; ITT population; superiority analysis.

Secondary endpointAnalysis populationMethodP-value
Objective Response Rate (ORR) ITT: 470 cabozantinib, 237 placebo Cochran-Mantel-Haenszel test 0.0086
Clinical Biostats interpretation

The reported P-value of 0.0086 is the formal comparison reported for ORR. It indicates evidence of a difference between the randomized treatment groups under the specified Cochran-Mantel-Haenszel analysis.

The ClinicalTrials.gov record does not provide the numerical response rate in each arm or a confidence interval for the response-rate difference. Therefore, this page does not infer or reconstruct those quantities.

This distinction is important statistically: a P-value can establish evidence against a null hypothesis without telling the reader the magnitude of the absolute difference. When an effect estimate and confidence interval are available, they provide additional information about the size and precision of the observed difference.

10. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected subjects over subjects at risk.

Safety measureCabozantinibPlacebo
Serious adverse events232 / 46787 / 237
Serious adverse events: affected / at risk
Cabozantinib
232 / 467
Placebo
87 / 237

The denominators in this safety measure are not identical to the randomized totals used in the ITT efficacy analysis. The reported serious-adverse-event figures therefore should be presented exactly as affected over at risk rather than silently treating 467 and 237 as the randomized sample sizes for every analysis.

Safety interpretation: the ClinicalTrials.gov record provides affected/at-risk counts but do not provide a formal between-arm statistical test or confidence interval for serious adverse events. The counts should therefore not be converted into an unreported inferential comparison.

11. Stratification and Why It Matters

The CELESTIAL OS and PFS analyses were stratified by three sets of factors: etiology of disease, geographic region, and presence of extrahepatic spread of disease and/or macrovascular invasion.

Stratification factorRole in reported analysis
Etiology of diseaseUsed to stratify the OS and PFS log-rank analyses.
Geographic regionUsed to stratify the OS and PFS log-rank analyses.
Extrahepatic spread and/or macrovascular invasionUsed to stratify the OS and PFS log-rank analyses.

Stratification does not mean that the treatment effect is estimated separately and independently within every subgroup. Instead, the analysis accounts for the prespecified strata when constructing the overall comparison.

This is especially useful in randomized trials when important prognostic characteristics may influence event timing. A stratified analysis can make the treatment comparison more closely aligned with the design used to randomize subjects.

12. Interim Analysis and Information Fraction

The primary OS analysis was based on a second planned interim analysis prespecified to occur at approximately the 75% information fraction, corresponding to approximately 466 deaths. The data cutoff was 01 June 2017.

Why use an interim analysis?

An event-driven trial can evaluate accumulating information before the final number of events has occurred. The interim analysis becomes part of the prespecified statistical design rather than an ad hoc examination of the data.

Why information fraction matters

For a time-to-event trial, statistical information is closely related to the number of observed events. The stated 75% information fraction therefore describes how much of the planned event information had accumulated at the interim analysis.

The registry data identify this as a planned interim analysis but do not provide a separate alpha-spending function or interim efficacy boundary in the ClinicalTrials.gov recordset. Accordingly, this page does not attribute a particular alpha-spending procedure or boundary to CELESTIAL.

Interpretation principle: interim monitoring must be considered when interpreting a P-value because repeatedly examining accumulating efficacy data can affect type I error if the analysis is not incorporated into the prespecified design. The registry-reported CELESTIAL data establish that the OS result came from a second planned interim analysis, but they do not provide enough information here to reconstruct the complete group-sequential error-control scheme.

13. Statistical Methods Explained

Why was a log-rank test used for OS and PFS?

OS and PFS are time-to-event endpoints. Subjects can experience events at different times, and some subjects may be censored before an event occurs. The log-rank test compares the treatment groups across the follow-up period while accounting for the timing of observed events. CELESTIAL used a stratified version based on the prespecified stratification factors.

What does an OS hazard ratio of 0.76 mean?

An HR of 0.76 means that the estimated instantaneous hazard of death in the cabozantinib group was 76% of that in the placebo group under the fitted time-to-event comparison. A useful relative interpretation is an estimated 24% lower hazard. It is not a 24-percentage-point reduction in mortality and does not imply the same effect for every patient.

Why is the confidence interval important?

The 95% CI of 0.63–0.92 places the estimated OS hazard ratio in the context of statistical uncertainty. The point estimate is 0.76, but the interval communicates that the data do not identify that value with infinite precision. It also shows that the entire reported interval lies below 1.

Why doesn't the P-value measure effect size?

The OS P-value of 0.0049 measures the evidence against the relevant null hypothesis under the statistical testing framework. It is not a scale for clinical magnitude. Two studies can have the same P-value with different effect sizes, and a very small P-value can occur with a modest effect when the information is sufficiently large.

Why was stratification used?

The log-rank analyses were stratified by disease etiology, geographic region, and the presence of extrahepatic spread and/or macrovascular invasion. Stratification incorporates these prespecified factors into the time-to-event comparison rather than ignoring the structure established in the trial design.

Why does the analysis population matter?

The OS and ORR analyses explicitly used the ITT population, while the serious-adverse-event ClinicalTrials.gov record use an affected/at-risk denominator of 467 for cabozantinib and 237 for placebo. A statistical estimate cannot be interpreted correctly without knowing which subjects were eligible for that particular analysis.

What is different about the ORR analysis?

ORR is a binary endpoint rather than a time-to-event endpoint. The registry reports a Cochran-Mantel-Haenszel analysis and a P-value of 0.0086. That is a different statistical framework from the stratified log-rank method used for OS and PFS because the underlying outcome structure is different.

14. Confidence Intervals and Effect Size

The CELESTIAL results illustrate why a complete statistical interpretation should present an effect estimate, its confidence interval, and the P-value together.

EndpointEffect95% CIP-valueWhat it communicates
Overall Survival HR 0.76 0.63–0.92 0.0049 Relative time-to-event effect, statistical uncertainty, and evidence against the null.
Progression-Free Survival HR 0.44 0.36–0.52 < 0.0001 Relative time-to-event effect, statistical uncertainty, and evidence against the null.
Objective Response Rate Not reported in the ClinicalTrials.gov record Not reported in the ClinicalTrials.gov record 0.0086 Evidence of a treatment-group difference under the reported categorical analysis.
A useful reading sequence

For a hazard-ratio result, first identify the direction of the effect, then examine the magnitude of the HR, then examine the confidence interval, and finally interpret the P-value in the context of the prespecified design. This sequence prevents the P-value from becoming the only statistic considered.

15. Time-to-Event Endpoints and Censoring

OS and PFS are fundamentally different from fixed-time binary outcomes because each subject contributes information over time. A subject who has not experienced the event by the end of observed follow-up can still contribute information before censoring.

Conceptual survival function
S(t) = P(T > t)

The survival function represents the probability that the event time T exceeds a given time t. Time-to-event methods use the observed event and censoring information rather than reducing the follow-up history to a single binary indicator.

For OS in CELESTIAL, the registry explicitly defines the event time as the time from randomization to death from any cause. This definition makes the randomized assignment and the time origin unambiguous for the primary analysis.

The ClinicalTrials.gov record does not provide individual censoring times or the underlying event-time dataset. Consequently, this page does not attempt to reconstruct Kaplan-Meier curves, median survival times, numbers at risk, or other quantities that require information beyond the reported summary statistics.

16. Primary Analysis Population

EndpointAnalysis populationSample sizes
Overall Survival ITT population; second interim analysis 707 randomized: 470 cabozantinib, 237 placebo
Progression-Free Survival First 707 randomized subjects 470 cabozantinib, 237 placebo
Objective Response Rate ITT population; all randomized subjects 470 cabozantinib, 237 placebo

The consistency of the randomized sample sizes across the reported efficacy analyses is useful because it makes clear that the principal efficacy results were not generated from a selectively treated subset. At the same time, the safety denominator for serious adverse events differs for cabozantinib, so efficacy and safety populations should not be conflated.

17. What the Hazard Ratios Do — and Do Not — Mean

Overall survival

The OS HR of 0.76 corresponds to an estimated 24% lower instantaneous hazard of death for cabozantinib relative to placebo under the reported model. It does not mean that 24% of subjects avoided death, that survival probability increased by 24 percentage points, or that the treatment effect was identical at every point in follow-up.

Progression-free survival

The PFS HR of 0.44 corresponds to an estimated 56% lower instantaneous hazard of the PFS event. It does not mean that 56% of subjects were free of progression or death, because a hazard ratio is not an absolute event-free proportion.

Confidence intervals

The OS confidence interval of 0.63–0.92 and the PFS confidence interval of 0.36–0.52 describe uncertainty around their respective HR estimates. They do not describe the range of treatment effects experienced by individual patients.

A complete interpretation therefore uses the hazard ratio as one component of the evidence. The endpoint definition, analysis population, stratification, interim-analysis context, confidence interval, and P-value all contribute to understanding what the reported estimate actually establishes.

18. Multiplicity and Multiple Endpoints

The ClinicalTrials.gov record identifies one primary endpoint, overall survival, and report progression-free survival and objective response rate as secondary analyses. The OS analysis also came from a planned second interim analysis.

EndpointRoleStatistical methodReported effect / result
Overall Survival Primary Stratified log-rank test HR 0.76; 95% CI 0.63–0.92; P = 0.0049
Progression-Free Survival Secondary Stratified log-rank test HR 0.44; 95% CI 0.36–0.52; P < 0.0001
Objective Response Rate Secondary Cochran-Mantel-Haenszel test P = 0.0086

The ClinicalTrials.gov record does not provide the complete multiplicity-adjustment procedure or an alpha-allocation scheme for the three posted analyses. Therefore, the reported P-values should be interpreted as the values attached to their respective registry analyses rather than reverse-engineering an unreported familywise-error procedure.

Why this matters: statistical significance is inseparable from the testing framework. An interim analysis, multiple endpoints, multiple looks at the data, and additional analyses can all affect how P-values should be interpreted. The available CELESTIAL data identify these design features but do not provide enough information to reconstruct every element of the trial's complete multiplicity strategy.

19. Missing Data and Analysis Assumptions

The ClinicalTrials.gov record identifies the ITT analysis populations and the statistical methods, but they do not provide a missing-data or imputation specification for the reported efficacy analyses.

For time-to-event endpoints, censoring is an intrinsic part of the analysis rather than ordinary missingness. The interpretation of Kaplan-Meier and hazard-ratio methods depends on the assumptions governing censoring and the adequacy of the time-to-event model. The ClinicalTrials.gov record does not provide individual-level follow-up information with which to evaluate those assumptions empirically.

For ORR, the registry specifies radiologic assessment every 8 weeks after randomization until disease progression or discontinuation and response determined by Investigator per RECIST 1.1. It does not provide an additional imputation rule in the ClinicalTrials.gov record.

Data discipline: the absence of a registry-reported imputation rule is not evidence that no rule existed in the full protocol or statistical analysis plan. It means only that this page does not attribute an unreported missing-data procedure to the trial.

20. Safety and Efficacy Use Different Statistical Questions

The efficacy analysis asks about randomized treatment assignment and clinical time-to-event outcomes. The serious-adverse-event summary instead reports affected subjects over subjects at risk. These are related components of the evidence but should not be collapsed into a single numerical treatment-effect measure.

Efficacy

OS and PFS were analyzed using randomized treatment groups and time-to-event methods, with ITT-based populations specified in the registry data.

Response

ORR was analyzed as a binary outcome using the Cochran-Mantel-Haenszel test in the ITT population.

Safety

Serious adverse events are reported as 232/467 for cabozantinib and 87/237 for placebo.

Interpretation

Each endpoint requires its own denominator, outcome definition, analysis population, and statistical method.

21. Limitations

22. Why This Trial Matters Statistically

CELESTIAL is a useful statistical teaching case because it combines randomized treatment allocation, quadruple masking, an event-driven primary endpoint, a planned interim analysis, stratified survival testing, hazard-ratio estimation, an ITT efficacy population, and a categorical response endpoint analyzed with a different statistical method.

ConceptHow it appears in CELESTIAL
Randomization707 subjects randomized to cabozantinib or placebo.
Quadruple maskingThe trial was registered as quadruple-masked.
ITT analysisOS and ORR analyses explicitly used the ITT population.
Time-to-event analysisOS was the primary endpoint and PFS was a secondary endpoint.
Log-rank testingOS and PFS were analyzed with stratified log-rank tests.
Hazard ratioHR 0.76 for OS and HR 0.44 for PFS.
Confidence intervalsTwo-sided 95% CIs were reported for both hazard ratios.
Stratified analysisOS and PFS analyses incorporated disease etiology, geographic region, and extrahepatic spread and/or macrovascular invasion.
Interim analysisOS was analyzed at a second planned interim analysis at approximately the 75% information fraction.
Categorical analysisORR was evaluated using the Cochran-Mantel-Haenszel test.
Safety denominatorsSerious adverse events are reported using affected/at-risk counts rather than the same denominator used for all efficacy analyses.

The most important statistical lesson is that the treatment effect cannot be separated from the structure of the analysis. A hazard ratio of 0.76 has meaning only when the reader also knows that it is an OS comparison, that the analysis was performed in the ITT population, that the log-rank test was stratified, and that the reported result came from a planned interim analysis.

23. Statistical Methods Explained: Reading the Results Together

Effect size

The OS HR of 0.76 and PFS HR of 0.44 quantify relative differences in event hazards between randomized groups.

Precision

The 95% CIs show how precisely the corresponding hazard ratios are estimated under the statistical framework.

Evidence

The P-values quantify evidence against the relevant null hypotheses but do not measure the magnitude of the effects.

Design context

The OS result was obtained at a planned interim analysis, making the timing of the analysis part of the statistical interpretation.

These four components should be read together. Looking only at the P-value loses information about magnitude and precision. Looking only at the hazard ratio loses information about uncertainty. Looking at either without the endpoint definition and analysis population risks attaching the statistic to the wrong clinical question.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Calculators

26. Sources

The quantitative trial statements on this page are restricted to the registry-reported CELESTIAL ClinicalTrials.gov data. The PubMed links are provided as source records associated with the trial; no additional numerical results from those publications are incorporated into this analysis.

Continue through the Clinical Biostats statistical learning pathway

Use the trial's endpoints and methods as a starting point for deeper study of survival analysis, stratified testing, confidence intervals, and clinical-trial methodology.

27. Record Summary

CELESTIAL provides a compact example of how a randomized phase 3 oncology trial can combine several statistical frameworks. The primary endpoint was overall survival, defined as time from randomization to death from any cause, and analyzed in the ITT population at a second planned interim analysis. The reported OS hazard ratio was 0.76 with a two-sided 95% CI of 0.63–0.92 and P = 0.0049. PFS was analyzed as a secondary time-to-event endpoint with HR 0.44, 95% CI 0.36–0.52, and P < 0.0001. ORR was analyzed using the Cochran-Mantel-Haenszel test, with P = 0.0086.

The statistical interpretation depends on more than those numbers. The analysis populations, stratification factors, interim-analysis timing, endpoint definitions, and distinction between time-to-event and binary outcomes determine what each statistic actually means. The serious-adverse-event data likewise require their own denominators and should not be treated as an efficacy comparison.

Clinical Biostats methodology: A trial-results page should separate reported evidence from statistical interpretation. For CELESTIAL, the ClinicalTrials.gov record supports detailed interpretation of the reported OS, PFS, and ORR analyses while leaving unreported quantities—such as survival medians, subgroup estimates, Kaplan-Meier curves, and complete multiplicity procedures—unreconstructed.