← Clinical Trials
Hepatocellular Carcinoma Phase 3 Time-to-Event NCT03298451

HIMALAYA: Complete Statistical Analysis of Durvalumab and Tremelimumab in Hepatocellular Carcinoma

An independent statistical review of the randomized phase 3 HIMALAYA trial evaluating durvalumab and tremelimumab regimens versus sorafenib as first-line treatment in patients with advanced hepatocellular carcinoma.

Trial start: 11 Oct 2017  ·  Primary completion: 27 Aug 2021  ·  Enrollment: 1,324
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

HIMALAYA was a randomized, parallel, open-label phase 3 study in hepatocellular carcinoma with four study arms. The ClinicalTrials.gov record reports an overall-survival primary comparison of tremelimumab 300 mg as a single dose plus durvalumab 1500 mg versus sorafenib 400 mg, and a secondary overall-survival comparison of durvalumab 1500 mg versus sorafenib 400 mg.

1,324
Enrolled
Phase 3
4
Study arms
Parallel design
0.78
Primary OS HR
95% CI 0.66–0.92
0.0035
Primary OS p-value
Stratified log-rank
FeatureHIMALAYA
Trial nameHIMALAYA
ClinicalTrials.gov identifierNCT03298451
PhasePhase 3
ConditionHepatocellular carcinoma
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment1,324
Study statusActive, not recruiting
Lead sponsorAstraZeneca
Sponsor typeIndustry

2. Clinical Question

The trial evaluates first-line treatment regimens for patients with hepatocellular carcinoma. The primary reported comparison asks whether tremelimumab 300 mg as a single dose plus durvalumab 1500 mg produces a different overall-survival outcome from sorafenib 400 mg. A secondary comparison evaluates durvalumab 1500 mg versus sorafenib 400 mg for overall survival, with a non-inferiority hypothesis followed sequentially by a superiority test.

Population

Patients with hepatocellular carcinoma enrolled in the phase 3 HIMALAYA study.

Intervention

Tremelimumab 300 mg x1 Dose + Durvalumab 1500 mg for the primary overall-survival comparison.

Comparator

Sorafenib 400 mg.

Primary question

How does overall survival compare between tremelimumab 300 mg x1 Dose + durvalumab 1500 mg and sorafenib 400 mg?

3. Trial Design

01
Randomize1,324 enrolled
02
Parallel arms4 study arms
03
Study treatmentDurvalumab / tremelimumab / sorafenib regimens
04
Follow-upOverall survival
05
AnalysisCox model + log-rank test
REGIMEN 1

Durvalumab 1500 mg

  • Durvalumab 1500 mg
REGIMEN 2

Tremelimumab 300 mg x1 Dose + Durvalumab 1500 mg

  • Tremelimumab 300 mg as a single dose
  • Durvalumab 1500 mg
REGIMEN 3

Tremelimumab 75 mg x4 Doses + Durvalumab 1500 mg

  • Tremelimumab 75 mg x4 doses
  • Durvalumab 1500 mg
  • This arm was closed during the study.
REGIMEN 4

Sorafenib 400 mg

  • Sorafenib 400 mg
Important design distinction: the registry data identify four study arms, but the statistical analyses posted on ClinicalTrials.gov focus on two specific comparisons. The primary efficacy comparison is tremelimumab 300 mg x1 Dose + durvalumab 1500 mg versus sorafenib 400 mg. The secondary analysis compares durvalumab 1500 mg versus sorafenib 400 mg. The tremelimumab 75 mg x4 Doses + durvalumab 1500 mg arm was closed during the study.

4. Analysis Population and Endpoint Framework

The reported efficacy analyses use the Full Analysis Set (FAS), defined in the ClinicalTrials.gov record as including all participants who were randomized. The registry also notes that the tremelimumab 75 mg x4 Doses + durvalumab 1500 mg arm was closed during the study.

Analysis populationRegistry definition / role
Full Analysis SetIncluded all participants who were randomized; used for the reported overall-survival analyses.
Tremelimumab 75 mg x4 Doses + Durvalumab armArm was closed during the study.

5. Primary Endpoint

EndpointRegistry definition / time frameStatistical analysis
Overall Survival (OS) - Treme 300 mg x1 Dose + Durva 1500 mg vs Sora 400 mg OS was defined as the time from the date of randomization until death due to any cause, regardless of whether the participant withdrew from randomized therapy or received another anticancer therapy. Any participant not known to have died at the DCO date was censored based on the last recorded date on which the participant was known to be alive. Time frame: from the date of randomization until death due to any cause, assessed up to the data cut-off date (27Aug2021, to a maximum of approximately 46 months). Cox proportional-hazards model and stratified log-rank test

This endpoint is a classic time-to-event outcome. The analysis therefore has to account for both the time until death and censoring of participants who were not known to have died by the data cutoff.

6. Statistical Methodology

Kaplan-Meier estimation

For an overall-survival endpoint, Kaplan-Meier estimation provides a nonparametric estimate of the probability of remaining alive over time. It is designed to incorporate right-censored observations, which occur when a participant reaches the data cutoff without a recorded death.

Conceptual survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of deaths at time ti, while ni is the number of participants at risk immediately before that time.

Stratified log-rank test

The registry reports a stratified log-rank test for the primary overall-survival comparison. The analysis adjusted for treatment, etiology of liver disease (HBV versus HCV versus others), ECOG (0 versus 1), and macro-vascular invasion (yes versus no). The registry text states that the values of these stratification factors were obtained from IVRS.

Cox proportional-hazards model

The primary hazard ratio was estimated using a Cox proportional-hazards model. The model adjusted for treatment, etiology of liver disease, ECOG, and macro-vascular invasion. This covariate adjustment changes the interpretation from a simple unadjusted comparison toward a model-based estimate accounting for the specified variables.

Hazard-ratio interpretation
HR = estimated hazard in treatment group ÷ estimated hazard in comparator group

An HR below 1 indicates a lower estimated instantaneous event rate in the first group relative to the comparator under the fitted model. It is not an absolute survival probability and does not directly state how many additional months an individual patient will live.

Covariate adjustment

The Cox model incorporated four prespecified adjustment dimensions reported in the registry analysis: treatment, etiology of liver disease, ECOG status, and macro-vascular invasion. Covariate adjustment can improve statistical efficiency and account for prognostic variables included in the model, but the resulting hazard ratio remains a model-based relative effect estimate.

Full Analysis Set

The reported analyses used the FAS, which included all randomized participants. This preserves the treatment assignment established by randomization rather than restricting efficacy analysis to participants who remained on treatment.

7. Primary Overall-Survival Result

The primary reported analysis compared tremelimumab 300 mg x1 Dose + durvalumab 1500 mg with sorafenib 400 mg. The analysis population was the FAS.

Hazard ratio for overall survival

0.78

95% CI: 0.66–0.92   ·   Two-sided CI   ·   Superiority hypothesis

Analysis: Cox proportional-hazards model with covariate adjustment.

Primary OS comparisonEstimate95% CIHypothesis
Treme 300 mg x1 Dose + Durva 1500 mg vs Sora 400 mgHR 0.780.66–0.92Superiority
Clinical Biostats interpretation

An HR of 0.78 means that, under the fitted Cox model, the estimated instantaneous hazard of death for the tremelimumab 300 mg x1 Dose + durvalumab 1500 mg group was approximately 22% lower than that for the sorafenib 400 mg group. The 22% figure is the simple relative interpretation of 1 − 0.78.

The HR does not mean that 22% of patients avoided death, that survival was extended by 22%, or that every participant experienced exactly a 22% reduction in risk. It is a relative, model-based time-to-event measure.

The 95% CI of 0.66–0.92 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of individual treatment effects among patients.

The confidence interval is particularly useful because it shows the precision of the estimate rather than simply whether a null value was crossed. A p-value, by contrast, addresses evidence against a specified null hypothesis; it does not measure the magnitude or clinical importance of the treatment effect.

Because the estimate comes from a Cox proportional-hazards model, interpretation also depends on the model's proportional-hazards framework. A single HR summarizes the relative hazard over the analyzed follow-up and should not automatically be interpreted as a constant relative risk at every time point.

Why the adjusted model matters

The reported Cox model did not simply compare the raw number or timing of deaths. It adjusted for etiology of liver disease, ECOG, and macro-vascular invasion in addition to treatment. These variables therefore form part of the statistical definition of the reported HR rather than being optional descriptive additions.

8. Primary Log-Rank Analysis

The same primary overall-survival comparison was evaluated with a stratified log-rank test.

Stratified log-rank test

P = 0.0035

Comparison: Treme 300 mg x1 Dose + Durva 1500 mg vs Sora 400 mg

Hypothesis type: Superiority

Clinical Biostats interpretation

The reported p-value of 0.0035 quantifies the evidence against the relevant null hypothesis under the specified stratified log-rank testing framework. It is not a measure of how large the treatment effect is.

The effect size is communicated by the HR of 0.78, while its uncertainty is communicated by the 95% CI of 0.66–0.92. The p-value and confidence interval therefore answer related but different statistical questions.

The stratified test accounts for the registry-specified stratification factors rather than treating all participants as belonging to a single homogeneous comparison. The corresponding Cox model provides the effect estimate, while the log-rank test provides the formal time-to-event comparison reported by the registry.

9. Secondary Overall-Survival Analysis: Non-Inferiority

A secondary overall-survival analysis compared durvalumab 1500 mg with sorafenib 400 mg. Unlike the primary superiority comparison, this analysis used a non-inferiority hypothesis and a prespecified non-inferiority margin.

Durvalumab vs sorafenib

HR 0.86

95.67% CI: 0.73–1.03

Non-inferiority margin: HR 1.08

Secondary OS comparisonEstimateConfidence intervalHypothesis
Durva 1500 mg vs Sora 400 mgHR 0.8695.67% CI 0.73–1.03Non-inferiority

The registry states that non-inferiority for the comparison of durvalumab 1500 mg versus sorafenib 400 mg is declared if the upper limit of the two-sided alpha-adjusted confidence interval for HR is less than the non-inferiority margin of 1.08.

Clinical Biostats interpretation

The HR of 0.86 is below 1, corresponding to an estimated 14% lower hazard of death under the Cox model. The primary statistical question here, however, is not simply whether the HR is below 1. It is whether the upper confidence limit remains below the prespecified non-inferiority margin of 1.08.

The reported 95.67% CI of 0.73–1.03 has an upper limit of 1.03, which is below the stated non-inferiority margin of 1.08. That is the comparison that matters for the reported non-inferiority criterion.

This illustrates why a non-inferiority trial should not be interpreted by applying the ordinary superiority rule of asking only whether the confidence interval excludes HR = 1. The relevant boundary is the non-inferiority margin.

The confidence interval also does not prove that the two treatments have identical effects. Rather, under the prespecified framework, it provides evidence that the upper bound remains within the amount of relative hazard permitted by the non-inferiority margin.

The registry specifies an alpha-adjusted confidence interval, and its unusual 95.67% level is therefore an important part of the formal testing framework rather than a number that should be silently replaced by 95%.

10. Sequential Superiority Testing for Durvalumab

After the non-inferiority criterion was satisfied, superiority was tested sequentially according to the protocol.

Stratified log-rank test

P = 0.0674

Comparison: Durva 1500 mg vs Sora 400 mg

Hypothesis type: Superiority

Clinical Biostats interpretation

The registry reports a p-value of 0.0674 for the superiority comparison after non-inferiority was satisfied. The sequential structure matters: superiority was not the first hypothesis tested in this comparison.

A p-value of 0.0674 does not measure the size of the observed treatment effect. The estimated effect remains the HR of 0.86, with the alpha-adjusted 95.67% CI of 0.73–1.03.

It is also important not to reinterpret the superiority p-value as a non-inferiority test. The two hypotheses use different statistical questions and different decision boundaries. Non-inferiority was assessed against the margin of 1.08; superiority was subsequently assessed as a separate sequential hypothesis.

11. Comparing the Two Statistical Questions

FeaturePrimary OS comparisonSecondary OS comparison
Treatment comparisonTreme 300 mg x1 Dose + Durva 1500 mg vs Sora 400 mgDurva 1500 mg vs Sora 400 mg
Effect measureHazard ratioHazard ratio
HR estimate0.780.86
Confidence interval95% CI 0.66–0.9295.67% CI 0.73–1.03
Primary hypothesisSuperiorityNon-inferiority
Non-inferiority marginNot applicable to this comparison1.08
Log-rank p-value0.00350.0674 for sequential superiority

This table illustrates an important statistical principle: the same effect measure can be used for different hypotheses. A hazard ratio of 0.86 is interpreted differently depending on whether the question is superiority or non-inferiority.

12. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

Overall survival records the time from randomization until death, with censoring for participants not known to have died by the data cutoff. The Cox model is designed for this type of time-to-event data and expresses the treatment comparison through a hazard ratio. In HIMALAYA, the reported model also adjusted for etiology of liver disease, ECOG, and macro-vascular invasion.

What does an HR of 0.78 mean?

An HR of 0.78 means that the fitted model estimates the instantaneous hazard of death to be about 78% as large in the tremelimumab 300 mg x1 Dose + durvalumab 1500 mg group as in the sorafenib 400 mg group, corresponding to a 22% lower estimated hazard. It does not mean that 22% of patients survived or that survival time increased by 22%.

Why is the confidence interval important?

The 95% CI of 0.66–0.92 shows the uncertainty around the HR estimate under the specified model and sampling framework. A confidence interval conveys substantially more information about precision than a p-value alone because it shows the range of effect estimates compatible with the statistical analysis at the stated confidence level.

Why is non-inferiority judged against a margin?

Non-inferiority asks whether the treatment's effect is not unacceptably worse than a prespecified reference boundary. For HIMALAYA's durvalumab-versus-sorafenib comparison, the registry specifies a margin of 1.08. The relevant question is therefore whether the upper limit of the alpha-adjusted confidence interval is below 1.08, rather than simply whether the interval excludes HR = 1.

Why does the unusual 95.67% confidence level matter?

The registry reports a 95.67% two-sided confidence interval for the secondary durvalumab-versus-sorafenib Cox analysis and states that it was alpha adjusted. This is a reminder that confidence levels in confirmatory trials can be modified to reflect the prespecified multiplicity and sequential-testing framework.

Why use both Cox regression and a log-rank test?

The two methods have complementary roles in the reported analysis. The Cox model supplies the hazard-ratio estimate and incorporates the specified covariate adjustment. The stratified log-rank test provides the formal comparison of survival distributions under the trial's stratification framework.

What does the FAS contribute to interpretation?

The Full Analysis Set includes all randomized participants according to the registry definition. An efficacy analysis based on randomized participants preserves the treatment comparison established by randomization and avoids redefining the primary comparison solely according to subsequent treatment exposure.

13. Stratification and Covariate Adjustment

The registry analysis identifies the same clinical variables in the survival analysis framework: etiology of liver disease (HBV versus HCV versus others), ECOG (0 versus 1), and macro-vascular invasion (yes versus no), together with treatment.

VariableReported categoriesRole in analysis
Etiology of liver diseaseHBV vs HCV vs othersCovariate adjustment / stratification
ECOG0 vs 1Covariate adjustment / stratification
Macro-vascular invasionYes vs noCovariate adjustment / stratification
TreatmentComparison-specific treatment groupsPrimary exposure of interest

The registry states that the values of the adjustment or stratification factors were obtained from IVRS. Statistically, this means the analysis was not based only on an unadjusted comparison of survival times.

Interpretive point: adjustment does not turn an observational analysis into a randomized trial. In this study, randomization supplies the principal basis for the treatment comparison; the Cox model then provides the specified adjusted effect estimate.

14. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm as the number affected over the number at risk.

Study armSerious adverse eventsAt risk
Durva 1500 mg115388
Treme 300 mg x1 Dose + Durva 1500 mg157388
Treme 75 mg x4 Doses + Durva 1500 mg52152
Sora 400 mg111374

The raw affected/at-risk counts show the number of participants with serious adverse events and the corresponding denominator reported by the registry. These counts should be kept separate from the efficacy analysis because safety and efficacy answer different questions and can use different analysis populations.

Serious adverse events: reported counts
Durva 1500 mg
115 / 388
Treme 300 + Durva
157 / 388
Treme 75 x4 + Durva
52 / 152
Sora 400 mg
111 / 374

The visualization is intentionally based on the reported affected counts, not a reconstructed percentage comparison. Because the denominators differ between arms, raw counts should not be interpreted as equivalent event rates.

15. Interim Analysis and Alpha Adjustment

The ClinicalTrials.gov record specifically identify interim analysis / alpha spending as a concept associated with the secondary durvalumab-versus-sorafenib Cox analysis. The registry also states that the 95.67% confidence interval was derived based upon the exact number of overall-survival events for each comparison using the Lan and DeMets approach that approximates the O'Brien Fleming spending function.

Why alpha adjustment matters

When a confirmatory analysis uses a prespecified testing sequence or multiple opportunities to evaluate hypotheses, the nominal confidence level may need adjustment so that the overall type I error remains consistent with the trial's design.

Why the CI is 95.67%

The reported secondary Cox analysis uses a 95.67% two-sided confidence interval rather than a conventional 95% interval. The registry explicitly identifies this as an alpha-adjusted interval.

The available registry text does not provide the complete Lan-DeMets specification or a full alpha-spending table. Accordingly, the statistical interpretation here is limited to the alpha-adjustment information explicitly reported in the ClinicalTrials.gov record.

16. Non-Inferiority Logic in HIMALAYA

The durvalumab 1500 mg versus sorafenib 400 mg comparison is particularly useful for understanding non-inferiority trial design.

Decision boundary
Upper CI limit < 1.08  →  non-inferiority criterion specified by the registry

The reported upper confidence limit is 1.03, while the prespecified non-inferiority margin is 1.08.

QuantityReported value
Hazard ratio0.86
Two-sided confidence level95.67%
Lower CI0.73
Upper CI1.03
Non-inferiority margin1.08
Subsequent superiority p-value0.0674

The key comparison is between 1.03 and 1.08. The confidence interval reaches above the conventional equality point of HR = 1, but its upper limit remains below the non-inferiority margin. That distinction is fundamental: a non-inferiority claim can remain statistically supported even when the confidence interval includes the null value of 1, provided the entire interval remains inside the prespecified non-inferiority boundary.

17. Overall-Survival Endpoint: Censoring and Interpretation

The registry definition of overall survival specifies that time is measured from randomization until death due to any cause. Participants not known to have died at the data cutoff are censored based on the last recorded date on which they were known to be alive.

Event

Death due to any cause.

Time origin

Date of randomization.

Censoring

Last recorded date on which a participant was known to be alive when death was not known at the data cutoff.

Intercurrent treatment

The endpoint definition states that OS remains measured regardless of withdrawal from randomized therapy or receipt of another anticancer therapy.

This is one reason overall survival is attractive as a clinical-trial endpoint: the event definition is death from any cause rather than a treatment-dependent assessment of disease progression. At the same time, because OS continues after treatment discontinuation or receipt of another anticancer therapy, it represents survival across the subsequent treatment pathway rather than only survival while receiving the assigned regimen.

18. Why a Hazard Ratio Is Not a Survival Probability

Statistical interpretation

The primary OS HR of 0.78 is a relative comparison of hazards. It does not tell us the probability that a particular participant will be alive at a particular calendar time.

A survival probability, such as S(t), answers a different question: the estimated probability of remaining alive beyond time t. The hazard ratio instead compares the instantaneous event rates between groups within the fitted Cox model.

Why absolute measures matter

Relative measures are useful for comparing treatment effects, but they do not by themselves reveal the absolute level of survival. A complete survival interpretation therefore normally considers Kaplan-Meier estimates or other absolute survival summaries alongside the hazard ratio.

Why proportional hazards matter

The Cox model expresses treatment effect through a hazard ratio under a proportional-hazards framework. If hazards vary substantially in their relative relationship over time, one HR can become an incomplete summary of the treatment difference. The ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption, so the HR should be interpreted within the stated model rather than as a guaranteed constant effect at every time point.

19. What the P-Values Do — and Do Not — Mean

Two p-values are reported for the survival comparisons in the ClinicalTrials.gov record: 0.0035 for the primary stratified log-rank comparison and 0.0674 for sequential superiority testing of durvalumab versus sorafenib.

p-valueComparisonStatistical role
0.0035Treme 300 mg x1 Dose + Durva 1500 mg vs Sora 400 mgPrimary superiority log-rank test
0.0674Durva 1500 mg vs Sora 400 mgSequential superiority test after non-inferiority was satisfied

A p-value is not the probability that the treatment works, nor is it the probability that the null hypothesis is true. It also cannot tell us whether an effect is clinically large. Those questions require the effect estimate, its confidence interval, and the clinical context of the endpoint.

20. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary overall-survival analysis reported an HR of 0.78 with a 95% CI of 0.66–0.92 and a stratified log-rank p-value of 0.0035. The secondary durvalumab comparison reported HR 0.86 with a 95.67% CI of 0.73–1.03 and a non-inferiority margin of 1.08.

Clinical interpretation

The statistical results describe relative differences in overall survival between specified randomized treatment groups. Clinical interpretation additionally requires considering the magnitude and precision of the effect, the endpoint definition, treatment exposure, safety, and the population represented by the trial.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

HIMALAYA is a useful statistical teaching case because its reported analyses combine several important clinical-trial concepts within the same overall-survival endpoint.

ConceptHow it appears in HIMALAYA
RandomizationParticipants were randomized in a phase 3 parallel trial.
Four-arm designThe study included four listed treatment/regimen arms.
Time-to-event endpointOverall survival was measured from randomization until death due to any cause.
CensoringParticipants not known to have died by the data cutoff were censored according to the last recorded date known alive.
Kaplan-Meier frameworkAppropriate for estimating survival distributions with right-censored observations.
Hazard ratioUsed as the reported effect measure for Cox survival analysis.
Cox regressionAdjusted for treatment, liver-disease etiology, ECOG, and macro-vascular invasion.
Log-rank testingUsed for formal time-to-event comparisons.
Stratified analysisSurvival testing incorporated the specified clinical stratification factors.
Non-inferiorityDurvalumab versus sorafenib was evaluated against an HR margin of 1.08.
Sequential testingSuperiority of durvalumab versus sorafenib was tested after non-inferiority was satisfied.
Alpha adjustmentThe secondary Cox analysis used a 95.67% alpha-adjusted confidence interval.
Safety analysisSerious adverse-event counts were reported separately by study arm.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue with Clinical Biostats statistical methods

Explore the statistical concepts that recur across randomized clinical trials, from survival analysis and hazard ratios to confidence intervals and non-inferiority design.

26. Record Summary

HIMALAYA provides a compact example of several important principles in clinical-trial statistics. The primary overall-survival comparison used a Cox proportional-hazards model and a stratified log-rank test, producing an HR of 0.78 with a 95% CI of 0.66–0.92 and a log-rank p-value of 0.0035. The secondary durvalumab-versus-sorafenib analysis illustrates a different inferential framework: a non-inferiority margin of 1.08, an alpha-adjusted 95.67% CI of 0.73–1.03, and sequential superiority testing with a p-value of 0.0674.

The most important statistical lesson is that these numbers cannot be interpreted independently of their hypotheses. The primary HR is a superiority effect estimate; the secondary HR is first evaluated against a non-inferiority margin and only then considered for sequential superiority. The confidence intervals quantify uncertainty, the p-values address hypothesis tests, and the Cox model defines how the reported hazard ratios are estimated.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical evidence from educational interpretation. For HIMALAYA, that means preserving the registry's endpoint definitions, analysis populations, effect measures, confidence levels, non-inferiority boundary, and sequential testing structure rather than reducing the trial to a single hazard ratio or p-value.