← Clinical Trials
Hepatocellular Carcinoma Phase 3 Recurrence-Free Survival NCT04102098

IMbrave050: Complete Statistical Analysis of Atezolizumab Plus Bevacizumab in Hepatocellular Carcinoma

An independent statistical review of the randomized phase 3 IMbrave050 trial evaluating atezolizumab plus bevacizumab versus active surveillance as adjuvant therapy in patients with hepatocellular carcinoma at high risk of recurrence after surgical resection or ablation.

Trial status: COMPLETED  ·  Enrollment: 668  ·  Primary completion: October 21, 2022
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

IMbrave050 was a randomized, parallel-group phase 3 trial comparing atezolizumab plus bevacizumab with active surveillance as adjuvant therapy in patients with hepatocellular carcinoma at high risk of recurrence after surgical resection or ablation.

668
Enrollment
Randomized trial
2
Arms
Parallel design
0.72
RFS HR
95% CI 0.56–0.93
0.0120
P-value
Two-sided
FeatureIMbrave050
Trial nameIMbrave050
NCT IDNCT04102098
PhasePhase 3
StatusCOMPLETED
ConditionCarcinoma, Hepatocellular
PopulationPatients with hepatocellular carcinoma at high risk of recurrence after surgical resection or ablation
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment668.0
InterventionAtezolizumab plus bevacizumab
ComparatorActive surveillance
Primary endpointRecurrence-Free Survival (RFS), as Determined by IRF
Primary endpoint typeTime-to-event
Primary analysis methodLog-rank test
Effect measureHazard ratio
Hypothesis typeSuperiority
Lead sponsorHoffmann-La Roche

2. Clinical Question

The primary statistical question was whether atezolizumab plus bevacizumab, compared with active surveillance, was associated with a difference in recurrence-free survival among patients with hepatocellular carcinoma at high risk of recurrence after surgical resection or ablation.

Population

Patients with hepatocellular carcinoma at high risk of recurrence after surgical resection or ablation.

Intervention

Atezolizumab plus bevacizumab.

Comparator

Active surveillance.

Primary question

Does atezolizumab plus bevacizumab improve recurrence-free survival relative to active surveillance?

3. Trial Design

IMbrave050 used a randomized, parallel-group phase 3 design with no masking. The trial enrolled 668 patients and compared two study groups: active surveillance and atezolizumab plus bevacizumab.

01
Randomize 668 patients
02
Two arms Active surveillance vs combination
03
Follow Baseline to approximately 33 months
04
Assess RFS IRF-determined recurrence or death
05
Compare Log-rank test and hazard ratio
ARM B · ACTIVE SURVEILLANCE

Active Surveillance

  • Active surveillance was the comparator arm.
  • Serious adverse events were reported for 34/330 affected/at risk.
ARM A · ATEZOLIZUMAB + BEVACIZUMAB

Atezolizumab Plus Bevacizumab

  • Atezolizumab and bevacizumab were the registered study interventions.
  • Serious adverse events were reported for 80/332 affected/at risk.
  • A separate crossover group is reported with 10/81 affected/at risk for serious adverse events.
What the design tells us: Randomization creates the framework for comparing recurrence-free survival between the assigned study groups. Because the trial was unmasked, the registry classifies the masking as none. The primary efficacy analysis nevertheless used the intention-to-treat population, preserving randomized treatment assignment as the basis for the comparison.

4. Trial Timeline

December 31, 2019

Trial start

The registry lists December 31, 2019 as the trial start date.

October 21, 2022

Primary completion

The registry lists October 21, 2022 as the primary completion date.

Approximately 33 months

Primary endpoint time frame

The registered recurrence-free survival endpoint is measured from baseline up to approximately 33 months.

5. Primary Endpoint

EndpointRegistry definitionTime frameEndpoint type
Recurrence-Free Survival (RFS), as Determined by IRF RFS is defined as the time from randomization to the first documented occurrence of intrahepatic or extrahepatic HCC as determined by an IRF, or death from any cause (whichever occurs first). Baseline up to approximately 33 months Time-to-event

This definition is important because recurrence-free survival is not simply a count of recurrent tumors. It is a time-to-event endpoint with a composite event definition: the first documented intrahepatic or extrahepatic HCC occurrence determined by an independent review function, or death from any cause, whichever occurs first.

Endpoint structure
RFS = time from randomization → first HCC recurrence or death

Patients who have not experienced the defined event by their last evaluable follow-up contribute information up to that point and are handled through time-to-event methods rather than being treated as simple binary responders or nonresponders.

6. Analysis Population and Stratification

The posted primary analysis used the intention-to-treat (ITT) population. The registry defines the ITT population as all randomized patients, whether or not the patient has received the assigned study treatment.

Analysis featureRegistry-supported description
Analysis population ITT population: all randomized patients, whether or not the patient has received the assigned study treatment.
Groups compared Arm B (Active Surveillance) vs Arm A (Atezolizumab Plus Bevacizumab)
Primary method Log Rank
Effect measure Hazard ratio
Hypothesis Superiority

The analysis also identifies stratified analysis as an analysis concept. The registry analysis notes specify the stratification factors as geographic region and high-risk features/curative procedure.

Stratification factorCategories
Geographic region Asia Pacific excluding Japan vs. rest of world
High risk features / curative procedure Ablation vs. resection with 1 high risk feature vs. resection with 2 or more high risk features

7. Statistical Methodology

Log-rank test

The primary comparison used a log-rank test. This is a standard method for comparing two time-to-event distributions while accounting for the timing of events and the fact that some patients may be censored before experiencing the event.

Conceptually, the log-rank test compares the observed number of events in each randomized group with the number expected under a null hypothesis of no difference between the groups. The test is therefore based on the ordering and timing of events rather than simply comparing the percentage of patients who eventually experienced recurrence.

Hazard ratio

The effect measure reported for the primary endpoint was a hazard ratio. The estimate was 0.72 for Arm B, active surveillance, versus Arm A, atezolizumab plus bevacizumab.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the numerator group

For this trial, the reported comparison is explicitly labeled Arm B (Active Surveillance) vs Arm A (Atezolizumab Plus Bevacizumab). Thus, careful attention to the group ordering is essential when interpreting the direction of the reported HR.

Stratified survival analysis

The registry analysis text identifies stratified analysis and provides the two sets of stratification factors. Stratification allows the treatment comparison to account for these prespecified factors rather than treating the trial population as completely homogeneous with respect to them.

Intention-to-treat analysis

The ITT principle means that randomized patients remain associated with their assigned study group for the primary efficacy analysis, whether or not they actually received the assigned treatment. This preserves the treatment comparison created by randomization.

Time-to-event analysis

RFS is inherently a time-to-event outcome. The statistical analysis therefore needs to distinguish between patients who experience recurrence or death, patients who remain event-free at their last assessment, and the timing at which those observations occur. The posted analysis uses a log-rank test and hazard ratio, both designed for this type of endpoint.

8. Primary Result: Recurrence-Free Survival

The registry reports a formal statistical analysis for the primary endpoint, Recurrence-Free Survival (RFS), as Determined by IRF, over the period from baseline up to approximately 33 months.

Hazard ratio for recurrence or death

0.72

95% CI: 0.56–0.93   ·   P = 0.0120

Comparison: Arm B (Active Surveillance) vs Arm A (Atezolizumab Plus Bevacizumab)

Primary endpointComparisonMethodEffect estimate95% CIP-value
Recurrence-Free Survival (RFS), as Determined by IRF Active Surveillance vs Atezolizumab Plus Bevacizumab Log-rank test HR 0.72 0.56–0.93 0.0120
Clinical Biostats interpretation

The reported hazard ratio of 0.72 means that, for the reported comparison of active surveillance versus atezolizumab plus bevacizumab, the estimated instantaneous rate of the defined RFS event in the numerator group was approximately 72% of that in the comparator group under the fitted time-to-event analysis. Equivalently, reversing the descriptive direction, the estimate corresponds to a hazard ratio of 1/0.72 only if one explicitly changes the group ordering; the registry's reported estimate should therefore be preserved as 0.72 rather than silently changing its direction.

The hazard ratio does not mean that 28% of patients avoided recurrence, that 28% of patients were cured, or that every patient experienced exactly a 28% reduction in recurrence risk. A hazard ratio is a relative time-to-event measure, not an absolute probability.

The 95% confidence interval of 0.56–0.93 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of outcomes that individual patients will experience.

The P-value of 0.0120 addresses evidence against the null hypothesis under the specified testing framework. It is not a measure of the magnitude or clinical importance of the treatment effect. Effect size and uncertainty are better conveyed by the hazard ratio and its confidence interval.

Because this is a time-to-event analysis, interpretation also depends on censoring and the assumptions underlying the hazard-ratio model. The ClinicalTrials.gov record does not provide a separate assessment of proportional hazards, so the reported HR should not be interpreted as proof that the hazard ratio was constant at every time point.

9. Understanding the Direction of the Hazard Ratio

The ordering of the groups is especially important in this registry analysis. The posted analysis explicitly lists Arm B (Active Surveillance) vs Arm A (Atezolizumab Plus Bevacizumab) and reports an HR of 0.72.

Reported direction

The registry reports the hazard ratio as Active Surveillance relative to Atezolizumab Plus Bevacizumab: HR 0.72.

What HR 0.72 is not

It is not an absolute recurrence probability, a difference in recurrence percentages, or a statement that 72% of patients experienced an event.

Why ordering matters

Changing the reference group changes the numerical reciprocal of a hazard ratio. Statistical interpretation must therefore retain the reported group order.

What the CI adds

The interval 0.56–0.93 communicates the uncertainty around the reported HR and is more informative than the point estimate alone.

This is a useful general lesson in clinical-trial reporting: a hazard ratio without its group ordering is incomplete. The same numerical effect can be expressed in the opposite direction when the reference group changes.

10. Why a Log-Rank Test Was Used

RFS is a time-to-event endpoint, so a simple comparison of the proportion of patients with recurrence would discard information about when recurrence or death occurred. The log-rank test instead uses the observed event times across follow-up.

Core statistical idea
Observed events − Expected events → evidence for separation of survival distributions

At each event time, the analysis considers how many patients remain at risk in each group and compares the observed pattern of events with the pattern expected under the null hypothesis of no treatment-group difference.

This approach is particularly appropriate when patients can have different follow-up durations. A patient who has not experienced recurrence by the end of observation is not treated as if the patient had been followed indefinitely without an event. Instead, the observation is censored at the appropriate time.

The log-rank test and the hazard ratio answer related but different statistical questions. The log-rank test provides a hypothesis test for differences between the time-to-event distributions, whereas the hazard ratio provides an estimate of the relative event rate under the fitted model.

11. Why the ITT Population Matters

The primary analysis population is explicitly defined as the ITT population: all randomized patients, whether or not they received the assigned study treatment.

Statistical interpretation

Randomization is the mechanism that establishes the treatment comparison. An ITT analysis maintains that randomized comparison rather than redefining groups according to treatment actually received. This reduces the risk that post-randomization treatment decisions systematically alter the original comparison.

ITT does not mean that every patient contributes an event. Patients can remain event-free at their last follow-up and therefore contribute censored observations. The analysis still uses the available follow-up time for each randomized patient.

For RFS, the combination of ITT analysis and time-to-event methodology allows the primary comparison to retain the randomized structure while appropriately incorporating different event and censoring times.

12. Stratified Analysis

The posted analysis identifies two stratification structures:

Stratification domainLevels specified in the registry analysis notes
Geographic region Asia Pacific excluding Japan vs. rest of world
High risk features / curative procedure Ablation vs. resection with 1 high risk feature vs. resection with 2 or more high risk features

Stratification is useful when important prognostic factors have been identified before the treatment comparison is made. Instead of assuming that all patients have identical underlying recurrence risk, the analysis can compare treatment groups while accounting for the prespecified strata.

Important distinction: stratification does not mean that each stratum is a separate randomized trial. The primary analysis remains a comparison within the overall randomized study, with the specified factors incorporated into the survival-analysis framework.

13. Statistical Methods Explained

Why was a log-rank test used?

Because the primary endpoint is recurrence-free survival, an outcome defined by the time from randomization until recurrence or death. The log-rank test is designed to compare time-to-event distributions while incorporating the timing of events and censored observations.

What does a hazard ratio of 0.72 mean?

It means that, for the reported comparison of active surveillance versus atezolizumab plus bevacizumab, the estimated instantaneous event rate was 0.72 times the corresponding rate in the comparison group under the fitted analysis. It does not mean that 72% of patients had an event or that every patient experienced a 28% reduction in risk.

Why is the confidence interval important?

The 95% confidence interval of 0.56–0.93 shows the statistical uncertainty around the estimated HR. A point estimate alone cannot communicate how precisely the treatment effect was estimated.

Why doesn't the P-value measure effect size?

The P-value of 0.0120 quantifies evidence against the null hypothesis under the specified statistical testing framework. It does not tell us how large the treatment effect is. The HR describes the estimated relative effect, while the confidence interval describes its uncertainty.

Why does the analysis use ITT?

The ITT population includes all randomized patients regardless of whether they received the assigned treatment. This preserves the comparison established by randomization and avoids redefining the efficacy population after randomization.

Why does the analysis use stratification?

The registry analysis identifies geographic region and high-risk features/curative procedure as stratification factors. Stratified analysis allows the treatment comparison to account for these prespecified categories when evaluating the time-to-event endpoint.

Why is RFS analyzed as a time-to-event endpoint rather than a simple percentage?

Because the registry definition is based on the time from randomization to the first documented recurrence or death. Patients can have different follow-up times, so a time-to-event analysis preserves information about when events occur and appropriately handles patients whose event status is not observed through the full follow-up period.

14. What the Confidence Interval Says

Primary RFS estimate

HR 0.72

95% CI: 0.56–0.93

Two-sided P = 0.0120

The confidence interval provides information that the point estimate alone cannot. The observed estimate is 0.72, but the analysis is subject to sampling uncertainty. The interval from 0.56 to 0.93 expresses that uncertainty under the specified statistical framework.

A confidence interval should not be interpreted as saying that there is a 95% probability that the true hazard ratio lies between 0.56 and 0.93. Rather, it is an interval produced by a procedure designed to have 95% coverage under repeated sampling when its assumptions hold.

The width of the interval also matters. A relatively narrow interval indicates greater precision than a very wide interval would, although precision and clinical importance remain distinct concepts.

Clinical Biostats interpretation

The most complete statistical statement is not simply "HR 0.72, P = 0.0120." It is that the trial reported a hazard ratio of 0.72, with a two-sided 95% CI of 0.56–0.93, using a log-rank analysis for the primary RFS endpoint in the ITT population, with the specified stratification factors incorporated into the analysis.

15. Primary Endpoint Interpretation

The primary endpoint combines two clinically important event types into one time-to-event outcome: documented intrahepatic or extrahepatic HCC recurrence and death from any cause. The event occurring first determines the RFS event.

Recurrence component

The first documented occurrence of intrahepatic or extrahepatic HCC determined by an IRF qualifies as an RFS event.

Death component

Death from any cause is also an RFS event, whichever occurs first.

Time origin

The registry defines RFS from randomization.

Assessment horizon

The registered time frame is baseline up to approximately 33 months.

This endpoint structure means that RFS should not be interpreted as recurrence alone. A treatment-group difference in RFS can arise through differences in the occurrence of documented recurrence, death before documented recurrence, or both.

16. Safety Results

The ClinicalTrials.gov record provides serious adverse events by arm as affected patients divided by the number at risk. These figures should be reported as reported in the registry rather than converted into a different metric.

GroupSerious adverse eventsReported affected / at risk
Active Surveillance Serious adverse events 34/330
Atezolizumab + Bevacizumab Serious adverse events 80/332
Crossover: Atezolizumab + Bevacizumab Serious adverse events 10/81
Safety interpretation: the ClinicalTrials.gov record reports serious adverse events as affected/at-risk counts. They do not provide a formal statistical comparison, confidence interval, or P-value for these safety figures. The counts therefore should not be presented as if a hypothesis test had been performed.

The safety and efficacy analyses also have different statistical purposes. The primary RFS analysis is explicitly based on the randomized ITT population, whereas the registry-reported serious-adverse-event data are presented as affected/at-risk counts by group. These should not be conflated into a single overall treatment-effect statistic.

17. Crossover and Its Statistical Implications

The ClinicalTrials.gov record identifies a crossover group for which serious adverse events were reported as 10/81. The ClinicalTrials.gov record does not provide the detailed crossover mechanism, timing, or an adjusted survival analysis specifically designed to estimate the effect of crossover.

From a statistical perspective, crossover is important because treatment received after randomization can complicate interpretation of an intention-to-treat comparison. The ITT comparison answers a question about the effect of being assigned to a strategy at randomization, whereas an analysis based on treatment actually received would answer a different question and can be affected by the reasons patients change treatment.

Interpretation rule: the presence of a crossover group does not justify automatically reanalyzing the primary endpoint according to treatment received. Any such analysis would require its own prespecified or appropriately justified methodology and assumptions.

18. Superiority Hypothesis

The registry classifies the primary hypothesis as superiority. This is distinct from a non-inferiority design.

FeatureIMbrave050
Hypothesis type Superiority
Primary endpoint Recurrence-Free Survival (RFS), as Determined by IRF
Primary endpoint type Time-to-event
Primary test Log-rank test
Effect measure Hazard ratio
Confidence interval 95%, two-sided

Because the ClinicalTrials.gov record identifies a superiority hypothesis, there is no non-inferiority margin to interpret here. The reported confidence interval and P-value should therefore be understood within a superiority-testing framework rather than through the logic of a prespecified non-inferiority boundary.

19. Multiplicity and Interim Analysis

The ClinicalTrials.gov record identifies one registered primary endpoint and one posted statistical analysis for that endpoint. They do not provide a multiplicity-adjustment procedure or an interim-analysis plan in the ClinicalTrials.gov record.

What is documented

One registered primary endpoint, one posted primary analysis, a log-rank method, a hazard ratio, a 95% two-sided confidence interval, and a superiority hypothesis.

What is not established here

The ClinicalTrials.gov record does not provide an alpha-spending scheme, multiplicity hierarchy, or interim-analysis boundary.

Accordingly, the reported P = 0.0120 should be described as the posted primary-analysis P-value. It should not be reinterpreted using an unreported alpha-spending or multiplicity procedure.

20. Missing Data, Censoring, and Time-to-Event Interpretation

Time-to-event analyses are designed to incorporate patients who have not experienced the event by their last available follow-up. Such observations contribute information until censoring rather than being classified simply as events or non-events.

The ClinicalTrials.gov record does not specify a detailed missing-data or imputation strategy for the RFS analysis. No particular imputation method should therefore be attributed to this trial on the basis of the ClinicalTrials.gov record.

Why censoring matters
Observed follow-up = event time, or time to last informative observation

The statistical analysis uses the timing of the event or censoring information rather than requiring every randomized patient to have a fully observed recurrence status at the end of the study.

A key assumption in conventional survival analysis is that the censoring mechanism is appropriately handled by the analysis framework. The ClinicalTrials.gov record does not provide enough information to assess that assumption empirically, so this page does not make a separate claim about the adequacy of the censoring process.

21. Proportional-Hazards Considerations

The primary effect measure is a hazard ratio. Hazard ratios are particularly useful for time-to-event comparisons, but their interpretation requires care when the relative event rates change substantially over time.

Clinical Biostats interpretation

An HR of 0.72 is a summary relative measure of the event rate under the fitted time-to-event analysis. It should not automatically be interpreted as saying that the treatment reduces the event probability by exactly 28% at every follow-up time.

If hazards are not approximately proportional, a single HR can summarize a more complex time-varying pattern imperfectly. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic or a time-varying treatment-effect analysis.

This distinction is especially important when explaining hazard ratios to readers who may otherwise interpret them as if they were ordinary risk ratios. A hazard ratio concerns the relative instantaneous event rate within the survival-analysis framework; it is not a direct statement about cumulative incidence at a particular time.

22. What This Result Does — and Does Not — Establish

It establishes a reported statistical comparison

The registry reports a formal log-rank analysis of the primary RFS endpoint, with HR 0.72, 95% CI 0.56–0.93, and P = 0.0120.

It does not establish an individual outcome

A population-level hazard ratio cannot determine whether a particular patient will experience recurrence or remain recurrence-free.

It does not equal an absolute risk reduction

The HR is a relative time-to-event measure and is not an absolute percentage-point difference in recurrence-free survival.

It does not prove a constant effect over time

The ClinicalTrials.gov record does not provide a separate proportional-hazards assessment or demonstrate that the HR is constant at every follow-up time.

23. Important Limitations and Interpretation Issues

24. Why This Trial Matters Statistically

IMbrave050 is a useful teaching case because its primary analysis brings together several core principles of modern clinical-trial biostatistics: randomization, an intention-to-treat population, a time-to-event endpoint, stratified analysis, a log-rank test, a hazard ratio, a confidence interval, and a superiority hypothesis.

ConceptHow it appears in IMbrave050
Randomization The trial is randomized with 668 enrolled patients.
Parallel design The design model is parallel with two arms.
Intention-to-treat analysis The primary analysis population includes all randomized patients, whether or not they received assigned study treatment.
Time-to-event endpoint RFS is measured from randomization to recurrence or death.
IRF determination HCC recurrence is determined by an IRF under the registered endpoint definition.
Log-rank test The posted primary analysis uses a log-rank method.
Hazard ratio The primary effect measure is HR 0.72.
Confidence interval The reported 95% two-sided CI is 0.56–0.93.
P-value The posted primary-analysis P-value is 0.0120.
Stratified analysis Geographic region and high-risk features/curative procedure are identified as stratification factors.
Superiority hypothesis The registry classifies the primary hypothesis as superiority.
Safety analysis Serious adverse events are reported as affected/at-risk counts by group.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Calculators

27. Sources

Continue with Clinical Biostats statistical methods

Explore the underlying biostatistical methods through tutorials, calculators, and additional clinical-trial analyses.

28. Record Summary

IMbrave050 provides a focused example of a randomized phase 3 time-to-event analysis. The trial enrolled 668 patients and evaluated atezolizumab plus bevacizumab versus active surveillance using a primary endpoint of Recurrence-Free Survival (RFS), as Determined by IRF. The registered RFS definition begins at randomization and ends at the first documented intrahepatic or extrahepatic HCC recurrence determined by an IRF, or death from any cause, whichever occurs first.

The posted primary analysis used the ITT population, a log-rank test, and a hazard ratio, with stratification by geographic region and high-risk features/curative procedure. The reported comparison of active surveillance versus atezolizumab plus bevacizumab produced an HR of 0.72, with a 95% two-sided CI of 0.56–0.93 and a P-value of 0.0120.

The central statistical lesson is that a time-to-event result should be interpreted as a complete package: the endpoint definition, randomized analysis population, group ordering, analysis method, effect estimate, confidence interval, and P-value all matter. The hazard ratio describes a relative event-rate measure, the confidence interval describes uncertainty around that estimate, and the P-value addresses evidence against the null hypothesis under the specified testing framework. None of these quantities alone describes the outcome of an individual patient.

Clinical Biostats methodology: This page separates the registry-reported statistical result from educational interpretation. Where the ClinicalTrials.gov record does not provide a design feature, adjustment procedure, subgroup result, or detailed missing-data method, no additional numerical or methodological claim has been added.