← Clinical Trial Results
Advanced Hepatocellular Carcinoma Phase 3 Randomized NCT00105443

SHARP: Complete Statistical Analysis of Sorafenib in Advanced Hepatocellular Carcinoma

An independent statistical review of the randomized phase 3 SHARP trial comparing sorafenib with placebo in patients with advanced hepatocellular carcinoma, focusing on overall survival, time to symptomatic progression, time to progression, disease control, and the statistical methods used to analyze them.

Phase 3  ·  Completed  ·  Enrollment 602  ·  Sponsor: Bayer
Scope of this record

This page separates reported trial results from statistical interpretation. Trial-specific numerical values on this page are restricted to the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SHARP was a randomized, double-blind, parallel phase 3 study evaluating sorafenib versus placebo in patients with advanced hepatocellular carcinoma. The trial enrolled 602 participants and had two prespecified primary time-to-event endpoints: overall survival and time to symptomatic progression.

602
Enrollment
2 treatment arms
2
Primary endpoints
OS + TTSP
0.6931
OS HR
95% CI 0.5549–0.8658
1.0764
TTSP HR
95% CI 0.8837–1.3110
FeatureSHARP
Trial nameSHARP
Brief titleA Phase III Study of Sorafenib in Patients With Advanced Hepatocellular Carcinoma
PhasePhase 3
ConditionCarcinoma, Hepatocellular
DesignRandomized, double-blind, parallel
Primary purposeTreatment
Enrollment602
Arms2
InterventionsSorafenib (Nexavar, BAY43-9006) and placebo
Start2005-03
Primary completion2008-11
StatusCompleted
Lead sponsorBayer
Sponsor typeIndustry

2. Clinical Question

The central statistical question was whether assignment to sorafenib, compared with placebo, was associated with a different time-to-event experience in patients with advanced hepatocellular carcinoma. The registry identifies superiority as the hypothesis type for the primary analyses.

Population

Patients enrolled in the phase 3 study of sorafenib in patients with advanced hepatocellular carcinoma.

Intervention

Sorafenib (Nexavar, BAY43-9006).

Comparator

Placebo.

Primary question

How does sorafenib compare with placebo for overall survival and time to symptomatic progression?

3. Trial Design

01
Randomization 602 participants
02
Double-blind 2 treatment arms
03
Follow-up Time-to-event outcomes
04
Interim analysis Planned for OS
05
Final assessment OS and TTSP
ARM A

Sorafenib

  • Sorafenib (Nexavar, BAY43-9006)
  • Randomized treatment assignment
  • Double-blind phase
ARM B

Placebo

  • Placebo
  • Randomized treatment assignment
  • Double-blind phase

The registry describes the study as randomized, double-blind, and parallel. These design features are important statistically because randomization establishes the intended comparison between treatment assignments, while double blinding is designed to reduce the influence of treatment knowledge on participant and investigator behavior and assessment.

4. Endpoints

EndpointRegistry definitionTime frameType
Overall Survival (OS) Overall Survival was defined as the time from date of starting treatment to death due to any cause. Subjects still alive at the time of analysis were censored at their last date of last contact. From randomization to death due to any cause until an average 7.2 months later up to the data cut-off date approximately 19 months after start of enrollment Time-to-event
Time to Symptomatic Progression (TTSP) TTSP was defined as the time from randomization to the first documented symptomatic progression. From randomization to the first documented symptomatic progression until an average 4.8 months later up to the data cut-off Time-to-event

The distinction between the two primary endpoints is statistically important. OS records death from any cause, whereas TTSP records the first documented symptomatic progression. They therefore represent different event processes and can legitimately produce different treatment-effect estimates.

5. Statistical Methodology

Intention-to-treat analysis

The registry states that overall survival was measured in the ITT population from randomization until death due to any cause. TTSP was also analyzed in the ITT population. This preserves the treatment comparison generated by randomization: participants remain associated with the treatment assignment to which they were randomized rather than being reclassified according to subsequent treatment exposure.

Stratified log-rank testing

The primary OS comparison used a log-rank test stratified by region, ECOG PS, and "tumor burden". The TTSP comparison used the same stratification factors. Stratification allows the comparison to account for prespecified factors that may influence the event-time distribution while maintaining a treatment comparison within the randomized framework.

Primary time-to-event framework
Randomization → follow-up → event or censoring → comparison of event-time distributions

The log-rank test compares the observed and expected event patterns between treatment groups over follow-up. A hazard ratio provides a relative summary of the treatment effect.

Hazard ratio

The registry reports hazard ratios as the effect measure for OS, TTSP, and TTP. A hazard ratio below 1 for sorafenib versus placebo indicates a lower estimated instantaneous event rate under the fitted time-to-event comparison; a hazard ratio above 1 indicates a higher estimated event rate.

Interpretation
HR < 1  →  lower estimated event hazard with sorafenib
HR = 1  →  no estimated relative hazard difference
HR > 1  →  higher estimated event hazard with sorafenib

A hazard ratio is not an absolute risk difference, a probability of survival, or the proportion of patients who benefit.

Cochran-Mantel-Haenszel analysis

Disease control was a binary endpoint. The registry reports a Cochran-Mantel-Haenszel test with adjustment for region, ECOG PS, and "tumor burden." The reported effect measure was a risk difference, defined in the registry as the sorafenib minus placebo difference.

Interim analysis and alpha spending

Two formal interim analyses of overall survival were planned. The registry states that an alpha-spending function was used so that the false-positive rate was less than or equal to 0.02 on a 1-sided basis. This is a group-sequential principle: because the accumulating data can be examined more than once, the statistical design must account for repeated opportunities to declare superiority.

Different alpha levels for different primary endpoints

The registry reports a 1-sided alpha of 0.02 for the OS comparison and a 1-sided alpha of 0.005 for TTSP. It also states that TTSP was analyzed only at the end of the study, so no alpha-spending adjustment was necessary for that endpoint.

EndpointTestDirectionAlpha information in registryStratification
Overall Survival Log-rank 1-sided Overall alpha ≤ 0.02 with alpha spending Region, ECOG PS, "tumor burden"
TTSP Log-rank 1-sided Alpha 0.005; no alpha spending necessary Region, ECOG PS, "tumor burden"
TTP Log-rank 1-sided Alpha 0.025 Region, ECOG PS, "tumor burden"
Disease Control Cochran-Mantel-Haenszel 1-sided Alpha 0.025 Region, ECOG PS, "tumor burden"

6. Results: Overall Survival

Overall survival was the first primary endpoint. The registry reports a log-rank comparison in the ITT population and a hazard ratio for sorafenib versus placebo.

Hazard ratio for overall survival

0.6931

95% CI: 0.5549–0.8658   ·   P = 0.00583

Analysis: 1-sided stratified log-rank test; superiority hypothesis.

Relative hazard interpretation
Sorafenib
0.6931
Placebo reference
1.0000
Clinical Biostats interpretation

An HR of 0.6931 means that the estimated instantaneous hazard of death under the fitted comparison was about 69.31% of the corresponding hazard under placebo. Equivalently, the estimate corresponds to an approximately 19 months after start of enrollment 30.69% lower estimated hazard for sorafenib relative to placebo.

That statement is about a relative hazard, not about an individual's probability of dying, the percentage of patients who benefit, or a 30.69% increase in survival time. It also does not mean that every participant experienced the same relative reduction.

The 95% confidence interval of 0.5549–0.8658 describes uncertainty around the estimated hazard ratio under the statistical framework. Its width matters: the point estimate alone gives only one estimate, whereas the interval communicates how precisely the treatment effect was estimated.

The P = 0.00583 value addresses evidence against the null hypothesis under the specified one-sided testing framework. It is not a measure of the size or clinical importance of the treatment effect. The magnitude of the effect is conveyed by the hazard ratio and its confidence interval.

Because this was a time-to-event analysis, interpretation also depends on censoring and the assumptions underlying the hazard-ratio summary. The ClinicalTrials.gov record does not provide enough detail to independently assess the proportional-hazards assumption from the underlying event-time data.

7. Results: Time to Symptomatic Progression

TTSP was the second primary endpoint. The registry defines it as the time from randomization to the first documented symptomatic progression and reports a stratified one-sided log-rank comparison in the ITT population.

Hazard ratio for TTSP

1.0764

95% CI: 0.8837–1.3110   ·   P = 0.7676

Analysis: 1-sided stratified log-rank test; superiority hypothesis.

Relative hazard interpretation
Sorafenib
1.0764
Placebo reference
1.0000
Clinical Biostats interpretation

An HR of 1.0764 is above 1, meaning that the estimated instantaneous hazard of symptomatic progression was approximately 7.64% higher for sorafenib relative to placebo under the reported hazard-ratio comparison. This point estimate should not be interpreted as evidence that sorafenib causes a 7.64% increase in symptomatic progression risk for every patient.

The 95% CI of 0.8837–1.3110 spans 1.0. The interval therefore includes values corresponding to a lower hazard, little relative difference, and a higher hazard. This indicates substantial uncertainty around the point estimate.

The P = 0.7676 value is very different from the question of effect size. It indicates that the observed test statistic was not unusual under the specified null hypothesis and one-sided testing framework. It does not prove that the two treatment groups are identical, nor does it quantify the probability that the null hypothesis is true.

TTSP also illustrates why multiple endpoints require careful interpretation. OS and TTSP were both primary endpoints, but they had different reported alpha levels and different interim-analysis rules. The statistical evidence for one endpoint should not simply be substituted for the evidence for the other.

8. Results: Secondary Time to Progression

Time to Progression (TTP) was analyzed as a secondary time-to-event endpoint. The registry states that the primary analysis for TTP used independent radiological review in the ITT population and that the comparison was stratified by region, ECOG PS, and "tumor burden."

Hazard ratio for time to progression

0.5764

95% CI: 0.4484–0.7410   ·   P = 0.000007

Analysis: 1-sided stratified log-rank test; alpha 0.025.

Statistical interpretation

The TTP hazard ratio of 0.5764 corresponds to an estimated instantaneous hazard approximately 57.64% of the placebo hazard, or an approximately 19 months after start of enrollment 42.36% lower estimated hazard for sorafenib relative to placebo under the reported model comparison.

The confidence interval, 0.4484–0.7410, remains below 1.0. The P-value of 0.000007 provides the reported statistical evidence under the specified one-sided testing framework, but it should not be confused with a measure of effect magnitude. The hazard ratio and its confidence interval carry the information about relative effect size and precision.

TTP and TTSP should not be treated as interchangeable endpoints. TTP is based on disease progression by radiological assessment, whereas TTSP is based on first documented symptomatic progression. Their event definitions and censoring rules therefore represent different clinical processes.

9. Results: Disease Control

Disease Control (DC) was a binary secondary endpoint. The registry states that disease control was determined in the ITT population by independent radiological review and investigator assessment, using response evaluation criteria in solid tumors. The reported statistical comparison used the Cochran-Mantel-Haenszel test.

Risk difference for disease control

−11.95

95% CI: −19.56 to −4.35   ·   P = 0.001641

Reported as the sorafenib minus placebo difference; 1-sided CMH test with alpha 0.025.

Clinical Biostats interpretation

The registry explicitly defines the reported effect as sorafenib minus placebo. The risk difference of −11.95 therefore represents the estimated difference in disease-control rates on that direction of subtraction. Because the estimate is negative, the reported sorafenib rate was lower than the placebo rate under this coding of the endpoint.

The 95% CI of −19.56 to −4.35 remains below zero, indicating that the interval does not include no difference under the reported scale. The P-value of 0.001641 describes evidence under the specified one-sided CMH testing framework; it does not itself describe how large the disease-control difference is.

Unlike a hazard ratio, a risk difference is on an absolute difference scale. It therefore answers a different statistical question from the time-to-event analyses. The two types of effect measure should not be converted into one another without the underlying event probabilities and an explicit statistical model.

10. Secondary Endpoint Summary

EndpointRoleMethodEffect measureEstimate95% CIP-value
Overall Survival Primary Log-rank Hazard ratio 0.6931 0.5549–0.8658 0.00583
Time to Symptomatic Progression Primary Log-rank Hazard ratio 1.0764 0.8837–1.3110 0.7676
Time to Progression Secondary Log-rank Hazard ratio 0.5764 0.4484–0.7410 0.000007
Disease Control Secondary Cochran-Mantel-Haenszel Risk difference −11.95 −19.56 to −4.35 0.001641

Putting the endpoints together requires preserving their definitions. OS is a death endpoint, TTSP is symptomatic progression, TTP is radiological progression, and disease control is binary. Their effect measures therefore answer different questions even though all four compare sorafenib with placebo.

11. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment phase and arm. Because the available data use several differently labelled risk sets, the figures should be reproduced as reported rather than collapsed into a single rate.

Reported group / phaseAffectedAt risk
Sorafenib (Nexavar) Arm — Double-Blind Phase 153 297
Placebo Arm — Double-Blind Phase (Interim) 164 302
Sorafenib (Nexavar) Arm — Complete 191 297
Placebo Arm — Complete Double-Blind Phase 185 302
Subjects Switching From Placebo to Sorafenib 26 47
Safety interpretation: these are serious-adverse-event counts reported by the registry data. The phase labels and denominators differ, so the figures should not be treated as though they represented one common analysis population. In particular, the "interim," "complete double-blind," and "switching" categories describe different populations or phases.

12. Statistical Methods Explained

Why was a log-rank test used for OS and TTSP?

Both OS and TTSP are time-to-event endpoints. A simple comparison of proportions at one arbitrary time point would discard information about when events occurred and would require a particular time point to be selected. The log-rank test instead compares the event experience over follow-up while accommodating right censoring.

What does an OS hazard ratio of 0.6931 mean?

The estimate indicates that the modeled instantaneous hazard of death for sorafenib relative to placebo was approximately 0.6931. A useful derived interpretation is that the estimated hazard was about 30.69% lower. This is a relative hazard interpretation, not a statement that 30.69% of patients avoided death or that survival time increased by 30.69%.

Why is the TTSP hazard ratio different from the OS hazard ratio?

They measure different events. OS records death from any cause. TTSP records the first documented symptomatic progression. Different event definitions naturally generate different event times, censoring patterns, and treatment-effect estimates. There is no statistical requirement that their hazard ratios be similar.

Why was the analysis stratified?

The registry reports stratification by region, ECOG PS, and "tumor burden." Stratification allows the primary comparison to account for these factors when comparing the event experience of the randomized groups. It is particularly useful when the factors are expected to be associated with prognosis or event timing.

Why was alpha spending used for overall survival?

The registry states that two formal interim OS analyses were planned in addition to the final analysis. If the same nominal significance threshold were used repeatedly without adjustment, the overall false-positive probability could increase. Alpha spending allocates the available type I error across the planned looks at the data.

What does a risk difference of −11.95 mean?

The registry defines the disease-control estimate as sorafenib minus placebo. A risk difference of −11.95 therefore indicates that the disease-control proportion was estimated to be 11.95 units lower in the sorafenib group on the reported scale. It is an absolute difference, unlike a hazard ratio, and it should be interpreted using the endpoint's exact coding and direction.

Why does a P-value not measure effect size?

A P-value describes how compatible the observed test statistic is with a specified null hypothesis under the stated testing procedure. It depends on both the estimated effect and the amount of information in the analysis. The hazard ratio or risk difference describes the estimated effect itself, while the confidence interval communicates its precision.

13. Primary Analysis Populations and Censoring

The registry specifies the ITT population for the primary efficacy analyses. For OS, patients were followed from randomization until death due to any cause, while patients who were alive at analysis were censored at their last date of last contact.

For TTSP, subjects who had not progressed symptomatically at the interim analysis were censored at the date of their last Functional Assessment of Cancer Therapy assessment according to the registry-reported analysis text. This illustrates an important feature of survival analysis: patients who have not experienced the endpoint by the analysis cutoff do not simply disappear. Their available follow-up contributes information up to the censoring time.

Right-censoring concept
Observed time = min(event time, censoring time)

A censored observation tells the analysis that the event had not been observed through the censoring time. It does not establish that the event would never occur.

The ClinicalTrials.gov record does not provide enough information to independently reconstruct the complete censoring distribution, evaluate the missing-data mechanism, or reproduce the Kaplan-Meier estimates from individual participant records. Those issues therefore should not be inferred beyond the registry's stated definitions.

14. Interim Analysis, Alpha Spending, and Multiplicity

The OS analysis has an unusually important design feature: the trial planned two formal interim analyses in addition to the final analysis. The registry states that an alpha-spending function was used to ensure that the false-positive rate was less than or equal to 0.02 on a one-sided basis.

Repeated looks at OS

Two formal interim analyses were planned before the final analysis. This creates the need for prespecified error control across information times.

Alpha spending

The alpha-spending function distributes the allowable type I error across the planned analyses rather than treating every interim P-value as though it were the only test.

TTSP

The registry states that TTSP was analyzed only at the end of the study, so no alpha-spending adjustment was necessary for that endpoint.

Different endpoint alpha

TTSP used a reported one-sided alpha of 0.005, while OS used an overall one-sided alpha of 0.02 with interim alpha spending.

These details matter when interpreting the reported P-values. A P-value cannot be interpreted correctly without knowing the testing framework that generated it. In SHARP, the registry explicitly identifies one-sided testing, stratification, endpoint-specific alpha information, and interim monitoring for OS.

15. Understanding the Confidence Intervals

EndpointEstimate95% CIWhat the interval indicates
Overall Survival HR 0.6931 0.5549–0.8658 Uncertainty around the estimated relative hazard of death.
TTSP HR 1.0764 0.8837–1.3110 Uncertainty includes values below and above a hazard ratio of 1.
TTP HR 0.5764 0.4484–0.7410 Uncertainty around the estimated relative hazard of radiological progression.
Disease Control RD −11.95 −19.56 to −4.35 Uncertainty around the sorafenib-minus-placebo absolute difference.
Why the interval matters

A point estimate is only one location on a range of statistically plausible values. For OS, the entire reported 95% confidence interval is below 1.0. For TTSP, the interval crosses 1.0. For disease control, the confidence interval is entirely below zero because the reported effect was defined as sorafenib minus placebo.

The confidence interval does not describe the range of individual patient outcomes. It quantifies uncertainty around the estimated treatment effect under the specified statistical framework.

16. Comparing Relative and Absolute Effect Measures

SHARP provides a useful example of why different effect measures should remain separate. The primary time-to-event analyses use hazard ratios, while the disease-control analysis uses a risk difference.

MeasureSHARP exampleCore question
Hazard ratio OS HR 0.6931 How does the estimated instantaneous event hazard compare between groups over follow-up?
Hazard ratio TTP HR 0.5764 How does the estimated instantaneous progression hazard compare?
Risk difference DC RD −11.95 How different are the disease-control proportions on an absolute difference scale?

A hazard ratio cannot be interpreted as though it were a risk difference. For example, HR 0.6931 does not mean that 69.31% of patients survived or that there was a 30.69-percentage-point improvement in survival. Likewise, RD −11.95 is not a hazard ratio and should not be described as a relative reduction in disease control.

17. Design Topics Supported by the Registry

Design topicWhat the ClinicalTrials.gov record supports
RandomizationThe study was randomized with two parallel treatment arms.
BlindingThe study was double-blind.
Intention-to-treatPrimary OS and TTSP analyses were conducted in the ITT population.
StratificationPrimary time-to-event analyses were stratified by region, ECOG PS, and "tumor burden."
Interim analysisTwo formal interim OS analyses were planned.
Alpha spendingAn alpha-spending function controlled the overall one-sided OS false-positive rate at ≤0.02.
SuperiorityThe primary analyses were specified as superiority hypotheses.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported in the ClinicalTrials.gov record.
Factorial designNot reported; the registry describes a two-arm parallel design.

18. Important Limitations and Interpretation Issues

19. Why This Trial Matters Statistically

SHARP is a useful teaching example because the trial combines randomized treatment assignment with multiple time-to-event endpoints, stratified testing, a binary secondary endpoint, and formal interim monitoring. It demonstrates why understanding the statistical design is essential before interpreting a P-value or effect estimate.

ConceptHow it appears in SHARP
Randomization Two-arm randomized parallel design comparing sorafenib with placebo.
Blinding Double-blind study design.
ITT analysis Primary OS and TTSP analyses used the ITT population.
Time-to-event endpoints OS, TTSP, and TTP were analyzed as time-to-event outcomes.
Log-rank test Used for OS, TTSP, and TTP comparisons.
Hazard ratio Reported as the effect measure for the time-to-event endpoints.
Confidence intervals Reported for all four statistical analyses reported.
Stratified analysis Region, ECOG PS, and "tumor burden" were used as stratification factors.
Interim analysis Two formal interim OS analyses were planned.
Alpha spending Used to control the overall one-sided OS false-positive rate.
Cochran-Mantel-Haenszel test Used for the binary disease-control endpoint.
Risk difference Disease control was summarized as sorafenib minus placebo.

20. Overall Statistical Interpretation

OS

The reported OS hazard ratio was 0.6931, with a 95% CI of 0.5549–0.8658 and P = 0.00583. The estimate is below 1, corresponding to an approximately 30.69% lower estimated hazard of death under the reported comparison.

TTSP

The reported TTSP hazard ratio was 1.0764, with a 95% CI of 0.8837–1.3110 and P = 0.7676. The estimate was close to 1 relative to the other reported effects, and the confidence interval included 1.

TTP

The reported TTP hazard ratio was 0.5764, with a 95% CI of 0.4484–0.7410 and P = 0.000007. This is an estimated approximately 42.36% lower instantaneous progression hazard under the reported hazard-ratio comparison.

Disease control

The reported disease-control risk difference was −11.95, defined as sorafenib minus placebo, with a 95% CI of −19.56 to −4.35 and P = 0.001641. This is an absolute difference rather than a relative hazard measure.

These results should be read endpoint by endpoint. The strongest statistical interpretation comes from retaining the distinction among event definitions, effect measures, confidence intervals, testing direction, and the prespecified interim-analysis framework rather than compressing all results into a single numerical conclusion.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Calculators

23. Sources

Continue with the statistical methods

Explore the survival-analysis and clinical-trial methods that provide the statistical framework for interpreting randomized time-to-event studies.

24. Record Summary

SHARP provides a compact but statistically rich example of randomized clinical-trial analysis. The trial used a double-blind parallel design, compared sorafenib with placebo, analyzed the primary efficacy endpoints in the ITT population, and used stratified one-sided log-rank testing for overall survival and time to symptomatic progression. Overall survival was monitored through planned interim analyses using alpha spending, while TTSP was analyzed at the end of the study without an alpha-spending adjustment.

The reported results also demonstrate why endpoint definitions and effect measures matter. OS produced a hazard ratio of 0.6931, TTSP produced a hazard ratio of 1.0764, TTP produced a hazard ratio of 0.5764, and disease control was summarized using a risk difference of −11.95. These estimates cannot be interpreted as though they were interchangeable because they represent different outcomes and different statistical scales.

Clinical Biostats methodology: A trial-results page should not merely repeat reported numbers. The statistical objective is to explain what each endpoint measures, why the selected analysis is appropriate, how effect estimates and confidence intervals should be interpreted, and which design features affect the meaning of the reported P-values.