← Clinical Trials
Hepatocellular Carcinoma Phase 3 Completed NCT01774344

RESORCE: Complete Statistical Analysis of Regorafenib in Hepatocellular Carcinoma

An independent statistical review of the randomized phase 3 RESORCE trial evaluating regorafenib versus placebo in patients with hepatocellular carcinoma, with emphasis on overall survival, time to progression, progression-free survival, tumor response, and the statistical methods used to analyze these endpoints.

Trial start: 2013-05-14  ·  Primary completion: 2016-02-29  ·  Enrollment: 573
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

RESORCE was a randomized, parallel-group, quadruple-masked phase 3 treatment trial evaluating regorafenib versus placebo in patients with hepatocellular carcinoma. The registry reports 573 enrolled participants, two study arms, one primary time-to-event endpoint, and posted statistical analyses using log-rank, stratified Cox, and Cochran-Mantel-Haenszel methods.

573
Enrolled
Phase 3
2
Arms
Parallel design
0.624
Primary OS HR
95% CI 0.498–0.782
0.000017
Primary OS P-value
Two-sided analysis
FeatureRESORCE
Trial nameRESORCE
Brief titleStudy of Regorafenib After Sorafenib in Patients With Hepatocellular Carcinoma
PhasePhase 3
StatusCOMPLETED
ConditionCarcinoma, Hepatocellular
AllocationRANDOMIZED
Design modelPARALLEL
MaskingQUADRUPLE
Primary purposeTREATMENT
Enrollment573.0
InterventionsRegorafenib (Stivarga, BAY73-4506); Placebo
Lead sponsorBayer
Sponsor typeINDUSTRY
Primary endpointOverall Survival (OS)
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Results postedYes
Statistical analyses posted15

2. Clinical Question

The central statistical question was whether randomized assignment to regorafenib could improve overall survival compared with placebo in patients with hepatocellular carcinoma, using a superiority framework.

Population

Patients enrolled in the RESORCE phase 3 study with the condition listed as carcinoma, hepatocellular.

Intervention

Regorafenib (Stivarga, BAY73-4506), including the 160 mg treatment specified in the posted statistical analyses.

Comparator

Placebo.

Primary question

Does regorafenib produce a different time-to-death distribution from placebo under the prespecified superiority comparison?

3. Trial Design

01
Randomize573 enrolled
02
Two armsRegorafenib or placebo
03
FollowTime-to-event outcomes
04
AssessOS, TTP, PFS, response
05
AnalyzeSurvival and categorical methods
INTERVENTION

Regorafenib

  • Regorafenib (Stivarga, BAY73-4506)
  • Posted primary and secondary analyses specify Regorafenib 160 mg (BAY73-4506)
  • Serious adverse events: 194 affected among 374 at risk
COMPARATOR

Placebo

  • Placebo
  • Compared directly with regorafenib in the posted analyses
  • Serious adverse events: 92 affected among 193 at risk

The registry identifies the design as randomized and parallel with quadruple masking. It therefore establishes a randomized treatment comparison rather than an observational comparison between patients who chose different therapies.

What the registry data do not establish: the ClinicalTrials.gov record does not provide a detailed randomization ratio, dosing schedule beyond the 160 mg designation in the statistical-analysis records, or a detailed sequence of treatment and follow-up procedures. Those details are therefore not added here.

4. Trial Timing and Registry Scope

2013-05-14

Trial start

The registry lists 2013-05-14 as the study start date.

2016-02-29

Primary completion

The registry lists 2016-02-29 as the primary completion date.

Completed

Current registry status

The ClinicalTrials.gov record identifies the study status as COMPLETED and reports results.

5. Endpoints

Primary endpoint

EndpointRegistry definitionTime frame
Overall Survival (OS) Overall Survival (OS) was defined as the time from date of randomization (Day 1) to death due to any cause. Subjects still alive at the time of analysis were censored at their last date of last contact. From randomization (Day 1) of the first subject until 419 days later

Secondary endpoints with posted analyses

EndpointTypeRegistry time frameMethod
Time to Progression (TTP) Time-to-event From date of randomization until 30 days after last study treatment (assessed every 6 weeks until PD; and after 8 cycle assessed every 12 weeks) (approximately 33 months) Log-rank test
Progression Free Survival (PFS) Time-to-event From date of randomization until 30 days after last study treatment (assessed every 6 weeks until PD; and after 8 cycle assessed every 12 weeks) (approximately 33 months) Log-rank test
Objective Tumor Response Rate (ORR) Binary From date of randomization until 30 days after last study treatment (assessed every 6 weeks until PD; and after 8 cycle assessed every 12 weeks) (approximately 33 months) Cochran-Mantel-Haenszel test
Disease Control Rate (DCR) Binary From date of randomization until 30 days after last study treatment (assessed every 6 weeks until PD; and after 8 cycle assessed every 12 weeks) (approximately 33 months) Cochran-Mantel-Haenszel test
Registry wording note: the ClinicalTrials.gov endpoint time-frame text ends with “after 8 cycle ”. That wording is reproduced rather than completed from outside information.

6. Statistical Methodology

The posted analyses use two principal statistical families. Time-to-event outcomes were analyzed with log-rank testing and Cox regression, with several analyses explicitly identified as stratified or unstratified sensitivity analyses. Binary response outcomes were analyzed with the Cochran-Mantel-Haenszel test.

Statistical componentHow it appears in RESORCEPurpose
Log-rank test Used for OS, TTP, and PFS analyses Compares time-to-event experience between randomized groups
Cox regression Used to calculate hazard ratios and confidence intervals Quantifies the relative event hazard between groups
Stratified analysis Used for the primary OS analysis and several secondary analyses Accounts for specified stratification factors in the survival comparison
Unstratified sensitivity analysis Reported for OS, TTP, and PFS Tests the robustness of the treatment estimate to the stratification specification
Cochran-Mantel-Haenszel test Used for ORR and DCR Compares binary outcomes using a stratified categorical-data framework

Stratified Cox model for overall survival

For the primary OS analysis, the registry states that the hazard ratio and 95% confidence interval were calculated for stratified IVRS using a Cox model. The model was stratified by geographic region (Asia or Rest of the World), ECOG-PS (0 versus 1), AFP level, presence versus absence of extrahepatic disease, and presence versus absence of macrovascular invasion.

Primary survival comparison
HR = estimated hazard of death under regorafenib relative to placebo

The registry reports a hazard ratio rather than a difference in median survival. The hazard ratio summarizes a relative event-rate comparison under the fitted Cox model and should not be interpreted as an absolute probability difference.

Kaplan-Meier estimation

The analysis notes state that Kaplan-Meier estimates for OS were used and that Kaplan-Meier survival curves were presented for each treatment. Kaplan-Meier estimation is appropriate for time-to-event data because it allows patients who remain event-free at the analysis point to contribute follow-up until censoring.

Conceptual Kaplan-Meier estimator
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at an event time and ni represents patients at risk immediately before that time. The ClinicalTrials.gov record does not provide the underlying event and censoring records needed to reconstruct the curve.

7. Primary Result: Overall Survival

Overall survival is the single registered primary endpoint. The ClinicalTrials.gov record contains three posted OS analyses. They represent the primary stratified analysis and two sensitivity analyses, rather than three separate primary endpoints.

AnalysisModel specificationHR95% CIP-value
Primary analysis Stratified Cox model; stratified IVRS 0.624 0.498–0.782 = 0.000017
Sensitivity analysis Stratified RAVE 0.660 0.527–0.828 = 0.000149
Sensitivity analysis Unstratified 0.674 0.546–0.831 = 0.000107

Primary OS estimate

HR 0.624

95% CI: 0.498–0.782   ·   P = 0.000017

Regorafenib 160 mg (BAY73-4506) versus placebo under the primary stratified analysis.

Clinical Biostats interpretation

An HR of 0.624 means that, under the fitted stratified Cox model, the estimated instantaneous hazard of death in the regorafenib group was 0.624 times that in the placebo group. Equivalently, 1 − 0.624 = 0.376, so the model-based relative hazard is approximately 37.6% lower under the comparison represented by the reported hazard ratio.

This does not mean that 37.6% of patients avoided death, that each individual had a 37.6% lower probability of dying, or that overall survival was extended by 37.6%. A hazard ratio is a relative time-to-event measure, not an absolute survival probability or a treatment-effect percentage for every patient.

The 95% CI of 0.498–0.782 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of effects experienced by individual patients.

The p-value of 0.000017 addresses evidence against the null hypothesis in the specified statistical test. It does not measure the size or clinical importance of the effect. Effect size and precision are better represented by the HR and its confidence interval.

The analysis is also dependent on the Cox-model framework, including the interpretation of a common hazard ratio over follow-up. The ClinicalTrials.gov record does not provide a formal assessment of the proportional-hazards assumption, so the HR should be understood as a model-based summary rather than a complete description of the survival experience.

Robustness across the posted OS analyses

The three posted estimates are 0.624, 0.660, and 0.674. Their confidence intervals all remain below 1.000, although they arise from different analysis specifications. The pattern is therefore useful statistically because the estimated treatment effect does not depend on one single reported Cox specification.

Stratified primary analysis

HR 0.624 with 95% CI 0.498–0.782 and P = 0.000017. The model incorporates the registry-specified stratification factors.

Stratified sensitivity analysis

HR 0.660 with 95% CI 0.527–0.828 and P = 0.000149. The registry identifies this as a stratified RAVE sensitivity analysis.

Unstratified sensitivity analysis

HR 0.674 with 95% CI 0.546–0.831 and P = 0.000107. This provides a comparison without the primary stratified specification.

Interpretive point

These are alternative analyses of the same primary endpoint, not three independent efficacy claims that should be counted as three separate primary hypotheses.

8. Secondary Results: Time to Progression

Time to Progression (TTP) is a time-to-event endpoint. The registry reports four statistical analyses: two based on stratified Cox regression and two based on unstratified Cox regression. The analysis notes state that one-sided p-values were calculated from the log-rank test and that 95% confidence intervals were computed using Kaplan-Meier estimates.

AnalysisSpecificationHR95% CIP-value
TTP analysis Stratified IVRS Cox regression 0.439 0.355–0.542 < 0.000001
TTP analysis Unstratified IVRS Cox regression 0.471 0.388–0.572 < 0.000001
TTP analysis Stratified IVRS Cox regression 0.412 0.334–0.509 < 0.000001
TTP analysis Unstratified IVRS Cox regression 0.444 0.365–0.539 < 0.000001

Reported TTP estimates

HR 0.412–0.471

All four posted analyses report P < 0.000001 with two-sided 95% confidence intervals.

The consistency of the TTP estimates across stratified and unstratified specifications is notable from a methodological perspective. The smallest reported HR is 0.412 and the largest is 0.471. These values should not be combined into a single estimate; they are separate analyses using different specifications.

How to interpret the TTP analyses

The hazard ratios are all below 1, indicating a lower estimated instantaneous progression hazard under the reported regorafenib comparison than under placebo. For example, an HR of 0.439 corresponds to a model-based estimated hazard approximately 43.9% as large as the comparator hazard, or approximately a 56.1% relative reduction in the estimated instantaneous hazard.

That interpretation does not imply that the probability of progression was reduced by exactly 56.1% for every patient. It also does not establish a particular median TTP because the ClinicalTrials.gov record contains no median estimates.

The p-values are extremely small under the reported testing framework, but they do not quantify the magnitude of the treatment effect. The HR and its confidence interval provide the effect-size and precision information.

Because the registry explicitly identifies one-sided testing for the p-value while reporting two-sided 95% confidence intervals, those quantities should not be treated as if they were generated under an identical one-sided/two-sided convention.

9. Secondary Results: Progression-Free Survival

Progression Free Survival (PFS) is also analyzed as a time-to-event endpoint. Four posted analyses are reported, again spanning stratified and unstratified Cox specifications.

AnalysisSpecificationHR95% CIP-value
PFS analysis Stratified IVRS Cox regression 0.453 0.369–0.555 < 0.000001
PFS analysis Unstratified Cox regression 0.480 0.397–0.580 < 0.000001
PFS analysis Stratified IVRS Cox regression 0.425 0.347–0.522 < 0.000001
PFS analysis Unstratified Cox regression 0.454 0.376–0.548 < 0.000001

Reported PFS estimates

HR 0.425–0.480

All four posted analyses report P < 0.000001 with two-sided 95% confidence intervals.

How to interpret the PFS analyses

The PFS hazard ratios range from 0.425 to 0.480. An HR of 0.453, for example, means that the fitted model estimates an instantaneous event hazard about 45.3% as large as the comparator hazard, corresponding to an approximately 54.7% lower estimated hazard.

The endpoint is still a time-to-event measure. A PFS hazard ratio does not provide a median PFS, a fixed-time PFS percentage, or an individual-level probability of remaining progression-free.

The confidence intervals provide the uncertainty around each estimate. The registry's analysis notes differ slightly across these records: some state that the 95% CI was computed using Cox regression, while another states Kaplan-Meier estimates. Those reported methodological distinctions should be preserved rather than silently harmonized.

The consistency of the estimates across stratified and unstratified models provides a useful sensitivity perspective, but these analyses should not be interpreted as independent opportunities to accumulate statistical significance.

10. Secondary Results: Objective Tumor Response Rate

Objective Tumor Response Rate (ORR) is a binary endpoint. The registry reports two Cochran-Mantel-Haenszel analyses with a reported difference as the effect measure.

AnalysisMethodReported difference95% CIP-value
ORR analysis Cochran-Mantel-Haenszel -6.88 -11.13 to -2.63 = 0.003650
ORR analysis Cochran-Mantel-Haenszel -4.15 -7.55 to -0.75 = 0.019991

The outcome unit is percentage of subjects. The registry comparison is ordered as “Placebo vs Regorafenib 160 mg (BAY73-4506),” and the effect measure is recorded as “Difference.” Because the ClinicalTrials.gov record does not explicitly state the subtraction convention used to construct that difference, the negative sign should not be independently translated into an arm-specific percentage-point advantage without assuming a formula that is not reported.

Statistical interpretation of ORR

The Cochran-Mantel-Haenszel test is appropriate for comparing binary response outcomes while accounting for stratification. The reported confidence intervals quantify uncertainty around the reported difference.

The two estimates, -6.88 and -4.15, are not interchangeable and should not be averaged. They represent separate posted analyses.

The p-values of 0.003650 and 0.019991 address the corresponding statistical tests. They do not measure the magnitude of the response difference, and they should not be interpreted without considering the multiplicity and analysis context of the full trial.

11. Secondary Results: Disease Control Rate

Disease Control Rate (DCR) is also reported as a binary endpoint and analyzed with the Cochran-Mantel-Haenszel test.

AnalysisMethodReported difference95% CIP-value
DCR analysis Cochran-Mantel-Haenszel -29.31 -37.52 to -21.11 < 0.000001
DCR analysis Cochran-Mantel-Haenszel -31.39 -39.57 to -23.22 < 0.000001

As with ORR, the ClinicalTrials.gov record specifies the comparison order and identify the effect measure as a difference, but do not provide an explicit formula defining which arm is subtracted from which. The safest interpretation is therefore to report the estimates exactly as registered rather than imposing an unreported sign convention.

Statistical interpretation of DCR

The two DCR estimates are -29.31 and -31.39, each accompanied by a 95% confidence interval that remains below zero and a p-value of < 0.000001.

These results indicate that the corresponding statistical comparisons were strongly separated under their reported analysis framework. They do not, by themselves, establish how long disease control lasted or how DCR translated into overall survival. Those are different clinical and statistical questions.

12. Secondary Results Summary

EndpointNumber of posted analysesEffect measures reportedStatistical method
Overall Survival 3 HR 0.624; 0.660; 0.674 Log-rank / Cox regression
Time to Progression 4 HR 0.439; 0.471; 0.412; 0.444 Log-rank / Cox regression
Progression Free Survival 4 HR 0.453; 0.480; 0.425; 0.454 Log-rank / Cox regression
Objective Tumor Response Rate 2 Difference -6.88; -4.15 Cochran-Mantel-Haenszel
Disease Control Rate 2 Difference -29.31; -31.39 Cochran-Mantel-Haenszel

The registry therefore provides a coherent statistical pattern: time-to-event endpoints use survival-analysis methods and binary endpoints use a stratified categorical-data method. The different effect measures should not be collapsed into a single summary statistic because they answer different questions.

13. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected patients divided by patients at risk.

ArmSerious adverse eventsAt risk
Placebo92193
Regorafenib (Stivarga, BAY73-4506)194374
Serious adverse events: affected / at risk
Placebo
92 / 193
Regorafenib
194 / 374

The displayed bars are simple visualizations of the registry-reported affected and at-risk counts. They are not a formal comparative safety analysis. In particular, the ClinicalTrials.gov record does not provide a confidence interval, p-value, exposure-adjusted incidence rate, seriousness definition beyond the category itself, or a formal between-arm statistical test for serious adverse events.

Safety denominator caution: the serious-adverse-event denominators are reported as 193 for placebo and 374 for regorafenib. They should not automatically be substituted for the overall enrollment number of 573 or treated as proof of the randomized allocation ratio.

14. Statistical Methods Explained

Why was a log-rank test used for overall survival?

Overall survival records the time from randomization to death and therefore contains both event timing and censoring information. A log-rank test is designed to compare survival distributions between groups across follow-up rather than reducing every patient to a simple binary outcome.

What does the primary hazard ratio of 0.624 mean?

It is a relative time-to-event estimate from the Cox model. Under the fitted model, the estimated instantaneous hazard of death in the regorafenib group is 0.624 times the comparator hazard. The corresponding relative reduction in estimated hazard is 37.6%. This is not an absolute survival difference and does not imply that every patient experiences the same proportional reduction.

Why was the OS analysis stratified?

The registry states that the primary OS Cox model was stratified by geographic region, ECOG-PS, AFP level, presence or absence of extrahepatic disease, and presence or absence of macrovascular invasion. Stratification allows the survival comparison to account for these prespecified factors without requiring the model to assign a single regression coefficient to each factor in the hazard-ratio estimate.

Why are there three OS hazard ratios?

The registry posts three analyses of the same primary endpoint: a stratified IVRS analysis, a stratified RAVE sensitivity analysis, and an unstratified sensitivity analysis. These are useful for assessing how the estimated treatment effect behaves under alternative analysis specifications. They should not be interpreted as three independent primary endpoints.

Why do the TTP and PFS analyses mention one-sided testing but report two-sided confidence intervals?

The registry explicitly states that one-sided p-values were calculated from the log-rank test while 95% confidence intervals were calculated separately. A p-value and a confidence interval are related but are not interchangeable summaries, and their sidedness should be reported exactly as specified rather than silently converting one to the other.

What does a confidence interval tell us?

A 95% confidence interval quantifies statistical uncertainty around an estimated effect under the relevant model and sampling framework. For the primary OS estimate, the interval is 0.498–0.782. It does not describe the probability that the true hazard ratio lies inside that particular interval, nor does it describe individual-patient treatment effects.

Why should the ORR and DCR differences not be interpreted like hazard ratios?

ORR and DCR are binary outcomes, so their reported effect measure is a difference rather than a hazard ratio. The numerical scale is therefore different. A difference should not be described as a relative reduction in hazard, and a hazard ratio should not be described as a percentage-point difference.

15. One-Sided Testing and Two-Sided Confidence Intervals

The registry provides an important methodological detail for TTP and PFS: the p-value was calculated using a one-sided log-rank test, while the reported confidence intervals are 95% and two-sided.

One-sided p-value

A one-sided test evaluates evidence in a prespecified direction. Its interpretation depends on the direction defined in the statistical hypothesis.

Two-sided 95% CI

A two-sided confidence interval communicates uncertainty on both sides of the estimated effect and is not simply a graphical restatement of a one-sided p-value.

This distinction matters when reading a registry table because it is possible to mistakenly assume that every reported inferential quantity uses the same sidedness. The ClinicalTrials.gov record explicitly say otherwise for the TTP and PFS analyses.

16. Stratified Analysis

Stratification is one of the recurring statistical themes in RESORCE. The primary OS analysis used a Cox model stratified by geographic region, ECOG-PS, AFP level, extrahepatic disease status, and macrovascular invasion status. Several TTP and PFS analyses are likewise identified as stratified or unstratified.

Stratification factorRegistry categories
Geographic regionAsia or Rest of the World (ROW)
ECOG-PS0 versus 1
AFP levelRegistry identifies AFP level as a stratification factor; the ClinicalTrials.gov record does not specify its categories
Extrahepatic diseasePresence versus absence
Macrovascular invasionPresence versus absence

A stratified analysis does not change the randomized treatment assignment. Instead, it changes how the statistical comparison uses the information associated with the specified strata. The availability of both stratified and unstratified analyses in the registry also creates a useful sensitivity-analysis framework.

17. Cochran-Mantel-Haenszel Analysis

The Cochran-Mantel-Haenszel test was used for ORR and DCR. This is a categorical-data method that can compare treatment groups while accounting for stratification. It is conceptually different from the survival-analysis machinery used for OS, TTP, and PFS.

Binary versus time-to-event endpoints
Binary outcome  →  CMH comparison   |   Time-to-event outcome  →  log-rank / Cox framework

The endpoint determines what information the statistical model must preserve. Binary response records whether an event occurred; time-to-event analysis additionally uses when the event occurred and handles censoring.

This distinction explains why the RESORCE statistical record contains both “Other difference” effect measures and hazard ratios. They are not alternative ways of reporting exactly the same estimand.

18. Multiplicity and the Interpretation of Multiple Analyses

The ClinicalTrials.gov record contains 15 statistical analyses: 3 for the primary OS endpoint, 4 for TTP, 4 for PFS, 2 for ORR, and 2 for DCR. The registry identifies OS as the single primary endpoint and the remaining outcomes as secondary endpoints.

Endpoint familyAnalyses postedRole
Overall Survival3Primary endpoint and sensitivity analyses
Time to Progression4Secondary endpoint
Progression Free Survival4Secondary endpoint
Objective Tumor Response Rate2Secondary endpoint
Disease Control Rate2Secondary endpoint

Multiple reported analyses do not automatically imply multiple independent confirmatory hypotheses. Sensitivity analyses are generally intended to evaluate robustness, while secondary endpoints address additional questions. Without a complete multiplicity strategy from the statistical analysis plan, the ClinicalTrials.gov record does not justify reconstructing a familywise-error hierarchy beyond what is explicitly reported.

Interpretive caution: very small p-values across several endpoints should not be treated as though they were independent replications. The endpoints are related, and several analyses are alternative specifications of the same endpoint.

19. Censoring and Time-to-Event Interpretation

The primary OS definition states that subjects who were alive at the time of analysis were censored at their last date of last contact. This is a central feature of survival analysis: a censored patient is not treated as having experienced the event at the censoring time.

Event

For OS, the event is death due to any cause.

Censoring

A subject still alive at analysis is censored at the last date of last contact.

Why this matters

Survival methods preserve follow-up information without pretending that an unobserved future event occurred at the last observed time.

Model dependence

The Cox hazard ratio adds a model-based summary on top of the underlying time-to-event data.

The ClinicalTrials.gov record does not provide the individual event and censoring records, so no independent Kaplan-Meier curve, median survival estimate, or reconstructed event count is generated on this page.

20. What the Primary Hazard Ratio Does — and Does Not — Mean

Effect size

The primary OS HR of 0.624 represents a relative hazard estimate. Numerically, it is 37.6% below 1.000, so it can be described as an approximately 37.6% lower estimated instantaneous hazard under the model.

Not an absolute risk reduction

The HR does not say that the absolute probability of death was reduced by 37.6 percentage points. No absolute survival probabilities are reported in the ClinicalTrials.gov record, so an absolute risk difference should not be invented.

Not a patient-level guarantee

A population-level hazard ratio does not imply that every individual experienced the same proportional change in risk. It summarizes the randomized groups statistically.

Confidence interval

The 95% CI of 0.498–0.782 shows the precision of the estimated HR under the specified analysis framework. It is not an interval containing individual treatment effects.

P-value

The p-value of 0.000017 describes the evidence against the null hypothesis under the specified test. It is not a probability that the null hypothesis is true and is not a measure of clinical magnitude.

21. Primary Analysis vs Sensitivity Analyses

The structure of the OS results is particularly useful for teaching sensitivity analysis. The registry reports a primary stratified analysis and two additional analyses using alternative specifications.

FeaturePrimary OS analysisSensitivity analyses
Analysis labelPrimaryStratified RAVE; unstratified
ModelCox modelCox model
StratificationSpecified clinical and geographic factorsStratified or unstratified depending on analysis
HR0.6240.660; 0.674
95% CI0.498–0.7820.527–0.828; 0.546–0.831
P-value= 0.000017= 0.000149; = 0.000107

The estimates are not identical, which is expected when the statistical specification changes. What matters for a sensitivity analysis is whether the substantive conclusion is highly dependent on one particular modeling choice. Within the ClinicalTrials.gov record, all three estimates are below 1 and all three reported confidence intervals remain below 1.

22. Limitations

23. Why This Trial Matters Statistically

RESORCE is a useful teaching case because the registry results show how a modern randomized oncology trial can combine several statistical frameworks within one study. The primary endpoint is a time-to-event outcome, while response outcomes are binary. The analysis also provides both stratified and unstratified sensitivity specifications.

ConceptHow it appears in RESORCE
RandomizationThe registry identifies the allocation as RANDOMIZED.
Parallel designThe design model is PARALLEL.
BlindingThe masking designation is QUADRUPLE.
Time-to-event endpointOverall Survival is the registered primary endpoint.
Kaplan-Meier estimationUsed for OS estimates and survival curves in the analysis notes.
Log-rank testUsed for OS, TTP, and PFS.
Hazard ratioUsed for OS, TTP, and PFS.
Stratified analysisUsed in the primary OS model and several secondary analyses.
Sensitivity analysisAlternative stratified and unstratified Cox specifications are posted.
Cochran-Mantel-Haenszel testUsed for ORR and DCR.
One-sided testingExplicitly reported for the TTP and PFS p-values.
Two-sided confidence intervalsReported for the hazard-ratio and binary difference estimates.
SuperiorityThe registry identifies the hypothesis type as SUPERIORITY.

The important lesson is that statistical interpretation should follow the endpoint. A hazard ratio is appropriate for describing the reported time-to-event comparison; a difference is the effect measure posted on ClinicalTrials.gov for ORR and DCR. Neither should be converted into the other merely for convenience.

24. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through the Clinical Biostats statistical pathway

Explore tutorials and statistical calculators related to survival analysis, hazard ratios, confidence intervals, stratified testing, and clinical-trial methods.

27. Record Summary

RESORCE provides a compact teaching example of how randomized clinical-trial evidence can be analyzed through multiple statistical lenses. The primary endpoint, overall survival, was analyzed with log-rank testing and a stratified Cox model, producing a primary hazard ratio of 0.624 with a 95% confidence interval of 0.498–0.782 and a p-value of 0.000017. Sensitivity analyses produced HRs of 0.660 and 0.674. Secondary time-to-event analyses for TTP and PFS also used log-rank and Cox methods, while ORR and DCR used Cochran-Mantel-Haenszel testing with reported differences.

The statistical story is therefore not simply the collection of p-values. It involves the choice of endpoint, the distinction between time-to-event and binary outcomes, stratification, censoring, model-based hazard ratios, confidence intervals, one-sided testing, and sensitivity analyses. The serious-adverse-event counts add a separate safety dimension, but the ClinicalTrials.gov record does not provide a formal comparative safety analysis.

Clinical Biostats methodology: A trial-results page should distinguish reported numerical evidence from statistical interpretation. The goal is to explain what each estimate means, what it does not mean, and how the analysis design affects interpretation without adding results that are not present in the underlying registry data.