This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record for RATIONALE-301.
1. Trial at a Glance
RATIONALE-301 was a randomized, parallel-group phase 3 study evaluating tislelizumab versus sorafenib in participants with unresectable hepatocellular carcinoma. The main study's prespecified primary endpoint was overall survival, analyzed in the intention-to-treat population using a non-inferiority framework followed by a superiority assessment.
| Feature | RATIONALE-301 |
|---|---|
| Phase | Phase 3 |
| Condition | Hepatocellular Carcinoma (HCC) |
| Population | Participants with unresectable hepatocellular carcinoma |
| Design | Randomized, parallel-group |
| Masking | None |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 684 |
| Primary endpoint for main study | Overall survival (OS) |
| Study start | December 18, 2017 |
| Primary completion | July 11, 2022 |
| Sponsor | BeiGene |
| ClinicalTrials.gov | NCT03412773 |
2. Clinical Question
The central statistical question was whether tislelizumab could demonstrate non-inferior overall survival compared with sorafenib in the main study. The registry analysis also describes a subsequent superiority test for overall survival if non-inferiority was demonstrated.
Population
Participants with unresectable hepatocellular carcinoma (HCC).
Intervention
Tislelizumab.
Comparator
Sorafenib.
Primary question
Is overall survival with tislelizumab non-inferior to sorafenib, with a subsequent test for superiority if the non-inferiority criterion is met?
3. Trial Design
Tislelizumab
- Tislelizumab was the study intervention in Arm A.
- The main efficacy comparison was against sorafenib.
- Primary overall survival analysis used the ITT analysis set.
Sorafenib
- Sorafenib was the comparator in Arm B.
- The main efficacy comparison was against tislelizumab.
- Primary overall survival analysis used the ITT analysis set.
4. Randomization, Stratification, and Analysis Populations
The ClinicalTrials.gov record shows that the primary OS Cox model and the corresponding survival analyses incorporated treatment together with prespecified clinical factors. For the Cox model, the covariates were geography, macrovascular invasion/extrahepatic spread, etiology, and ECOG score. The survival comparisons were also described as stratified by these factors.
| Analysis population | Definition / role in the ClinicalTrials.gov record |
|---|---|
| Intent-to-treat (ITT) | All randomized participants in the main study; used for the primary overall survival analysis and the registry-reported efficacy analyses. |
| ITT with paired quality-of-life observations | For the Cycle 4 and Cycle 6 quality-of-life analyses, only participants with data at both baseline and the specified cycle were included. |
| Sub-study Safety Analysis Set | Reported as part of the analysis population for the investigator-assessed ORR and PFS outcomes at the later data cutoff. |
Covariate adjustment in the primary OS model
The registry-reported model specification used geography (Asia vs. EU/US), macrovascular invasion/extrahepatic spread (present vs. absent), etiology (HCV vs. other), and ECOG (0 vs. 1) as covariates.
This is different from simply comparing two unadjusted survival curves. The hazard ratio reported from the Cox model is a model-based estimate that incorporates the listed covariates, while the log-rank framework provides the time-to-event hypothesis-testing component.
5. Primary Endpoints
The registry data identify three primary endpoints overall: two in the safety run-in sub-study and one in the main study. The registry-reported formal statistical analyses include the main-study overall survival comparison.
| Primary endpoint | Time frame | Type | Statistical information reported |
|---|---|---|---|
| Safety Run-in Sub-study: Number of Participants With Treatment-emergent Adverse Events (TEAEs) | From the first dose to 30 days after the last dose, new anticancer therapy, or the analysis cutoff of December 14th, 2023 (a maximum of 64 months) | Binary | Results posted; no formal statistical analysis record reported |
| Safety Run-in Sub-study: Serum Concentration of Tislelizumab | Cycle 1 and Cycle 5 at end of infusion, 24 and 72 hours post-dose, and 8 days and 15 days post-dose | Other / unclear | Results posted; no formal statistical analysis record reported |
| Main Study: Overall Survival (OS) | Through the primary analysis data cut-off date of July 11th, 2022 (up to approximately 55 months) | Time-to-event | Formal Cox and log-rank analyses reported |
Main-study overall survival definition
Overall survival was defined as the time from the date of randomization to the date of death due to any cause. The registry states that median OS was estimated using Kaplan-Meier methodology.
6. Primary Results: Overall Survival
The primary efficacy result posted on ClinicalTrials.gov for RATIONALE-301 is the main-study overall survival comparison between tislelizumab and sorafenib in the ITT analysis set. The primary analysis data cutoff was July 11th, 2022, with follow-up of up to approximately 55 months.
Hazard ratio for overall survival
95.003% CI: 0.712–1.019 · two-sided
Primary OS comparison: tislelizumab vs sorafenib in the ITT analysis set.
| Feature | Reported result |
|---|---|
| Endpoint | Main Study: Overall Survival (OS) |
| Analysis population | Intent-To-Treat (ITT) analysis set |
| Comparison | Arm A: Tislelizumab vs Arm B: Sorafenib |
| Cox model HR | 0.85 |
| Cox model CI | 95.003% CI 0.712–1.019 |
| Log-rank HR | 0.85 |
| Log-rank CI | 95% CI 0.712–1.019 |
| Reported p-value | 0.0398 |
| Primary hypothesis framework | Non-inferiority, followed by superiority if non-inferiority was demonstrated |
The reported hazard ratio of 0.85 means that, under the fitted Cox model, the estimated hazard of death for tislelizumab relative to sorafenib was 0.85. Expressed as a simple relative interpretation, this corresponds to a 15% lower estimated hazard for tislelizumab relative to sorafenib.
The hazard ratio does not mean that 15% of patients benefited, that individual patients experienced exactly a 15% reduction in risk, or that overall survival probabilities were 15% higher. A hazard ratio is a relative time-to-event measure, not an absolute survival probability.
The confidence interval provides information about statistical precision. The registry-reported 95.003% CI is 0.712 to 1.019, so the interval extends on both sides of 1.00. In a conventional superiority interpretation, that means the interval includes the possibility of no hazard-ratio difference. In this trial, however, the primary question was structured first as a non-inferiority question, so the relevant comparison is not simply whether the interval excludes 1.00.
The p-value of 0.0398 is a reported two-sided p-value. A p-value quantifies compatibility with a specified null hypothesis under the statistical model; it does not measure the size, clinical importance, or certainty of the treatment effect. The registry separately describes a one-sided superiority boundary of 0.0223 after non-inferiority is demonstrated. The ClinicalTrials.gov record should therefore not be treated as though the two-sided p-value and the one-sided superiority boundary were interchangeable.
Finally, the HR depends on the Cox model and its assumptions, including the proportional-hazards framework. Censoring, the ITT population, the covariate specification, and the prespecified non-inferiority logic all matter when interpreting the estimate.
Non-inferiority framework
The registry-reported analysis text states that the null hypothesis for non-inferiority assumed the hazard ratio for tislelizumab versus sorafenib was greater than or equal to 1.08. The alternative hypothesis was therefore directed toward a hazard ratio below that non-inferiority boundary.
For this framework, the clinically relevant reference is the prespecified non-inferiority margin of 1.08, rather than 1.00 alone. A hazard ratio below 1.08 can satisfy the statistical criterion for non-inferiority even when the estimated treatment effect is not necessarily demonstrated to be superior under the separate superiority test.
Superiority was a separate question
The registry analysis states that superiority of tislelizumab over sorafenib for OS was tested using a stratified log-rank test in the ITT analysis set only when non-inferiority was demonstrated. It describes a one-sided superiority boundary of 0.0223, with superiority declared if the one-sided p-value crossed that boundary in favor of Arm A.
7. Primary Safety Run-in Endpoints
The ClinicalTrials.gov record identifies two safety run-in endpoints as primary endpoints for the sub-study. Results were posted, but no formal statistical-analysis records for these two endpoints were included in the statistical analyses posted on ClinicalTrials.gov.
Treatment-emergent adverse events
The endpoint was the Number of Participants With Treatment-emergent Adverse Events (TEAEs), measured from the first dose to 30 days after the last dose, new anticancer therapy, or the analysis cutoff of December 14th, 2023 (a maximum of 64 months).
The registry definition posted on ClinicalTrials.gov for an adverse event describes an unfavorable or unintended sign, symptom, or disease temporally associated with study drug use, regardless of causality. A serious adverse event was defined using criteria including death, life-threatening events, hospitalization or prolonged hospitalization, disability/incapacity, and congenital anomaly/birth defect.
Serum concentration of tislelizumab
Serum concentration of tislelizumab was a pre-specified primary endpoint for the safety run-in sub-study only. Measurements were scheduled in Cycle 1 and Cycle 5 at end of infusion, 24 and 72 hours post-dose, and 8 days and 15 days post-dose, with each cycle described as 3 weeks in the ClinicalTrials.gov record.
8. Secondary Efficacy Results
The ClinicalTrials.gov record includes overall response rate, progression-free survival, and time to progression. These endpoints illustrate why a clinical-trial statistical analysis should not treat every outcome as though it answered the same question.
Overall Response Rate by Blinded Independent Review Committee
Overall response rate (ORR) assessed by the Blinded Independent Review Committee (BIRC) was analyzed through the July 11th, 2022 primary analysis cutoff, up to approximately 55 months. The analysis used the ITT analysis set and a Cochran-Mantel-Haenszel test.
CMH ORR difference
95% CI: 3.85–12.70 · P = 0.0003
Reported as a difference in percentage of participants: tislelizumab vs sorafenib.
The reported 8.28 is an estimated between-arm ORR difference expressed in percentage-point units. It is not a hazard ratio and should not be interpreted as a relative increase in response probability.
The 95% CI of 3.85 to 12.70 quantifies uncertainty around the estimated difference under the specified analysis. The interval remains above zero, which is consistent with the direction of the reported hypothesis that ORR was higher in the tislelizumab group.
The reported p = 0.0003 addresses the null hypothesis that ORR is equal between groups under the registry-reported CMH framework. It does not measure the size of the response difference; the effect estimate and confidence interval provide that information.
Because the CMH analysis accounts for the trial's stratified categorical-data structure, the result should not be casually equated with a simple unadjusted subtraction of two percentages unless the same estimand and weighting structure are being used.
Investigator-Assessed Overall Response Rate
The later investigator-assessed ORR analysis used the study completion data cutoff of December 14th, 2023, up to approximately 65 months. The registry-reported analysis used the Cochran-Mantel-Haenszel test.
Later ORR difference
95% CI: 4.71–13.78 · P < 0.0001
Reported as a Cochran-Mantel-Haenszel ORR difference: tislelizumab vs sorafenib.
The later analysis estimates a 9.24 percentage-point difference in investigator-assessed ORR. The confidence interval of 4.71 to 13.78 describes the uncertainty around that estimate.
This is a different assessment source and data cutoff from the BIRC ORR result. The two estimates therefore should not be silently combined as though they were measurements from the same assessment process at the same time.
The very small reported p-value is evidence against the registry-reported equality null hypothesis under the CMH analysis. It is not itself a measure of clinical magnitude or treatment benefit.
Progression-Free Survival by BIRC
BIRC-assessed progression-free survival was evaluated through July 11th, 2022, up to approximately 55 months. The ITT analysis set was used, with a one-sided log-rank test and a Cox-derived hazard ratio.
PFS hazard ratio
95% CI: 0.92–1.33 · P = 0.1364
Reported for tislelizumab vs sorafenib.
A hazard ratio of 1.11 means the fitted model estimated a higher instantaneous rate of progression or death for tislelizumab relative to sorafenib over the analyzed period. It does not mean that progression occurred 11% more often in individual patients.
The 95% CI of 0.92 to 1.33 spans 1.00, so the interval includes both a possible modest lower hazard and a possible higher hazard under the model. The p-value of 0.1364 is a hypothesis-test result, not a measure of the clinical importance of the difference.
The registry-reported analysis describes the Cox model as stratified by geography, macrovascular invasion/extrahepatic spread, etiology, and ECOG. The proportional-hazards assumption remains relevant to interpretation of a single HR.
Investigator-Assessed Progression-Free Survival
Investigator-assessed PFS was analyzed through December 14th, 2023, up to approximately 65 months, using the ITT analysis set and a one-sided log-rank test.
Later PFS hazard ratio
95% CI: 0.90–1.26 · P = 0.2622
Reported for tislelizumab vs sorafenib.
The later PFS estimate of 1.06 is close to 1.00. The 95% CI of 0.90 to 1.26 includes 1.00 and therefore encompasses no estimated hazard difference as well as modest effects in either direction.
The later investigator assessment and the BIRC assessment should be interpreted as distinct analyses because the assessment source and data cutoff differ. Neither should be used to overwrite the other.
Time to Progression by BIRC
BIRC-assessed time to progression was evaluated through July 11th, 2022, up to approximately 55 months, in the ITT analysis set.
TTP hazard ratio
95% CI: 0.94–1.38 · P = 0.0859
Reported for tislelizumab vs sorafenib.
The estimated HR of 1.14 indicates a higher estimated instantaneous rate of progression for tislelizumab relative to sorafenib under the fitted model. The 95% CI of 0.94 to 1.38 includes 1.00, so the estimate has substantial uncertainty around the direction and magnitude of the difference.
TTP differs conceptually from PFS because TTP focuses on progression rather than the combined event of progression or death. These endpoints therefore should not be treated as interchangeable measures.
Investigator-Assessed Time to Progression
Investigator-assessed TTP was evaluated through December 14th, 2023, up to approximately 65 months, using the ITT analysis set.
Later TTP hazard ratio
95% CI: 0.94–1.34 · P = 0.1182
Reported for tislelizumab vs sorafenib.
The later TTP estimate of 1.12 is directionally similar to the BIRC TTP estimate of 1.14. However, the confidence interval of 0.94 to 1.34 includes 1.00, and the reported p-value is 0.1182.
The similarity of two point estimates does not by itself establish reproducibility or treatment-effect consistency. Their assessment source, cutoff date, and statistical uncertainty must also be considered.
9. Quality-of-Life Results
The registry analyses included change from baseline in EORTC QLQ HCC 18 and EORTC QLQ-C30 scores. These are continuous longitudinal outcomes rather than time-to-event endpoints, so the statistical model differs from the Cox and log-rank methods used for OS and PFS.
EORTC QLQ HCC 18 Index Score — Cycle 4
LS mean difference
95% CI: -3.8 to -0.8 · P = 0.0033
Change from baseline to Cycle 4; each cycle was 21 days.
The reported least-squares mean difference of -2.3 represents the model-based between-arm difference in change from baseline at Cycle 4, with the treatment contrast defined as tislelizumab versus sorafenib.
The 95% CI of -3.8 to -0.8 quantifies uncertainty around the model-based difference. Because the interval is entirely below zero, the estimated difference is directionally negative under the registry-reported model.
The p-value of 0.0033 tests the corresponding model-based null hypothesis. It does not establish that a particular magnitude is clinically important; clinical meaning depends on the scale, context, and prespecified interpretation of the questionnaire.
EORTC QLQ HCC 18 Index Score — Cycle 6
LS mean difference
95% CI: -4.7 to -0.7 · P = 0.0096
Change from baseline to Cycle 6; each cycle was 21 days.
The Cycle 6 estimate is a model-based LS mean difference of -2.7, with a 95% CI from -4.7 to -0.7. The confidence interval remains below zero.
The analysis population was the ITT analysis set, but only participants with data at both baseline and Cycle 6 were included. This restriction is important: the result is not identical in interpretation to a simple comparison of all randomized participants regardless of questionnaire availability.
EORTC QLQ-C30 Global Health Status / Quality of Life — Cycle 4
LS mean difference
95% CI: 1.4–7.3 · P = 0.0037
Change from baseline to Cycle 4; each cycle was 21 days.
The Cycle 4 global health status/quality-of-life estimate is a positive LS mean difference of 4.3 for tislelizumab versus sorafenib. The 95% CI of 1.4 to 7.3 indicates uncertainty around that model-based estimate while remaining above zero.
Because this is a continuous score rather than a time-to-event endpoint, the result should not be translated into a hazard ratio or a percentage reduction in risk.
EORTC QLQ-C30 Global Health Status / Quality of Life — Cycle 6
LS mean difference
95% CI: 1.8–8.2 · P = 0.0022
Change from baseline to Cycle 6; each cycle was 21 days.
The Cycle 6 estimate is a positive LS mean difference of 5.0. The 95% CI ranges from 1.8 to 8.2, and the registry-reported p-value is 0.0022.
The analysis is longitudinal in structure and uses a mixed model for repeated measures. The estimate therefore reflects the specified model rather than an unadjusted arithmetic difference between two raw means.
10. Statistical Methodology
Kaplan-Meier estimation
The registry defines OS as a time-to-event endpoint and states that median OS was estimated using Kaplan-Meier methodology. Kaplan-Meier estimation is appropriate for right-censored survival data because participants can contribute follow-up information even when they have not experienced the event by the end of observation.
Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.
The ClinicalTrials.gov record does not provide median OS values, Kaplan-Meier event counts, or reconstructed survival curves. Accordingly, this page does not manufacture those quantities from the reported hazard ratio.
Log-rank test
The main-study OS analysis included a log-rank comparison. The ClinicalTrials.gov record identifies a stratified log-rank framework for superiority testing after non-inferiority was demonstrated. PFS and TTP were also analyzed using one-sided log-rank tests in their registry-reported secondary analyses.
The log-rank test compares the observed and expected event patterns between groups over follow-up. It is therefore a time-to-event hypothesis test rather than a test of a single fixed-time survival percentage.
Cox proportional-hazards model
The primary OS analysis used a Cox proportional-hazards model. The model included treatment, geography, macrovascular invasion/extrahepatic spread, etiology, and ECOG as covariates. For the secondary PFS and TTP analyses, the registry-reported analysis notes describe treatment as a covariate with stratification by the same clinical factors.
This is a relative model-based comparison. It is not an absolute risk difference, a probability of benefit, or a statement about an individual participant's outcome.
Cochran-Mantel-Haenszel test
The BIRC and investigator-assessed ORR analyses used the Cochran-Mantel-Haenszel test. This method is useful for comparing categorical outcomes across treatment groups while accounting for stratification structure.
For RATIONALE-301, the reported estimand was an ORR difference rather than a hazard ratio. That distinction matters because a categorical response endpoint and a time-to-event endpoint contain different information and require different statistical models.
Mixed-effects model for repeated measures
The quality-of-life outcomes used a mixed model for repeated measures. The registry-reported model included baseline scores, stratification factors, treatment arms, visits, and treatment-arm-by-visit interaction as fixed effects. Visit was treated as a repeated measure with an unstructured covariance matrix based on the missing-at-random assumption.
This structure allows the analysis to model repeated observations from participants rather than treating each visit as an unrelated cross-sectional comparison.
11. Statistical Methods Explained
Why was a non-inferiority framework used for overall survival?
A non-inferiority design asks whether the experimental treatment preserves an acceptable amount of the comparator's efficacy. Here, the registry-reported null hypothesis was HR ≥ 1.08, with the alternative directed toward HR < 1.08. This means the trial's first question was not simply whether the HR was below 1.00; it was whether the upper side of the uncertainty was sufficiently favorable relative to the prespecified 1.08 boundary.
Why does the number 1.08 matter more than 1.00 for the primary non-inferiority question?
A value of 1.00 represents equality of hazards. A non-inferiority margin of 1.08 represents the largest unfavorable hazard ratio that the prespecified design was prepared to regard as compatible with non-inferiority. Therefore, the inferential reference for the first hypothesis is 1.08, not merely 1.00.
What does an OS hazard ratio of 0.85 mean?
It means the fitted Cox model estimated the instantaneous hazard of death in the tislelizumab group at 85% of that in the sorafenib group. The simple derived interpretation is a 15% lower estimated hazard. It does not mean 15% fewer deaths at every time point or a 15-percentage-point increase in survival.
Why can the PFS hazard ratio be above 1 while the OS hazard ratio is below 1?
OS and PFS are different endpoints. OS records death from any cause, while PFS records progression or death. Their event processes, censoring patterns, and subsequent treatment pathways can differ. A treatment therefore need not have the same hazard ratio for every time-to-event endpoint.
Why use a Cochran-Mantel-Haenszel test for ORR?
ORR is a categorical outcome, unlike OS, PFS, and TTP. The registry-reported analysis used the Cochran-Mantel-Haenszel framework to compare response between treatment groups while accounting for the stratified structure. The resulting effect measure was an ORR difference, not a hazard ratio.
Why was a mixed model used for quality-of-life outcomes?
Quality-of-life measurements were collected repeatedly over visits. A mixed model for repeated measures can account for within-participant correlation across visits while estimating treatment, visit, and treatment-by-visit effects. The registry-reported model also incorporated baseline scores and stratification factors and used an unstructured covariance matrix under a missing-at-random assumption.
Why is the p-value not the same thing as the effect size?
The p-value describes evidence against a specified null hypothesis under the assumed statistical model. It does not tell the reader how large the treatment effect is. For RATIONALE-301, the HR, ORR difference, LS mean difference, and their confidence intervals are the quantities that describe estimated effect magnitude and precision.
12. Missing Data and Analysis Assumptions
The registry-reported quality-of-life analyses explicitly state that only participants with data at both baseline and the relevant cycle were included. The mixed model for repeated measures was based on a missing-at-random assumption.
What missing at random means here
The model relies on an assumption about the relationship between missingness and the observed information included in the analysis. It is not equivalent to assuming that no observations are missing.
Why it matters
If the missing-data mechanism differs materially from the assumed structure, the estimated treatment difference in quality-of-life outcomes can be affected.
The ClinicalTrials.gov record does not describe a separate imputation procedure for the OS, PFS, or TTP analyses. Those time-to-event analyses instead operate through their specified event and censoring framework.
13. Safety Results
The registry-reported serious adverse-event data provide affected participants and participants at risk for the safety run-in sub-study, the two main-study treatment arms, and all enrolled participants during the screening period.
| Safety group | Serious adverse events | At risk | Affected / at risk |
|---|---|---|---|
| Safety Run-In Sub-study | 1 | 10 | 1/10 |
| Arm A: Tislelizumab | 104 | 338 | 104/338 |
| Arm B: Sorafenib | 91 | 324 | 91/324 |
| Screening Period: All Enrolled Participants | 7 | 684 | 7/684 |
These counts should be kept separate from the efficacy analysis. A serious adverse event is a safety outcome defined by prespecified clinical criteria; it is not an endpoint that can be directly combined with OS, ORR, PFS, or quality-of-life scores into a single statistical effect.
14. Non-Inferiority, Superiority, and the Logic of the Primary Analysis
RATIONALE-301 is statistically instructive because its primary OS analysis is not a single yes-or-no superiority test. The registry text describes a sequential inferential structure.
| Step | Question | Supplied statistical specification |
|---|---|---|
| 1 | Non-inferiority | Null hypothesis assumes HR ≥ 1.08; alternative favors HR < 1.08. |
| 2 | Superiority, if non-inferiority demonstrated | Stratified log-rank test in the ITT analysis set. |
| 3 | Superiority boundary | One-sided p-value boundary of 0.0223 in favor of Arm A. |
This hierarchy changes how the HR and confidence interval should be read. A confidence interval that includes 1.00 is not automatically evidence against non-inferiority, because the relevant margin is 1.08. Conversely, demonstrating non-inferiority is not identical to demonstrating superiority over the comparator.
The primary result is best understood as a margin-based hypothesis test followed, conditionally, by a separate superiority question. This is fundamentally different from asking only whether a 95% confidence interval excludes 1.00.
The ClinicalTrials.gov record reports an OS HR of 0.85 with a confidence interval extending to 1.019. Because 1.019 is below the stated non-inferiority boundary of 1.08, the interval is compatible with the non-inferiority framework described in the registry. The separate superiority assessment has its own one-sided boundary and should not be inferred by simply relabeling the reported two-sided p-value.
15. Stratified Analysis and Covariate Adjustment
The registry-reported Cox analyses repeatedly identify the same four clinical factors: geography, macrovascular invasion/extrahepatic spread, etiology, and ECOG score. Understanding their role is important because a stratified analysis is not the same thing as reporting an unadjusted treatment contrast.
| Factor | Categories reported in the analysis |
|---|---|
| Geography | Asia vs. EU/US |
| Macrovascular invasion / extrahepatic spread | Present vs. absent |
| Etiology | HCV vs. other |
| ECOG | 0 vs. 1 |
For the primary Cox analysis, these variables were included as covariates. For the registry-reported secondary survival analyses, the analysis notes describe treatment as a covariate with stratification by these factors. The distinction matters because the model specification determines how the treatment effect is estimated.
16. Censoring and Time-to-Event Interpretation
OS, PFS, and TTP are time-to-event outcomes. Participants may have different amounts of observed follow-up, and some participants may not experience the event before the analysis cutoff. Such observations are handled through censoring rather than being treated as though the event occurred at the last observation.
OS event
Death due to any cause, measured from randomization.
PFS event
Progression or death, according to the specified assessment.
TTP event
Progression, with the registry-reported endpoint distinguished from PFS.
Censoring
Allows participants without an observed event at the relevant observation limit to contribute partial follow-up information.
Because hazard ratios summarize relative instantaneous event rates over follow-up, they should not be interpreted as fixed-time risk ratios. A complete survival interpretation is strongest when the hazard ratio is considered together with Kaplan-Meier estimates and clinically meaningful time points; the ClinicalTrials.gov record does not provide those additional numerical survival estimates.
17. Comparing the Time-to-Event Results
| Endpoint | Assessment | HR | 95% CI | P-value |
|---|---|---|---|---|
| Overall Survival | Main study | 0.85 | 0.712–1.019 | 0.0398 |
| Progression-Free Survival | BIRC | 1.11 | 0.92–1.33 | 0.1364 |
| Progression-Free Survival | Investigator | 1.06 | 0.90–1.26 | 0.2622 |
| Time to Progression | BIRC | 1.14 | 0.94–1.38 | 0.0859 |
| Time to Progression | Investigator | 1.12 | 0.94–1.34 | 0.1182 |
The pattern illustrates why a trial should not be summarized by a single statistic. The OS estimate is below 1.00, whereas the registry-reported PFS and TTP estimates are above 1.00. The endpoints have different definitions, different event processes, and different inferential roles.
18. Multiplicity and Multiple Endpoints
the ClinicalTrials.gov record reports 3 primary endpoints overall and 24 posted outcome measures, with 12 statistical analyses posted. The primary endpoints span safety, serum concentration, and overall survival, while the secondary analyses cover response, progression, and quality-of-life measures.
Multiple endpoints create a statistical interpretation issue because each additional hypothesis test can contribute to the overall chance of observing a small p-value by chance alone. The ClinicalTrials.gov record explicitly specify a sequential non-inferiority/superiority framework for the main-study OS endpoint, but they do not provide a complete multiplicity-adjustment scheme for every secondary endpoint in the registry-reported material.
19. Analysis Population and the Meaning of ITT
The primary OS analysis used the Intent-To-Treat analysis set, defined in the ClinicalTrials.gov record as all randomized participants in the main study.
ITT analysis preserves the treatment assignment created by randomization. This is especially important in a comparative trial because post-randomization treatment changes, discontinuation, or other deviations can otherwise make treatment groups less comparable.
For RATIONALE-301, the registry-reported primary OS analysis explicitly uses the ITT set. The quality-of-life analyses have an additional requirement for observed baseline and Cycle 4 or Cycle 6 data.
20. Limitations
- Non-inferiority requires margin-based interpretation: the primary OS question is not adequately described by simply asking whether the HR confidence interval excludes 1.00. The registry-reported non-inferiority boundary is 1.08.
- Superiority is a separate hypothesis: the registry-reported analysis describes a one-sided superiority boundary of 0.0223 after non-inferiority is demonstrated. The reported OS p-value is reported as two-sided 0.0398 and should not be silently converted.
- Hazard-ratio assumptions: Cox HRs are model-based and rely on the proportional-hazards framework. A single HR does not describe the entire survival experience if hazards change materially over time.
- Endpoint differences: OS, PFS, TTP, ORR, and quality-of-life scores measure different clinical phenomena and cannot be collapsed into one common effect measure.
- Assessment source: BIRC and investigator assessments are different sources of outcome classification. Later results also use a different data cutoff.
- Quality-of-life missingness: the registry-reported mixed model uses a missing-at-random assumption, and the Cycle 4 and Cycle 6 analyses include participants with data at both baseline and the specified cycle.
- Multiple endpoints: the trial has several primary and secondary outcomes. The registry-reported material does not provide a complete multiplicity framework for every reported secondary p-value.
- Third arm: the registry profile reports 3 arms, but the registry-reported main-study statistical analyses specify only Arm A versus Arm B. No unsupported interpretation of the third arm is made here.
- Limited numerical survival detail in the ClinicalTrials.gov record: median survival values and time-specific survival probabilities are not included in the ClinicalTrials.gov record, so they are not reported here.
- Safety denominators: serious adverse-event figures are posted on ClinicalTrials.gov for different analysis groups and should not automatically be compared as if they shared identical exposure definitions.
21. Why This Trial Matters Statistically
RATIONALE-301 is a useful statistical teaching case because it combines a non-inferiority primary time-to-event question with a conditional superiority assessment and several distinct secondary endpoint types.
| Concept | How it appears in RATIONALE-301 |
|---|---|
| Randomization | Randomized, parallel-group phase 3 design. |
| Intention-to-treat analysis | Primary OS analysis includes all randomized participants in the main study. |
| Non-inferiority | Primary OS null hypothesis uses HR ≥ 1.08. |
| Superiority | Separate OS superiority assessment after non-inferiority, with a one-sided boundary of 0.0223. |
| Kaplan-Meier estimation | Registry states that median OS was estimated using Kaplan-Meier methodology. |
| Cox proportional-hazards model | Used for the primary OS HR and secondary PFS/TTP HRs. |
| Log-rank test | Used for OS, PFS, and TTP time-to-event comparisons. |
| Cochran-Mantel-Haenszel test | Used for BIRC and investigator-assessed ORR. |
| Covariate adjustment | Primary OS model includes geography, macrovascular invasion/extrahepatic spread, etiology, and ECOG. |
| Stratified analysis | Survival analyses account for the specified clinical stratification factors. |
| Mixed-effects model | Repeated-measures analysis of EORTC QLQ HCC 18 and QLQ-C30 outcomes. |
| Missing-data assumptions | Quality-of-life mixed model uses a missing-at-random assumption. |
| Multiple endpoints | Three registered primary endpoints and multiple secondary outcome measures. |
22. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The main-study OS analysis estimated an HR of 0.85 with a 95.003% CI of 0.712–1.019. The registry-reported analysis used an ITT population, a Cox model, and a non-inferiority margin of 1.08, with a separate superiority framework.
Clinical interpretation
The statistical estimates describe differences between randomized treatment groups for several distinct outcomes. Their clinical meaning depends on the endpoint, magnitude, uncertainty, follow-up, assessment method, and safety context rather than on any single p-value.
23. Interpreting Confidence Intervals Correctly
Confidence intervals are especially important in this trial because the endpoint types use different effect measures.
| Effect measure | Example from RATIONALE-301 | What the interval describes |
|---|---|---|
| Hazard ratio | OS HR 0.85; 95.003% CI 0.712–1.019 | Uncertainty around the model-based relative hazard estimate. |
| ORR difference | 8.28; 95% CI 3.85–12.70 | Uncertainty around the estimated between-arm response difference. |
| LS mean difference | Cycle 6 QLQ-C30: 5.0; 95% CI 1.8–8.2 | Uncertainty around the model-based difference in longitudinal change. |
These intervals should not be interpreted as ranges containing most individual patient outcomes. They quantify uncertainty about the estimated treatment contrast under the corresponding statistical model and inferential framework.
24. What the Primary OS Result Does — and Does Not — Mean
The OS HR of 0.85 means that the fitted Cox model estimated the hazard of death under tislelizumab at 85% of the hazard under sorafenib, corresponding to a simple derived interpretation of a 15% lower estimated hazard.
The reported 95.003% CI of 0.712–1.019 describes uncertainty around the estimated hazard ratio. It does not describe the range of effects experienced by individual participants.
The primary non-inferiority question uses 1.08 as the stated margin. Thus, the relevant inferential issue is whether the uncertainty around the estimated HR is sufficiently favorable relative to 1.08, not merely whether it is below 1.00.
The registry-reported OS p-value is 0.0398 and is reported as two-sided. It does not measure the size of the treatment effect and should not be substituted for the separate one-sided superiority boundary of 0.0223 described in the registry analysis.
25. Related Tutorials
Learn more about the statistical methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: RATIONALE-301, NCT03412773.
- PubMed: PMID 30969136.
- PubMed: PMID 37796513.
- PubMed: PMID 39435268.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts behind randomized trials, survival analysis, non-inferiority testing, categorical endpoints, and longitudinal models.
28. Record Summary
RATIONALE-301 provides a useful example of how the statistical interpretation of a phase 3 oncology trial depends on the exact hypothesis and endpoint rather than on a single p-value. The main-study OS analysis used an ITT population, a Cox proportional-hazards model, a reported HR of 0.85, and a 95.003% CI of 0.712–1.019. Its primary framework was non-inferiority against a stated HR margin of 1.08, followed by a separate superiority assessment with a one-sided boundary of 0.0223 if non-inferiority was demonstrated.
The secondary results illustrate the breadth of modern clinical-trial analysis: ORR was evaluated with the Cochran-Mantel-Haenszel test; PFS and TTP used log-rank and Cox methods; and quality-of-life outcomes used mixed models for repeated measures. The appropriate interpretation therefore combines effect estimates, confidence intervals, endpoint definitions, analysis populations, model assumptions, and the prespecified hypothesis structure.