This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for RATIONALE-303 and are presented at the reported precision.
1. Trial at a Glance
RATIONALE-303 was a randomized, open-label, parallel phase 3 trial comparing tislelizumab with docetaxel in participants with non-small cell lung cancer in the second- or third-line setting. The registry reports two co-primary overall-survival endpoints: one in all participants and one in participants whose tumors were PD-L1 positive.
| Feature | RATIONALE-303 |
|---|---|
| Phase | Phase 3 |
| Condition | Non-small cell lung cancer |
| Design | Randomized, parallel |
| Allocation | Randomized |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 805 |
| Interventions | Tislelizumab and docetaxel |
| Primary endpoints | Overall survival in all participants; overall survival in PD-L1-positive participants |
| Trial status | Completed |
| Lead sponsor | BeiGene |
| ClinicalTrials.gov | NCT03358875 |
2. Clinical Question
The central statistical question was whether treatment with tislelizumab differed from docetaxel with respect to overall survival in participants with non-small cell lung cancer treated in the second- or third-line setting.
Population
Participants with non-small cell lung cancer enrolled in the second- or third-line treatment setting.
Intervention
Tislelizumab.
Comparator
Docetaxel.
Primary question
Does the randomized comparison show a difference in overall survival between tislelizumab and docetaxel, both in all participants and in the prespecified PD-L1-positive population?
3. Trial Design
Tislelizumab
- Drug intervention.
- Compared with docetaxel.
- Primary efficacy comparison based on overall survival.
Docetaxel
- Drug intervention.
- Comparator to tislelizumab.
- Used as the reference group for reported effect measures.
4. Endpoints
The registry identifies two co-primary endpoints, both based on overall survival. The first applies to all participants; the second applies to participants whose tumors were PD-L1 positive.
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Overall Survival (OS) in All Participants (Co-primary Endpoint) | OS was defined as the time from randomization to death from any cause. Median OS was calculated using the Kaplan-Meier method. Data for participants who were not reported as having died at the time of analysis were censored at the date they were last known to be alive. Data for participants who did not have postbaseline information were censored at the date of randomization. | From randomization to the data cutoff date of 10 August 2020; up to 32.4 months |
| Overall Survival (OS) in PD-L1-Positive Participants (Co-primary Endpoint) | OS was defined as the time from randomization to death from any cause. Median OS was calculated using the Kaplan-Meier method. Data for participants who were not reported as having died at the time of analysis were censored at the date they were last known to be alive. Data for participants who did not have postbaseline information were censored at the date of randomization. | From randomization up to the final efficacy analysis data cut-off date of 15 July 2021; up to 43 months |
PD-L1-positive population
The registry defines the PD-L1-Positive Analysis Set as all randomized patients whose tumors were PD-L1 positive, with PD-L1 positivity defined as ≥ 25% of tumor cells with PD-L1 membrane staining.
5. Statistical Methodology
Primary time-to-event analysis
The registry reports a 1-sided log-rank test for each co-primary overall-survival endpoint. The treatment effect is reported as a hazard ratio, with a two-sided 95% confidence interval.
For the all-participant OS analysis, the comparison was stratified by histology, line of therapy, and PD-L1 expression. The registry specifies histology as squamous versus non-squamous, line of therapy as second versus third, and PD-L1 expression as ≥25% versus <25% tumor cells.
For the PD-L1-positive OS analysis, the comparison was stratified by histology and line of therapy: squamous versus non-squamous and second versus third line.
The registry states that a stratified Cox proportional hazards model with Efron's method of tie handling was used to determine the hazard ratio and its 95% confidence interval for the all-participant OS analysis. The PD-L1-positive OS analysis is described as stratified by histology and line of therapy.
Secondary endpoint methodology
The secondary analyses span three statistical families: categorical response outcomes, time-to-event outcomes, and longitudinal quality-of-life outcomes.
| Endpoint family | Registry method | Effect measure |
|---|---|---|
| Objective response rate | Cochran-Mantel-Haenszel test | Odds ratio |
| Duration of response | Log-rank test | Hazard ratio |
| Progression-free survival | Log-rank test | Hazard ratio |
| Health-related quality of life | Mixed-effects model / MMRM | Least-squares mean difference |
Mixed-effects model for repeated measures
The quality-of-life analyses used a linear mixed-effect model for repeated measures. The registry description specifies baseline score, stratification factors, treatment arm, visit, and treatment-arm-by-visit interaction as fixed effects, with visit as a repeated measure.
This is important because the quality-of-life endpoint is not simply a comparison of two independent means at Cycle 6. The model incorporates the longitudinal structure of repeated assessments and explicitly models treatment-by-visit differences.
6. Results: Overall Survival in All Participants
The first co-primary endpoint was overall survival in all participants. The analysis used the Intent-to-Treat Analysis Set, defined in the registry as including all randomized patients.
Hazard ratio for overall survival
95% CI: 0.527–0.778 · 1-sided log-rank test P < 0.0001
Data cutoff: 10 August 2020; up to 32.4 months
| Element | Reported result |
|---|---|
| Analysis population | Intent-to-Treat Analysis Set; all randomized patients |
| Comparison | Tislelizumab vs Docetaxel |
| Method | 1-sided Log Rank Test |
| Effect measure | Hazard Ratio |
| Estimate | 0.64 |
| 95% CI | 0.527–0.778 |
| P-value | <0.0001 |
| Hypothesis type | Superiority |
| Stratification | Histology, line of therapy, and PD-L1 expression |
An HR of 0.64 means that, under the reported time-to-event model, the estimated hazard of death in the tislelizumab group was 0.64 times the estimated hazard in the docetaxel group over the analyzed follow-up. Put as a simple relative interpretation, this corresponds to an estimated 36% lower hazard of death because 1 − 0.64 = 0.36.
The HR does not mean that 36% of participants avoided death, that survival increased by 36%, or that every individual experienced the same reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute survival probability.
The 95% CI of 0.527–0.778 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of treatment effects that must occur in individual patients.
The P < 0.0001 result addresses evidence against the null hypothesis under the reported test. It does not measure the size or clinical importance of the effect. Effect size is conveyed by the HR and its confidence interval, while absolute survival measures would provide a complementary clinical perspective.
The analysis was stratified, and the registry also specifies a stratified Cox model for HR estimation. Interpretation therefore depends on the assumptions of the survival-analysis framework, including the adequacy of the model and the handling of censoring. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.
7. Results: Overall Survival in PD-L1-Positive Participants
The second co-primary endpoint evaluated overall survival among randomized participants whose tumors were PD-L1 positive, defined as ≥25% of tumor cells with PD-L1 membrane staining.
Hazard ratio for overall survival in PD-L1-positive participants
95% CI: 0.407–0.702 · 1-sided log-rank test P < 0.0001
Final efficacy analysis data cutoff: 15 July 2021; up to 43 months
| Element | Reported result |
|---|---|
| Analysis population | PD-L1-Positive Analysis Set; all randomized patients whose tumors were PD-L1 positive |
| PD-L1 definition | ≥25% of tumor cells with PD-L1 membrane staining |
| Comparison | Tislelizumab vs Docetaxel |
| Method | 1-sided Log Rank Test |
| Effect measure | Hazard Ratio |
| Estimate | 0.53 |
| 95% CI | 0.407–0.702 |
| P-value | <0.0001 |
| Hypothesis type | Superiority |
| Stratification | Histology and line of therapy |
An HR of 0.53 means that the estimated hazard of death was 0.53 times the estimated hazard in the docetaxel group for the PD-L1-positive analysis population. As a simple derived interpretation, this corresponds to an estimated 47% lower hazard because 1 − 0.53 = 0.47.
The estimate should not be read as a 47% increase in survival probability or as evidence that 47% of patients benefited. It is a relative time-to-event measure describing the fitted comparison between treatment groups.
The 95% CI of 0.407–0.702 gives the statistical uncertainty around the estimated HR. Its width also illustrates why an estimate should be considered together with its interval rather than treated as a single exact number.
The reported P < 0.0001 is evidence against the null under the stated one-sided log-rank test. A P-value is not an effect-size metric and does not tell us how clinically meaningful an observed difference is.
This analysis also uses a different data cutoff and analysis population from the all-participant OS result. The two HRs should therefore be compared descriptively, not treated as if they were estimates from precisely the same population and follow-up window.
8. Secondary Efficacy Results
The registry contains formal statistical analyses for objective response rate, duration of response, and progression-free survival. These results complement the two co-primary overall-survival analyses but answer different statistical questions.
Objective Response Rate in All Participants
Odds ratio for objective response
95% CI: 2.336–6.393 · P < 0.0001
Final efficacy analysis data cutoff: 15 July 2021; up to 43 months
The analysis used the Intent-to-Treat Analysis Set. The registry reports a Cochran-Mantel-Haenszel analysis stratified by histology, line of therapy, and PD-L1 expression.
An odds ratio of 3.86 means that the estimated odds of objective response were 3.86 times as high with tislelizumab as with docetaxel under the reported stratified analysis. This is an odds ratio, not a risk ratio or a difference in response percentages.
The 95% CI of 2.336–6.393 quantifies uncertainty around the estimated odds ratio. The P < 0.0001 value assesses evidence under the reported superiority test; it does not measure the magnitude of the response difference.
Objective Response Rate in PD-L1-Positive Participants
Odds ratio for objective response
95% CI: 3.721–17.379 · P < 0.0001
Final efficacy analysis data cutoff: 15 July 2021; up to 43 months
The PD-L1-positive analysis used the PD-L1-Positive Analysis Set and was stratified by histology and line of therapy.
An odds ratio of 8.04 indicates that the estimated odds of objective response were 8.04 times as high with tislelizumab as with docetaxel in the reported PD-L1-positive analysis. The estimate is not itself a probability and should not be interpreted as meaning that eight times as many participants responded.
The 95% CI of 3.721–17.379 indicates considerable uncertainty around the size of the odds ratio even though the entire reported interval is above 1. The P-value of <0.0001 addresses statistical evidence, not effect magnitude.
Duration of Response in All Responders
Hazard ratio for duration of response
95% CI: 0.176–0.536 · P < 0.0001
This analysis included participants in the Intent-to-Treat Analysis Set who had an objective response. The registry used a log-rank test and a stratified Cox proportional hazards model with Efron's method of tie handling for the HR and 95% CI. Stratification included histology, line of therapy, and PD-L1 expression.
An HR of 0.31 means that the estimated hazard of the response-ending event was 0.31 times that in the comparator group under the reported duration-of-response analysis. As a simple derived interpretation, this corresponds to an estimated 69% lower hazard because 1 − 0.31 = 0.69.
The analysis is conditional on having an objective response, so it addresses durability among responders rather than the probability of achieving a response in the first place. That distinction is important when interpreting it alongside ORR.
Duration of Response in PD-L1-Positive Responders
Hazard ratio for duration of response
95% CI: 0.066–0.370 · P < 0.0001
The PD-L1-positive duration-of-response analysis included participants in the PD-L1-Positive Analysis Set who had an objective response. The comparison was stratified by histology and line of therapy.
An HR of 0.16 corresponds to an estimated 84% lower hazard of the response-ending event because 1 − 0.16 = 0.84. It does not mean that 84% of responders maintained their response or that 84% were cured.
The 95% CI of 0.066–0.370 indicates uncertainty around the estimated relative hazard. Because this analysis is restricted to responders, it should not be substituted for the overall treatment comparison.
Progression-Free Survival in All Participants
Hazard ratio for progression-free survival
95% CI: 0.528–0.745 · P < 0.0001
The PFS analysis used the Intent-to-Treat Analysis Set. The registry reports a log-rank test and a stratified Cox proportional hazards model, with stratification by histology, line of therapy, and PD-L1 expression.
An HR of 0.63 means that the estimated hazard of the PFS event was 0.63 times that in the docetaxel group. As a simple derived interpretation, this corresponds to an estimated 37% lower hazard because 1 − 0.63 = 0.37.
The PFS hazard ratio does not tell us the absolute proportion of participants who remained progression-free at a particular time. It is a relative time-to-event estimate and should be interpreted together with its confidence interval and the definition of the PFS event.
Progression-Free Survival in PD-L1-Positive Participants
Hazard ratio for progression-free survival
95% CI: 0.285–0.494 · P < 0.0001
The PD-L1-positive PFS analysis used the PD-L1 Positive Analysis Set and was stratified by histology and line of therapy. The registry reports a stratified Cox proportional hazards model for the HR and 95% CI.
An HR of 0.38 corresponds to an estimated 62% lower hazard of the PFS event because 1 − 0.38 = 0.62. This is a relative hazard interpretation, not a statement that 62% of participants avoided progression or death.
The 95% CI of 0.285–0.494 quantifies uncertainty around the estimate, while P < 0.0001 addresses the statistical test rather than the clinical magnitude of the effect.
9. Quality-of-Life Results
The registry reports three formal mixed-model analyses of change from baseline in EORTC quality-of-life measures at Cycle 6, with each cycle defined as 3 weeks. These endpoints were analyzed in the HRQoL Analysis Set.
| Endpoint | Estimate | 95% CI | P-value |
|---|---|---|---|
| EORTC QLQ-C30 Global Health Status / Quality of Life Score | LS Mean Difference 5.7 | 2.38–9.07 | 0.0008 |
| QLQ-LC13 Coughing Scale | LS Mean Difference -8.3 | -13.02 to -3.51 | 0.0007 |
| QLQ-LC13 Dyspnoea Scale | LS Mean Difference -3.2 | -6.52 to 0.11 | 0.0579 |
| QLQ-LC13 Chest Pain Scale | LS Mean Difference -2.2 | -6.05 to 1.56 | 0.2472 |
The ClinicalTrials.gov record describes these as mixed-model analyses. The linear mixed-effect model for repeated measures includes baseline score, stratification factors, treatment arm, visit, and treatment-arm-by-visit interaction as fixed effects, with visit as a repeated measure.
Global health status / QOL
The reported LS mean difference was 5.7, with a 95% CI of 2.38–9.07 and P = 0.0008.
Coughing
The reported LS mean difference was -8.3, with a 95% CI of -13.02 to -3.51 and P = 0.0007.
Dyspnoea
The reported LS mean difference was -3.2, with a 95% CI of -6.52 to 0.11 and P = 0.0579.
Chest pain
The reported LS mean difference was -2.2, with a 95% CI of -6.05 to 1.56 and P = 0.2472.
10. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by the corresponding at-risk population.
| Treatment arm | Serious adverse events |
|---|---|
| Tislelizumab | 192/534 |
| Docetaxel | 84/258 |
The bars above are a visual representation of the reported affected counts, not a percentage comparison. The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates, so no P-value, risk ratio, odds ratio, or confidence interval is inferred here.
11. Statistical Methods Explained
Why was a log-rank test used for overall survival?
Overall survival is a time-to-event endpoint. Some participants die before the analysis cutoff, while others remain alive and are therefore censored. A log-rank test is designed to compare survival experience between groups while incorporating the timing of events and censoring rather than reducing each participant to a simple yes/no outcome.
What does an HR of 0.64 mean?
It means that the estimated hazard of death in the tislelizumab group was 0.64 times that in the docetaxel group under the reported survival model. The simple derived expression, 1 − 0.64 = 0.36, gives a 36% relative reduction in estimated hazard. It does not represent a 36-percentage-point improvement in survival probability.
Why was the analysis stratified?
The registry reports stratification by factors including histology, line of therapy, and PD-L1 expression. Stratification allows the time-to-event comparison to account for these prespecified factors rather than treating all participants as if they came from one completely homogeneous risk set.
Why is the odds ratio for ORR different from the hazard ratio for OS?
ORR is a binary outcome: a participant either meets the response definition or does not. OS is a time-to-event outcome: both whether and when death occurs matter. An odds ratio therefore summarizes relative odds of response, while a hazard ratio summarizes relative instantaneous event rates over time. They cannot be interpreted as interchangeable effect measures.
Why is the quality-of-life analysis different from the survival analysis?
The quality-of-life endpoints are continuous longitudinal measures observed at baseline and Cycle 6. The registry therefore uses a mixed-effects model for repeated measures rather than a survival model. The specified MMRM includes baseline score, treatment, visit, stratification factors, and treatment-by-visit interaction, with visit treated as a repeated measure.
Why should the two co-primary endpoints not be treated as one ordinary endpoint?
The registry identifies OS in all participants and OS in PD-L1-positive participants as separate co-primary endpoints. They also use different analysis populations and different data cutoffs. The ClinicalTrials.gov record does not specify an alpha-allocation or detailed multiplicity procedure, so no additional error-control rule is inferred. The two reported analyses should therefore be described using the registry's co-primary designation rather than creating an unreported hierarchy.
12. Understanding the Hazard Ratio
An HR below 1 indicates a lower estimated event hazard in the tislelizumab group under the fitted model. An HR of 1 would indicate no relative hazard difference. The hazard ratio is not an absolute risk difference and is not itself a survival probability.
OS HR 0.64
Estimated 36% lower hazard of death relative to the comparator under the reported model.
PD-L1-positive OS HR 0.53
Estimated 47% lower hazard of death in the PD-L1-positive analysis.
PFS HR 0.63
Estimated 37% lower hazard of the PFS event in all participants.
PD-L1-positive PFS HR 0.38
Estimated 62% lower hazard of the PFS event in the PD-L1-positive analysis.
These derived percentage statements are mathematical interpretations of the reported hazard ratios. They should not be confused with absolute percentage-point differences in survival or progression-free survival.
13. Confidence Intervals and P-values
| Analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| All-participant OS | HR 0.64 | 0.527–0.778 | <0.0001 |
| PD-L1-positive OS | HR 0.53 | 0.407–0.702 | <0.0001 |
| All-participant ORR | OR 3.86 | 2.336–6.393 | <0.0001 |
| PD-L1-positive ORR | OR 8.04 | 3.721–17.379 | <0.0001 |
| All-responder DOR | HR 0.31 | 0.176–0.536 | <0.0001 |
| PD-L1-positive responder DOR | HR 0.16 | 0.066–0.370 | <0.0001 |
| All-participant PFS | HR 0.63 | 0.528–0.745 | <0.0001 |
| PD-L1-positive PFS | HR 0.38 | 0.285–0.494 | <0.0001 |
| GHS/QOL change | LS Mean Difference 5.7 | 2.38–9.07 | 0.0008 |
| Coughing change | LS Mean Difference -8.3 | -13.02 to -3.51 | 0.0007 |
| Dyspnoea change | LS Mean Difference -3.2 | -6.52 to 0.11 | 0.0579 |
| Chest pain change | LS Mean Difference -2.2 | -6.05 to 1.56 | 0.2472 |
A confidence interval provides information about the precision of an estimate. It should be read together with the point estimate rather than as a substitute for it. A P-value answers a different question: it quantifies evidence under the specified null hypothesis and testing framework. Neither measure tells us whether an effect is clinically important by itself.
14. Analysis Populations
| Population | Definition / use |
|---|---|
| Intent-to-Treat Analysis Set | All randomized patients. Used for the all-participant OS, ORR, DOR, and PFS analyses reported in the ClinicalTrials.gov record. |
| PD-L1-Positive Analysis Set | All randomized patients whose tumors were PD-L1 positive, defined as ≥25% of tumor cells with PD-L1 membrane staining. |
| HRQoL Analysis Set | All randomized participants who received ≥1 dose of study drug and completed ≥1 HRQoL assessment; the registry-reported description states that participants with available data at baseline and Cycle 6 are included. |
| Responder populations | For DOR, participants in the relevant analysis population who had an objective response. |
The use of different analysis populations is not a technical footnote. It changes the question being answered. ITT preserves the randomized comparison for efficacy, whereas the HRQoL analysis requires treatment exposure and quality-of-life assessments, and DOR necessarily conditions on having responded.
15. Stratified Analysis
Stratification appears repeatedly in the RATIONALE-303 statistical analyses. For the all-participant primary OS endpoint, the registry identifies three stratification factors:
Histology
Squamous versus non-squamous.
Line of therapy
Second versus third.
PD-L1 expression
≥25% versus <25% tumor cells for the all-participant OS analysis.
PD-L1-positive analyses
The PD-L1-positive OS, ORR, DOR and PFS analyses are stratified by histology and line of therapy.
Stratification is particularly relevant in a randomized trial because it connects the statistical analysis to factors used to structure the comparison. It does not mean that the treatment effect is separately estimated within every stratum and then simply averaged. The reported analysis uses a stratified comparison framework.
16. Kaplan-Meier Estimation and Censoring
The registry explicitly states that median OS was calculated using the Kaplan-Meier method. It also specifies how participants who had not died at the time of analysis were handled: they were censored at the date they were last known to be alive.
At each observed event time, the estimated survival function is updated according to the number of events and the number of participants at risk immediately before that time.
The registry also states that participants without postbaseline information were censored at the date of randomization. This is an important detail because censoring rules determine how available follow-up contributes to the survival estimate.
17. The One-Sided Test and Two-Sided Confidence Interval
Both primary analyses are reported with a 1-sided log-rank test and a two-sided 95% confidence interval. This distinction is worth preserving rather than silently converting the analysis into a generic two-sided test.
One-sided P-value
The reported log-rank P-value evaluates the prespecified directional superiority framework recorded in the registry.
Two-sided CI
The reported confidence interval is two-sided and gives an interval estimate around the hazard ratio.
The ClinicalTrials.gov record does not report the complete multiplicity or alpha-allocation procedure associated with the two co-primary endpoints. Consequently, this page does not infer a particular familywise error allocation or rewrite the reported P-values under an unreported procedure.
18. Multiplicity and Multiple Endpoints
RATIONALE-303 has two registered co-primary endpoints, both overall survival endpoints but in different analysis populations and at different data cutoffs. The ClinicalTrials.gov record also contain numerous secondary analyses covering response, duration of response, progression-free survival, and quality of life.
| Endpoint role | Examples in the ClinicalTrials.gov record | Statistical issue |
|---|---|---|
| Co-primary | OS in all participants; OS in PD-L1-positive participants | Multiple confirmatory questions require the prespecified trial testing framework. |
| Secondary efficacy | ORR, DOR, PFS | These answer different questions and are not interchangeable with OS. |
| HRQoL | GHS/QOL, coughing, dyspnoea, chest pain | Multiple longitudinal outcomes and repeated measurements create additional interpretation considerations. |
19. Missing Data and Censoring
The primary OS endpoint has explicit censoring rules in the registry definition. Participants not reported as having died at analysis were censored at the date they were last known to be alive, while participants without postbaseline information were censored at randomization.
The ClinicalTrials.gov record does not specify an imputation method for missing quality-of-life scores. The MMRM specification is reported, but no additional missing-data sensitivity analysis or imputation strategy is provided in the ClinicalTrials.gov record.
20. Bayesian Methods, Non-Inferiority, and Crossover
Several common clinical-trial design topics are not supported by the registry-reported RATIONALE-303 data and are therefore not presented as trial features.
| Topic | What the ClinicalTrials.gov record supports |
|---|---|
| Bayesian methods | No Bayesian analysis is reported. |
| Non-inferiority margin | No non-inferiority margin is reported. The primary analyses are identified as superiority analyses. |
| Crossover | No crossover analysis or crossover treatment rule is reported in the ClinicalTrials.gov record. |
| Factorial design | The design is reported as parallel with two arms, not factorial. |
| Interim analysis | The ClinicalTrials.gov record does not describe an interim-analysis schedule or stopping boundary. |
This distinction is important for an independent statistical analysis: absence of a reported design feature is not evidence that the feature occurred. The analysis therefore stays within the ClinicalTrials.gov record.
21. Statistical Comparison of the Major Efficacy Measures
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| OS, all participants | Hazard ratio | 0.64 | 0.527–0.778 | <0.0001 |
| OS, PD-L1-positive | Hazard ratio | 0.53 | 0.407–0.702 | <0.0001 |
| ORR, all participants | Odds ratio | 3.86 | 2.336–6.393 | <0.0001 |
| ORR, PD-L1-positive | Odds ratio | 8.04 | 3.721–17.379 | <0.0001 |
| DOR, all responders | Hazard ratio | 0.31 | 0.176–0.536 | <0.0001 |
| DOR, PD-L1-positive responders | Hazard ratio | 0.16 | 0.066–0.370 | <0.0001 |
| PFS, all participants | Hazard ratio | 0.63 | 0.528–0.745 | <0.0001 |
| PFS, PD-L1-positive | Hazard ratio | 0.38 | 0.285–0.494 | <0.0001 |
The important statistical point is that these are different estimands. OS concerns death from any cause. PFS concerns the progression-free survival event. ORR concerns the odds of objective response. DOR concerns the time-to-event experience among responders. A single numerical ranking across these measures would obscure rather than clarify their different meanings.
22. What the Secondary Endpoints Add
ORR
Provides information about the probability of achieving an objective response and uses an odds ratio as the effect measure.
DOR
Adds information about durability among participants who achieved an objective response.
PFS
Captures a time-to-event outcome occurring before or at death and therefore complements OS.
HRQoL
Examines changes in patient-reported quality-of-life measures over the specified baseline-to-Cycle-6 interval.
Taken together, these endpoints illustrate why a clinical-trial statistical analysis should not collapse every result into one number. Different endpoints provide different evidence about different dimensions of the treatment comparison.
23. Limitations
- Different data cutoffs: the two co-primary OS analyses use different cutoff dates and follow-up frames.
- Different analysis populations: the all-participant and PD-L1-positive analyses do not use identical populations.
- Different effect measures: hazard ratios, odds ratios, and least-squares mean differences answer different statistical questions and should not be directly compared as though they were the same measure.
- Proportional-hazards assumption: the Cox model provides hazard ratios, but the ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.
- Censoring: survival analysis depends on the specified censoring rules and the assumptions underlying the treatment of censored observations.
- Responder conditioning: DOR analyses are restricted to participants who had an objective response and therefore address durability among responders rather than the overall randomized population.
- HRQoL population: the HRQoL Analysis Set requires treatment exposure and HRQoL assessment, so it is not identical to the ITT population used for primary efficacy analyses.
- Multiplicity: the ClinicalTrials.gov record identifies co-primary and multiple secondary endpoints but do not provide the complete multiplicity-adjustment framework.
- Missing-data information: the ClinicalTrials.gov record specifies the MMRM structure but do not provide a detailed imputation or sensitivity-analysis strategy for HRQoL.
- Safety comparison: serious adverse-event counts are provided by arm, but the ClinicalTrials.gov record does not provide a formal comparative statistical analysis.
- Unreported design details: the ClinicalTrials.gov record does not establish a non-inferiority margin, Bayesian method, crossover rule, or interim-analysis boundary.
24. Why This Trial Matters Statistically
RATIONALE-303 is a useful statistical teaching case because it combines randomized treatment comparison, stratified survival analysis, a co-primary endpoint structure, categorical response analysis, duration-of-response analysis, progression-free survival, and longitudinal patient-reported outcomes.
| Concept | How it appears in RATIONALE-303 |
|---|---|
| Randomization | 805 participants were enrolled in a randomized two-arm phase 3 parallel trial. |
| Intention-to-treat analysis | The all-participant primary OS analysis used all randomized patients. |
| Kaplan-Meier estimation | The registry specifies Kaplan-Meier calculation of median OS. |
| Log-rank testing | Primary OS and several secondary time-to-event endpoints used log-rank methods. |
| Hazard ratio | OS, PFS and DOR effects were reported using hazard ratios. |
| Stratified analysis | Histology, line of therapy and PD-L1 expression were used as stratification factors for specified analyses. |
| Cochran-Mantel-Haenszel test | ORR was analyzed using a stratified CMH approach. |
| Odds ratio | ORR treatment effects were reported as odds ratios. |
| Mixed-effects model | HRQoL changes were analyzed using a linear mixed-effect model for repeated measures. |
| Confidence intervals | All posted formal statistical analyses reported here include 95% confidence intervals. |
| P-values | Primary survival analyses use one-sided log-rank P-values; reported confidence intervals are two-sided. |
| Multiple endpoints | Two co-primary OS endpoints are accompanied by multiple secondary efficacy and HRQoL analyses. |
25. A Practical Reading Sequence for This Trial
A useful way to read the RATIONALE-303 results is to move from design to estimand, then from effect estimate to uncertainty, and finally to endpoint-specific limitations.
Start with randomization
Identify the randomized comparison, the two treatment arms, the phase 3 parallel design, and the ITT principle.
Define the OS questions
Separate OS in all participants from OS in PD-L1-positive participants because the analysis populations and data cutoffs differ.
Read the hazard ratio
Interpret HR 0.64 and HR 0.53 as relative hazard measures rather than absolute survival probabilities.
Read the confidence interval
Assess how much statistical uncertainty surrounds each point estimate rather than focusing only on the P-value.
Move to ORR, DOR, PFS and HRQoL
Interpret each endpoint according to its own estimand and statistical method.
26. Statistical Methods Explained: A Deeper View
Why does the ITT population matter for the primary OS analysis?
Using all randomized patients preserves the treatment assignment created by randomization. This means that the primary efficacy comparison remains tied to the randomized groups rather than being restricted to patients who completed treatment or had a particular post-randomization experience.
Why does the PD-L1-positive OS analysis have a different stratification scheme?
Once the analysis is restricted to participants meeting the PD-L1-positive definition, PD-L1 expression is no longer used as the same binary stratification factor across the analyzed population. The registry instead specifies stratification by histology and line of therapy for the PD-L1-positive analysis.
Why is an odds ratio of 8.04 not the same as an eightfold response rate?
An odds ratio compares odds, not probabilities. If a response probability is p, its odds are p/(1-p). The odds ratio compares those odds between treatment groups. Because probability and odds are nonlinear transformations of one another, an OR of 8.04 cannot be converted into an eightfold probability without knowing the underlying response probabilities.
Why does DOR use a responder population?
Duration of response is only defined for participants who experience an objective response. The analysis therefore asks a conditional question: among participants who responded, how does the time-to-loss-of-response experience compare between randomized groups? This differs from ORR, which asks how frequently response occurs in the broader analysis population.
What does an MMRM treatment-by-visit interaction do?
The treatment-by-visit interaction allows the estimated difference between treatment groups to vary across visits. That is important for longitudinal quality-of-life analysis because the treatment contrast at one visit does not necessarily have to equal the contrast at another visit.
Why can a statistically significant result still require clinical interpretation?
A P-value describes evidence against a null hypothesis under the stated statistical framework. It does not quantify the size of an effect, its practical importance, or its relevance to an individual patient. Those questions require examination of the effect estimate, confidence interval, endpoint scale, timing, and clinical context.
27. Overall Statistical Interpretation
The ClinicalTrials.gov record reports HR 0.64 for overall survival in all participants and HR 0.53 for overall survival in PD-L1-positive participants, with corresponding 95% confidence intervals of 0.527–0.778 and 0.407–0.702, respectively. Both analyses report P < 0.0001 from a one-sided log-rank test and are identified as superiority analyses.
The secondary analyses report HR 0.63 for PFS in all participants, HR 0.38 for PFS in PD-L1-positive participants, OR 3.86 for ORR in all participants, and OR 8.04 for ORR in PD-L1-positive participants. Duration-of-response analyses report HR 0.31 among all responders and HR 0.16 among PD-L1-positive responders.
The HRQoL analyses use an MMRM framework. The reported LS mean differences were 5.7 for global health status/QOL, -8.3 for coughing, -3.2 for dyspnoea, and -2.2 for chest pain, with the corresponding confidence intervals and P-values reported above.
The statistical story is therefore broader than a single hazard ratio. RATIONALE-303 combines two co-primary time-to-event analyses with categorical response measures, duration-of-response analyses, PFS, and longitudinal quality-of-life outcomes. Each result requires interpretation according to its endpoint definition, analysis population, effect measure, confidence interval, and testing framework.
28. Related Tutorials
Learn more about the methods used in this trial:
29. Related Calculators
30. Sources
- ClinicalTrials.gov: NCT03358875 — RATIONALE-303.
- PubMed: PMID 36184068.
- PubMed: PMID 37587845.
- PubMed: PMID 42529140.
Continue through the Clinical Biostats statistical library
Explore the survival, categorical-data, longitudinal, and clinical-trial methods that appear in RATIONALE-303.
31. Record Summary
RATIONALE-303 provides a useful example of a modern randomized phase 3 analysis in which the primary question is expressed through time-to-event methodology and supported by several complementary endpoint families. The ClinicalTrials.gov record reports two co-primary overall-survival analyses, both using stratified log-rank testing and hazard-ratio estimation, together with secondary analyses using the Cochran-Mantel-Haenszel test, log-rank methods, and a linear mixed-effects model for repeated measures.
The primary all-participant OS analysis reports HR 0.64 with a 95% CI of 0.527–0.778 and P < 0.0001. The PD-L1-positive OS analysis reports HR 0.53 with a 95% CI of 0.407–0.702 and P < 0.0001. Secondary analyses report effects for ORR, DOR, PFS, and HRQoL using their corresponding statistical estimands.
The most important statistical discipline is to keep those estimands separate. A hazard ratio is not an odds ratio; a response analysis is not a survival analysis; a responder-only DOR analysis is not an ITT analysis; and a Cycle 6 HRQoL comparison is not interchangeable with an OS endpoint. Reading the trial correctly therefore requires attention to the endpoint definition, analysis population, data cutoff, effect measure, confidence interval, P-value, and assumptions of the underlying model.