← Clinical Trials
Non-Small Cell Lung Cancer Phase 3 Completed NCT03358875

RATIONALE-303: Complete Statistical Analysis of Tislelizumab in Non-Small Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 RATIONALE-303 trial comparing tislelizumab with docetaxel as treatment in the second- or third-line setting in participants with non-small cell lung cancer.

Trial start: 30 November 2017  ·  Primary completion: 15 July 2021  ·  Enrollment: 805
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for RATIONALE-303 and are presented at the reported precision.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

RATIONALE-303 was a randomized, open-label, parallel phase 3 trial comparing tislelizumab with docetaxel in participants with non-small cell lung cancer in the second- or third-line setting. The registry reports two co-primary overall-survival endpoints: one in all participants and one in participants whose tumors were PD-L1 positive.

805
Enrollment
Randomized trial
2
Arms
Tislelizumab vs docetaxel
0.64
OS HR
All participants
0.53
OS HR
PD-L1-positive participants
FeatureRATIONALE-303
PhasePhase 3
ConditionNon-small cell lung cancer
DesignRandomized, parallel
AllocationRandomized
MaskingNone
Primary purposeTreatment
Enrollment805
InterventionsTislelizumab and docetaxel
Primary endpointsOverall survival in all participants; overall survival in PD-L1-positive participants
Trial statusCompleted
Lead sponsorBeiGene
ClinicalTrials.govNCT03358875

2. Clinical Question

The central statistical question was whether treatment with tislelizumab differed from docetaxel with respect to overall survival in participants with non-small cell lung cancer treated in the second- or third-line setting.

Population

Participants with non-small cell lung cancer enrolled in the second- or third-line treatment setting.

Intervention

Tislelizumab.

Comparator

Docetaxel.

Primary question

Does the randomized comparison show a difference in overall survival between tislelizumab and docetaxel, both in all participants and in the prespecified PD-L1-positive population?

3. Trial Design

01
Randomize805 participants
02
Two armsTislelizumab vs docetaxel
03
FollowSurvival and response outcomes
04
AssessEfficacy, HRQoL and safety
05
AnalyzeTime-to-event and other endpoints
Allocation
Randomized allocation was used to compare the two treatment groups.
Design model
Parallel design with two treatment arms.
Masking
The trial was unmasked; the registry lists masking as none.
Primary purpose
Treatment.
ARM A

Tislelizumab

  • Drug intervention.
  • Compared with docetaxel.
  • Primary efficacy comparison based on overall survival.
ARM B

Docetaxel

  • Drug intervention.
  • Comparator to tislelizumab.
  • Used as the reference group for reported effect measures.
What the ClinicalTrials.gov record does not establish: the ClinicalTrials.gov record does not provide treatment-arm enrollment counts, dosing schedules, treatment duration, crossover information, interim-analysis rules, missing-data imputation rules, or a non-inferiority margin. Those details are therefore not inferred here.

4. Endpoints

The registry identifies two co-primary endpoints, both based on overall survival. The first applies to all participants; the second applies to participants whose tumors were PD-L1 positive.

EndpointRegistry definitionTime frame
Overall Survival (OS) in All Participants (Co-primary Endpoint) OS was defined as the time from randomization to death from any cause. Median OS was calculated using the Kaplan-Meier method. Data for participants who were not reported as having died at the time of analysis were censored at the date they were last known to be alive. Data for participants who did not have postbaseline information were censored at the date of randomization. From randomization to the data cutoff date of 10 August 2020; up to 32.4 months
Overall Survival (OS) in PD-L1-Positive Participants (Co-primary Endpoint) OS was defined as the time from randomization to death from any cause. Median OS was calculated using the Kaplan-Meier method. Data for participants who were not reported as having died at the time of analysis were censored at the date they were last known to be alive. Data for participants who did not have postbaseline information were censored at the date of randomization. From randomization up to the final efficacy analysis data cut-off date of 15 July 2021; up to 43 months

PD-L1-positive population

The registry defines the PD-L1-Positive Analysis Set as all randomized patients whose tumors were PD-L1 positive, with PD-L1 positivity defined as ≥ 25% of tumor cells with PD-L1 membrane staining.

Two different data cutoffs matter. The all-participant co-primary endpoint uses a data cutoff of 10 August 2020, with a stated follow-up frame of up to 32.4 months. The PD-L1-positive co-primary endpoint uses the final efficacy analysis data cutoff of 15 July 2021, with a stated follow-up frame of up to 43 months. These are not interchangeable analyses and should not be treated as though they arose from one identical follow-up period.

5. Statistical Methodology

Primary time-to-event analysis

The registry reports a 1-sided log-rank test for each co-primary overall-survival endpoint. The treatment effect is reported as a hazard ratio, with a two-sided 95% confidence interval.

For the all-participant OS analysis, the comparison was stratified by histology, line of therapy, and PD-L1 expression. The registry specifies histology as squamous versus non-squamous, line of therapy as second versus third, and PD-L1 expression as ≥25% versus <25% tumor cells.

For the PD-L1-positive OS analysis, the comparison was stratified by histology and line of therapy: squamous versus non-squamous and second versus third line.

Primary survival framework
Randomization → time-to-death → censoring where applicable → stratified log-rank comparison → hazard-ratio estimation

The registry states that a stratified Cox proportional hazards model with Efron's method of tie handling was used to determine the hazard ratio and its 95% confidence interval for the all-participant OS analysis. The PD-L1-positive OS analysis is described as stratified by histology and line of therapy.

Secondary endpoint methodology

The secondary analyses span three statistical families: categorical response outcomes, time-to-event outcomes, and longitudinal quality-of-life outcomes.

Endpoint familyRegistry methodEffect measure
Objective response rateCochran-Mantel-Haenszel testOdds ratio
Duration of responseLog-rank testHazard ratio
Progression-free survivalLog-rank testHazard ratio
Health-related quality of lifeMixed-effects model / MMRMLeast-squares mean difference

Mixed-effects model for repeated measures

The quality-of-life analyses used a linear mixed-effect model for repeated measures. The registry description specifies baseline score, stratification factors, treatment arm, visit, and treatment-arm-by-visit interaction as fixed effects, with visit as a repeated measure.

This is important because the quality-of-life endpoint is not simply a comparison of two independent means at Cycle 6. The model incorporates the longitudinal structure of repeated assessments and explicitly models treatment-by-visit differences.

6. Results: Overall Survival in All Participants

The first co-primary endpoint was overall survival in all participants. The analysis used the Intent-to-Treat Analysis Set, defined in the registry as including all randomized patients.

Hazard ratio for overall survival

0.64

95% CI: 0.527–0.778   ·   1-sided log-rank test P < 0.0001

Data cutoff: 10 August 2020; up to 32.4 months

ElementReported result
Analysis populationIntent-to-Treat Analysis Set; all randomized patients
ComparisonTislelizumab vs Docetaxel
Method1-sided Log Rank Test
Effect measureHazard Ratio
Estimate0.64
95% CI0.527–0.778
P-value<0.0001
Hypothesis typeSuperiority
StratificationHistology, line of therapy, and PD-L1 expression
Clinical Biostats interpretation

An HR of 0.64 means that, under the reported time-to-event model, the estimated hazard of death in the tislelizumab group was 0.64 times the estimated hazard in the docetaxel group over the analyzed follow-up. Put as a simple relative interpretation, this corresponds to an estimated 36% lower hazard of death because 1 − 0.64 = 0.36.

The HR does not mean that 36% of participants avoided death, that survival increased by 36%, or that every individual experienced the same reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute survival probability.

The 95% CI of 0.527–0.778 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of treatment effects that must occur in individual patients.

The P < 0.0001 result addresses evidence against the null hypothesis under the reported test. It does not measure the size or clinical importance of the effect. Effect size is conveyed by the HR and its confidence interval, while absolute survival measures would provide a complementary clinical perspective.

The analysis was stratified, and the registry also specifies a stratified Cox model for HR estimation. Interpretation therefore depends on the assumptions of the survival-analysis framework, including the adequacy of the model and the handling of censoring. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.

7. Results: Overall Survival in PD-L1-Positive Participants

The second co-primary endpoint evaluated overall survival among randomized participants whose tumors were PD-L1 positive, defined as ≥25% of tumor cells with PD-L1 membrane staining.

Hazard ratio for overall survival in PD-L1-positive participants

0.53

95% CI: 0.407–0.702   ·   1-sided log-rank test P < 0.0001

Final efficacy analysis data cutoff: 15 July 2021; up to 43 months

ElementReported result
Analysis populationPD-L1-Positive Analysis Set; all randomized patients whose tumors were PD-L1 positive
PD-L1 definition≥25% of tumor cells with PD-L1 membrane staining
ComparisonTislelizumab vs Docetaxel
Method1-sided Log Rank Test
Effect measureHazard Ratio
Estimate0.53
95% CI0.407–0.702
P-value<0.0001
Hypothesis typeSuperiority
StratificationHistology and line of therapy
Clinical Biostats interpretation

An HR of 0.53 means that the estimated hazard of death was 0.53 times the estimated hazard in the docetaxel group for the PD-L1-positive analysis population. As a simple derived interpretation, this corresponds to an estimated 47% lower hazard because 1 − 0.53 = 0.47.

The estimate should not be read as a 47% increase in survival probability or as evidence that 47% of patients benefited. It is a relative time-to-event measure describing the fitted comparison between treatment groups.

The 95% CI of 0.407–0.702 gives the statistical uncertainty around the estimated HR. Its width also illustrates why an estimate should be considered together with its interval rather than treated as a single exact number.

The reported P < 0.0001 is evidence against the null under the stated one-sided log-rank test. A P-value is not an effect-size metric and does not tell us how clinically meaningful an observed difference is.

This analysis also uses a different data cutoff and analysis population from the all-participant OS result. The two HRs should therefore be compared descriptively, not treated as if they were estimates from precisely the same population and follow-up window.

8. Secondary Efficacy Results

The registry contains formal statistical analyses for objective response rate, duration of response, and progression-free survival. These results complement the two co-primary overall-survival analyses but answer different statistical questions.

Objective Response Rate in All Participants

Odds ratio for objective response

3.86

95% CI: 2.336–6.393   ·   P < 0.0001

Final efficacy analysis data cutoff: 15 July 2021; up to 43 months

The analysis used the Intent-to-Treat Analysis Set. The registry reports a Cochran-Mantel-Haenszel analysis stratified by histology, line of therapy, and PD-L1 expression.

Interpretation

An odds ratio of 3.86 means that the estimated odds of objective response were 3.86 times as high with tislelizumab as with docetaxel under the reported stratified analysis. This is an odds ratio, not a risk ratio or a difference in response percentages.

The 95% CI of 2.336–6.393 quantifies uncertainty around the estimated odds ratio. The P < 0.0001 value assesses evidence under the reported superiority test; it does not measure the magnitude of the response difference.

Objective Response Rate in PD-L1-Positive Participants

Odds ratio for objective response

8.04

95% CI: 3.721–17.379   ·   P < 0.0001

Final efficacy analysis data cutoff: 15 July 2021; up to 43 months

The PD-L1-positive analysis used the PD-L1-Positive Analysis Set and was stratified by histology and line of therapy.

Interpretation

An odds ratio of 8.04 indicates that the estimated odds of objective response were 8.04 times as high with tislelizumab as with docetaxel in the reported PD-L1-positive analysis. The estimate is not itself a probability and should not be interpreted as meaning that eight times as many participants responded.

The 95% CI of 3.721–17.379 indicates considerable uncertainty around the size of the odds ratio even though the entire reported interval is above 1. The P-value of <0.0001 addresses statistical evidence, not effect magnitude.

Duration of Response in All Responders

Hazard ratio for duration of response

0.31

95% CI: 0.176–0.536   ·   P < 0.0001

This analysis included participants in the Intent-to-Treat Analysis Set who had an objective response. The registry used a log-rank test and a stratified Cox proportional hazards model with Efron's method of tie handling for the HR and 95% CI. Stratification included histology, line of therapy, and PD-L1 expression.

Interpretation

An HR of 0.31 means that the estimated hazard of the response-ending event was 0.31 times that in the comparator group under the reported duration-of-response analysis. As a simple derived interpretation, this corresponds to an estimated 69% lower hazard because 1 − 0.31 = 0.69.

The analysis is conditional on having an objective response, so it addresses durability among responders rather than the probability of achieving a response in the first place. That distinction is important when interpreting it alongside ORR.

Duration of Response in PD-L1-Positive Responders

Hazard ratio for duration of response

0.16

95% CI: 0.066–0.370   ·   P < 0.0001

The PD-L1-positive duration-of-response analysis included participants in the PD-L1-Positive Analysis Set who had an objective response. The comparison was stratified by histology and line of therapy.

Interpretation

An HR of 0.16 corresponds to an estimated 84% lower hazard of the response-ending event because 1 − 0.16 = 0.84. It does not mean that 84% of responders maintained their response or that 84% were cured.

The 95% CI of 0.066–0.370 indicates uncertainty around the estimated relative hazard. Because this analysis is restricted to responders, it should not be substituted for the overall treatment comparison.

Progression-Free Survival in All Participants

Hazard ratio for progression-free survival

0.63

95% CI: 0.528–0.745   ·   P < 0.0001

The PFS analysis used the Intent-to-Treat Analysis Set. The registry reports a log-rank test and a stratified Cox proportional hazards model, with stratification by histology, line of therapy, and PD-L1 expression.

Interpretation

An HR of 0.63 means that the estimated hazard of the PFS event was 0.63 times that in the docetaxel group. As a simple derived interpretation, this corresponds to an estimated 37% lower hazard because 1 − 0.63 = 0.37.

The PFS hazard ratio does not tell us the absolute proportion of participants who remained progression-free at a particular time. It is a relative time-to-event estimate and should be interpreted together with its confidence interval and the definition of the PFS event.

Progression-Free Survival in PD-L1-Positive Participants

Hazard ratio for progression-free survival

0.38

95% CI: 0.285–0.494   ·   P < 0.0001

The PD-L1-positive PFS analysis used the PD-L1 Positive Analysis Set and was stratified by histology and line of therapy. The registry reports a stratified Cox proportional hazards model for the HR and 95% CI.

Interpretation

An HR of 0.38 corresponds to an estimated 62% lower hazard of the PFS event because 1 − 0.38 = 0.62. This is a relative hazard interpretation, not a statement that 62% of participants avoided progression or death.

The 95% CI of 0.285–0.494 quantifies uncertainty around the estimate, while P < 0.0001 addresses the statistical test rather than the clinical magnitude of the effect.

9. Quality-of-Life Results

The registry reports three formal mixed-model analyses of change from baseline in EORTC quality-of-life measures at Cycle 6, with each cycle defined as 3 weeks. These endpoints were analyzed in the HRQoL Analysis Set.

EndpointEstimate95% CIP-value
EORTC QLQ-C30 Global Health Status / Quality of Life Score LS Mean Difference 5.7 2.38–9.07 0.0008
QLQ-LC13 Coughing Scale LS Mean Difference -8.3 -13.02 to -3.51 0.0007
QLQ-LC13 Dyspnoea Scale LS Mean Difference -3.2 -6.52 to 0.11 0.0579
QLQ-LC13 Chest Pain Scale LS Mean Difference -2.2 -6.05 to 1.56 0.2472

The ClinicalTrials.gov record describes these as mixed-model analyses. The linear mixed-effect model for repeated measures includes baseline score, stratification factors, treatment arm, visit, and treatment-arm-by-visit interaction as fixed effects, with visit as a repeated measure.

Global health status / QOL

The reported LS mean difference was 5.7, with a 95% CI of 2.38–9.07 and P = 0.0008.

Coughing

The reported LS mean difference was -8.3, with a 95% CI of -13.02 to -3.51 and P = 0.0007.

Dyspnoea

The reported LS mean difference was -3.2, with a 95% CI of -6.52 to 0.11 and P = 0.0579.

Chest pain

The reported LS mean difference was -2.2, with a 95% CI of -6.05 to 1.56 and P = 0.2472.

Interpretation caution: the sign of a least-squares mean difference depends on how the score and treatment contrast are coded. The registry provides the estimates and confidence intervals, but does not provide a separate clinical-importance threshold in the ClinicalTrials.gov record. Statistical significance therefore should not be equated automatically with clinical importance.

10. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by the corresponding at-risk population.

Treatment armSerious adverse events
Tislelizumab192/534
Docetaxel84/258
Serious adverse events by arm
Tislelizumab
192/534
Docetaxel
84/258

The bars above are a visual representation of the reported affected counts, not a percentage comparison. The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates, so no P-value, risk ratio, odds ratio, or confidence interval is inferred here.

Important denominator distinction: the serious-adverse-event denominators of 534 and 258 are not used here to reconstruct randomized-arm sample sizes. The ClinicalTrials.gov record states only that the overall trial enrollment was 805 and that the ITT set included all randomized patients.

11. Statistical Methods Explained

Why was a log-rank test used for overall survival?

Overall survival is a time-to-event endpoint. Some participants die before the analysis cutoff, while others remain alive and are therefore censored. A log-rank test is designed to compare survival experience between groups while incorporating the timing of events and censoring rather than reducing each participant to a simple yes/no outcome.

What does an HR of 0.64 mean?

It means that the estimated hazard of death in the tislelizumab group was 0.64 times that in the docetaxel group under the reported survival model. The simple derived expression, 1 − 0.64 = 0.36, gives a 36% relative reduction in estimated hazard. It does not represent a 36-percentage-point improvement in survival probability.

Why was the analysis stratified?

The registry reports stratification by factors including histology, line of therapy, and PD-L1 expression. Stratification allows the time-to-event comparison to account for these prespecified factors rather than treating all participants as if they came from one completely homogeneous risk set.

Why is the odds ratio for ORR different from the hazard ratio for OS?

ORR is a binary outcome: a participant either meets the response definition or does not. OS is a time-to-event outcome: both whether and when death occurs matter. An odds ratio therefore summarizes relative odds of response, while a hazard ratio summarizes relative instantaneous event rates over time. They cannot be interpreted as interchangeable effect measures.

Why is the quality-of-life analysis different from the survival analysis?

The quality-of-life endpoints are continuous longitudinal measures observed at baseline and Cycle 6. The registry therefore uses a mixed-effects model for repeated measures rather than a survival model. The specified MMRM includes baseline score, treatment, visit, stratification factors, and treatment-by-visit interaction, with visit treated as a repeated measure.

Why should the two co-primary endpoints not be treated as one ordinary endpoint?

The registry identifies OS in all participants and OS in PD-L1-positive participants as separate co-primary endpoints. They also use different analysis populations and different data cutoffs. The ClinicalTrials.gov record does not specify an alpha-allocation or detailed multiplicity procedure, so no additional error-control rule is inferred. The two reported analyses should therefore be described using the registry's co-primary designation rather than creating an unreported hierarchy.

12. Understanding the Hazard Ratio

Conceptual interpretation
HR = estimated hazard in tislelizumab group ÷ estimated hazard in docetaxel group

An HR below 1 indicates a lower estimated event hazard in the tislelizumab group under the fitted model. An HR of 1 would indicate no relative hazard difference. The hazard ratio is not an absolute risk difference and is not itself a survival probability.

OS HR 0.64

Estimated 36% lower hazard of death relative to the comparator under the reported model.

PD-L1-positive OS HR 0.53

Estimated 47% lower hazard of death in the PD-L1-positive analysis.

PFS HR 0.63

Estimated 37% lower hazard of the PFS event in all participants.

PD-L1-positive PFS HR 0.38

Estimated 62% lower hazard of the PFS event in the PD-L1-positive analysis.

These derived percentage statements are mathematical interpretations of the reported hazard ratios. They should not be confused with absolute percentage-point differences in survival or progression-free survival.

13. Confidence Intervals and P-values

AnalysisEstimate95% CIP-value
All-participant OSHR 0.640.527–0.778<0.0001
PD-L1-positive OSHR 0.530.407–0.702<0.0001
All-participant ORROR 3.862.336–6.393<0.0001
PD-L1-positive ORROR 8.043.721–17.379<0.0001
All-responder DORHR 0.310.176–0.536<0.0001
PD-L1-positive responder DORHR 0.160.066–0.370<0.0001
All-participant PFSHR 0.630.528–0.745<0.0001
PD-L1-positive PFSHR 0.380.285–0.494<0.0001
GHS/QOL changeLS Mean Difference 5.72.38–9.070.0008
Coughing changeLS Mean Difference -8.3-13.02 to -3.510.0007
Dyspnoea changeLS Mean Difference -3.2-6.52 to 0.110.0579
Chest pain changeLS Mean Difference -2.2-6.05 to 1.560.2472
How to read this table

A confidence interval provides information about the precision of an estimate. It should be read together with the point estimate rather than as a substitute for it. A P-value answers a different question: it quantifies evidence under the specified null hypothesis and testing framework. Neither measure tells us whether an effect is clinically important by itself.

14. Analysis Populations

PopulationDefinition / use
Intent-to-Treat Analysis Set All randomized patients. Used for the all-participant OS, ORR, DOR, and PFS analyses reported in the ClinicalTrials.gov record.
PD-L1-Positive Analysis Set All randomized patients whose tumors were PD-L1 positive, defined as ≥25% of tumor cells with PD-L1 membrane staining.
HRQoL Analysis Set All randomized participants who received ≥1 dose of study drug and completed ≥1 HRQoL assessment; the registry-reported description states that participants with available data at baseline and Cycle 6 are included.
Responder populations For DOR, participants in the relevant analysis population who had an objective response.

The use of different analysis populations is not a technical footnote. It changes the question being answered. ITT preserves the randomized comparison for efficacy, whereas the HRQoL analysis requires treatment exposure and quality-of-life assessments, and DOR necessarily conditions on having responded.

15. Stratified Analysis

Stratification appears repeatedly in the RATIONALE-303 statistical analyses. For the all-participant primary OS endpoint, the registry identifies three stratification factors:

Histology

Squamous versus non-squamous.

Line of therapy

Second versus third.

PD-L1 expression

≥25% versus <25% tumor cells for the all-participant OS analysis.

PD-L1-positive analyses

The PD-L1-positive OS, ORR, DOR and PFS analyses are stratified by histology and line of therapy.

Stratification is particularly relevant in a randomized trial because it connects the statistical analysis to factors used to structure the comparison. It does not mean that the treatment effect is separately estimated within every stratum and then simply averaged. The reported analysis uses a stratified comparison framework.

16. Kaplan-Meier Estimation and Censoring

The registry explicitly states that median OS was calculated using the Kaplan-Meier method. It also specifies how participants who had not died at the time of analysis were handled: they were censored at the date they were last known to be alive.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

At each observed event time, the estimated survival function is updated according to the number of events and the number of participants at risk immediately before that time.

The registry also states that participants without postbaseline information were censored at the date of randomization. This is an important detail because censoring rules determine how available follow-up contributes to the survival estimate.

No median survival is added here. Although the registry describes median OS as a calculated measure, the ClinicalTrials.gov record does not provide median OS estimates. They are therefore not reconstructed or imported from an external publication.

17. The One-Sided Test and Two-Sided Confidence Interval

Both primary analyses are reported with a 1-sided log-rank test and a two-sided 95% confidence interval. This distinction is worth preserving rather than silently converting the analysis into a generic two-sided test.

One-sided P-value

The reported log-rank P-value evaluates the prespecified directional superiority framework recorded in the registry.

Two-sided CI

The reported confidence interval is two-sided and gives an interval estimate around the hazard ratio.

The ClinicalTrials.gov record does not report the complete multiplicity or alpha-allocation procedure associated with the two co-primary endpoints. Consequently, this page does not infer a particular familywise error allocation or rewrite the reported P-values under an unreported procedure.

18. Multiplicity and Multiple Endpoints

RATIONALE-303 has two registered co-primary endpoints, both overall survival endpoints but in different analysis populations and at different data cutoffs. The ClinicalTrials.gov record also contain numerous secondary analyses covering response, duration of response, progression-free survival, and quality of life.

Endpoint roleExamples in the ClinicalTrials.gov recordStatistical issue
Co-primaryOS in all participants; OS in PD-L1-positive participantsMultiple confirmatory questions require the prespecified trial testing framework.
Secondary efficacyORR, DOR, PFSThese answer different questions and are not interchangeable with OS.
HRQoLGHS/QOL, coughing, dyspnoea, chest painMultiple longitudinal outcomes and repeated measurements create additional interpretation considerations.
Registry-data limitation: the ClinicalTrials.gov record identifies the endpoints, methods, and hypothesis types but do not provide a detailed multiplicity-adjustment procedure or alpha allocation. It would therefore be inappropriate to invent a hierarchy, adjusted significance threshold, or claim of familywise-error control beyond what the registry data explicitly support.

19. Missing Data and Censoring

The primary OS endpoint has explicit censoring rules in the registry definition. Participants not reported as having died at analysis were censored at the date they were last known to be alive, while participants without postbaseline information were censored at randomization.

The ClinicalTrials.gov record does not specify an imputation method for missing quality-of-life scores. The MMRM specification is reported, but no additional missing-data sensitivity analysis or imputation strategy is provided in the ClinicalTrials.gov record.

Why this matters: time-to-event censoring and missing longitudinal measurements are different statistical problems. Censoring determines how survival follow-up is incorporated, while missing HRQoL observations affect estimation in the repeated-measures analysis. They should not be described as the same type of missingness.

20. Bayesian Methods, Non-Inferiority, and Crossover

Several common clinical-trial design topics are not supported by the registry-reported RATIONALE-303 data and are therefore not presented as trial features.

TopicWhat the ClinicalTrials.gov record supports
Bayesian methodsNo Bayesian analysis is reported.
Non-inferiority marginNo non-inferiority margin is reported. The primary analyses are identified as superiority analyses.
CrossoverNo crossover analysis or crossover treatment rule is reported in the ClinicalTrials.gov record.
Factorial designThe design is reported as parallel with two arms, not factorial.
Interim analysisThe ClinicalTrials.gov record does not describe an interim-analysis schedule or stopping boundary.

This distinction is important for an independent statistical analysis: absence of a reported design feature is not evidence that the feature occurred. The analysis therefore stays within the ClinicalTrials.gov record.

21. Statistical Comparison of the Major Efficacy Measures

EndpointEffect measureEstimate95% CIP-value
OS, all participantsHazard ratio0.640.527–0.778<0.0001
OS, PD-L1-positiveHazard ratio0.530.407–0.702<0.0001
ORR, all participantsOdds ratio3.862.336–6.393<0.0001
ORR, PD-L1-positiveOdds ratio8.043.721–17.379<0.0001
DOR, all respondersHazard ratio0.310.176–0.536<0.0001
DOR, PD-L1-positive respondersHazard ratio0.160.066–0.370<0.0001
PFS, all participantsHazard ratio0.630.528–0.745<0.0001
PFS, PD-L1-positiveHazard ratio0.380.285–0.494<0.0001

The important statistical point is that these are different estimands. OS concerns death from any cause. PFS concerns the progression-free survival event. ORR concerns the odds of objective response. DOR concerns the time-to-event experience among responders. A single numerical ranking across these measures would obscure rather than clarify their different meanings.

22. What the Secondary Endpoints Add

ORR

Provides information about the probability of achieving an objective response and uses an odds ratio as the effect measure.

DOR

Adds information about durability among participants who achieved an objective response.

PFS

Captures a time-to-event outcome occurring before or at death and therefore complements OS.

HRQoL

Examines changes in patient-reported quality-of-life measures over the specified baseline-to-Cycle-6 interval.

Taken together, these endpoints illustrate why a clinical-trial statistical analysis should not collapse every result into one number. Different endpoints provide different evidence about different dimensions of the treatment comparison.

23. Limitations

24. Why This Trial Matters Statistically

RATIONALE-303 is a useful statistical teaching case because it combines randomized treatment comparison, stratified survival analysis, a co-primary endpoint structure, categorical response analysis, duration-of-response analysis, progression-free survival, and longitudinal patient-reported outcomes.

ConceptHow it appears in RATIONALE-303
Randomization805 participants were enrolled in a randomized two-arm phase 3 parallel trial.
Intention-to-treat analysisThe all-participant primary OS analysis used all randomized patients.
Kaplan-Meier estimationThe registry specifies Kaplan-Meier calculation of median OS.
Log-rank testingPrimary OS and several secondary time-to-event endpoints used log-rank methods.
Hazard ratioOS, PFS and DOR effects were reported using hazard ratios.
Stratified analysisHistology, line of therapy and PD-L1 expression were used as stratification factors for specified analyses.
Cochran-Mantel-Haenszel testORR was analyzed using a stratified CMH approach.
Odds ratioORR treatment effects were reported as odds ratios.
Mixed-effects modelHRQoL changes were analyzed using a linear mixed-effect model for repeated measures.
Confidence intervalsAll posted formal statistical analyses reported here include 95% confidence intervals.
P-valuesPrimary survival analyses use one-sided log-rank P-values; reported confidence intervals are two-sided.
Multiple endpointsTwo co-primary OS endpoints are accompanied by multiple secondary efficacy and HRQoL analyses.

25. A Practical Reading Sequence for This Trial

A useful way to read the RATIONALE-303 results is to move from design to estimand, then from effect estimate to uncertainty, and finally to endpoint-specific limitations.

Step 1 · Design

Start with randomization

Identify the randomized comparison, the two treatment arms, the phase 3 parallel design, and the ITT principle.

Step 2 · Primary estimands

Define the OS questions

Separate OS in all participants from OS in PD-L1-positive participants because the analysis populations and data cutoffs differ.

Step 3 · Effect size

Read the hazard ratio

Interpret HR 0.64 and HR 0.53 as relative hazard measures rather than absolute survival probabilities.

Step 4 · Precision

Read the confidence interval

Assess how much statistical uncertainty surrounds each point estimate rather than focusing only on the P-value.

Step 5 · Supporting endpoints

Move to ORR, DOR, PFS and HRQoL

Interpret each endpoint according to its own estimand and statistical method.

26. Statistical Methods Explained: A Deeper View

Why does the ITT population matter for the primary OS analysis?

Using all randomized patients preserves the treatment assignment created by randomization. This means that the primary efficacy comparison remains tied to the randomized groups rather than being restricted to patients who completed treatment or had a particular post-randomization experience.

Why does the PD-L1-positive OS analysis have a different stratification scheme?

Once the analysis is restricted to participants meeting the PD-L1-positive definition, PD-L1 expression is no longer used as the same binary stratification factor across the analyzed population. The registry instead specifies stratification by histology and line of therapy for the PD-L1-positive analysis.

Why is an odds ratio of 8.04 not the same as an eightfold response rate?

An odds ratio compares odds, not probabilities. If a response probability is p, its odds are p/(1-p). The odds ratio compares those odds between treatment groups. Because probability and odds are nonlinear transformations of one another, an OR of 8.04 cannot be converted into an eightfold probability without knowing the underlying response probabilities.

Why does DOR use a responder population?

Duration of response is only defined for participants who experience an objective response. The analysis therefore asks a conditional question: among participants who responded, how does the time-to-loss-of-response experience compare between randomized groups? This differs from ORR, which asks how frequently response occurs in the broader analysis population.

What does an MMRM treatment-by-visit interaction do?

The treatment-by-visit interaction allows the estimated difference between treatment groups to vary across visits. That is important for longitudinal quality-of-life analysis because the treatment contrast at one visit does not necessarily have to equal the contrast at another visit.

Why can a statistically significant result still require clinical interpretation?

A P-value describes evidence against a null hypothesis under the stated statistical framework. It does not quantify the size of an effect, its practical importance, or its relevance to an individual patient. Those questions require examination of the effect estimate, confidence interval, endpoint scale, timing, and clinical context.

27. Overall Statistical Interpretation

Primary OS findings

The ClinicalTrials.gov record reports HR 0.64 for overall survival in all participants and HR 0.53 for overall survival in PD-L1-positive participants, with corresponding 95% confidence intervals of 0.527–0.778 and 0.407–0.702, respectively. Both analyses report P < 0.0001 from a one-sided log-rank test and are identified as superiority analyses.

Supporting efficacy findings

The secondary analyses report HR 0.63 for PFS in all participants, HR 0.38 for PFS in PD-L1-positive participants, OR 3.86 for ORR in all participants, and OR 8.04 for ORR in PD-L1-positive participants. Duration-of-response analyses report HR 0.31 among all responders and HR 0.16 among PD-L1-positive responders.

Patient-reported outcomes

The HRQoL analyses use an MMRM framework. The reported LS mean differences were 5.7 for global health status/QOL, -8.3 for coughing, -3.2 for dyspnoea, and -2.2 for chest pain, with the corresponding confidence intervals and P-values reported above.

The statistical story is therefore broader than a single hazard ratio. RATIONALE-303 combines two co-primary time-to-event analyses with categorical response measures, duration-of-response analyses, PFS, and longitudinal quality-of-life outcomes. Each result requires interpretation according to its endpoint definition, analysis population, effect measure, confidence interval, and testing framework.

28. Related Tutorials

Learn more about the methods used in this trial:

29. Related Calculators

30. Sources

Continue through the Clinical Biostats statistical library

Explore the survival, categorical-data, longitudinal, and clinical-trial methods that appear in RATIONALE-303.

31. Record Summary

RATIONALE-303 provides a useful example of a modern randomized phase 3 analysis in which the primary question is expressed through time-to-event methodology and supported by several complementary endpoint families. The ClinicalTrials.gov record reports two co-primary overall-survival analyses, both using stratified log-rank testing and hazard-ratio estimation, together with secondary analyses using the Cochran-Mantel-Haenszel test, log-rank methods, and a linear mixed-effects model for repeated measures.

The primary all-participant OS analysis reports HR 0.64 with a 95% CI of 0.527–0.778 and P < 0.0001. The PD-L1-positive OS analysis reports HR 0.53 with a 95% CI of 0.407–0.702 and P < 0.0001. Secondary analyses report effects for ORR, DOR, PFS, and HRQoL using their corresponding statistical estimands.

The most important statistical discipline is to keep those estimands separate. A hazard ratio is not an odds ratio; a response analysis is not a survival analysis; a responder-only DOR analysis is not an ITT analysis; and a Cycle 6 HRQoL comparison is not interchangeable with an OS endpoint. Reading the trial correctly therefore requires attention to the endpoint definition, analysis population, data cutoff, effect measure, confidence interval, P-value, and assumptions of the underlying model.

Clinical Biostats methodology: This page uses only the trial information posted on ClinicalTrials.gov for RATIONALE-303. Where the ClinicalTrials.gov record does not report a numerical result, design parameter, multiplicity rule, or additional analysis detail, that information has not been reconstructed or imported from another source.