← Clinical Trials
Refractory Metastatic Colorectal Cancer Phase 3 Time-to-Event Analysis NCT04737187

SUNLIGHT: Complete Statistical Analysis of Trifluridine/Tipiracil With Bevacizumab in Refractory Metastatic Colorectal Cancer

An independent statistical review of the randomized phase 3 SUNLIGHT trial evaluating trifluridine/tipiracil with bevacizumab versus trifluridine/tipiracil alone in patients with refractory metastatic colorectal cancer.

Trial status: Completed  ·  Enrollment: 492  ·  Primary completion: July 19, 2022
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

SUNLIGHT was a randomized, parallel-group, open-label phase 3 study evaluating trifluridine/tipiracil with and without bevacizumab in patients with refractory metastatic colorectal cancer. The registry reports 492 enrolled participants, two treatment arms, time-to-event primary endpoints, and a superiority framework using a stratified log-rank test.

492
Enrolled
Phase 3
2
Treatment arms
Parallel design
0.61
OS HR
95% CI 0.49–0.77
0.44
PFS HR
95% CI 0.36–0.54
FeatureSUNLIGHT
Trial nameSUNLIGHT
PhasePhase 3
PopulationRefractory metastatic colorectal cancer
DesignRandomized, parallel
MaskingNone
AllocationRandomized
Primary purposeTreatment
Enrollment492
Primary endpoint typeTime-to-event
Primary hypothesisSuperiority
Primary analysis methodStratified log-rank test
Lead sponsorTaiho Oncology, Inc.
Sponsor typeIndustry
ClinicalTrials.govNCT04737187

2. Clinical Question

The central question was whether adding bevacizumab to trifluridine/tipiracil improved time-to-event outcomes compared with trifluridine/tipiracil alone in patients with refractory metastatic colorectal cancer.

Population

Patients with refractory metastatic colorectal cancer.

Intervention

Trifluridine/tipiracil with bevacizumab.

Comparator

Trifluridine/tipiracil without bevacizumab.

Primary question

Does adding bevacizumab improve the randomized treatment comparison for overall survival and related registered time-to-event endpoints?

3. Trial Design

01
Randomize492 enrolled
02
Two armsCombination vs trifluridine/tipiracil
03
FollowTime-to-event outcomes
04
AnalyzeStratified log-rank test
05
EstimateHazard ratios and 95% CIs
ARM A · 246 at risk for reported serious-AE analysis

Trifluridine/Tipiracil + Bevacizumab

  • Trifluridine/tipiracil
  • Bevacizumab
  • Randomized comparison against trifluridine/tipiracil alone
ARM B · 246 at risk for reported serious-AE analysis

Trifluridine/Tipiracil

  • Trifluridine/tipiracil
  • No bevacizumab
  • Randomized comparator arm
Design interpretation: The registry describes the study as randomized, parallel, and unmasked. Because the allocation was randomized, the principal efficacy comparison is anchored to the assigned treatment groups rather than to treatment actually received after randomization.

4. Trial Timeline

November 25, 2020

Study start

The registered study start date was November 25, 2020.

July 19, 2022

Primary completion

The registered primary completion date was July 19, 2022.

Completed

Registry status

The trial is listed as completed, with results posted on ClinicalTrials.gov.

5. Primary Endpoints

The registry identifies four primary endpoints. All four are time-to-event endpoints related to overall survival, with the overall survival analysis providing the formal statistical analysis reported in the ClinicalTrials.gov record.

EndpointRegistered time frameEndpoint typeFormal analysis posted
Overall Survival (OS) From date of randomization to the death due to any cause or cut-off date, whichever comes first (maximum duration: up to 20 months) Time-to-event Yes
Survival Probability at 6 Months From date of randomization until 6 months post treatment Time-to-event Yes
Survival Probability at 12 Months From date of randomization until 12 months post treatment Time-to-event Yes
Survival Probability at 18 Months From date of randomization until 18 months post treatment Time-to-event Yes

Overall survival definition

The registry defines overall survival as the observed time elapsed between the date of randomization and the date of death due to any cause. The primary estimand was defined to assess the effect of randomized treatments on survival duration in all participants regardless of whether or not intercurrent events had occurred, using a treatment-policy strategy.

Registry wording: the registry-reported endpoint record states that the analysis was performed using Kaplan-Meier methods. The formal statistical analysis additionally reports a stratified log-rank test and a stratified Cox proportional-hazards model for the hazard ratio.

6. Secondary Endpoint

The statistical analyses posted on ClinicalTrials.gov include one secondary time-to-event endpoint: progression-free survival.

EndpointRegistered time frameAnalysis populationMethodEffect measure
Progression Free Survival (PFS) From randomization to the date of radiological tumour progression or death due to any cause or data cut-off date whichever comes first FAS Stratified log-rank test Hazard ratio

7. Primary Result: Overall Survival

The registry reports a formal primary analysis of overall survival in the full analysis set (FAS). The comparison was trifluridine/tipiracil plus bevacizumab versus trifluridine/tipiracil alone.

Hazard ratio for death

0.61

95% CI: 0.49–0.77   ·   P < 0.001

Stratified log-rank test; superiority hypothesis; two-sided 95% confidence interval.

Primary endpointAnalysis populationComparisonMethodHR95% CIP-value
Overall Survival (OS) FAS Trifluridine/Tipiracil + Bevacizumab vs Trifluridine/Tipiracil Stratified log-rank test; stratified Cox proportional-hazards model 0.61 0.49–0.77 < 0.001
Clinical Biostats interpretation

The estimated hazard ratio of 0.61 means that, under the fitted time-to-event model, the estimated instantaneous hazard of death was 0.61 times that of the trifluridine/tipiracil comparator group. Equivalently, 1 − 0.61 = 0.39, so the estimate corresponds to an approximately 39% lower estimated hazard of death for the combination relative to the comparator.

The HR does not mean that 39% of patients avoided death, that survival time increased by 39%, or that every individual patient experienced a 39% reduction in risk. It is a relative time-to-event measure derived from the statistical model.

The 95% confidence interval of 0.49–0.77 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not a range containing the treatment effect for 95% of individual patients.

The P < 0.001 result addresses the strength of evidence against the relevant null hypothesis under the prespecified testing framework. A p-value does not measure the magnitude of the treatment effect and should not be interpreted as the probability that the treatment has no effect.

Because the effect is expressed as a hazard ratio from a Cox model, interpretation also depends on the model's proportional-hazards framework. A single HR summarizes the relative event rate over the analyzed follow-up; it does not describe the complete survival curve by itself.

8. Secondary Result: Progression-Free Survival

The registry reports a formal secondary analysis of progression-free survival in the FAS population. The endpoint was defined from randomization to radiological tumour progression, death due to any cause, or the data cut-off date, whichever came first.

Hazard ratio for progression or death

0.44

95% CI: 0.36–0.54   ·   P < 0.001

Stratified log-rank test; two-sided 95% confidence interval.

Secondary endpointAnalysis populationComparisonMethodHR95% CIP-value
Progression Free Survival (PFS) FAS Trifluridine/Tipiracil + Bevacizumab vs Trifluridine/Tipiracil Stratified log-rank test; stratified Cox proportional-hazards model 0.44 0.36–0.54 < 0.001
Clinical Biostats interpretation

The estimated hazard ratio of 0.44 corresponds to an estimated instantaneous hazard of progression or death that was 0.44 times the comparator hazard under the fitted model. Equivalently, the point estimate corresponds to an approximately 56% lower estimated hazard of progression or death.

This does not mean that 56% of patients were progression-free, nor does it mean that each patient's individual probability of progression was reduced by exactly 56%. The hazard ratio is a relative model-based time-to-event measure.

The 95% confidence interval of 0.36–0.54 indicates the statistical precision of the estimated relative effect. The interval is substantially below 1, while still expressing uncertainty about the exact magnitude of the hazard ratio.

The P < 0.001 result provides evidence against the relevant null hypothesis under the stated testing framework. It should not be confused with a measure of effect size: the magnitude is conveyed by the HR and its confidence interval.

The registry analysis notes that a hierarchical testing method was used for the key secondary endpoint. Testing of the key secondary endpoint followed significance of the primary outcome measure, with statistical significance assessed at the 0.05 level. This hierarchy matters because it helps control the interpretation of multiple confirmatory tests.

9. Comparing the Two Time-to-Event Effects

EndpointHR95% CIP-valueEstimated hazard reduction
Overall Survival 0.61 0.49–0.77 < 0.001 Approximately 39%
Progression Free Survival 0.44 0.36–0.54 < 0.001 Approximately 56%

The two estimates answer different clinical questions. Overall survival measures time to death from any cause, whereas progression-free survival measures time to radiological tumour progression or death. A treatment can therefore produce different relative effects for these endpoints without the difference itself constituting a statistical contradiction.

Do not compare hazard ratios as if they were simple percentages of patients benefiting. The OS HR of 0.61 and PFS HR of 0.44 are estimates from separate time-to-event analyses with different event definitions. Their numerical difference should not be interpreted as a direct measure of how much more important one endpoint is than the other.

10. Statistical Methodology

Kaplan-Meier estimation

The registry describes Kaplan-Meier analysis for overall survival. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not experience the event during the observed follow-up. Such participants can contribute information until their censoring time rather than being treated as if the event had occurred at that point.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here di represents the number of events at time ti, while ni represents the number of participants at risk immediately before that event time.

Stratified log-rank test

The registered and reported primary method was the stratified log-rank test. A log-rank test compares the observed and expected pattern of events between randomized groups over follow-up. Stratification extends that comparison by accounting for prespecified strata when constructing the test statistic.

The ClinicalTrials.gov record specifically identify stratified analysis as an analysis concept. The exact stratification variables are not included in the ClinicalTrials.gov record, so this page does not infer or add them.

Stratified Cox proportional-hazards model

The formal OS analysis notes that the hazard ratio and 95% confidence interval were estimated with a stratified Cox proportional-hazards model. This provides a model-based estimate of the relative event hazard while retaining the stratified analysis framework.

Hazard ratio interpretation
HR < 1  →  lower estimated instantaneous event hazard in the treatment group

An HR below 1 favors the numerator treatment group in a comparison expressed as trifluridine/tipiracil plus bevacizumab versus trifluridine/tipiracil. The HR is not an absolute risk difference and is not itself a probability.

Full analysis set

Both reported statistical analyses were performed on the FAS population. The ClinicalTrials.gov record does not provide a separate definition of the FAS, so this page does not impose an additional definition beyond identifying it as the analysis population reported by the registry.

11. Statistical Methods Explained

Why use a stratified log-rank test?

Time-to-event outcomes contain information not only about whether an event occurred but also about when it occurred. The log-rank test uses this event-time information across follow-up. Stratification allows the comparison to account for the trial's prespecified analysis strata rather than treating the entire study population as a single unstructured group.

What does an OS hazard ratio of 0.61 mean?

An HR of 0.61 indicates an estimated instantaneous hazard of death that is 61% of the comparator hazard under the fitted Cox model. The corresponding relative reduction in estimated hazard is 39%. It does not mean that 39% of participants survived or that survival time increased by 39%.

What does the 95% confidence interval of 0.49–0.77 tell us?

The interval describes uncertainty around the estimated OS hazard ratio. The point estimate is 0.61, while the interval extends from 0.49 to 0.77 under the specified 95% confidence level and analysis framework. Precision is therefore better communicated by considering the interval together with the point estimate rather than reporting the HR alone.

Why is the PFS hazard ratio different from the OS hazard ratio?

PFS and OS use different event definitions. PFS counts radiological tumour progression or death as an event, whereas OS counts death from any cause. Because the event processes differ, their estimated hazard ratios need not be the same.

Why does the p-value not measure treatment effect size?

A p-value summarizes the compatibility of the observed data with a specified null hypothesis under the statistical testing framework. It is affected by both the magnitude of an effect and the amount of information in the data. The hazard ratio describes relative effect magnitude, while its confidence interval describes statistical precision.

Why does the hierarchical test matter for PFS?

The registry-reported PFS analysis notes that a hierarchical testing method was used to control type I error and handle the key secondary endpoint. The primary outcome measure was tested first, and testing of the key secondary outcome measure was then performed sequentially after significance of the primary outcome measure. This means the PFS p-value should be interpreted in the context of the prespecified testing sequence rather than as an isolated test selected after looking at the results.

What assumption should be considered when interpreting a Cox hazard ratio?

The Cox proportional-hazards interpretation relies on a proportional-hazards framework. A single HR is most naturally interpreted as a relative hazard measure over the analyzed follow-up. If relative hazards change materially over time, one HR may not fully describe the evolving treatment difference. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic, so no such diagnostic is inferred here.

12. Multiplicity and Hierarchical Testing

Multiplicity is an important feature of the SUNLIGHT statistical analysis because the registry identifies multiple primary endpoints and a key secondary endpoint. Testing several hypotheses without an appropriate strategy can increase the probability of at least one false-positive conclusion.

Endpoint / analysisRole in the ClinicalTrials.gov recordMultiplicity information
Overall Survival Primary endpoint Formal primary analysis reported
Survival Probability at 6 Months Primary endpoint Results posted in the registry; no separate formal estimate is included in the registry-reported statistical-analyses record
Survival Probability at 12 Months Primary endpoint Results posted in the registry; no separate formal estimate is included in the registry-reported statistical-analyses record
Survival Probability at 18 Months Primary endpoint Results posted in the registry; no separate formal estimate is included in the registry-reported statistical-analyses record
Progression Free Survival Secondary endpoint Hierarchical testing used; sequential testing followed significance of the primary outcome measure
Important distinction: the ClinicalTrials.gov record identifies four registered primary endpoints but provide a formal effect estimate and p-value for only one primary endpoint, Overall Survival. This page therefore does not manufacture separate HRs, confidence intervals, or p-values for the three survival-probability endpoints.

13. Survival Probability Endpoints

The registry lists survival probability at 6, 12, and 18 months as primary endpoints. These are time-specific summaries of the overall-survival process rather than alternative definitions of the underlying death event.

Registered endpointTime frameStatistical interpretation
Survival Probability at 6 Months From date of randomization until 6 months post treatment Time-specific survival probability derived from the overall-survival process
Survival Probability at 12 Months From date of randomization until 12 months post treatment Time-specific survival probability derived from the overall-survival process
Survival Probability at 18 Months From date of randomization until 18 months post treatment Time-specific survival probability derived from the overall-survival process

For a time-specific survival probability, the Kaplan-Meier estimator provides a natural framework because it accounts for censoring before the specified time point. The registry-reported statistical-analyses data do not provide separate numerical estimates, confidence intervals, or p-values for these three endpoints, so none are added here.

14. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These counts are presented as affected participants over participants at risk.

Safety measureTrifluridine/Tipiracil + BevacizumabTrifluridine/Tipiracil
Serious adverse events 66/246 79/246
Reported serious adverse events by arm
+ Bevacizumab
66/246
Trifluridine/Tipiracil
79/246

The ClinicalTrials.gov record supports reporting the number of affected participants and the corresponding at-risk denominator. They do not provide a formal hypothesis test, confidence interval, severity breakdown, or exposure-adjusted incidence rate for serious adverse events. Those quantities are therefore not inferred.

15. Interpreting the Analysis Population and Censoring

Both formal analyses in the registry-reported statistical-analyses record were performed on the FAS population. This is important because the population used for an efficacy analysis determines which randomized participants contribute to the treatment comparison.

Time-to-event analysis also differs fundamentally from a simple comparison of event proportions. Participants who have not experienced the event by the end of their observable follow-up can be censored, allowing their partial follow-up information to contribute to the Kaplan-Meier and Cox analyses.

Censoring

A participant without an observed event at the relevant follow-up endpoint does not automatically count as an event. Their observed follow-up contributes information up to the censoring time.

Analysis population

The reported OS and PFS analyses were performed on the FAS population, as stated in the registry analysis records.

16. Why the Confidence Intervals Matter

Overall survival

The OS estimate is HR 0.61 with a 95% CI of 0.49–0.77. The interval communicates uncertainty around the point estimate and should be considered whenever the magnitude of the treatment effect is described.

Progression-free survival

The PFS estimate is HR 0.44 with a 95% CI of 0.36–0.54. The interval gives a statistical range around the estimated relative hazard and is more informative than the point estimate alone.

Confidence intervals and p-values serve different purposes. The confidence interval focuses attention on the estimated effect and its precision, while the p-value addresses the compatibility of the data with a specified null hypothesis. Neither quantity describes the distribution of benefit across individual patients.

17. What the P-Values Do — and Do Not — Tell Us

EndpointP-valueWhat it addressesWhat it does not provide
Overall Survival < 0.001 Evidence against the relevant null hypothesis under the stated primary testing framework It is not the size of the treatment effect or the probability that the treatment works
Progression Free Survival < 0.001 Evidence under the hierarchical testing framework reported for the key secondary endpoint It is not a measure of how much longer PFS lasted or the proportion of patients benefiting

A very small p-value can coexist with an imprecisely estimated effect if the sample information and model produce such a result; conversely, a clinically meaningful point estimate can be accompanied by a wide confidence interval. The HR, confidence interval, and testing framework therefore need to be read together.

18. Design Features That Are Not Supported by the Supplied Data

Several common clinical-trial statistical topics are not described in the registry-reported SUNLIGHT data. They should not be reconstructed from assumptions about phase 3 oncology trials.

TopicWhat can be stated from the ClinicalTrials.gov record
Non-inferiority marginNot applicable to the registry-reported superiority hypothesis description; no non-inferiority margin is reported.
CrossoverNo crossover information is reported.
Factorial designThe design model is reported as parallel; no factorial design is reported.
Bayesian methodsNo Bayesian method is reported.
Interim analysisNo interim-analysis procedure is reported.
Missing-data imputationNo imputation method is posted on ClinicalTrials.gov for the reported time-to-event analyses.
Stratification variablesStratified analysis is reported, but the ClinicalTrials.gov record does not identify the individual stratification variables.
Why this matters: absence of a detail in the ClinicalTrials.gov record is not evidence that the procedure was absent from the actual protocol or statistical analysis plan. This page simply avoids attributing an unreported method to the trial.

19. Statistical Interpretation of the Primary Finding

Relative effect

The primary OS analysis produced an HR of 0.61. In the direction of the comparison specified in the registry, this corresponds to an estimated 39% lower hazard of death for trifluridine/tipiracil plus bevacizumab relative to trifluridine/tipiracil alone.

Precision

The corresponding 95% CI was 0.49–0.77. The confidence interval provides information about the statistical precision of the estimated treatment effect and should be reported alongside the HR rather than omitted.

Statistical evidence

The reported p-value was < 0.001. This is evidence against the relevant null hypothesis under the trial's stated superiority framework, but the p-value does not quantify the clinical magnitude of the treatment effect.

Secondary endpoint

The PFS analysis produced an HR of 0.44 with a 95% CI of 0.36–0.54 and P < 0.001. The registry notes that this key secondary endpoint was handled using hierarchical testing after significance of the primary outcome measure.

20. Limitations

21. Why This Trial Matters Statistically

SUNLIGHT is a useful teaching example because the ClinicalTrials.gov record connects a randomized superiority comparison with several core survival-analysis concepts: Kaplan-Meier estimation, stratified log-rank testing, Cox proportional-hazards modeling, hazard ratios, confidence intervals, and hierarchical control of type I error.

ConceptHow it appears in SUNLIGHT
RandomizationThe study is registered as randomized.
Parallel designThe design model is registered as parallel.
Time-to-event endpointsThe four registered primary endpoints are time-to-event outcomes.
Kaplan-Meier estimationKaplan-Meier analysis is identified in the OS endpoint definition.
Stratified log-rank testUsed for the reported OS and PFS comparisons.
Hazard ratioUsed to quantify relative treatment effects for OS and PFS.
Confidence interval95% two-sided intervals are reported for both formal HR estimates.
Superiority testingThe registered hypothesis type is superiority.
Hierarchical testingUsed for the key secondary PFS endpoint according to the registry-reported analysis notes.
Full analysis setBoth registry-reported formal statistical analyses were performed on the FAS population.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Statistical Calculators

24. Sources

Continue with the underlying statistical methods

Explore the survival-analysis, confidence-interval, multiplicity, randomization, and stratified-analysis concepts that support interpretation of randomized clinical trials.

25. Record Summary

SUNLIGHT provides a clear example of a randomized phase 3 time-to-event analysis in which the principal treatment comparison is expressed through hazard ratios and evaluated with stratified log-rank testing. The ClinicalTrials.gov record reports an OS HR of 0.61 (95% CI 0.49–0.77; P < 0.001) and a PFS HR of 0.44 (95% CI 0.36–0.54; P < 0.001). Both analyses were performed in the FAS population. The PFS analysis additionally used hierarchical testing as described in the registry analysis notes.

The statistical interpretation is strongest when the numerical effect estimates, confidence intervals, p-values, endpoint definitions, analysis population, and testing hierarchy are considered together. The hazard ratios quantify relative time-to-event effects; the confidence intervals communicate statistical precision; and the p-values provide evidence relative to the specified null hypotheses. None of these measures alone describes the complete clinical experience of individual participants.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. For SUNLIGHT, the ClinicalTrials.gov record supports a detailed analysis of randomized time-to-event methods while leaving unreported design details, endpoint-specific estimates, and safety statistics unaltered rather than inferred.