This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
SUNLIGHT was a randomized, parallel-group, open-label phase 3 study evaluating trifluridine/tipiracil with and without bevacizumab in patients with refractory metastatic colorectal cancer. The registry reports 492 enrolled participants, two treatment arms, time-to-event primary endpoints, and a superiority framework using a stratified log-rank test.
| Feature | SUNLIGHT |
|---|---|
| Trial name | SUNLIGHT |
| Phase | Phase 3 |
| Population | Refractory metastatic colorectal cancer |
| Design | Randomized, parallel |
| Masking | None |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 492 |
| Primary endpoint type | Time-to-event |
| Primary hypothesis | Superiority |
| Primary analysis method | Stratified log-rank test |
| Lead sponsor | Taiho Oncology, Inc. |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT04737187 |
2. Clinical Question
The central question was whether adding bevacizumab to trifluridine/tipiracil improved time-to-event outcomes compared with trifluridine/tipiracil alone in patients with refractory metastatic colorectal cancer.
Population
Patients with refractory metastatic colorectal cancer.
Intervention
Trifluridine/tipiracil with bevacizumab.
Comparator
Trifluridine/tipiracil without bevacizumab.
Primary question
Does adding bevacizumab improve the randomized treatment comparison for overall survival and related registered time-to-event endpoints?
3. Trial Design
Trifluridine/Tipiracil + Bevacizumab
- Trifluridine/tipiracil
- Bevacizumab
- Randomized comparison against trifluridine/tipiracil alone
Trifluridine/Tipiracil
- Trifluridine/tipiracil
- No bevacizumab
- Randomized comparator arm
4. Trial Timeline
Study start
The registered study start date was November 25, 2020.
Primary completion
The registered primary completion date was July 19, 2022.
Registry status
The trial is listed as completed, with results posted on ClinicalTrials.gov.
5. Primary Endpoints
The registry identifies four primary endpoints. All four are time-to-event endpoints related to overall survival, with the overall survival analysis providing the formal statistical analysis reported in the ClinicalTrials.gov record.
| Endpoint | Registered time frame | Endpoint type | Formal analysis posted |
|---|---|---|---|
| Overall Survival (OS) | From date of randomization to the death due to any cause or cut-off date, whichever comes first (maximum duration: up to 20 months) | Time-to-event | Yes |
| Survival Probability at 6 Months | From date of randomization until 6 months post treatment | Time-to-event | Yes |
| Survival Probability at 12 Months | From date of randomization until 12 months post treatment | Time-to-event | Yes |
| Survival Probability at 18 Months | From date of randomization until 18 months post treatment | Time-to-event | Yes |
Overall survival definition
The registry defines overall survival as the observed time elapsed between the date of randomization and the date of death due to any cause. The primary estimand was defined to assess the effect of randomized treatments on survival duration in all participants regardless of whether or not intercurrent events had occurred, using a treatment-policy strategy.
6. Secondary Endpoint
The statistical analyses posted on ClinicalTrials.gov include one secondary time-to-event endpoint: progression-free survival.
| Endpoint | Registered time frame | Analysis population | Method | Effect measure |
|---|---|---|---|---|
| Progression Free Survival (PFS) | From randomization to the date of radiological tumour progression or death due to any cause or data cut-off date whichever comes first | FAS | Stratified log-rank test | Hazard ratio |
7. Primary Result: Overall Survival
The registry reports a formal primary analysis of overall survival in the full analysis set (FAS). The comparison was trifluridine/tipiracil plus bevacizumab versus trifluridine/tipiracil alone.
Hazard ratio for death
95% CI: 0.49–0.77 · P < 0.001
Stratified log-rank test; superiority hypothesis; two-sided 95% confidence interval.
| Primary endpoint | Analysis population | Comparison | Method | HR | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Overall Survival (OS) | FAS | Trifluridine/Tipiracil + Bevacizumab vs Trifluridine/Tipiracil | Stratified log-rank test; stratified Cox proportional-hazards model | 0.61 | 0.49–0.77 | < 0.001 |
The estimated hazard ratio of 0.61 means that, under the fitted time-to-event model, the estimated instantaneous hazard of death was 0.61 times that of the trifluridine/tipiracil comparator group. Equivalently, 1 − 0.61 = 0.39, so the estimate corresponds to an approximately 39% lower estimated hazard of death for the combination relative to the comparator.
The HR does not mean that 39% of patients avoided death, that survival time increased by 39%, or that every individual patient experienced a 39% reduction in risk. It is a relative time-to-event measure derived from the statistical model.
The 95% confidence interval of 0.49–0.77 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not a range containing the treatment effect for 95% of individual patients.
The P < 0.001 result addresses the strength of evidence against the relevant null hypothesis under the prespecified testing framework. A p-value does not measure the magnitude of the treatment effect and should not be interpreted as the probability that the treatment has no effect.
Because the effect is expressed as a hazard ratio from a Cox model, interpretation also depends on the model's proportional-hazards framework. A single HR summarizes the relative event rate over the analyzed follow-up; it does not describe the complete survival curve by itself.
8. Secondary Result: Progression-Free Survival
The registry reports a formal secondary analysis of progression-free survival in the FAS population. The endpoint was defined from randomization to radiological tumour progression, death due to any cause, or the data cut-off date, whichever came first.
Hazard ratio for progression or death
95% CI: 0.36–0.54 · P < 0.001
Stratified log-rank test; two-sided 95% confidence interval.
| Secondary endpoint | Analysis population | Comparison | Method | HR | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Progression Free Survival (PFS) | FAS | Trifluridine/Tipiracil + Bevacizumab vs Trifluridine/Tipiracil | Stratified log-rank test; stratified Cox proportional-hazards model | 0.44 | 0.36–0.54 | < 0.001 |
The estimated hazard ratio of 0.44 corresponds to an estimated instantaneous hazard of progression or death that was 0.44 times the comparator hazard under the fitted model. Equivalently, the point estimate corresponds to an approximately 56% lower estimated hazard of progression or death.
This does not mean that 56% of patients were progression-free, nor does it mean that each patient's individual probability of progression was reduced by exactly 56%. The hazard ratio is a relative model-based time-to-event measure.
The 95% confidence interval of 0.36–0.54 indicates the statistical precision of the estimated relative effect. The interval is substantially below 1, while still expressing uncertainty about the exact magnitude of the hazard ratio.
The P < 0.001 result provides evidence against the relevant null hypothesis under the stated testing framework. It should not be confused with a measure of effect size: the magnitude is conveyed by the HR and its confidence interval.
The registry analysis notes that a hierarchical testing method was used for the key secondary endpoint. Testing of the key secondary endpoint followed significance of the primary outcome measure, with statistical significance assessed at the 0.05 level. This hierarchy matters because it helps control the interpretation of multiple confirmatory tests.
9. Comparing the Two Time-to-Event Effects
| Endpoint | HR | 95% CI | P-value | Estimated hazard reduction |
|---|---|---|---|---|
| Overall Survival | 0.61 | 0.49–0.77 | < 0.001 | Approximately 39% |
| Progression Free Survival | 0.44 | 0.36–0.54 | < 0.001 | Approximately 56% |
The two estimates answer different clinical questions. Overall survival measures time to death from any cause, whereas progression-free survival measures time to radiological tumour progression or death. A treatment can therefore produce different relative effects for these endpoints without the difference itself constituting a statistical contradiction.
10. Statistical Methodology
Kaplan-Meier estimation
The registry describes Kaplan-Meier analysis for overall survival. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not experience the event during the observed follow-up. Such participants can contribute information until their censoring time rather than being treated as if the event had occurred at that point.
Here di represents the number of events at time ti, while ni represents the number of participants at risk immediately before that event time.
Stratified log-rank test
The registered and reported primary method was the stratified log-rank test. A log-rank test compares the observed and expected pattern of events between randomized groups over follow-up. Stratification extends that comparison by accounting for prespecified strata when constructing the test statistic.
The ClinicalTrials.gov record specifically identify stratified analysis as an analysis concept. The exact stratification variables are not included in the ClinicalTrials.gov record, so this page does not infer or add them.
Stratified Cox proportional-hazards model
The formal OS analysis notes that the hazard ratio and 95% confidence interval were estimated with a stratified Cox proportional-hazards model. This provides a model-based estimate of the relative event hazard while retaining the stratified analysis framework.
An HR below 1 favors the numerator treatment group in a comparison expressed as trifluridine/tipiracil plus bevacizumab versus trifluridine/tipiracil. The HR is not an absolute risk difference and is not itself a probability.
Full analysis set
Both reported statistical analyses were performed on the FAS population. The ClinicalTrials.gov record does not provide a separate definition of the FAS, so this page does not impose an additional definition beyond identifying it as the analysis population reported by the registry.
11. Statistical Methods Explained
Why use a stratified log-rank test?
Time-to-event outcomes contain information not only about whether an event occurred but also about when it occurred. The log-rank test uses this event-time information across follow-up. Stratification allows the comparison to account for the trial's prespecified analysis strata rather than treating the entire study population as a single unstructured group.
What does an OS hazard ratio of 0.61 mean?
An HR of 0.61 indicates an estimated instantaneous hazard of death that is 61% of the comparator hazard under the fitted Cox model. The corresponding relative reduction in estimated hazard is 39%. It does not mean that 39% of participants survived or that survival time increased by 39%.
What does the 95% confidence interval of 0.49–0.77 tell us?
The interval describes uncertainty around the estimated OS hazard ratio. The point estimate is 0.61, while the interval extends from 0.49 to 0.77 under the specified 95% confidence level and analysis framework. Precision is therefore better communicated by considering the interval together with the point estimate rather than reporting the HR alone.
Why is the PFS hazard ratio different from the OS hazard ratio?
PFS and OS use different event definitions. PFS counts radiological tumour progression or death as an event, whereas OS counts death from any cause. Because the event processes differ, their estimated hazard ratios need not be the same.
Why does the p-value not measure treatment effect size?
A p-value summarizes the compatibility of the observed data with a specified null hypothesis under the statistical testing framework. It is affected by both the magnitude of an effect and the amount of information in the data. The hazard ratio describes relative effect magnitude, while its confidence interval describes statistical precision.
Why does the hierarchical test matter for PFS?
The registry-reported PFS analysis notes that a hierarchical testing method was used to control type I error and handle the key secondary endpoint. The primary outcome measure was tested first, and testing of the key secondary outcome measure was then performed sequentially after significance of the primary outcome measure. This means the PFS p-value should be interpreted in the context of the prespecified testing sequence rather than as an isolated test selected after looking at the results.
What assumption should be considered when interpreting a Cox hazard ratio?
The Cox proportional-hazards interpretation relies on a proportional-hazards framework. A single HR is most naturally interpreted as a relative hazard measure over the analyzed follow-up. If relative hazards change materially over time, one HR may not fully describe the evolving treatment difference. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic, so no such diagnostic is inferred here.
12. Multiplicity and Hierarchical Testing
Multiplicity is an important feature of the SUNLIGHT statistical analysis because the registry identifies multiple primary endpoints and a key secondary endpoint. Testing several hypotheses without an appropriate strategy can increase the probability of at least one false-positive conclusion.
| Endpoint / analysis | Role in the ClinicalTrials.gov record | Multiplicity information |
|---|---|---|
| Overall Survival | Primary endpoint | Formal primary analysis reported |
| Survival Probability at 6 Months | Primary endpoint | Results posted in the registry; no separate formal estimate is included in the registry-reported statistical-analyses record |
| Survival Probability at 12 Months | Primary endpoint | Results posted in the registry; no separate formal estimate is included in the registry-reported statistical-analyses record |
| Survival Probability at 18 Months | Primary endpoint | Results posted in the registry; no separate formal estimate is included in the registry-reported statistical-analyses record |
| Progression Free Survival | Secondary endpoint | Hierarchical testing used; sequential testing followed significance of the primary outcome measure |
13. Survival Probability Endpoints
The registry lists survival probability at 6, 12, and 18 months as primary endpoints. These are time-specific summaries of the overall-survival process rather than alternative definitions of the underlying death event.
| Registered endpoint | Time frame | Statistical interpretation |
|---|---|---|
| Survival Probability at 6 Months | From date of randomization until 6 months post treatment | Time-specific survival probability derived from the overall-survival process |
| Survival Probability at 12 Months | From date of randomization until 12 months post treatment | Time-specific survival probability derived from the overall-survival process |
| Survival Probability at 18 Months | From date of randomization until 18 months post treatment | Time-specific survival probability derived from the overall-survival process |
For a time-specific survival probability, the Kaplan-Meier estimator provides a natural framework because it accounts for censoring before the specified time point. The registry-reported statistical-analyses data do not provide separate numerical estimates, confidence intervals, or p-values for these three endpoints, so none are added here.
14. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. These counts are presented as affected participants over participants at risk.
| Safety measure | Trifluridine/Tipiracil + Bevacizumab | Trifluridine/Tipiracil |
|---|---|---|
| Serious adverse events | 66/246 | 79/246 |
The ClinicalTrials.gov record supports reporting the number of affected participants and the corresponding at-risk denominator. They do not provide a formal hypothesis test, confidence interval, severity breakdown, or exposure-adjusted incidence rate for serious adverse events. Those quantities are therefore not inferred.
15. Interpreting the Analysis Population and Censoring
Both formal analyses in the registry-reported statistical-analyses record were performed on the FAS population. This is important because the population used for an efficacy analysis determines which randomized participants contribute to the treatment comparison.
Time-to-event analysis also differs fundamentally from a simple comparison of event proportions. Participants who have not experienced the event by the end of their observable follow-up can be censored, allowing their partial follow-up information to contribute to the Kaplan-Meier and Cox analyses.
Censoring
A participant without an observed event at the relevant follow-up endpoint does not automatically count as an event. Their observed follow-up contributes information up to the censoring time.
Analysis population
The reported OS and PFS analyses were performed on the FAS population, as stated in the registry analysis records.
16. Why the Confidence Intervals Matter
The OS estimate is HR 0.61 with a 95% CI of 0.49–0.77. The interval communicates uncertainty around the point estimate and should be considered whenever the magnitude of the treatment effect is described.
The PFS estimate is HR 0.44 with a 95% CI of 0.36–0.54. The interval gives a statistical range around the estimated relative hazard and is more informative than the point estimate alone.
Confidence intervals and p-values serve different purposes. The confidence interval focuses attention on the estimated effect and its precision, while the p-value addresses the compatibility of the data with a specified null hypothesis. Neither quantity describes the distribution of benefit across individual patients.
17. What the P-Values Do — and Do Not — Tell Us
| Endpoint | P-value | What it addresses | What it does not provide |
|---|---|---|---|
| Overall Survival | < 0.001 | Evidence against the relevant null hypothesis under the stated primary testing framework | It is not the size of the treatment effect or the probability that the treatment works |
| Progression Free Survival | < 0.001 | Evidence under the hierarchical testing framework reported for the key secondary endpoint | It is not a measure of how much longer PFS lasted or the proportion of patients benefiting |
A very small p-value can coexist with an imprecisely estimated effect if the sample information and model produce such a result; conversely, a clinically meaningful point estimate can be accompanied by a wide confidence interval. The HR, confidence interval, and testing framework therefore need to be read together.
18. Design Features That Are Not Supported by the Supplied Data
Several common clinical-trial statistical topics are not described in the registry-reported SUNLIGHT data. They should not be reconstructed from assumptions about phase 3 oncology trials.
| Topic | What can be stated from the ClinicalTrials.gov record |
|---|---|
| Non-inferiority margin | Not applicable to the registry-reported superiority hypothesis description; no non-inferiority margin is reported. |
| Crossover | No crossover information is reported. |
| Factorial design | The design model is reported as parallel; no factorial design is reported. |
| Bayesian methods | No Bayesian method is reported. |
| Interim analysis | No interim-analysis procedure is reported. |
| Missing-data imputation | No imputation method is posted on ClinicalTrials.gov for the reported time-to-event analyses. |
| Stratification variables | Stratified analysis is reported, but the ClinicalTrials.gov record does not identify the individual stratification variables. |
19. Statistical Interpretation of the Primary Finding
The primary OS analysis produced an HR of 0.61. In the direction of the comparison specified in the registry, this corresponds to an estimated 39% lower hazard of death for trifluridine/tipiracil plus bevacizumab relative to trifluridine/tipiracil alone.
The corresponding 95% CI was 0.49–0.77. The confidence interval provides information about the statistical precision of the estimated treatment effect and should be reported alongside the HR rather than omitted.
The reported p-value was < 0.001. This is evidence against the relevant null hypothesis under the trial's stated superiority framework, but the p-value does not quantify the clinical magnitude of the treatment effect.
The PFS analysis produced an HR of 0.44 with a 95% CI of 0.36–0.54 and P < 0.001. The registry notes that this key secondary endpoint was handled using hierarchical testing after significance of the primary outcome measure.
20. Limitations
- Registry-level reporting: the analysis presented here is constrained to the ClinicalTrials.gov record and does not reconstruct information that was not provided.
- Incomplete endpoint-specific numerical reporting: four primary endpoints are registered, but the registry-reported formal statistical-analysis record provides a complete estimate, confidence interval, and p-value for only overall survival.
- Time-to-event assumptions: Cox-model hazard ratios are model-based and should be interpreted within the proportional-hazards framework.
- Unknown stratification variables: stratified analysis is reported, but the ClinicalTrials.gov record does not identify the variables used for stratification.
- Limited safety detail: serious adverse events are reported by arm, but the ClinicalTrials.gov record does not include a formal comparative safety analysis.
- No reconstructed survival curves: the registry-reported summary data are insufficient to recreate patient-level Kaplan-Meier curves without making unsupported assumptions.
- No unreported design details: the ClinicalTrials.gov record does not establish an interim-analysis procedure, missing-data strategy, crossover policy, or Bayesian analysis.
- Hierarchical testing: interpretation of the PFS p-value depends on the reported testing sequence rather than treating it as an isolated hypothesis test.
21. Why This Trial Matters Statistically
SUNLIGHT is a useful teaching example because the ClinicalTrials.gov record connects a randomized superiority comparison with several core survival-analysis concepts: Kaplan-Meier estimation, stratified log-rank testing, Cox proportional-hazards modeling, hazard ratios, confidence intervals, and hierarchical control of type I error.
| Concept | How it appears in SUNLIGHT |
|---|---|
| Randomization | The study is registered as randomized. |
| Parallel design | The design model is registered as parallel. |
| Time-to-event endpoints | The four registered primary endpoints are time-to-event outcomes. |
| Kaplan-Meier estimation | Kaplan-Meier analysis is identified in the OS endpoint definition. |
| Stratified log-rank test | Used for the reported OS and PFS comparisons. |
| Hazard ratio | Used to quantify relative treatment effects for OS and PFS. |
| Confidence interval | 95% two-sided intervals are reported for both formal HR estimates. |
| Superiority testing | The registered hypothesis type is superiority. |
| Hierarchical testing | Used for the key secondary PFS endpoint according to the registry-reported analysis notes. |
| Full analysis set | Both registry-reported formal statistical analyses were performed on the FAS population. |
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: SUNLIGHT, NCT04737187.
- PubMed: PMID 37133585.
- PubMed: PMID 38953855.
Continue with the underlying statistical methods
Explore the survival-analysis, confidence-interval, multiplicity, randomization, and stratified-analysis concepts that support interpretation of randomized clinical trials.
25. Record Summary
SUNLIGHT provides a clear example of a randomized phase 3 time-to-event analysis in which the principal treatment comparison is expressed through hazard ratios and evaluated with stratified log-rank testing. The ClinicalTrials.gov record reports an OS HR of 0.61 (95% CI 0.49–0.77; P < 0.001) and a PFS HR of 0.44 (95% CI 0.36–0.54; P < 0.001). Both analyses were performed in the FAS population. The PFS analysis additionally used hierarchical testing as described in the registry analysis notes.
The statistical interpretation is strongest when the numerical effect estimates, confidence intervals, p-values, endpoint definitions, analysis population, and testing hierarchy are considered together. The hazard ratios quantify relative time-to-event effects; the confidence intervals communicate statistical precision; and the p-values provide evidence relative to the specified null hypotheses. None of these measures alone describes the complete clinical experience of individual participants.