This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information contained in the ClinicalTrials.gov record.
1. Trial at a Glance
BREAKWATER is a randomized, open-label, parallel phase 3 study in oncology. The registry describes 841 participants, 7 arms, 4 registered primary endpoints, and posted statistical analyses using the log-rank test and Cochran-Mantel-Haenszel test.
| Feature | BREAKWATER |
|---|---|
| Phase | Phase 3 |
| Population | People with previously untreated metastatic colorectal cancer |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Allocation | Randomized |
| Enrollment | 841 |
| Arms | 7 |
| Primary endpoint types | Binary; Time-to-event |
| Hypothesis type | Superiority |
| Results posted | Yes |
| Lead sponsor | Pfizer |
| Trial status | Active, not recruiting |
2. Clinical Question
The statistical question in BREAKWATER is framed around randomized comparisons of treatment strategies in people with previously untreated metastatic colorectal cancer. The registry reports primary comparisons involving encorafenib plus cetuximab with mFOLFOX6 versus standard-of-care chemotherapy, and encorafenib plus cetuximab with FOLFIRI versus FOLFIRI with or without bevacizumab.
Population
People with previously untreated metastatic colorectal cancer, as described in the trial's brief title.
Interventions
The registered interventions include encorafenib, cetuximab, oxaliplatin, irinotecan, leucovorin, 5-FU, capecitabine, and bevacizumab.
Comparator
The primary statistical comparisons reported include standard-of-care chemotherapy in Phase 3 and FOLFIRI with or without bevacizumab in Cohort 3.
Primary question
Under the registry's superiority framework, how do the prespecified treatment groups compare for progression-free survival and objective response rate?
3. Trial Design
The presence of 7 arms is important statistically. A multi-arm trial can contain several clinically distinct comparisons, but the interpretation of any one estimate remains tied to the exact groups and analysis population specified for that endpoint. BREAKWATER's posted primary analyses are not all comparisons across all 7 arms; they are specific prespecified comparisons.
Reported Phase 3 and Cohort 3 comparisons
Encorafenib + cetuximab
The registry identifies Arm A as EC and reports 46 serious adverse events among 153 participants at risk.
Encorafenib + cetuximab + mFOLFOX6
Arm B is the treatment group in the primary Phase 3 comparisons with Arm C. The serious adverse event count is 107 among 232 participants at risk.
Standard-of-care chemotherapy
Arm C is the control arm for the primary Phase 3 comparison with Arm B. The serious adverse event count is 89 among 229 participants at risk.
FOLFIRI comparison
Arm D is EC + FOLFIRI, while Arm E is FOLFIRI with or without bevacizumab. Serious adverse events were reported as 28/71 and 25/68, respectively.
4. Endpoints
| Registered primary endpoint | Time frame | Type | Analysis reported? |
|---|---|---|---|
| SLI: Number of Participants With Dose Limiting Toxicity (DLTs) | Cycle 1 (28 days) | Binary | Endpoint results posted; no formal statistical comparison reported in the statistical analyses dataset |
| Phase 3: Progression Free Survival (PFS) as Assessed by BICR for Arm B vs Arm C - FAS | From date of randomization to earliest documentation of PD by BICR or death or censoring date, whichever occurred first (maximum up to 37.25 months) | Time-to-event | Yes |
| Phase 3: Objective Response Rate (ORR) as Assessed by BICR for Arm B vs Arm C - FAS ORR Subset | From date of randomization until documented PD by BICR, or start of subsequent anticancer therapy or death, whichever occurred first (maximum up to 24.71 months) | Binary | Yes |
| Cohort 3: ORR as Assessed by BICR for Arm D vs Arm E - FAS | From date of randomization until documented PD by BICR, or start of subsequent anticancer therapy or death, whichever occurred first (maximum up to 24.71 months) | Binary | Yes |
The PFS definition specifies time from randomization to the earliest documented disease progression according to RECIST version 1.1 or death due to any cause, as assessed by blinded independent central review. The ORR definitions specify confirmed complete response or partial response according to RECIST version 1.1 as assessed by BICR.
The registry's primary endpoint list therefore combines two different statistical structures: a binary toxicity endpoint and response endpoints, plus a censored time-to-event endpoint. That distinction determines which statistical methods are appropriate.
5. Statistical Methodology
Full Analysis Set
The Phase 3 Full Analysis Set included all participants randomized in the Phase 3 portion of the study. The Phase 3 ORR analysis used the first 110 participants randomized in each of Arm B and Arm C. The Cohort 3 FAS included all participants randomized in the Cohort 3 portion of the study.
The analysis population matters because an effect estimate is meaningful only in the population to which it was actually applied. The denominator and eligibility for a response analysis can differ from those of a time-to-event analysis.
Log-rank test
The registry reports the log-rank test for the primary Phase 3 PFS comparison and the secondary Phase 3 OS comparison. The log-rank framework compares survival experience over time while accounting for right-censored observations.
The test is based on observed and expected event patterns over follow-up. It is not a test of whether the numerical hazard ratio is exactly a particular value.
Stratified Cox proportional-hazards model
For the reported Phase 3 PFS and OS analyses, the hazard ratio was based on a stratified Cox proportional-hazards model. The registry analysis notes identify the log-rank p-value as a 1-sided p-value and the confidence interval as two-sided.
A hazard ratio is a relative, model-based measure of event rates over the analyzed follow-up. It is not a percentage of patients who benefit and is not an absolute difference in survival probability.
Cochran-Mantel-Haenszel test
The Phase 3 ORR comparison and Cohort 3 ORR comparison used the Cochran-Mantel-Haenszel test. The registry also states that the odds ratio was estimated using the stratified Cochran-Mantel-Haenszel method.
This approach is particularly useful when a binary endpoint is compared across randomized groups while accounting for prespecified stratification. Instead of simply pooling all response counts and ignoring strata, the CMH framework combines evidence across strata.
Odds ratio
An odds ratio above 1 indicates higher estimated odds of the binary event in the first-listed group relative to the comparator. It is not identical to a risk ratio or a difference in response probabilities.
6. Results
The ClinicalTrials.gov record contains four formal statistical analyses: three primary endpoint analyses and one secondary endpoint analysis. Three of the analyses include an effect estimate and confidence interval.
6.1 Phase 3 Progression-Free Survival: Arm B vs Arm C
Hazard ratio for progression or death
95% CI: 0.407–0.677 · P < 0.0001
Arm B: EC + mFOLFOX6 vs Arm C: Standard of Care Chemotherapy
| Feature | Reported result |
|---|---|
| Endpoint | Progression Free Survival (PFS) as assessed by BICR for Arm B vs Arm C - FAS |
| Endpoint type | Time-to-event |
| Analysis population | Phase 3 Full Analysis Set |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.53 |
| 95% CI | 0.407–0.677 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The estimated hazard ratio of 0.53 means that, under the stratified Cox model used for this analysis, the estimated instantaneous rate of progression or death in Arm B was approximately 53% of the corresponding rate in Arm C. Expressed as a simple relative interpretation, this corresponds to an estimated 47% lower hazard in Arm B under that model.
The HR does not mean that 47% of participants avoided progression, that median PFS differed by 47%, or that each individual experienced a 47% reduction in risk. It is a relative time-to-event measure.
The two-sided 95% confidence interval of 0.407–0.677 describes the statistical uncertainty around the estimated hazard ratio under the specified analysis framework. It does not describe the range of individual patient outcomes.
The reported P < 0.0001 addresses the evidence against the relevant null hypothesis under the reported testing framework. A p-value does not measure the size of the treatment effect, clinical importance, or probability that the treatment hypothesis is true.
Because the hazard ratio was obtained from a Cox proportional-hazards model, its interpretation also depends on the appropriateness of the model for the observed time-to-event data. The registry supplies a hazard ratio and stratified log-rank analysis, but does not provide enough information in the ClinicalTrials.gov recordset to independently assess proportional hazards.
6.2 Phase 3 Objective Response Rate: Arm B vs Arm C
Odds ratio for objective response
99.8% CI: 0.989–6.089 · P = 0.0008
Arm B: EC + mFOLFOX6 vs Arm C: Standard of Care Chemotherapy
| Feature | Reported result |
|---|---|
| Endpoint | Objective Response Rate (ORR) as assessed by BICR for Arm B vs Arm C - FAS ORR Subset |
| Endpoint type | Binary |
| Analysis population | First 110 participants randomized in each of Arm B and Arm C |
| Method | Cochran-Mantel-Haenszel test |
| Effect measure | Odds ratio |
| Estimate | 2.443 |
| 99.8% CI | 0.989–6.089 |
| P-value | =0.0008 |
| Hypothesis | Superiority |
The odds ratio of 2.443 indicates that the estimated odds of objective response were 2.443 times as high in Arm B as in Arm C under the stratified Cochran-Mantel-Haenszel analysis.
An odds ratio is not the same as saying that the probability of response was 2.443 times as high. The odds scale and probability scale are different, particularly when response is not a rare event. The ClinicalTrials.gov record does not include the arm-level response counts needed to translate this odds ratio into response probabilities.
The unusually specific 99.8% confidence interval of 0.989–6.089 reflects the confidence level reported for this formal analysis. The interval is relatively broad compared with the point estimate, illustrating that a point estimate alone should not be treated as a complete description of statistical precision.
The reported P = 0.0008 is the registry's stated p-value for the superiority comparison. The analysis notes specify that this is a 1-sided p-value from the stratified CMH test, whereas the confidence interval is two-sided. These are different conventions and should not be silently treated as though they were generated by the same tail specification.
The ORR analysis also used a specific subset: the first 110 participants randomized in each of Arms B and C. That population restriction is important when interpreting the estimate and prevents it from automatically being treated as an estimate based on every participant enrolled in the overall trial.
6.3 Cohort 3 Objective Response Rate: Arm D vs Arm E
Odds ratio for objective response
95% CI: 1.420–5.348 · P = 0.0011
Arm D: EC + FOLFIRI vs Arm E: FOLFIRI With or Without Bevacizumab
| Feature | Reported result |
|---|---|
| Endpoint | Cohort 3 ORR as assessed by BICR for Arm D vs Arm E - FAS |
| Endpoint type | Binary |
| Analysis population | Cohort 3 Full Analysis Set |
| Method | Cochran-Mantel-Haenszel test |
| Effect measure | Odds ratio |
| Estimate | 2.756 |
| 95% CI | 1.420–5.348 |
| P-value | =0.0011 |
| Hypothesis | Superiority |
The odds ratio of 2.756 indicates that the estimated odds of objective response were 2.756 times as high in Arm D as in Arm E under the stratified CMH analysis.
The 95% confidence interval of 1.420–5.348 quantifies uncertainty around the odds-ratio estimate under the analysis framework. It does not provide a range of response probabilities for individual patients.
The reported P = 0.0011 is evidence against the null hypothesis specified for the superiority analysis under the reported testing framework. It does not quantify the magnitude of benefit, and it should not be interpreted as the probability that the observed effect occurred by chance.
The registry does not supply arm-level response counts in the ClinicalTrials.gov record. Consequently, the odds ratio should be interpreted on its reported odds scale rather than converted into an unreported response-rate difference or risk ratio.
6.4 SLI Dose-Limiting Toxicity Endpoint
The registry lists “Number of Participants With Dose Limiting Toxicity (DLTs)” as a primary endpoint for the SLI portion of the study, with a time frame of Cycle 1 (28 days). The registry-reported statistical-analyses dataset does not contain a formal effect estimate, confidence interval, or p-value for this endpoint.
For a binary toxicity endpoint such as DLT occurrence, a typical statistical description would begin with the number and proportion of participants experiencing a DLT within the prespecified assessment window. If a comparative hypothesis were specified, an appropriate comparison could use a two-group categorical method or an exact method depending on sample size and the design. The ClinicalTrials.gov record does not provide a formal comparative analysis for the DLT endpoint, so no statistical comparison is inferred here.
7. Secondary Endpoint Result: Overall Survival
The ClinicalTrials.gov record contains one secondary statistical analysis: overall survival for Phase 3 Arm B versus Arm C.
Hazard ratio for overall survival
95% CI: 0.375–0.632 · P < 0.0001
Arm B: EC + mFOLFOX6 vs Arm C: Standard of Care Chemotherapy
| Feature | Reported result |
|---|---|
| Endpoint | Phase 3 OS for Arm B vs Arm C - FAS |
| Time frame | From date of first dose to death due to any cause or censoring date, whichever occurred first (maximum up to 37.25 months) |
| Endpoint type | Time-to-event |
| Analysis population | Phase 3 Full Analysis Set |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.49 |
| 95% CI | 0.375–0.632 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The OS hazard ratio of 0.49 means that, under the stratified Cox model used for this analysis, the estimated instantaneous rate of death in Arm B was approximately 49% of that in Arm C. As a simple relative interpretation, that corresponds to an estimated 51% lower hazard under the model.
This does not mean that 51% of participants survived because of treatment, nor does it mean that every participant experienced a 51% reduction in their individual probability of death. A hazard ratio summarizes a relative event-rate relationship over time.
The 95% confidence interval of 0.375–0.632 describes uncertainty around the estimated hazard ratio. It is not a prediction interval for individual survival times and does not state that every plausible patient-level treatment effect lies inside this interval.
The reported P < 0.0001 addresses statistical evidence under the reported one-sided stratified log-rank test. It should not be confused with effect size. The magnitude of the estimated association is communicated by the hazard ratio and its confidence interval.
As with the PFS analysis, interpretation of the Cox hazard ratio depends on the model and censoring framework. The ClinicalTrials.gov record does not provide enough information to evaluate the proportional-hazards assumption independently.
8. Comparing the Time-to-Event and Binary Analyses
| Feature | PFS / OS | ORR |
|---|---|---|
| Data structure | Time-to-event | Binary |
| Primary method | Log-rank test | Cochran-Mantel-Haenszel test |
| Main effect measure | Hazard ratio | Odds ratio |
| Censoring | Yes, as specified by the endpoint definition | Endpoint assessed through the specified response window |
| Clinical question | How does the occurrence of progression/death or death evolve over time? | How do the odds of achieving a confirmed response compare? |
This distinction is central to interpreting BREAKWATER. PFS and OS incorporate the timing of events and censoring, whereas ORR reduces the response assessment to a binary outcome within the prespecified assessment framework. Consequently, a hazard ratio and an odds ratio cannot be compared numerically as though they were interchangeable effect measures.
9. What the Hazard Ratios Mean
The estimated instantaneous rate of progression or death in Arm B relative to Arm C was approximately 53% under the reported stratified Cox model. A value below 1 indicates a lower estimated event rate for the first-listed group.
The estimated instantaneous rate of death in Arm B relative to Arm C was approximately 49% under the reported stratified Cox model. This is a relative time-to-event measure, not an absolute survival probability.
10. What the Odds Ratios Mean
The estimated odds of objective response were 2.443 times those in the comparator group under the stratified CMH analysis. The confidence interval and p-value should be read alongside the point estimate rather than treating the point estimate as exact.
The estimated odds of objective response were 2.756 times those in the comparator group under the stratified CMH analysis. This does not establish that the response probability itself was 2.756 times higher.
An odds ratio becomes a risk or response-rate comparison only after specifying the underlying event probabilities. Because the ClinicalTrials.gov record does not include the corresponding arm-level response counts, the statistically appropriate interpretation here is to remain on the odds scale reported by the registry.
11. Confidence Intervals and Precision
BREAKWATER provides confidence intervals for all three formal primary endpoint effect estimates. These intervals are an important part of the evidence because they show how much uncertainty surrounds each point estimate.
| Endpoint | Effect | Confidence interval | Confidence level |
|---|---|---|---|
| Phase 3 PFS | HR 0.53 | 0.407–0.677 | 95% |
| Phase 3 ORR | OR 2.443 | 0.989–6.089 | 99.8% |
| Cohort 3 ORR | OR 2.756 | 1.420–5.348 | 95% |
| Phase 3 OS | HR 0.49 | 0.375–0.632 | 95% |
The Phase 3 ORR analysis is especially useful for teaching the distinction between the confidence level and the p-value. Its confidence interval is reported at 99.8%, while the analysis note identifies a 1-sided p-value from the stratified CMH test. The confidence interval should therefore be interpreted according to its stated two-sided 99.8% construction rather than reverse-engineering a different interval from the p-value.
12. Statistical Methods Explained
Why was a log-rank test used for PFS and OS?
PFS and OS are time-to-event endpoints. Participants can be followed for different lengths of time, and some observations can be censored before an event occurs. The log-rank test is designed to compare the survival experience of randomized groups while incorporating the timing of observed events and censoring.
What does a hazard ratio of 0.53 mean?
It means that the estimated instantaneous rate of progression or death in Arm B was approximately 53% of the rate in Arm C under the reported Cox model. It does not mean that exactly 53% of participants experienced the event or that the median PFS was 53% as long.
Why was the Cochran-Mantel-Haenszel test used for ORR?
ORR is binary: a participant either achieves the prespecified response category or does not. The CMH method provides a way to compare treatment groups while accounting for stratification rather than treating all observations as though no stratifying structure existed.
What does an odds ratio of 2.443 mean?
An OR of 2.443 means that the estimated odds of response in Arm B were 2.443 times the estimated odds in Arm C. Odds are not probabilities, so the number cannot be interpreted as a 144.3% increase in the response rate without additional information.
Why is the analysis population important?
The Phase 3 PFS analysis used the Phase 3 FAS, while the Phase 3 ORR analysis used the first 110 participants randomized in each of Arms B and C. Those are different analysis populations. An estimate from the ORR subset should not automatically be generalized to the entire randomized Phase 3 population.
Why should the p-value not be used as an effect-size measure?
A p-value quantifies statistical evidence under a specified null hypothesis and testing framework. It does not describe how large the treatment effect is. For effect magnitude, the relevant quantities here are the hazard ratios or odds ratios, while their confidence intervals describe uncertainty.
Why does stratification matter?
The registry-reported analysis notes identify stratified log-rank testing for the Phase 3 PFS and OS analyses and stratified CMH methods for ORR. Stratified analysis preserves information about the trial's comparison structure and can improve the alignment between the statistical analysis and the way randomization or comparison strata were defined.
13. One-Sided Tests and Two-Sided Confidence Intervals
The registry's analysis notes contain an important detail that is easy to miss: the PFS analysis reports a 1-sided p-value from a stratified log-rank test, while the hazard ratio confidence interval is explicitly identified as 95% two-sided. The Phase 3 ORR analysis similarly reports a 1-sided p-value from the stratified CMH test alongside a 99.8% two-sided confidence interval.
One-sided p-value
The test asks whether the evidence supports the prespecified superiority direction rather than allocating the testing probability symmetrically to both directions.
Two-sided confidence interval
The interval expresses uncertainty on both sides of the estimated effect according to the stated confidence level.
These are complementary but not interchangeable summaries. A reader should preserve the registry's reported tail specification rather than assuming that every p-value and interval were generated under identical conventions.
14. Serious Adverse Events by Arm
The ClinicalTrials.gov record reports serious adverse events as affected participants divided by participants at risk for seven listed arms or cohorts. These figures are presented exactly as reported in the registry and are not converted into additional percentages.
| Study portion / arm | Serious adverse events affected / at risk |
|---|---|
| SLI: Cohort 1 [EC + FOLFIRI] | 14/30 |
| SLI: Cohort 2 [EC + mFOLFOX6] | 12/27 |
| Phase 3: Arm A [EC] | 46/153 |
| Phase 3: Arm B [EC + mFOLFOX6] | 107/232 |
| Phase 3: Arm C [Standard of Care Chemotherapy, Control Arm] | 89/229 |
| Cohort 3: Arm D [EC + FOLFIRI] | 28/71 |
| Cohort 3: Arm E [FOLFIRI With or Without Bevacizumab] | 25/68 |
These are safety counts, not efficacy effect estimates. They should not be compared using the same interpretive framework as the PFS hazard ratio or ORR odds ratios. In particular, the ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates by arm.
15. Design Topics Supported by the Registry Data
| Topic | What the ClinicalTrials.gov record establishes |
|---|---|
| Randomization | The study is randomized. |
| Parallel design | The design model is parallel. |
| Masking | None. |
| Superiority | The reported formal hypothesis type is superiority. |
| Stratified analysis | Identified for the reported log-rank and CMH analyses. |
| Time-to-event analysis | Used for PFS and OS. |
| Binary analysis | Used for DLT and ORR endpoints. |
| Bayesian methods | Not reported in the ClinicalTrials.gov record. |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Factorial design | Not reported; the registry identifies the design model as parallel. |
| Missing-data imputation | Not reported in the ClinicalTrials.gov record. |
| Interim analysis / alpha spending | Not reported in the ClinicalTrials.gov record. |
This separation is deliberate. Statistical methods that are not present in the ClinicalTrials.gov record should not be attributed to the trial merely because they are common in other phase 3 studies.
16. Limitations and Interpretation Issues
- Multiple arms and comparisons: BREAKWATER contains 7 arms, but the formal analyses reported here concern specific comparisons rather than every possible pair of arms.
- Different analysis populations: the Phase 3 PFS and OS analyses use the Phase 3 FAS, whereas the Phase 3 ORR analysis uses the first 110 randomized participants in each of Arms B and C.
- Different effect measures: PFS and OS are summarized with hazard ratios, while ORR is summarized with odds ratios. These measures should not be treated as interchangeable.
- Confidence-level differences: the Phase 3 ORR analysis uses a 99.8% confidence interval, while the other registry-reported formal primary analyses use 95% confidence intervals.
- One-sided testing: the analysis notes identify 1-sided p-values for the Phase 3 PFS and ORR analyses. This should be retained when interpreting statistical evidence.
- Proportional-hazards assumption: the PFS and OS hazard ratios are based on stratified Cox proportional-hazards models. The ClinicalTrials.gov record does not provide sufficient information to independently evaluate that assumption.
- Incomplete endpoint-level results: the SLI DLT endpoint is listed as a primary endpoint and has results posted, but the registry-reported statistical-analyses dataset does not provide a formal statistical comparison for it.
- Safety interpretation: serious adverse-event counts are reported by arm, but no formal comparative safety analysis is included in the ClinicalTrials.gov record.
- Limited reconstruction: the ClinicalTrials.gov record does not include arm-level response counts, Kaplan-Meier coordinates, median PFS, median OS, subgroup estimates, or other quantities needed for additional quantitative reconstruction.
17. Why This Trial Matters Statistically
BREAKWATER is a useful teaching example because the ClinicalTrials.gov record connect several fundamental clinical-trial methods within one randomized study. The trial contains both binary and time-to-event primary endpoints, multiple treatment arms, stratified analyses, different effect measures, and different analysis populations.
| Concept | How it appears in BREAKWATER |
|---|---|
| Randomization | The study uses randomized allocation. |
| Parallel design | The registry identifies a parallel design model. |
| Full Analysis Set | The Phase 3 PFS and OS analyses use the Phase 3 FAS. |
| Time-to-event endpoints | PFS and OS are analyzed using the log-rank framework. |
| Hazard ratio | PFS HR 0.53 and OS HR 0.49 are reported for Arm B versus Arm C. |
| Cox model | The reported PFS and OS HRs are based on stratified Cox proportional-hazards models. |
| Binary endpoint | ORR is analyzed as a binary outcome. |
| Cochran-Mantel-Haenszel test | Used for the reported Phase 3 and Cohort 3 ORR comparisons. |
| Odds ratio | OR 2.443 for Phase 3 ORR and OR 2.756 for Cohort 3 ORR. |
| Stratified analysis | Identified in the analysis notes for the reported formal comparisons. |
| Confidence intervals | Reported for all three registry-reported primary endpoint effect estimates. |
| One-sided p-values | Specified in the analysis notes for the Phase 3 PFS and ORR tests. |
| Safety analysis | Serious adverse events are reported as affected / at-risk counts by arm. |
18. Interpreting the Trial as a Statistical Story
The most important statistical feature of BREAKWATER is not any single number. It is the way different endpoint types require different analyses and different interpretations.
PFS
PFS is a time-to-event endpoint. Its analysis accounts for when progression or death occurs and for censoring. The primary effect measure is the hazard ratio.
ORR
ORR is binary. The registry-reported analysis uses a stratified CMH test and reports an odds ratio rather than a hazard ratio.
OS
OS is another time-to-event endpoint. The secondary analysis uses the same general log-rank/Cox framework as the reported PFS analysis.
Safety
Serious adverse events are reported as counts relative to participants at risk. These data are descriptive in the ClinicalTrials.gov record rather than a formal efficacy-style comparison.
This structure illustrates why clinical-trial statistics should not be reduced to a single p-value. The endpoint definition, analysis population, effect measure, confidence interval, testing direction, and statistical model all contribute to the interpretation.
19. Results Summary
| Endpoint | Comparison | Method | Effect | Confidence interval | P-value |
|---|---|---|---|---|---|
| PFS | Arm B vs Arm C | Log-rank | HR 0.53 | 95% CI 0.407–0.677 | <0.0001 |
| ORR | Arm B vs Arm C | CMH | OR 2.443 | 99.8% CI 0.989–6.089 | =0.0008 |
| ORR | Arm D vs Arm E | CMH | OR 2.756 | 95% CI 1.420–5.348 | =0.0011 |
| OS | Arm B vs Arm C | Log-rank | HR 0.49 | 95% CI 0.375–0.632 | <0.0001 |
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Statistical Calculators
22. Sources
- ClinicalTrials.gov: NCT04607421 — BREAKWATER.
- Linked publication: PubMed PMID 41230651.
- Linked publication: PubMed PMID 40444708.
- Linked publication: PubMed PMID 39863775.
- Linked publication: PubMed PMID 36763936.
Continue through the Clinical Biostats statistical tutorials
Explore the statistical methods behind randomized clinical trials, survival analysis, categorical endpoints, confidence intervals, and treatment-effect estimation.
23. Record Summary
BREAKWATER provides a useful statistical case study because the ClinicalTrials.gov record combines randomized parallel treatment allocation with both binary and time-to-event primary endpoints. The Phase 3 PFS comparison of Arm B versus Arm C reports a hazard ratio of 0.53 with a 95% CI of 0.407–0.677 and a p-value of <0.0001. The Phase 3 ORR comparison reports an odds ratio of 2.443 with a 99.8% CI of 0.989–6.089 and a p-value of =0.0008. The Cohort 3 ORR comparison reports an odds ratio of 2.756 with a 95% CI of 1.420–5.348 and a p-value of =0.0011. The secondary Phase 3 OS analysis reports a hazard ratio of 0.49 with a 95% CI of 0.375–0.632 and a p-value of <0.0001.
The statistical interpretation depends on preserving the distinctions among these analyses: PFS and OS are time-to-event endpoints analyzed with log-rank testing and stratified Cox models, while ORR is a binary endpoint analyzed with the stratified Cochran-Mantel-Haenszel method. The analysis populations also differ, particularly for the Phase 3 ORR endpoint. Finally, the registry's the ClinicalTrials.gov record are descriptive serious-adverse-event counts by arm rather than formal comparative efficacy-style analyses.