← Clinical Trials
Metastatic Colorectal Cancer Phase 3 Superiority NCT04607421

BREAKWATER: Complete Statistical Analysis of Encorafenib Plus Cetuximab in Metastatic Colorectal Cancer

An independent statistical analysis of the randomized phase 3 BREAKWATER trial evaluating encorafenib plus cetuximab with or without chemotherapy in people with previously untreated metastatic colorectal cancer.

Trial start: 2020-12-21  ·  Primary completion: 2025-03-01  ·  Lead sponsor: Pfizer
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information contained in the ClinicalTrials.gov record.

1. Trial at a Glance

BREAKWATER is a randomized, open-label, parallel phase 3 study in oncology. The registry describes 841 participants, 7 arms, 4 registered primary endpoints, and posted statistical analyses using the log-rank test and Cochran-Mantel-Haenszel test.

841
Enrollment
Registry enrollment
7
Arms
Parallel trial design
4
Primary endpoints
Binary + time-to-event
4
Statistical analyses
3 primary + 1 secondary
FeatureBREAKWATER
PhasePhase 3
PopulationPeople with previously untreated metastatic colorectal cancer
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
AllocationRandomized
Enrollment841
Arms7
Primary endpoint typesBinary; Time-to-event
Hypothesis typeSuperiority
Results postedYes
Lead sponsorPfizer
Trial statusActive, not recruiting

2. Clinical Question

The statistical question in BREAKWATER is framed around randomized comparisons of treatment strategies in people with previously untreated metastatic colorectal cancer. The registry reports primary comparisons involving encorafenib plus cetuximab with mFOLFOX6 versus standard-of-care chemotherapy, and encorafenib plus cetuximab with FOLFIRI versus FOLFIRI with or without bevacizumab.

Population

People with previously untreated metastatic colorectal cancer, as described in the trial's brief title.

Interventions

The registered interventions include encorafenib, cetuximab, oxaliplatin, irinotecan, leucovorin, 5-FU, capecitabine, and bevacizumab.

Comparator

The primary statistical comparisons reported include standard-of-care chemotherapy in Phase 3 and FOLFIRI with or without bevacizumab in Cohort 3.

Primary question

Under the registry's superiority framework, how do the prespecified treatment groups compare for progression-free survival and objective response rate?

3. Trial Design

01
Randomize841 enrolled
02
7 armsParallel allocation
03
TreatmentOpen-label
04
AssessmentPFS / ORR / safety
05
Follow-upTime-to-event outcomes
Allocation
Randomized allocation.
Design model
Parallel.
Masking
None.
Primary purpose
Treatment.

The presence of 7 arms is important statistically. A multi-arm trial can contain several clinically distinct comparisons, but the interpretation of any one estimate remains tied to the exact groups and analysis population specified for that endpoint. BREAKWATER's posted primary analyses are not all comparisons across all 7 arms; they are specific prespecified comparisons.

Reported Phase 3 and Cohort 3 comparisons

PHASE 3 · ARM A

Encorafenib + cetuximab

The registry identifies Arm A as EC and reports 46 serious adverse events among 153 participants at risk.

PHASE 3 · ARM B

Encorafenib + cetuximab + mFOLFOX6

Arm B is the treatment group in the primary Phase 3 comparisons with Arm C. The serious adverse event count is 107 among 232 participants at risk.

PHASE 3 · ARM C

Standard-of-care chemotherapy

Arm C is the control arm for the primary Phase 3 comparison with Arm B. The serious adverse event count is 89 among 229 participants at risk.

COHORT 3 · ARMS D / E

FOLFIRI comparison

Arm D is EC + FOLFIRI, while Arm E is FOLFIRI with or without bevacizumab. Serious adverse events were reported as 28/71 and 25/68, respectively.

Important design distinction. The registry also identifies SLI Cohort 1 groups involving EC + FOLFIRI and EC + mFOLFOX6. The registered SLI primary endpoint is dose-limiting toxicity during Cycle 1, whereas the formal statistical analyses in the ClinicalTrials.gov record concern the Phase 3 PFS and ORR comparisons, the Cohort 3 ORR comparison, and Phase 3 OS.

4. Endpoints

Registered primary endpointTime frameTypeAnalysis reported?
SLI: Number of Participants With Dose Limiting Toxicity (DLTs) Cycle 1 (28 days) Binary Endpoint results posted; no formal statistical comparison reported in the statistical analyses dataset
Phase 3: Progression Free Survival (PFS) as Assessed by BICR for Arm B vs Arm C - FAS From date of randomization to earliest documentation of PD by BICR or death or censoring date, whichever occurred first (maximum up to 37.25 months) Time-to-event Yes
Phase 3: Objective Response Rate (ORR) as Assessed by BICR for Arm B vs Arm C - FAS ORR Subset From date of randomization until documented PD by BICR, or start of subsequent anticancer therapy or death, whichever occurred first (maximum up to 24.71 months) Binary Yes
Cohort 3: ORR as Assessed by BICR for Arm D vs Arm E - FAS From date of randomization until documented PD by BICR, or start of subsequent anticancer therapy or death, whichever occurred first (maximum up to 24.71 months) Binary Yes

The PFS definition specifies time from randomization to the earliest documented disease progression according to RECIST version 1.1 or death due to any cause, as assessed by blinded independent central review. The ORR definitions specify confirmed complete response or partial response according to RECIST version 1.1 as assessed by BICR.

The registry's primary endpoint list therefore combines two different statistical structures: a binary toxicity endpoint and response endpoints, plus a censored time-to-event endpoint. That distinction determines which statistical methods are appropriate.

5. Statistical Methodology

Full Analysis Set

The Phase 3 Full Analysis Set included all participants randomized in the Phase 3 portion of the study. The Phase 3 ORR analysis used the first 110 participants randomized in each of Arm B and Arm C. The Cohort 3 FAS included all participants randomized in the Cohort 3 portion of the study.

Analysis populations reported by the registry
Phase 3 PFS: FAS  ·  Phase 3 ORR: first 110 randomized in Arm B and Arm C  ·  Cohort 3 ORR: Cohort 3 FAS

The analysis population matters because an effect estimate is meaningful only in the population to which it was actually applied. The denominator and eligibility for a response analysis can differ from those of a time-to-event analysis.

Log-rank test

The registry reports the log-rank test for the primary Phase 3 PFS comparison and the secondary Phase 3 OS comparison. The log-rank framework compares survival experience over time while accounting for right-censored observations.

Conceptual survival comparison
H0: survival distributions are equivalent between the specified randomized groups

The test is based on observed and expected event patterns over follow-up. It is not a test of whether the numerical hazard ratio is exactly a particular value.

Stratified Cox proportional-hazards model

For the reported Phase 3 PFS and OS analyses, the hazard ratio was based on a stratified Cox proportional-hazards model. The registry analysis notes identify the log-rank p-value as a 1-sided p-value and the confidence interval as two-sided.

Hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the first-listed treatment group

A hazard ratio is a relative, model-based measure of event rates over the analyzed follow-up. It is not a percentage of patients who benefit and is not an absolute difference in survival probability.

Cochran-Mantel-Haenszel test

The Phase 3 ORR comparison and Cohort 3 ORR comparison used the Cochran-Mantel-Haenszel test. The registry also states that the odds ratio was estimated using the stratified Cochran-Mantel-Haenszel method.

This approach is particularly useful when a binary endpoint is compared across randomized groups while accounting for prespecified stratification. Instead of simply pooling all response counts and ignoring strata, the CMH framework combines evidence across strata.

Odds ratio

Conceptual form
OR = (odds of response in treatment group) / (odds of response in comparator group)

An odds ratio above 1 indicates higher estimated odds of the binary event in the first-listed group relative to the comparator. It is not identical to a risk ratio or a difference in response probabilities.

6. Results

The ClinicalTrials.gov record contains four formal statistical analyses: three primary endpoint analyses and one secondary endpoint analysis. Three of the analyses include an effect estimate and confidence interval.

6.1 Phase 3 Progression-Free Survival: Arm B vs Arm C

Hazard ratio for progression or death

0.53

95% CI: 0.407–0.677   ·   P < 0.0001

Arm B: EC + mFOLFOX6   vs   Arm C: Standard of Care Chemotherapy

FeatureReported result
EndpointProgression Free Survival (PFS) as assessed by BICR for Arm B vs Arm C - FAS
Endpoint typeTime-to-event
Analysis populationPhase 3 Full Analysis Set
MethodLog-rank test
Effect measureHazard ratio
Estimate0.53
95% CI0.407–0.677
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The estimated hazard ratio of 0.53 means that, under the stratified Cox model used for this analysis, the estimated instantaneous rate of progression or death in Arm B was approximately 53% of the corresponding rate in Arm C. Expressed as a simple relative interpretation, this corresponds to an estimated 47% lower hazard in Arm B under that model.

The HR does not mean that 47% of participants avoided progression, that median PFS differed by 47%, or that each individual experienced a 47% reduction in risk. It is a relative time-to-event measure.

The two-sided 95% confidence interval of 0.407–0.677 describes the statistical uncertainty around the estimated hazard ratio under the specified analysis framework. It does not describe the range of individual patient outcomes.

The reported P < 0.0001 addresses the evidence against the relevant null hypothesis under the reported testing framework. A p-value does not measure the size of the treatment effect, clinical importance, or probability that the treatment hypothesis is true.

Because the hazard ratio was obtained from a Cox proportional-hazards model, its interpretation also depends on the appropriateness of the model for the observed time-to-event data. The registry supplies a hazard ratio and stratified log-rank analysis, but does not provide enough information in the ClinicalTrials.gov recordset to independently assess proportional hazards.

6.2 Phase 3 Objective Response Rate: Arm B vs Arm C

Odds ratio for objective response

2.443

99.8% CI: 0.989–6.089   ·   P = 0.0008

Arm B: EC + mFOLFOX6   vs   Arm C: Standard of Care Chemotherapy

FeatureReported result
EndpointObjective Response Rate (ORR) as assessed by BICR for Arm B vs Arm C - FAS ORR Subset
Endpoint typeBinary
Analysis populationFirst 110 participants randomized in each of Arm B and Arm C
MethodCochran-Mantel-Haenszel test
Effect measureOdds ratio
Estimate2.443
99.8% CI0.989–6.089
P-value=0.0008
HypothesisSuperiority
Clinical Biostats interpretation

The odds ratio of 2.443 indicates that the estimated odds of objective response were 2.443 times as high in Arm B as in Arm C under the stratified Cochran-Mantel-Haenszel analysis.

An odds ratio is not the same as saying that the probability of response was 2.443 times as high. The odds scale and probability scale are different, particularly when response is not a rare event. The ClinicalTrials.gov record does not include the arm-level response counts needed to translate this odds ratio into response probabilities.

The unusually specific 99.8% confidence interval of 0.989–6.089 reflects the confidence level reported for this formal analysis. The interval is relatively broad compared with the point estimate, illustrating that a point estimate alone should not be treated as a complete description of statistical precision.

The reported P = 0.0008 is the registry's stated p-value for the superiority comparison. The analysis notes specify that this is a 1-sided p-value from the stratified CMH test, whereas the confidence interval is two-sided. These are different conventions and should not be silently treated as though they were generated by the same tail specification.

The ORR analysis also used a specific subset: the first 110 participants randomized in each of Arms B and C. That population restriction is important when interpreting the estimate and prevents it from automatically being treated as an estimate based on every participant enrolled in the overall trial.

6.3 Cohort 3 Objective Response Rate: Arm D vs Arm E

Odds ratio for objective response

2.756

95% CI: 1.420–5.348   ·   P = 0.0011

Arm D: EC + FOLFIRI   vs   Arm E: FOLFIRI With or Without Bevacizumab

FeatureReported result
EndpointCohort 3 ORR as assessed by BICR for Arm D vs Arm E - FAS
Endpoint typeBinary
Analysis populationCohort 3 Full Analysis Set
MethodCochran-Mantel-Haenszel test
Effect measureOdds ratio
Estimate2.756
95% CI1.420–5.348
P-value=0.0011
HypothesisSuperiority
Clinical Biostats interpretation

The odds ratio of 2.756 indicates that the estimated odds of objective response were 2.756 times as high in Arm D as in Arm E under the stratified CMH analysis.

The 95% confidence interval of 1.420–5.348 quantifies uncertainty around the odds-ratio estimate under the analysis framework. It does not provide a range of response probabilities for individual patients.

The reported P = 0.0011 is evidence against the null hypothesis specified for the superiority analysis under the reported testing framework. It does not quantify the magnitude of benefit, and it should not be interpreted as the probability that the observed effect occurred by chance.

The registry does not supply arm-level response counts in the ClinicalTrials.gov record. Consequently, the odds ratio should be interpreted on its reported odds scale rather than converted into an unreported response-rate difference or risk ratio.

6.4 SLI Dose-Limiting Toxicity Endpoint

The registry lists “Number of Participants With Dose Limiting Toxicity (DLTs)” as a primary endpoint for the SLI portion of the study, with a time frame of Cycle 1 (28 days). The registry-reported statistical-analyses dataset does not contain a formal effect estimate, confidence interval, or p-value for this endpoint.

For a binary toxicity endpoint such as DLT occurrence, a typical statistical description would begin with the number and proportion of participants experiencing a DLT within the prespecified assessment window. If a comparative hypothesis were specified, an appropriate comparison could use a two-group categorical method or an exact method depending on sample size and the design. The ClinicalTrials.gov record does not provide a formal comparative analysis for the DLT endpoint, so no statistical comparison is inferred here.

Interpretation boundary: the absence of a formal statistical analysis in the ClinicalTrials.gov recordset is not evidence of absence of DLTs. It means only that the ClinicalTrials.gov record does not provide an estimate, confidence interval, or p-value for this endpoint.

7. Secondary Endpoint Result: Overall Survival

The ClinicalTrials.gov record contains one secondary statistical analysis: overall survival for Phase 3 Arm B versus Arm C.

Hazard ratio for overall survival

0.49

95% CI: 0.375–0.632   ·   P < 0.0001

Arm B: EC + mFOLFOX6   vs   Arm C: Standard of Care Chemotherapy

FeatureReported result
EndpointPhase 3 OS for Arm B vs Arm C - FAS
Time frameFrom date of first dose to death due to any cause or censoring date, whichever occurred first (maximum up to 37.25 months)
Endpoint typeTime-to-event
Analysis populationPhase 3 Full Analysis Set
MethodLog-rank test
Effect measureHazard ratio
Estimate0.49
95% CI0.375–0.632
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The OS hazard ratio of 0.49 means that, under the stratified Cox model used for this analysis, the estimated instantaneous rate of death in Arm B was approximately 49% of that in Arm C. As a simple relative interpretation, that corresponds to an estimated 51% lower hazard under the model.

This does not mean that 51% of participants survived because of treatment, nor does it mean that every participant experienced a 51% reduction in their individual probability of death. A hazard ratio summarizes a relative event-rate relationship over time.

The 95% confidence interval of 0.375–0.632 describes uncertainty around the estimated hazard ratio. It is not a prediction interval for individual survival times and does not state that every plausible patient-level treatment effect lies inside this interval.

The reported P < 0.0001 addresses statistical evidence under the reported one-sided stratified log-rank test. It should not be confused with effect size. The magnitude of the estimated association is communicated by the hazard ratio and its confidence interval.

As with the PFS analysis, interpretation of the Cox hazard ratio depends on the model and censoring framework. The ClinicalTrials.gov record does not provide enough information to evaluate the proportional-hazards assumption independently.

8. Comparing the Time-to-Event and Binary Analyses

FeaturePFS / OSORR
Data structureTime-to-eventBinary
Primary methodLog-rank testCochran-Mantel-Haenszel test
Main effect measureHazard ratioOdds ratio
CensoringYes, as specified by the endpoint definitionEndpoint assessed through the specified response window
Clinical questionHow does the occurrence of progression/death or death evolve over time?How do the odds of achieving a confirmed response compare?

This distinction is central to interpreting BREAKWATER. PFS and OS incorporate the timing of events and censoring, whereas ORR reduces the response assessment to a binary outcome within the prespecified assessment framework. Consequently, a hazard ratio and an odds ratio cannot be compared numerically as though they were interchangeable effect measures.

9. What the Hazard Ratios Mean

PFS HR = 0.53

The estimated instantaneous rate of progression or death in Arm B relative to Arm C was approximately 53% under the reported stratified Cox model. A value below 1 indicates a lower estimated event rate for the first-listed group.

OS HR = 0.49

The estimated instantaneous rate of death in Arm B relative to Arm C was approximately 49% under the reported stratified Cox model. This is a relative time-to-event measure, not an absolute survival probability.

Do not equate HR 0.49 with a 51% absolute improvement. A hazard ratio is neither an absolute risk difference nor a difference in survival percentages. Absolute interpretation requires time-specific survival estimates or other directly reported measures, which are not included in the ClinicalTrials.gov record.

10. What the Odds Ratios Mean

Phase 3 ORR OR = 2.443

The estimated odds of objective response were 2.443 times those in the comparator group under the stratified CMH analysis. The confidence interval and p-value should be read alongside the point estimate rather than treating the point estimate as exact.

Cohort 3 ORR OR = 2.756

The estimated odds of objective response were 2.756 times those in the comparator group under the stratified CMH analysis. This does not establish that the response probability itself was 2.756 times higher.

An odds ratio becomes a risk or response-rate comparison only after specifying the underlying event probabilities. Because the ClinicalTrials.gov record does not include the corresponding arm-level response counts, the statistically appropriate interpretation here is to remain on the odds scale reported by the registry.

11. Confidence Intervals and Precision

BREAKWATER provides confidence intervals for all three formal primary endpoint effect estimates. These intervals are an important part of the evidence because they show how much uncertainty surrounds each point estimate.

EndpointEffectConfidence intervalConfidence level
Phase 3 PFSHR 0.530.407–0.67795%
Phase 3 ORROR 2.4430.989–6.08999.8%
Cohort 3 ORROR 2.7561.420–5.34895%
Phase 3 OSHR 0.490.375–0.63295%

The Phase 3 ORR analysis is especially useful for teaching the distinction between the confidence level and the p-value. Its confidence interval is reported at 99.8%, while the analysis note identifies a 1-sided p-value from the stratified CMH test. The confidence interval should therefore be interpreted according to its stated two-sided 99.8% construction rather than reverse-engineering a different interval from the p-value.

12. Statistical Methods Explained

Why was a log-rank test used for PFS and OS?

PFS and OS are time-to-event endpoints. Participants can be followed for different lengths of time, and some observations can be censored before an event occurs. The log-rank test is designed to compare the survival experience of randomized groups while incorporating the timing of observed events and censoring.

What does a hazard ratio of 0.53 mean?

It means that the estimated instantaneous rate of progression or death in Arm B was approximately 53% of the rate in Arm C under the reported Cox model. It does not mean that exactly 53% of participants experienced the event or that the median PFS was 53% as long.

Why was the Cochran-Mantel-Haenszel test used for ORR?

ORR is binary: a participant either achieves the prespecified response category or does not. The CMH method provides a way to compare treatment groups while accounting for stratification rather than treating all observations as though no stratifying structure existed.

What does an odds ratio of 2.443 mean?

An OR of 2.443 means that the estimated odds of response in Arm B were 2.443 times the estimated odds in Arm C. Odds are not probabilities, so the number cannot be interpreted as a 144.3% increase in the response rate without additional information.

Why is the analysis population important?

The Phase 3 PFS analysis used the Phase 3 FAS, while the Phase 3 ORR analysis used the first 110 participants randomized in each of Arms B and C. Those are different analysis populations. An estimate from the ORR subset should not automatically be generalized to the entire randomized Phase 3 population.

Why should the p-value not be used as an effect-size measure?

A p-value quantifies statistical evidence under a specified null hypothesis and testing framework. It does not describe how large the treatment effect is. For effect magnitude, the relevant quantities here are the hazard ratios or odds ratios, while their confidence intervals describe uncertainty.

Why does stratification matter?

The registry-reported analysis notes identify stratified log-rank testing for the Phase 3 PFS and OS analyses and stratified CMH methods for ORR. Stratified analysis preserves information about the trial's comparison structure and can improve the alignment between the statistical analysis and the way randomization or comparison strata were defined.

13. One-Sided Tests and Two-Sided Confidence Intervals

The registry's analysis notes contain an important detail that is easy to miss: the PFS analysis reports a 1-sided p-value from a stratified log-rank test, while the hazard ratio confidence interval is explicitly identified as 95% two-sided. The Phase 3 ORR analysis similarly reports a 1-sided p-value from the stratified CMH test alongside a 99.8% two-sided confidence interval.

One-sided p-value

The test asks whether the evidence supports the prespecified superiority direction rather than allocating the testing probability symmetrically to both directions.

Two-sided confidence interval

The interval expresses uncertainty on both sides of the estimated effect according to the stated confidence level.

These are complementary but not interchangeable summaries. A reader should preserve the registry's reported tail specification rather than assuming that every p-value and interval were generated under identical conventions.

14. Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events as affected participants divided by participants at risk for seven listed arms or cohorts. These figures are presented exactly as reported in the registry and are not converted into additional percentages.

Study portion / armSerious adverse events affected / at risk
SLI: Cohort 1 [EC + FOLFIRI]14/30
SLI: Cohort 2 [EC + mFOLFOX6]12/27
Phase 3: Arm A [EC]46/153
Phase 3: Arm B [EC + mFOLFOX6]107/232
Phase 3: Arm C [Standard of Care Chemotherapy, Control Arm]89/229
Cohort 3: Arm D [EC + FOLFIRI]28/71
Cohort 3: Arm E [FOLFIRI With or Without Bevacizumab]25/68

These are safety counts, not efficacy effect estimates. They should not be compared using the same interpretive framework as the PFS hazard ratio or ORR odds ratios. In particular, the ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates by arm.

15. Design Topics Supported by the Registry Data

TopicWhat the ClinicalTrials.gov record establishes
RandomizationThe study is randomized.
Parallel designThe design model is parallel.
MaskingNone.
SuperiorityThe reported formal hypothesis type is superiority.
Stratified analysisIdentified for the reported log-rank and CMH analyses.
Time-to-event analysisUsed for PFS and OS.
Binary analysisUsed for DLT and ORR endpoints.
Bayesian methodsNot reported in the ClinicalTrials.gov record.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designNot reported; the registry identifies the design model as parallel.
Missing-data imputationNot reported in the ClinicalTrials.gov record.
Interim analysis / alpha spendingNot reported in the ClinicalTrials.gov record.

This separation is deliberate. Statistical methods that are not present in the ClinicalTrials.gov record should not be attributed to the trial merely because they are common in other phase 3 studies.

16. Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

BREAKWATER is a useful teaching example because the ClinicalTrials.gov record connect several fundamental clinical-trial methods within one randomized study. The trial contains both binary and time-to-event primary endpoints, multiple treatment arms, stratified analyses, different effect measures, and different analysis populations.

ConceptHow it appears in BREAKWATER
RandomizationThe study uses randomized allocation.
Parallel designThe registry identifies a parallel design model.
Full Analysis SetThe Phase 3 PFS and OS analyses use the Phase 3 FAS.
Time-to-event endpointsPFS and OS are analyzed using the log-rank framework.
Hazard ratioPFS HR 0.53 and OS HR 0.49 are reported for Arm B versus Arm C.
Cox modelThe reported PFS and OS HRs are based on stratified Cox proportional-hazards models.
Binary endpointORR is analyzed as a binary outcome.
Cochran-Mantel-Haenszel testUsed for the reported Phase 3 and Cohort 3 ORR comparisons.
Odds ratioOR 2.443 for Phase 3 ORR and OR 2.756 for Cohort 3 ORR.
Stratified analysisIdentified in the analysis notes for the reported formal comparisons.
Confidence intervalsReported for all three registry-reported primary endpoint effect estimates.
One-sided p-valuesSpecified in the analysis notes for the Phase 3 PFS and ORR tests.
Safety analysisSerious adverse events are reported as affected / at-risk counts by arm.

18. Interpreting the Trial as a Statistical Story

The most important statistical feature of BREAKWATER is not any single number. It is the way different endpoint types require different analyses and different interpretations.

PFS

PFS is a time-to-event endpoint. Its analysis accounts for when progression or death occurs and for censoring. The primary effect measure is the hazard ratio.

ORR

ORR is binary. The registry-reported analysis uses a stratified CMH test and reports an odds ratio rather than a hazard ratio.

OS

OS is another time-to-event endpoint. The secondary analysis uses the same general log-rank/Cox framework as the reported PFS analysis.

Safety

Serious adverse events are reported as counts relative to participants at risk. These data are descriptive in the ClinicalTrials.gov record rather than a formal efficacy-style comparison.

This structure illustrates why clinical-trial statistics should not be reduced to a single p-value. The endpoint definition, analysis population, effect measure, confidence interval, testing direction, and statistical model all contribute to the interpretation.

19. Results Summary

EndpointComparisonMethodEffectConfidence intervalP-value
PFS Arm B vs Arm C Log-rank HR 0.53 95% CI 0.407–0.677 <0.0001
ORR Arm B vs Arm C CMH OR 2.443 99.8% CI 0.989–6.089 =0.0008
ORR Arm D vs Arm E CMH OR 2.756 95% CI 1.420–5.348 =0.0011
OS Arm B vs Arm C Log-rank HR 0.49 95% CI 0.375–0.632 <0.0001
Reading the table correctly: the hazard ratios summarize time-to-event comparisons, while the odds ratios summarize binary response comparisons. The p-values communicate statistical evidence under their respective testing frameworks; they are not interchangeable measures of effect size.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical tutorials

Explore the statistical methods behind randomized clinical trials, survival analysis, categorical endpoints, confidence intervals, and treatment-effect estimation.

23. Record Summary

BREAKWATER provides a useful statistical case study because the ClinicalTrials.gov record combines randomized parallel treatment allocation with both binary and time-to-event primary endpoints. The Phase 3 PFS comparison of Arm B versus Arm C reports a hazard ratio of 0.53 with a 95% CI of 0.407–0.677 and a p-value of <0.0001. The Phase 3 ORR comparison reports an odds ratio of 2.443 with a 99.8% CI of 0.989–6.089 and a p-value of =0.0008. The Cohort 3 ORR comparison reports an odds ratio of 2.756 with a 95% CI of 1.420–5.348 and a p-value of =0.0011. The secondary Phase 3 OS analysis reports a hazard ratio of 0.49 with a 95% CI of 0.375–0.632 and a p-value of <0.0001.

The statistical interpretation depends on preserving the distinctions among these analyses: PFS and OS are time-to-event endpoints analyzed with log-rank testing and stratified Cox models, while ORR is a binary endpoint analyzed with the stratified Cochran-Mantel-Haenszel method. The analysis populations also differ, particularly for the Phase 3 ORR endpoint. Finally, the registry's the ClinicalTrials.gov record are descriptive serious-adverse-event counts by arm rather than formal comparative efficacy-style analyses.

Clinical Biostats methodology: A trial-results page should separate reported evidence from statistical interpretation. The goal is to explain what each estimate means, what it does not mean, and how the endpoint definition, analysis population, effect measure, confidence interval, and testing framework determine the appropriate interpretation.