← Clinical Trials
Metastatic Colorectal Cancer Phase 3 Time-to-Event Analysis NCT00700102

ML18147: Complete Statistical Analysis of Bevacizumab in Metastatic Colorectal Cancer

An independent statistical review of the randomized phase 3 ML18147 trial evaluating chemotherapy with versus without bevacizumab in patients with metastatic colorectal cancer, focusing on overall survival, progression-free survival, response rate, and the statistical methods used to compare the treatment groups.

Trial start: February 2006  ·  Primary completion: May 2013  ·  Enrollment: 820
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

ML18147 was a randomized, open-label, parallel-group phase 3 trial in metastatic colorectal cancer. The study enrolled 820 participants and compared chemotherapy alone with chemotherapy plus bevacizumab.

820
Enrollment
Randomized phase 3 trial
2
Arms
Parallel treatment groups
0.81
Overall Survival HR
95% CI 0.69–0.94
0.68
PFS HR
95% CI 0.59–0.78
FeatureML18147
Trial nameML18147
ClinicalTrials.gov identifierNCT00700102
PhasePhase 3
StatusCompleted
ConditionColorectal Cancer
Enrollment820
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
InterventionsChemotherapy; Bevacizumab
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry
Primary endpoint typeTime-to-event

2. Clinical Question

The central statistical question was whether chemotherapy plus bevacizumab differed from chemotherapy alone with respect to overall survival, defined as time from randomization to death from any cause.

Population

Participants enrolled in the phase 3 study of metastatic colorectal cancer.

Intervention

Chemotherapy plus bevacizumab.

Comparator

Chemotherapy.

Primary question

Does chemotherapy plus bevacizumab improve overall survival relative to chemotherapy alone?

The registered primary hypothesis type was superiority. That matters statistically: the objective was to determine whether the randomized groups differed in the specified direction of the treatment comparison rather than to demonstrate that the two regimens were sufficiently similar.

3. Trial Design

01
Randomize820 participants
02
Two armsChemotherapy vs combination
03
FollowTime-to-event outcomes
04
CompareLog-rank and categorical methods
05
InterpretEffect estimates and uncertainty
ARM 1

Chemotherapy

  • Chemotherapy was the comparator intervention.
  • Serious adverse events: 137 affected participants among 409 at risk.
ARM 2

Chemotherapy + Bevacizumab

  • Chemotherapy was administered with bevacizumab.
  • Serious adverse events: 130 affected participants among 401 at risk.
Crossover is an important design caveat. The registry notes that patients randomized to receive bevacizumab could continue to receive bevacizumab following discontinuation of chemotherapy. According to the registry, this was intended to minimize potential bias introduced by differential follow-up time between the treatment arms.

4. Trial Timing and Registry Record

February 2006

Trial start

The registry lists February 2006 as the study start.

May 2013

Primary completion

The registry lists May 2013 as the primary completion date.

Completed

Final registry status

The ClinicalTrials.gov record identifies the study as completed and reports statistical results for the primary and secondary outcomes listed below.

5. Primary Endpoint

EndpointRegistry definitionTime frameAnalysis
Overall Survival Time From Randomization to Death From Any Cause within 6.5 years Log-rank test; hazard ratio

The primary endpoint is a classic time-to-event endpoint. Rather than reducing each participant to a simple yes/no outcome, the analysis uses both the event status and the amount of observed follow-up. Participants who remain alive at the end of their observed follow-up can contribute information through their censoring time.

6. Results: Overall Survival

The primary analysis compared chemotherapy with chemotherapy plus bevacizumab in the intention-to-treat population. The reported method was the log-rank test, with a hazard ratio as the effect measure.

Hazard ratio for death

0.81

95% CI: 0.69–0.94   ·   P = 0.0062

Analysis population: intention to treat  ·  Hypothesis type: superiority

Clinical Biostats interpretation

A hazard ratio of 0.81 means that, within the time-to-event framework used for this comparison, the estimated hazard of death in the chemotherapy + bevacizumab group was 81% of the estimated hazard in the chemotherapy group. Equivalently, the estimate corresponds to a 19% lower estimated hazard of death relative to chemotherapy alone.

The hazard ratio does not mean that 19% of patients avoided death, that survival was extended by exactly 19%, or that every participant experienced the same reduction in risk. It is a relative time-to-event effect estimate.

The 95% confidence interval of 0.69–0.94 describes the statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects experienced by individual patients.

The P = 0.0062 value addresses evidence against the null hypothesis under the specified statistical test. It does not measure the magnitude or clinical importance of the treatment effect. Effect size and uncertainty are better described by the hazard ratio and confidence interval.

Because this is a time-to-event analysis, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio framework. The ClinicalTrials.gov record reports the log-rank method and hazard ratio, but do not provide a separate proportional-hazards diagnostic.

Why the log-rank test fits this endpoint

The log-rank test is designed to compare survival distributions between groups while accounting for the timing of events and right-censoring. That makes it fundamentally different from a simple comparison of the proportion of participants who had died by a fixed date.

Conceptual interpretation
HR = 0.81  →  estimated hazard in combination group / estimated hazard in chemotherapy group

The ratio is relative and time-to-event based. It should not be interpreted as an absolute survival probability or as a percentage of participants benefiting.

7. Secondary Endpoint: Overall Survival From First-Line Therapy

A second overall-survival analysis measured survival from the time of first-line therapy rather than from randomization. The registry gives a time frame of within approximately 9.6 years and identifies Kaplan-Meier estimation in the analysis notes.

EndpointChemotherapyChemotherapy + BevacizumabEffect estimateP-value
Overall Survival: Months From Time of First Line Therapy Comparator Intervention HR 0.90
95% CI 0.77–1.05
0.1713

Hazard ratio from first-line therapy

0.90

95% CI: 0.77–1.05   ·   P = 0.1713

Time frame: within approximately 9.6 years  ·  Analysis note: Kaplan Meier Estimate

Clinical Biostats interpretation

The estimated hazard ratio of 0.90 corresponds to an estimated hazard of death that is 90% of that in the chemotherapy group for this analysis, or approximately a 10% lower estimated hazard.

The 95% confidence interval of 0.77–1.05 includes 1.00. In statistical terms, the interval therefore includes the value corresponding to no relative difference in hazard.

The reported P = 0.1713 does not provide conventional evidence against the null hypothesis at commonly used two-sided significance levels. Importantly, the P-value does not establish that the treatments are equivalent and does not quantify the size of any possible effect.

The analysis is also distinct from the primary endpoint because its time origin is first-line therapy, not randomization. Changing the time origin changes the estimand and therefore the interpretation of the resulting hazard ratio.

8. Secondary Endpoint: Progression-Free Survival

Progression-free survival was evaluated as a time-to-event endpoint within 6.5 years. The registry analysis specifies the unstratified intention-to-treat population and reports a log-rank analysis.

EndpointAnalysis populationMethodEffect measureEstimate95% CIP-value
Progression Free Survival: Time to Event Unstratified intention to treat population Log-rank Hazard ratio 0.68 0.59–0.78 <.0001

Hazard ratio for progression-free survival

0.68

95% CI: 0.59–0.78   ·   P <.0001

Analysis population: unstratified intention to treat population

Clinical Biostats interpretation

A hazard ratio of 0.68 corresponds to an estimated hazard of the progression-free-survival event that is 68% of the hazard in the chemotherapy group. Expressed as a relative hazard difference, the estimate corresponds to a 32% lower estimated hazard for the chemotherapy + bevacizumab group.

The 95% confidence interval of 0.59–0.78 describes uncertainty around that estimated relative effect. The interval lies below 1.00, indicating that the reported data are inconsistent with a null hazard ratio of 1.00 under the stated statistical framework.

The reported P <.0001 indicates strong statistical evidence against the null hypothesis for this comparison. It is not a measure of how large the treatment effect is; the hazard ratio and its confidence interval provide that information.

The registry identifies the analysis as unstratified ITT while also listing stratified analysis among the other concepts in the analysis record. Those details should not be silently converted into a claim that a stratified primary PFS model was used. The result reported here follows the specific statistical-analysis record posted on ClinicalTrials.gov for this endpoint.

9. Secondary Endpoint: Response Rate

Response rate was defined as the percentage of participants with best overall response, defined as confirmed complete response (CR) or partial response (PR) according to RECIST criteria. The time frame was within 6.5 years, and the analysis population was participants with measurable disease.

EndpointAnalysis populationMethodEffect measureEstimate95% CIP-value
Response Rate Participants with measurable disease Chi-squared Mean Difference (Final Values) 1.50 -1.5–4.5 0.3113
Response Rate Participants with measurable disease Cochran-Mantel-Haenszel Not reported Not reported Not reported 0.4315
Clinical Biostats interpretation

The registry reports a mean difference of 1.50 with a 95% confidence interval of -1.5 to 4.5 for the chi-squared analysis. Because the reported confidence interval spans 0, it includes the value corresponding to no difference on the reported difference scale.

The reported P = 0.3113 does not provide conventional statistical evidence against the null hypothesis for this comparison. The P-value should not be interpreted as the probability that the treatment effect is zero, nor does it measure the clinical importance of the observed response-rate difference.

A second response-rate analysis used the Cochran-Mantel-Haenszel test and reported P = 0.4315, but the ClinicalTrials.gov record does not provide an effect estimate or confidence interval for that analysis. Accordingly, no additional numerical treatment effect is inferred here.

Why response rate uses a different statistical framework

Response rate is binary: a participant either meets the registered definition of confirmed CR or PR or does not. Unlike overall survival and progression-free survival, the endpoint is not inherently a time-to-event variable. That is why the registry reports categorical-data methods such as the chi-squared test and Cochran-Mantel-Haenszel test.

10. Statistical Methodology

Log-rank test

The log-rank test is the principal reported method for the overall-survival and progression-free-survival analyses. It compares the event-time distributions of the treatment groups while accounting for the ordering and timing of observed events and for censored observations.

Time-to-event comparison
Observed events + risk-set information + censoring  →  comparison of survival distributions

The method uses information accumulated over follow-up rather than treating every participant as simply having experienced or not experienced an event.

Hazard ratio

The hazard ratio is the effect measure reported for the overall-survival and progression-free-survival analyses. A hazard ratio below 1 indicates a lower estimated instantaneous event hazard in the first-listed treatment group relative to the comparator, under the model and analysis framework used.

Interpretation of the reported hazard ratios
OS HR 0.81  →  approximately 19% lower estimated hazard
PFS HR 0.68  →  approximately 32% lower estimated hazard

These are relative hazard interpretations. They are not equivalent to absolute risk reductions, median-survival differences, or probabilities that an individual patient will benefit.

Intention-to-treat analysis

The primary overall-survival analysis was conducted in the intention-to-treat population. ITT analysis preserves the randomized treatment assignment as the basis for comparison. This is important because treatment discontinuation, crossover, and subsequent treatment can occur after randomization without changing which group a participant belongs to for the primary randomized comparison.

Kaplan-Meier estimation

The secondary overall-survival analysis from first-line therapy includes the analysis note Kaplan Meier Estimate. Kaplan-Meier estimation is commonly used to estimate the survival function over time while accounting for right-censoring.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at an event time and ni represents participants at risk immediately before that time.

Chi-squared test

The response-rate analysis used a chi-squared test. This is appropriate to the general structure of a binary categorical endpoint because it compares the distribution of response classifications across treatment groups. The registry reports the effect measure for this analysis as a mean difference in final values.

Cochran-Mantel-Haenszel test

A second response-rate analysis used the Cochran-Mantel-Haenszel test. This family of methods is useful when categorical treatment comparisons are evaluated across strata. The ClinicalTrials.gov record identifies the method but does not provide the stratification variable or an effect estimate for this particular analysis, so those details are not inferred.

Confidence intervals

Confidence intervals quantify statistical uncertainty around an estimated effect. For the hazard-ratio analyses, the interval is on the ratio scale; for the response-rate chi-squared analysis, the registry-reported effect measure is a mean difference and the interval is therefore interpreted on that difference scale.

11. Statistical Methods Explained

Why was a log-rank test used for overall survival?

Overall survival records the time from randomization until death, with participants who have not died by the end of observed follow-up being censored. The log-rank test is specifically designed for comparing such time-to-event distributions. A simple chi-squared test would discard the timing of events and the information contributed by different follow-up durations.

What does an overall-survival hazard ratio of 0.81 mean?

It means the estimated hazard of death for chemotherapy plus bevacizumab was 81% of the estimated hazard for chemotherapy in the reported primary analysis. As a relative interpretation, that corresponds to a 19% lower estimated hazard. It does not mean that 19% more participants survived, nor does it specify an absolute difference in survival probability.

What does the 95% CI of 0.69–0.94 tell us?

The interval describes uncertainty around the estimated hazard ratio. Values inside the interval are compatible with the statistical uncertainty represented by the analysis, subject to the assumptions of the method. Because the interval does not contain 1.00, it does not include the no-difference hazard ratio.

Why is the P-value not an effect size?

A P-value measures how compatible the observed data are with a specified null hypothesis under the statistical test. It depends on both the magnitude of an observed difference and the amount of information available. The hazard ratio and confidence interval are therefore necessary to understand the estimated size and precision of a time-to-event effect.

Why does the response-rate analysis use a chi-squared test?

The registered response endpoint is binary: confirmed CR or PR according to RECIST criteria versus not meeting that response definition. A chi-squared test evaluates whether the categorical response distribution differs between the treatment groups.

Why are there two statistical analyses for response rate?

The ClinicalTrials.gov record reports both a chi-squared analysis and a Cochran-Mantel-Haenszel analysis for response rate. These are distinct statistical procedures. The registry reports P = 0.3113 for the chi-squared analysis and P = 0.4315 for the Cochran-Mantel-Haenszel analysis, but it does not supply an effect estimate and confidence interval for the latter.

Why does the time origin matter?

The primary endpoint starts at randomization, whereas the secondary overall-survival analysis is defined as months from the time of first-line therapy. These are different time origins and therefore represent different statistical estimands. Hazard ratios from the two analyses should not be treated as interchangeable measurements.

12. Primary and Secondary Results at a Glance

EndpointRoleMethodEffect95% CIP-value
Overall Survival: Time From Randomization to Death From Any Cause Primary Log-rank HR 0.81 0.69–0.94 0.0062
Overall Survival: Months From Time of First Line Therapy Secondary Log-rank HR 0.90 0.77–1.05 0.1713
Progression Free Survival: Time to Event Secondary Log-rank HR 0.68 0.59–0.78 <.0001
Response Rate Secondary Chi-squared Mean difference 1.50 -1.5–4.5 0.3113
Response Rate Secondary Cochran-Mantel-Haenszel Not reported Not reported 0.4315

This table illustrates why a clinical-trial results page should not collapse every endpoint into a single significance statement. The endpoint definition, analysis population, time origin, statistical method, effect measure, confidence interval, and P-value all contribute to the interpretation.

13. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. The denominator is described as participants at risk, so the figures are presented exactly in that form rather than converted into percentages.

Safety measureChemotherapyChemotherapy + Bevacizumab
Serious adverse events 137/409 130/401
Serious adverse events: affected / at risk
Chemotherapy
137/409
Chemotherapy + Bevacizumab
130/401

The registry data do not provide a formal comparative P-value, confidence interval, or effect estimate for serious adverse events in the ClinicalTrials.gov record. Accordingly, the safety data are described rather than converted into an inferred treatment comparison.

14. Crossover and Follow-Up Considerations

The registry contains an explicit caveat concerning treatment after discontinuation of chemotherapy. Patients randomized to receive bevacizumab could continue to receive bevacizumab following discontinuation of chemotherapy.

Why this matters statistically: the registry states that continued bevacizumab after chemotherapy discontinuation could minimize potential bias introduced by differential follow-up time between the treatment arms. This is a design and interpretation issue rather than a separate efficacy endpoint, and it should be considered when interpreting the randomized overall-survival comparison.

The distinction between randomized assignment and subsequent treatment exposure is fundamental in clinical-trial analysis. The ITT framework retains the randomized comparison, while post-randomization treatment patterns can influence the observed outcome trajectories.

15. Multiplicity and Multiple Analyses

The ClinicalTrials.gov record identifies one registered primary endpoint and five statistical analyses in total: one primary analysis and four secondary analyses. The primary endpoint is overall survival from randomization to death from any cause.

Analysis familyRegistry roleNumber of registry-reported analyses
Primary endpoint analysisOverall survival from randomization1
Secondary analysesOverall survival from first-line therapy, progression-free survival, and response rate analyses4
Total statistical analyses postedPrimary + secondary analyses5

The ClinicalTrials.gov record identifies the hypothesis type as superiority, but they do not provide an alpha-allocation scheme, hierarchical testing procedure, multiplicity-adjustment procedure, or interim-analysis plan. Those features are therefore not inferred.

Interpretation principle: multiple reported endpoints should not automatically be treated as if every P-value had the same confirmatory status as the prespecified primary endpoint. The registry's endpoint role and analysis designation should remain explicit.

16. Stratified Analysis and the Cochran-Mantel-Haenszel Method

Stratification appears in the registry's posted analyses through the Cochran-Mantel-Haenszel test used for one response-rate comparison. The specific strata used for each analysis are not provided in the ClinicalTrials.gov record.

Why stratification can help

Stratified methods can account for categorical factors that define clinically or statistically meaningful groups, reducing the influence of imbalance across those strata on an overall comparison.

What cannot be inferred

The ClinicalTrials.gov record does not identify the specific stratification variables or provide stratum-specific estimates. Those details should therefore not be reconstructed from general knowledge of the trial.

17. Missing Data, Censoring, and Analysis Assumptions

Time-to-event analyses inherently involve censoring because some participants may remain free of the event when their available follow-up ends. The log-rank analysis and Kaplan-Meier framework are designed to use the available event and follow-up information rather than requiring an observed event for every participant.

The ClinicalTrials.gov record does not specify a missing-data imputation method, a prespecified rule for handling missing response assessments, or a detailed censoring algorithm. Those methods are therefore not added to this analysis.

Similarly, the ClinicalTrials.gov record identifies hazard ratios and log-rank tests but do not provide a formal assessment of the proportional-hazards assumption. A hazard ratio should therefore be interpreted as the reported relative time-to-event measure, without claiming that proportional hazards were empirically demonstrated.

18. What the Hazard Ratio Does — and Does Not — Mean

Primary overall-survival interpretation

The primary OS hazard ratio of 0.81 indicates an estimated instantaneous hazard of death that was approximately 81% as large in the chemotherapy + bevacizumab group as in the chemotherapy group. As a relative hazard interpretation, this corresponds to approximately a 19% lower estimated hazard.

It does not mean that 19% of participants were saved, that each patient lived 19% longer, or that the absolute probability of survival increased by 19 percentage points.

Precision

The 95% CI of 0.69–0.94 shows the statistical uncertainty surrounding the reported estimate. It is not an interval containing the outcomes that individual patients might experience.

P-value

The primary P-value of 0.0062 quantifies evidence against the specified null hypothesis under the log-rank testing framework. It does not quantify the magnitude of the treatment effect. A smaller P-value is not automatically a larger treatment effect.

19. Comparing the Time-to-Event Results

Time-to-event endpointHazard ratio95% CIP-valueRelative interpretation
Overall survival from randomization 0.81 0.69–0.94 0.0062 Approximately 19% lower estimated hazard
Overall survival from first-line therapy 0.90 0.77–1.05 0.1713 Approximately 10% lower estimated hazard
Progression-free survival 0.68 0.59–0.78 <.0001 Approximately 32% lower estimated hazard

These estimates should be read as separate analyses rather than as interchangeable measures of a single treatment effect. Overall survival from randomization, overall survival from first-line therapy, and progression-free survival have different endpoint definitions and, in the first two cases, different time origins.

20. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary randomized overall-survival comparison produced a hazard ratio of 0.81 with a 95% CI of 0.69–0.94 and P = 0.0062. The secondary progression-free-survival analysis produced a hazard ratio of 0.68 with a 95% CI of 0.59–0.78 and P <.0001.

What remains separate

Response rate was analyzed as a categorical endpoint, with the registry-reported chi-squared analysis reporting a mean difference of 1.50 and P = 0.3113. Safety is separately described by serious adverse events rather than combined into an efficacy statistic.

This distinction is important because a clinical trial contains multiple estimands. Survival endpoints address time until an event, response rate addresses a binary tumor-response classification, and safety endpoints address adverse outcomes. No single statistic summarizes all of these dimensions.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

ML18147 is a useful teaching case because its registry results bring together several fundamental clinical-trial methods: randomized treatment comparison, intention-to-treat analysis, time-to-event endpoints, log-rank testing, hazard ratios, confidence intervals, Kaplan-Meier estimation, categorical response analysis, and stratified methods.

ConceptHow it appears in ML18147
RandomizationThe trial used randomized allocation in a parallel-group phase 3 design.
Intention-to-treat analysisThe primary overall-survival analysis used the intention-to-treat population.
Time-to-event endpointOverall survival was defined as time from randomization to death from any cause.
Kaplan-Meier estimationThe secondary overall-survival analysis from first-line therapy includes Kaplan-Meier estimation.
Log-rank testingLog-rank analysis was reported for overall survival and progression-free survival.
Hazard ratioHazard ratios were reported for the time-to-event comparisons.
Confidence interval95% two-sided confidence intervals were reported for the primary OS, secondary OS, and PFS hazard ratios.
Chi-squared testResponse rate was analyzed using a chi-squared test.
Cochran-Mantel-Haenszel testA second response-rate analysis used the Cochran-Mantel-Haenszel method.
Stratified analysisStratified analysis is identified among the concepts associated with the PFS analysis.
Safety analysisSerious adverse events are reported by treatment arm as affected/at-risk counts.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Connect the trial's endpoints and methods to deeper statistical tutorials and practical analysis tools.

26. Record Summary

ML18147 provides a useful example of how different endpoint types require different statistical methods. The primary endpoint was overall survival from randomization to death from any cause and was analyzed using a log-rank test in the intention-to-treat population, producing a hazard ratio of 0.81 with a 95% CI of 0.69–0.94 and P = 0.0062. Progression-free survival was also evaluated with a log-rank analysis and produced a hazard ratio of 0.68 with a 95% CI of 0.59–0.78 and P <.0001.

The secondary overall-survival analysis used a different time origin and produced an HR of 0.90 with a 95% CI of 0.77–1.05 and P = 0.1713. Response rate was analyzed as a categorical endpoint using both chi-squared and Cochran-Mantel-Haenszel methods, with the registry-reported chi-squared analysis reporting a mean difference of 1.50 and a 95% CI of -1.5–4.5.

The statistical lesson is that the interpretation of a clinical-trial result depends on more than whether a P-value crosses a threshold. The endpoint definition, time origin, analysis population, statistical method, effect measure, confidence interval, censoring, and post-randomization treatment all determine what a reported number actually means.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. The goal is to explain the statistical structure of the trial without adding estimates, analyses, or design features that are not supported by the ClinicalTrials.gov record.