This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
ML18147 was a randomized, open-label, parallel-group phase 3 trial in metastatic colorectal cancer. The study enrolled 820 participants and compared chemotherapy alone with chemotherapy plus bevacizumab.
| Feature | ML18147 |
|---|---|
| Trial name | ML18147 |
| ClinicalTrials.gov identifier | NCT00700102 |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Colorectal Cancer |
| Enrollment | 820 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Interventions | Chemotherapy; Bevacizumab |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | Industry |
| Primary endpoint type | Time-to-event |
2. Clinical Question
The central statistical question was whether chemotherapy plus bevacizumab differed from chemotherapy alone with respect to overall survival, defined as time from randomization to death from any cause.
Population
Participants enrolled in the phase 3 study of metastatic colorectal cancer.
Intervention
Chemotherapy plus bevacizumab.
Comparator
Chemotherapy.
Primary question
Does chemotherapy plus bevacizumab improve overall survival relative to chemotherapy alone?
The registered primary hypothesis type was superiority. That matters statistically: the objective was to determine whether the randomized groups differed in the specified direction of the treatment comparison rather than to demonstrate that the two regimens were sufficiently similar.
3. Trial Design
Chemotherapy
- Chemotherapy was the comparator intervention.
- Serious adverse events: 137 affected participants among 409 at risk.
Chemotherapy + Bevacizumab
- Chemotherapy was administered with bevacizumab.
- Serious adverse events: 130 affected participants among 401 at risk.
4. Trial Timing and Registry Record
Trial start
The registry lists February 2006 as the study start.
Primary completion
The registry lists May 2013 as the primary completion date.
Final registry status
The ClinicalTrials.gov record identifies the study as completed and reports statistical results for the primary and secondary outcomes listed below.
5. Primary Endpoint
| Endpoint | Registry definition | Time frame | Analysis |
|---|---|---|---|
| Overall Survival | Time From Randomization to Death From Any Cause | within 6.5 years | Log-rank test; hazard ratio |
The primary endpoint is a classic time-to-event endpoint. Rather than reducing each participant to a simple yes/no outcome, the analysis uses both the event status and the amount of observed follow-up. Participants who remain alive at the end of their observed follow-up can contribute information through their censoring time.
6. Results: Overall Survival
The primary analysis compared chemotherapy with chemotherapy plus bevacizumab in the intention-to-treat population. The reported method was the log-rank test, with a hazard ratio as the effect measure.
Hazard ratio for death
95% CI: 0.69–0.94 · P = 0.0062
Analysis population: intention to treat · Hypothesis type: superiority
A hazard ratio of 0.81 means that, within the time-to-event framework used for this comparison, the estimated hazard of death in the chemotherapy + bevacizumab group was 81% of the estimated hazard in the chemotherapy group. Equivalently, the estimate corresponds to a 19% lower estimated hazard of death relative to chemotherapy alone.
The hazard ratio does not mean that 19% of patients avoided death, that survival was extended by exactly 19%, or that every participant experienced the same reduction in risk. It is a relative time-to-event effect estimate.
The 95% confidence interval of 0.69–0.94 describes the statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects experienced by individual patients.
The P = 0.0062 value addresses evidence against the null hypothesis under the specified statistical test. It does not measure the magnitude or clinical importance of the treatment effect. Effect size and uncertainty are better described by the hazard ratio and confidence interval.
Because this is a time-to-event analysis, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio framework. The ClinicalTrials.gov record reports the log-rank method and hazard ratio, but do not provide a separate proportional-hazards diagnostic.
Why the log-rank test fits this endpoint
The log-rank test is designed to compare survival distributions between groups while accounting for the timing of events and right-censoring. That makes it fundamentally different from a simple comparison of the proportion of participants who had died by a fixed date.
The ratio is relative and time-to-event based. It should not be interpreted as an absolute survival probability or as a percentage of participants benefiting.
7. Secondary Endpoint: Overall Survival From First-Line Therapy
A second overall-survival analysis measured survival from the time of first-line therapy rather than from randomization. The registry gives a time frame of within approximately 9.6 years and identifies Kaplan-Meier estimation in the analysis notes.
| Endpoint | Chemotherapy | Chemotherapy + Bevacizumab | Effect estimate | P-value |
|---|---|---|---|---|
| Overall Survival: Months From Time of First Line Therapy | Comparator | Intervention | HR 0.90 95% CI 0.77–1.05 |
0.1713 |
Hazard ratio from first-line therapy
95% CI: 0.77–1.05 · P = 0.1713
Time frame: within approximately 9.6 years · Analysis note: Kaplan Meier Estimate
The estimated hazard ratio of 0.90 corresponds to an estimated hazard of death that is 90% of that in the chemotherapy group for this analysis, or approximately a 10% lower estimated hazard.
The 95% confidence interval of 0.77–1.05 includes 1.00. In statistical terms, the interval therefore includes the value corresponding to no relative difference in hazard.
The reported P = 0.1713 does not provide conventional evidence against the null hypothesis at commonly used two-sided significance levels. Importantly, the P-value does not establish that the treatments are equivalent and does not quantify the size of any possible effect.
The analysis is also distinct from the primary endpoint because its time origin is first-line therapy, not randomization. Changing the time origin changes the estimand and therefore the interpretation of the resulting hazard ratio.
8. Secondary Endpoint: Progression-Free Survival
Progression-free survival was evaluated as a time-to-event endpoint within 6.5 years. The registry analysis specifies the unstratified intention-to-treat population and reports a log-rank analysis.
| Endpoint | Analysis population | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Progression Free Survival: Time to Event | Unstratified intention to treat population | Log-rank | Hazard ratio | 0.68 | 0.59–0.78 | <.0001 |
Hazard ratio for progression-free survival
95% CI: 0.59–0.78 · P <.0001
Analysis population: unstratified intention to treat population
A hazard ratio of 0.68 corresponds to an estimated hazard of the progression-free-survival event that is 68% of the hazard in the chemotherapy group. Expressed as a relative hazard difference, the estimate corresponds to a 32% lower estimated hazard for the chemotherapy + bevacizumab group.
The 95% confidence interval of 0.59–0.78 describes uncertainty around that estimated relative effect. The interval lies below 1.00, indicating that the reported data are inconsistent with a null hazard ratio of 1.00 under the stated statistical framework.
The reported P <.0001 indicates strong statistical evidence against the null hypothesis for this comparison. It is not a measure of how large the treatment effect is; the hazard ratio and its confidence interval provide that information.
The registry identifies the analysis as unstratified ITT while also listing stratified analysis among the other concepts in the analysis record. Those details should not be silently converted into a claim that a stratified primary PFS model was used. The result reported here follows the specific statistical-analysis record posted on ClinicalTrials.gov for this endpoint.
9. Secondary Endpoint: Response Rate
Response rate was defined as the percentage of participants with best overall response, defined as confirmed complete response (CR) or partial response (PR) according to RECIST criteria. The time frame was within 6.5 years, and the analysis population was participants with measurable disease.
| Endpoint | Analysis population | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Response Rate | Participants with measurable disease | Chi-squared | Mean Difference (Final Values) | 1.50 | -1.5–4.5 | 0.3113 |
| Response Rate | Participants with measurable disease | Cochran-Mantel-Haenszel | Not reported | Not reported | Not reported | 0.4315 |
The registry reports a mean difference of 1.50 with a 95% confidence interval of -1.5 to 4.5 for the chi-squared analysis. Because the reported confidence interval spans 0, it includes the value corresponding to no difference on the reported difference scale.
The reported P = 0.3113 does not provide conventional statistical evidence against the null hypothesis for this comparison. The P-value should not be interpreted as the probability that the treatment effect is zero, nor does it measure the clinical importance of the observed response-rate difference.
A second response-rate analysis used the Cochran-Mantel-Haenszel test and reported P = 0.4315, but the ClinicalTrials.gov record does not provide an effect estimate or confidence interval for that analysis. Accordingly, no additional numerical treatment effect is inferred here.
Why response rate uses a different statistical framework
Response rate is binary: a participant either meets the registered definition of confirmed CR or PR or does not. Unlike overall survival and progression-free survival, the endpoint is not inherently a time-to-event variable. That is why the registry reports categorical-data methods such as the chi-squared test and Cochran-Mantel-Haenszel test.
10. Statistical Methodology
Log-rank test
The log-rank test is the principal reported method for the overall-survival and progression-free-survival analyses. It compares the event-time distributions of the treatment groups while accounting for the ordering and timing of observed events and for censored observations.
The method uses information accumulated over follow-up rather than treating every participant as simply having experienced or not experienced an event.
Hazard ratio
The hazard ratio is the effect measure reported for the overall-survival and progression-free-survival analyses. A hazard ratio below 1 indicates a lower estimated instantaneous event hazard in the first-listed treatment group relative to the comparator, under the model and analysis framework used.
PFS HR 0.68 → approximately 32% lower estimated hazard
These are relative hazard interpretations. They are not equivalent to absolute risk reductions, median-survival differences, or probabilities that an individual patient will benefit.
Intention-to-treat analysis
The primary overall-survival analysis was conducted in the intention-to-treat population. ITT analysis preserves the randomized treatment assignment as the basis for comparison. This is important because treatment discontinuation, crossover, and subsequent treatment can occur after randomization without changing which group a participant belongs to for the primary randomized comparison.
Kaplan-Meier estimation
The secondary overall-survival analysis from first-line therapy includes the analysis note Kaplan Meier Estimate. Kaplan-Meier estimation is commonly used to estimate the survival function over time while accounting for right-censoring.
Here, di represents events at an event time and ni represents participants at risk immediately before that time.
Chi-squared test
The response-rate analysis used a chi-squared test. This is appropriate to the general structure of a binary categorical endpoint because it compares the distribution of response classifications across treatment groups. The registry reports the effect measure for this analysis as a mean difference in final values.
Cochran-Mantel-Haenszel test
A second response-rate analysis used the Cochran-Mantel-Haenszel test. This family of methods is useful when categorical treatment comparisons are evaluated across strata. The ClinicalTrials.gov record identifies the method but does not provide the stratification variable or an effect estimate for this particular analysis, so those details are not inferred.
Confidence intervals
Confidence intervals quantify statistical uncertainty around an estimated effect. For the hazard-ratio analyses, the interval is on the ratio scale; for the response-rate chi-squared analysis, the registry-reported effect measure is a mean difference and the interval is therefore interpreted on that difference scale.
11. Statistical Methods Explained
Why was a log-rank test used for overall survival?
Overall survival records the time from randomization until death, with participants who have not died by the end of observed follow-up being censored. The log-rank test is specifically designed for comparing such time-to-event distributions. A simple chi-squared test would discard the timing of events and the information contributed by different follow-up durations.
What does an overall-survival hazard ratio of 0.81 mean?
It means the estimated hazard of death for chemotherapy plus bevacizumab was 81% of the estimated hazard for chemotherapy in the reported primary analysis. As a relative interpretation, that corresponds to a 19% lower estimated hazard. It does not mean that 19% more participants survived, nor does it specify an absolute difference in survival probability.
What does the 95% CI of 0.69–0.94 tell us?
The interval describes uncertainty around the estimated hazard ratio. Values inside the interval are compatible with the statistical uncertainty represented by the analysis, subject to the assumptions of the method. Because the interval does not contain 1.00, it does not include the no-difference hazard ratio.
Why is the P-value not an effect size?
A P-value measures how compatible the observed data are with a specified null hypothesis under the statistical test. It depends on both the magnitude of an observed difference and the amount of information available. The hazard ratio and confidence interval are therefore necessary to understand the estimated size and precision of a time-to-event effect.
Why does the response-rate analysis use a chi-squared test?
The registered response endpoint is binary: confirmed CR or PR according to RECIST criteria versus not meeting that response definition. A chi-squared test evaluates whether the categorical response distribution differs between the treatment groups.
Why are there two statistical analyses for response rate?
The ClinicalTrials.gov record reports both a chi-squared analysis and a Cochran-Mantel-Haenszel analysis for response rate. These are distinct statistical procedures. The registry reports P = 0.3113 for the chi-squared analysis and P = 0.4315 for the Cochran-Mantel-Haenszel analysis, but it does not supply an effect estimate and confidence interval for the latter.
Why does the time origin matter?
The primary endpoint starts at randomization, whereas the secondary overall-survival analysis is defined as months from the time of first-line therapy. These are different time origins and therefore represent different statistical estimands. Hazard ratios from the two analyses should not be treated as interchangeable measurements.
12. Primary and Secondary Results at a Glance
| Endpoint | Role | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| Overall Survival: Time From Randomization to Death From Any Cause | Primary | Log-rank | HR 0.81 | 0.69–0.94 | 0.0062 |
| Overall Survival: Months From Time of First Line Therapy | Secondary | Log-rank | HR 0.90 | 0.77–1.05 | 0.1713 |
| Progression Free Survival: Time to Event | Secondary | Log-rank | HR 0.68 | 0.59–0.78 | <.0001 |
| Response Rate | Secondary | Chi-squared | Mean difference 1.50 | -1.5–4.5 | 0.3113 |
| Response Rate | Secondary | Cochran-Mantel-Haenszel | Not reported | Not reported | 0.4315 |
This table illustrates why a clinical-trial results page should not collapse every endpoint into a single significance statement. The endpoint definition, analysis population, time origin, statistical method, effect measure, confidence interval, and P-value all contribute to the interpretation.
13. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. The denominator is described as participants at risk, so the figures are presented exactly in that form rather than converted into percentages.
| Safety measure | Chemotherapy | Chemotherapy + Bevacizumab |
|---|---|---|
| Serious adverse events | 137/409 | 130/401 |
The registry data do not provide a formal comparative P-value, confidence interval, or effect estimate for serious adverse events in the ClinicalTrials.gov record. Accordingly, the safety data are described rather than converted into an inferred treatment comparison.
14. Crossover and Follow-Up Considerations
The registry contains an explicit caveat concerning treatment after discontinuation of chemotherapy. Patients randomized to receive bevacizumab could continue to receive bevacizumab following discontinuation of chemotherapy.
The distinction between randomized assignment and subsequent treatment exposure is fundamental in clinical-trial analysis. The ITT framework retains the randomized comparison, while post-randomization treatment patterns can influence the observed outcome trajectories.
15. Multiplicity and Multiple Analyses
The ClinicalTrials.gov record identifies one registered primary endpoint and five statistical analyses in total: one primary analysis and four secondary analyses. The primary endpoint is overall survival from randomization to death from any cause.
| Analysis family | Registry role | Number of registry-reported analyses |
|---|---|---|
| Primary endpoint analysis | Overall survival from randomization | 1 |
| Secondary analyses | Overall survival from first-line therapy, progression-free survival, and response rate analyses | 4 |
| Total statistical analyses posted | Primary + secondary analyses | 5 |
The ClinicalTrials.gov record identifies the hypothesis type as superiority, but they do not provide an alpha-allocation scheme, hierarchical testing procedure, multiplicity-adjustment procedure, or interim-analysis plan. Those features are therefore not inferred.
16. Stratified Analysis and the Cochran-Mantel-Haenszel Method
Stratification appears in the registry's posted analyses through the Cochran-Mantel-Haenszel test used for one response-rate comparison. The specific strata used for each analysis are not provided in the ClinicalTrials.gov record.
Why stratification can help
Stratified methods can account for categorical factors that define clinically or statistically meaningful groups, reducing the influence of imbalance across those strata on an overall comparison.
What cannot be inferred
The ClinicalTrials.gov record does not identify the specific stratification variables or provide stratum-specific estimates. Those details should therefore not be reconstructed from general knowledge of the trial.
17. Missing Data, Censoring, and Analysis Assumptions
Time-to-event analyses inherently involve censoring because some participants may remain free of the event when their available follow-up ends. The log-rank analysis and Kaplan-Meier framework are designed to use the available event and follow-up information rather than requiring an observed event for every participant.
The ClinicalTrials.gov record does not specify a missing-data imputation method, a prespecified rule for handling missing response assessments, or a detailed censoring algorithm. Those methods are therefore not added to this analysis.
Similarly, the ClinicalTrials.gov record identifies hazard ratios and log-rank tests but do not provide a formal assessment of the proportional-hazards assumption. A hazard ratio should therefore be interpreted as the reported relative time-to-event measure, without claiming that proportional hazards were empirically demonstrated.
18. What the Hazard Ratio Does — and Does Not — Mean
The primary OS hazard ratio of 0.81 indicates an estimated instantaneous hazard of death that was approximately 81% as large in the chemotherapy + bevacizumab group as in the chemotherapy group. As a relative hazard interpretation, this corresponds to approximately a 19% lower estimated hazard.
It does not mean that 19% of participants were saved, that each patient lived 19% longer, or that the absolute probability of survival increased by 19 percentage points.
The 95% CI of 0.69–0.94 shows the statistical uncertainty surrounding the reported estimate. It is not an interval containing the outcomes that individual patients might experience.
The primary P-value of 0.0062 quantifies evidence against the specified null hypothesis under the log-rank testing framework. It does not quantify the magnitude of the treatment effect. A smaller P-value is not automatically a larger treatment effect.
19. Comparing the Time-to-Event Results
| Time-to-event endpoint | Hazard ratio | 95% CI | P-value | Relative interpretation |
|---|---|---|---|---|
| Overall survival from randomization | 0.81 | 0.69–0.94 | 0.0062 | Approximately 19% lower estimated hazard |
| Overall survival from first-line therapy | 0.90 | 0.77–1.05 | 0.1713 | Approximately 10% lower estimated hazard |
| Progression-free survival | 0.68 | 0.59–0.78 | <.0001 | Approximately 32% lower estimated hazard |
These estimates should be read as separate analyses rather than as interchangeable measures of a single treatment effect. Overall survival from randomization, overall survival from first-line therapy, and progression-free survival have different endpoint definitions and, in the first two cases, different time origins.
20. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary randomized overall-survival comparison produced a hazard ratio of 0.81 with a 95% CI of 0.69–0.94 and P = 0.0062. The secondary progression-free-survival analysis produced a hazard ratio of 0.68 with a 95% CI of 0.59–0.78 and P <.0001.
What remains separate
Response rate was analyzed as a categorical endpoint, with the registry-reported chi-squared analysis reporting a mean difference of 1.50 and P = 0.3113. Safety is separately described by serious adverse events rather than combined into an efficacy statistic.
This distinction is important because a clinical trial contains multiple estimands. Survival endpoints address time until an event, response rate addresses a binary tumor-response classification, and safety endpoints address adverse outcomes. No single statistic summarizes all of these dimensions.
21. Important Limitations and Interpretation Issues
- Crossover and treatment continuation: the registry specifically notes that patients randomized to bevacizumab could continue bevacizumab after discontinuation of chemotherapy, with the stated purpose of minimizing potential bias from differential follow-up time.
- Different time origins: the primary OS endpoint starts at randomization, whereas one secondary OS analysis is measured from first-line therapy. The two hazard ratios therefore represent different analyses.
- Analysis-population differences: the primary OS analysis used intention to treat, while the PFS analysis is explicitly described as using the unstratified intention-to-treat population and the response analysis used participants with measurable disease.
- Response-rate effect measure: the registry reports a mean difference for the chi-squared response-rate analysis rather than a risk ratio or odds ratio. The interpretation should therefore remain on the reported difference scale.
- Incomplete method detail: the ClinicalTrials.gov record does not provide a detailed SAP, exact stratification variables, missing-data rules, censoring rules, or a proportional-hazards diagnostic.
- No inferred median survival: the ClinicalTrials.gov record provides hazard ratios and confidence intervals but do not provide median survival estimates, so none are reported here.
- No inferred subgroup results: the ClinicalTrials.gov record does not provide subgroup-specific effect estimates, so no subgroup conclusions are drawn.
- Safety comparison: serious adverse events are reported as affected/at-risk counts, but no formal comparative effect estimate or P-value is posted on ClinicalTrials.gov for that endpoint.
22. Why This Trial Matters Statistically
ML18147 is a useful teaching case because its registry results bring together several fundamental clinical-trial methods: randomized treatment comparison, intention-to-treat analysis, time-to-event endpoints, log-rank testing, hazard ratios, confidence intervals, Kaplan-Meier estimation, categorical response analysis, and stratified methods.
| Concept | How it appears in ML18147 |
|---|---|
| Randomization | The trial used randomized allocation in a parallel-group phase 3 design. |
| Intention-to-treat analysis | The primary overall-survival analysis used the intention-to-treat population. |
| Time-to-event endpoint | Overall survival was defined as time from randomization to death from any cause. |
| Kaplan-Meier estimation | The secondary overall-survival analysis from first-line therapy includes Kaplan-Meier estimation. |
| Log-rank testing | Log-rank analysis was reported for overall survival and progression-free survival. |
| Hazard ratio | Hazard ratios were reported for the time-to-event comparisons. |
| Confidence interval | 95% two-sided confidence intervals were reported for the primary OS, secondary OS, and PFS hazard ratios. |
| Chi-squared test | Response rate was analyzed using a chi-squared test. |
| Cochran-Mantel-Haenszel test | A second response-rate analysis used the Cochran-Mantel-Haenszel method. |
| Stratified analysis | Stratified analysis is identified among the concepts associated with the PFS analysis. |
| Safety analysis | Serious adverse events are reported by treatment arm as affected/at-risk counts. |
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT00700102 — ML18147.
- PubMed: PMID 38186268.
- PubMed: PMID 23852309.
- PubMed: PMID 23168366.
Continue through the Clinical Biostats statistical pathway
Connect the trial's endpoints and methods to deeper statistical tutorials and practical analysis tools.
26. Record Summary
ML18147 provides a useful example of how different endpoint types require different statistical methods. The primary endpoint was overall survival from randomization to death from any cause and was analyzed using a log-rank test in the intention-to-treat population, producing a hazard ratio of 0.81 with a 95% CI of 0.69–0.94 and P = 0.0062. Progression-free survival was also evaluated with a log-rank analysis and produced a hazard ratio of 0.68 with a 95% CI of 0.59–0.78 and P <.0001.
The secondary overall-survival analysis used a different time origin and produced an HR of 0.90 with a 95% CI of 0.77–1.05 and P = 0.1713. Response rate was analyzed as a categorical endpoint using both chi-squared and Cochran-Mantel-Haenszel methods, with the registry-reported chi-squared analysis reporting a mean difference of 1.50 and a 95% CI of -1.5–4.5.
The statistical lesson is that the interpretation of a clinical-trial result depends on more than whether a P-value crosses a threshold. The endpoint definition, time origin, analysis population, statistical method, effect measure, confidence interval, censoring, and post-randomization treatment all determine what a reported number actually means.