This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
MONARCH 2 was a randomized, double-blind, parallel phase 3 trial evaluating abemaciclib combined with fulvestrant against placebo combined with fulvestrant in women with hormone receptor positive, HER2 negative breast cancer. The registered primary endpoint was progression-free survival (PFS), a time-to-event endpoint analyzed using a stratified log-rank test.
| Feature | MONARCH 2 |
|---|---|
| Trial name | MONARCH 2 |
| Phase | Phase 3 |
| Population | Women with hormone receptor positive, HER2 negative breast cancer |
| Design | Randomized, double-blind, parallel |
| Primary purpose | Treatment |
| Enrollment | 669 |
| Arms | 2 |
| Primary endpoint | Progression-Free Survival (PFS) |
| Primary endpoint type | Time-to-event |
| Primary analysis | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Trial status | Active, not recruiting |
| Lead sponsor | Eli Lilly and Company |
| ClinicalTrials.gov | NCT02107703 |
2. Clinical Question
The central statistical question was whether treatment with abemaciclib plus fulvestrant produced a different progression-free survival experience from placebo plus fulvestrant in women with hormone receptor positive, HER2 negative breast cancer.
Population
Women with hormone receptor positive, HER2 negative breast cancer.
Intervention
Abemaciclib combined with fulvestrant.
Comparator
Placebo combined with fulvestrant.
Primary question
How does progression-free survival compare between the randomized treatment groups?
Because PFS is a time-to-event endpoint, the statistical question is not simply whether more participants eventually progressed. It concerns the distribution of time from randomization to the first qualifying event, with participants who have not experienced an event by their available follow-up contributing censored information.
3. Trial Design
Abemaciclib + Fulvestrant
- Abemaciclib
- Fulvestrant
Placebo + Fulvestrant
- Placebo
- Fulvestrant
The randomized and double-blind structure is important statistically because treatment assignment is established before the outcome is observed, while blinding is intended to reduce the potential for knowledge of assignment to influence trial conduct or assessment.
4. Endpoints
| Endpoint | Registered definition / time frame | Type |
|---|---|---|
| Progression-Free Survival (PFS) | From Date of Randomization until Disease Progression or Death Due to Any Cause (Up To 31 Months) | Time-to-event |
The registry definition specifies PFS as the time from randomization to the first evidence of disease progression according to RECIST v1.1 or death from any cause. The registry text further defines progressive disease as at least a 20% increase in the sum of the diameters of target lesions, with the smallest sum on study as the reference and an absolute increase of at least 5 mm, followed by the registry-reported definition text.
The registered time frame is therefore an explicitly time-based endpoint: participants contribute follow-up from randomization until progression, death, or censoring according to the applicable analysis rules.
5. Statistical Methodology
Primary analysis: stratified log-rank test
The registry reports a log-rank test as the primary analysis method. The analysis notes specify that the log-rank test was stratified by endocrine sensitivity and nature of disease as determined by the interactive web response system (IWRS).
The log-rank procedure compares the observed pattern of events over follow-up between treatment groups rather than comparing only the proportion of participants who have progressed at one fixed time point.
Hazard ratio as the effect measure
The registry reports the treatment effect as a hazard ratio (HR). The primary estimate was 0.553 with a two-sided 95% confidence interval of 0.449 to 0.681.
An HR below 1 indicates a lower estimated instantaneous event rate in the abemaciclib group under the time-to-event model. It is not an absolute risk difference, a probability of benefit for an individual participant, or a statement that every participant experiences the same relative change.
Analysis population and censoring
The posted analysis population was all randomized participants. The analysis notes identify 224 censored participants in the abemaciclib group. Censoring is a fundamental feature of time-to-event analysis: a participant can contribute information up to the point at which their event status is no longer observed according to the analysis definition.
The important distinction is that censoring does not mean that a participant is treated as having experienced progression at the censoring time. Instead, the participant's observed follow-up contributes information to the survival analysis up to that point.
Planned event information and statistical power
The final analysis was planned at 378 PFS events. The registry states that this would provide approximately 90% power assuming a hazard ratio of 0.703 at a one-sided α of 0.025.
Planning assumption
The HR of 0.703 was a design assumption used for planning the event-driven analysis; it is not the observed treatment-effect estimate.
Observed estimate
The reported primary analysis estimate was HR 0.553 with a two-sided 95% CI of 0.449–0.681.
6. Results
Primary Endpoint: Progression-Free Survival
Hazard ratio for progression or death
95% CI: 0.449–0.681 · P < 0.0000001
Analysis: stratified log-rank test · Population: all randomized participants
| Endpoint | Comparison | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|
| Progression-Free Survival | Abemaciclib + Fulvestrant vs Placebo + Fulvestrant | HR 0.553 | 0.449–0.681 | <0.0000001 |
The reported estimate of 0.553 means that the estimated instantaneous rate of the PFS event—disease progression or death—was 55.3% of the corresponding rate in the placebo-plus-fulvestrant group under the reported analysis framework. Expressed as a simple relative-hazard interpretation, 1 − 0.553 = 0.447, so the estimate corresponds to an approximately 44.7% lower estimated hazard.
What the estimate means: The HR of 0.553 is a relative time-to-event measure. It summarizes the estimated difference in the instantaneous rate of progression or death between the randomized groups over the analyzed follow-up.
What it does not mean: It does not mean that 44.7% of participants avoided progression, that an individual participant has exactly a 44.7% lower probability of progression, or that median PFS was reduced or increased by 44.7%. No median PFS is reported in the ClinicalTrials.gov record used for this page.
Confidence interval: The two-sided 95% CI of 0.449–0.681 describes uncertainty around the estimated hazard ratio under the statistical framework. It is an interval for the treatment-effect parameter, not a range containing 95% of individual patient outcomes.
P-value: The P-value of <0.0000001 quantifies the incompatibility of the observed data with the null hypothesis under the specified testing framework. It does not measure the magnitude of the treatment effect. The magnitude is described by the HR and its confidence interval.
Important cautions: The analysis is time-to-event based and includes censoring. The registry identifies a stratified analysis and a one-sided α of 0.025 for the event-driven design. The reported confidence interval is two-sided at 95%, so its interpretation should not be confused with the one-sided design alpha. As with hazard-ratio methods generally, interpretation of a single HR is most direct when the relative hazards are reasonably represented by the model over time.
7. Understanding the PFS Result
The PFS analysis combines three pieces of information that should be read together: the effect estimate, the confidence interval, and the hypothesis-test result.
Effect size
HR 0.553 describes the relative treatment effect on the hazard scale. Its distance from 1 is more informative about effect magnitude than the P-value alone.
Precision
The 95% CI of 0.449–0.681 provides a range of values reflecting statistical uncertainty around the estimated HR.
Evidence against the null
The reported P-value of <0.0000001 indicates very strong statistical evidence against the null hypothesis used for the reported test.
Time-to-event context
PFS incorporates when events occur and also uses information from participants who are censored rather than simply classifying everyone as event or no event.
A useful statistical distinction is that effect size and statistical evidence are not interchangeable. With a sufficiently large amount of event information, even a relatively modest effect can produce a small P-value. Conversely, an important estimated effect can have a wide confidence interval when information is limited. The MONARCH 2 registry result supplies all three quantities, allowing them to be interpreted together.
8. Statistical Methods Explained
Why was a log-rank test used?
PFS is a time-to-event endpoint. Participants can experience progression or death at different times, and some participants can be censored. A log-rank test is designed to compare survival or event-time distributions between groups while incorporating the timing of observed events rather than reducing follow-up to a single binary outcome.
What does an HR of 0.553 mean?
An HR of 0.553 means that the estimated instantaneous event rate in the abemaciclib-plus-fulvestrant group was 0.553 times that in the placebo-plus-fulvestrant group under the reported analysis. The corresponding arithmetic expression, 1 − 0.553, is 0.447, or 44.7%. That does not mean that exactly 44.7% of participants benefited.
Why is the confidence interval important?
The point estimate is only one estimate of the underlying treatment effect. The 95% CI of 0.449–0.681 communicates statistical precision around the HR. A narrower interval generally indicates greater precision than a wider interval, all else equal. The interval should not be interpreted as a distribution of individual treatment effects.
Why doesn't the P-value measure effect size?
The P-value depends on both the observed treatment contrast and the amount of statistical information available. It addresses compatibility with the null hypothesis under the specified testing framework. The HR describes the estimated relative treatment effect, while the confidence interval describes uncertainty around that estimate.
Why was the analysis stratified?
The registry analysis notes specify stratification by endocrine sensitivity and nature of disease as determined by IWRS. Stratification allows the time-to-event comparison to account for prespecified groups rather than treating all participants as though they came from a single unstructured population.
Why does the one-sided alpha differ from the two-sided confidence interval?
The trial planning note specifies a one-sided α of 0.025 for the event-driven final analysis, whereas the posted effect estimate is accompanied by a two-sided 95% confidence interval. These are related but distinct inferential quantities. A one-sided testing threshold concerns the direction and error rate of a hypothesis test; a two-sided confidence interval expresses uncertainty around the parameter in both directions.
9. Randomization and Blinding
MONARCH 2 is described in the ClinicalTrials.gov record as randomized and double-blind. These are design features rather than statistical results, but they are central to interpretation.
These features do not themselves establish the magnitude of the treatment effect. Instead, they define the structure within which the primary PFS comparison was made.
10. Planned Analysis and Event-Driven Design
The registry states that the final analysis was planned at 378 PFS events. The event-driven framework is important because the amount of information in a time-to-event trial depends substantially on the number of observed events, not simply the number of enrolled participants.
These quantities describe the statistical design and planning assumptions. The observed HR of 0.553 is a result from the posted analysis and should not be substituted into the original power calculation.
This distinction between design assumptions and observed results is fundamental. A trial can be designed around one anticipated treatment effect and ultimately observe a different effect. The planned HR of 0.703 is therefore not a claim about what the trial ultimately demonstrated.
11. Stratified Analysis
The primary analysis was not described simply as an unstratified log-rank test. The registry specifically states that the log-rank test was stratified by:
- Endocrine sensitivity
- Nature of disease by interactive web response system (IWRS)
| Feature | Registry-supported detail |
|---|---|
| Analysis | Log Rank |
| Stratification | Endocrine sensitivity and nature of disease by IWRS |
| Effect measure | Hazard ratio |
| Analysis population | All randomized participants |
Stratification is useful when clinically relevant or design-defined characteristics are expected to influence the event process. Rather than discarding those factors, the analysis compares treatment groups while respecting the prespecified strata.
12. Censoring and the Analysis Population
The primary PFS analysis used all randomized participants. The registry specifically reports 224 censored participants in the abemaciclib group.
Why randomization matters
Analyzing participants according to randomized assignment preserves the treatment-comparison framework created at randomization.
Why censoring matters
A censored participant contributes follow-up information until the censoring time but is not counted as having experienced the PFS event at that time.
Time-to-event methods are particularly useful for clinical trials because follow-up is rarely identical for every participant. Some participants experience progression or death during observation, while others remain event-free when their usable follow-up ends.
13. Safety
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by the corresponding at-risk count.
| Safety measure | Affected / at risk |
|---|---|
| Abemaciclib + Fulvestrant | 129 / 441 |
| Placebo + Fulvestrant | 33 / 223 |
These figures should be read as serious adverse events by arm using the affected/at-risk format reported in the ClinicalTrials.gov record. They are not the same statistical endpoint as PFS and should not be combined with the efficacy estimate into a single numerical measure.
The denominators also should not be silently replaced by the overall enrollment of 669. The registry specifically supplies 441 and 223 as the at-risk denominators for these serious-adverse-event figures, so those values are retained here exactly as reported.
14. What the Hazard Ratio Does — and Does Not — Mean
The PFS HR of 0.553 indicates a lower estimated instantaneous rate of progression or death in the abemaciclib-plus-fulvestrant group relative to the placebo-plus-fulvestrant group under the reported time-to-event analysis.
The arithmetic complement, 44.7%, is a convenient way to describe the estimated relative reduction in hazard. It is not a statement that 44.7% of patients avoided progression and does not describe an individual's probability of benefit.
The 95% CI of 0.449–0.681 communicates uncertainty around the estimated HR. It provides a statistical interval for the treatment-effect parameter rather than a range of outcomes that individual participants can be expected to experience.
The reported P < 0.0000001 is evidence against the null hypothesis under the reported testing framework. It should not be interpreted as the probability that the null hypothesis is true, nor as a measure of how large the treatment effect is.
The ClinicalTrials.gov record does not report median PFS or fixed-time PFS rates. Consequently, this page does not create those quantities from the HR. The available result is a relative time-to-event effect and should be interpreted as such.
15. Important Limitations and Interpretation Issues
- Limited reported endpoint information: the ClinicalTrials.gov record contains one formal statistical analysis, for PFS. They do not provide additional formal statistical analyses for other posted outcome measures.
- No median PFS in the ClinicalTrials.gov record: the HR and confidence interval cannot be converted into a median PFS without additional survival information.
- No reconstructed Kaplan-Meier curve: a valid Kaplan-Meier reconstruction requires underlying event and censoring information or appropriately detailed source data.
- Censoring: 224 participants are identified as censored in the abemaciclib group. Censoring is therefore an important component of the reported time-to-event analysis.
- Hazard-ratio interpretation: a single HR is a model-based relative measure. Its interpretation is most straightforward when the relative event rates are reasonably represented over time by the underlying survival-analysis framework.
- Stratification: the reported log-rank analysis was stratified by endocrine sensitivity and nature of disease by IWRS. An unstratified analysis should not be substituted for the reported primary method when interpreting the posted result.
- One-sided versus two-sided inference: the design note specifies a one-sided α of 0.025, while the reported confidence interval is two-sided at 95%. These should not be treated as interchangeable quantities.
- Analysis population: the reported PFS analysis used all randomized participants, which is distinct from an analysis restricted to participants who remained on treatment.
- Safety and efficacy: serious adverse-event counts and PFS measure different aspects of the trial and require separate interpretation.
16. Why This Trial Matters Statistically
MONARCH 2 provides a compact teaching example of how a randomized phase 3 trial can use a time-to-event endpoint, stratified hypothesis testing, hazard ratios, confidence intervals, event-driven planning, and censoring within one coherent analysis framework.
| Concept | How it appears in MONARCH 2 |
|---|---|
| Randomization | The trial uses randomized allocation with two parallel arms. |
| Blinding | The registry describes the trial as double-blind. |
| Time-to-event endpoint | PFS is defined from randomization until disease progression or death. |
| Log-rank test | The registry reports the log-rank test as the primary statistical method. |
| Stratified analysis | The log-rank test was stratified by endocrine sensitivity and nature of disease by IWRS. |
| Hazard ratio | PFS treatment effect is reported as HR 0.553. |
| Confidence interval | The HR has a two-sided 95% CI of 0.449–0.681. |
| P-value | The reported P-value is <0.0000001. |
| Censoring | The analysis identifies 224 censored participants in the abemaciclib group. |
| Event-driven planning | The final analysis was planned at 378 PFS events. |
| Statistical power | The registry states approximately 90% power under an assumed HR of 0.703 and one-sided α of 0.025. |
The trial is especially useful for understanding why clinical-trial survival analysis is not simply a comparison of percentages. The timing of events, censoring, prespecified strata, and the number of observed events all influence how the treatment comparison is estimated and tested.
17. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
18. Related Statistical Calculators
19. Sources
- ClinicalTrials.gov: NCT02107703 — MONARCH 2.
- PubMed: PMID 34425869.
- PubMed: PMID 33797023.
- PubMed: PMID 33392835.
- PubMed: PMID 31563959.
- PubMed: PMID 28580882.
Continue with the statistical methods behind MONARCH 2
Explore the broader Clinical Biostats tutorials and calculators for survival analysis, hazard ratios, confidence intervals, log-rank testing, and clinical-trial methodology.
20. Record Summary
MONARCH 2 provides a focused example of randomized time-to-event analysis. The trial enrolled 669 participants in a randomized, double-blind, parallel phase 3 design comparing abemaciclib plus fulvestrant with placebo plus fulvestrant. Its registered primary endpoint was PFS, defined from randomization until disease progression or death due to any cause, with a time frame of up to 31 months.
The posted primary analysis used a stratified log-rank test in all randomized participants. The reported PFS hazard ratio was 0.553, with a two-sided 95% CI of 0.449–0.681 and a P-value of <0.0000001. The registry specifies stratification by endocrine sensitivity and nature of disease by IWRS. The final analysis was planned at 378 PFS events, with approximately 90% power under an assumed HR of 0.703 and one-sided α of 0.025.
Statistically, the most important lesson is that the result should be read as a complete set of quantities rather than as a P-value alone: the hazard ratio describes the estimated relative event rate, the confidence interval describes statistical uncertainty around that estimate, and the P-value addresses the evidence against the null hypothesis under the specified testing framework. Censoring, stratification, randomization, and the event-driven design are integral parts of that interpretation.