← Clinical Trials
Breast Neoplasms Phase 3 Progression-Free Survival NCT02107703

MONARCH 2: Complete Statistical Analysis of Abemaciclib in Hormone Receptor Positive HER2 Negative Breast Cancer

An independent statistical analysis of the randomized, double-blind phase 3 MONARCH 2 trial evaluating abemaciclib plus fulvestrant versus placebo plus fulvestrant in women with hormone receptor positive, HER2 negative breast cancer.

Trial period: 2014-07-22 to 2017-02-14  ·  Enrollment: 669  ·  Primary endpoint: PFS
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

MONARCH 2 was a randomized, double-blind, parallel phase 3 trial evaluating abemaciclib combined with fulvestrant against placebo combined with fulvestrant in women with hormone receptor positive, HER2 negative breast cancer. The registered primary endpoint was progression-free survival (PFS), a time-to-event endpoint analyzed using a stratified log-rank test.

669
Enrolled
Randomized trial
2
Arms
Parallel design
0.553
PFS HR
95% CI 0.449–0.681
<0.0000001
P-value
Stratified log-rank
FeatureMONARCH 2
Trial nameMONARCH 2
PhasePhase 3
PopulationWomen with hormone receptor positive, HER2 negative breast cancer
DesignRandomized, double-blind, parallel
Primary purposeTreatment
Enrollment669
Arms2
Primary endpointProgression-Free Survival (PFS)
Primary endpoint typeTime-to-event
Primary analysisStratified log-rank test
Effect measureHazard ratio
Trial statusActive, not recruiting
Lead sponsorEli Lilly and Company
ClinicalTrials.govNCT02107703

2. Clinical Question

The central statistical question was whether treatment with abemaciclib plus fulvestrant produced a different progression-free survival experience from placebo plus fulvestrant in women with hormone receptor positive, HER2 negative breast cancer.

Population

Women with hormone receptor positive, HER2 negative breast cancer.

Intervention

Abemaciclib combined with fulvestrant.

Comparator

Placebo combined with fulvestrant.

Primary question

How does progression-free survival compare between the randomized treatment groups?

Because PFS is a time-to-event endpoint, the statistical question is not simply whether more participants eventually progressed. It concerns the distribution of time from randomization to the first qualifying event, with participants who have not experienced an event by their available follow-up contributing censored information.

3. Trial Design

01
Randomize669 enrolled
02
Assign2 parallel arms
03
BlindDouble-blind design
04
AssessPFS over time
05
AnalyzeStratified log-rank
Allocation
Randomized
Design model
Parallel
Masking
Double-blind
Primary purpose
Treatment
ARM A

Abemaciclib + Fulvestrant

  • Abemaciclib
  • Fulvestrant
ARM B

Placebo + Fulvestrant

  • Placebo
  • Fulvestrant

The randomized and double-blind structure is important statistically because treatment assignment is established before the outcome is observed, while blinding is intended to reduce the potential for knowledge of assignment to influence trial conduct or assessment.

4. Endpoints

EndpointRegistered definition / time frameType
Progression-Free Survival (PFS) From Date of Randomization until Disease Progression or Death Due to Any Cause (Up To 31 Months) Time-to-event

The registry definition specifies PFS as the time from randomization to the first evidence of disease progression according to RECIST v1.1 or death from any cause. The registry text further defines progressive disease as at least a 20% increase in the sum of the diameters of target lesions, with the smallest sum on study as the reference and an absolute increase of at least 5 mm, followed by the registry-reported definition text.

The registered time frame is therefore an explicitly time-based endpoint: participants contribute follow-up from randomization until progression, death, or censoring according to the applicable analysis rules.

5. Statistical Methodology

Primary analysis: stratified log-rank test

The registry reports a log-rank test as the primary analysis method. The analysis notes specify that the log-rank test was stratified by endocrine sensitivity and nature of disease as determined by the interactive web response system (IWRS).

Primary analysis framework
H0: no difference in the PFS experience between randomized groups

The log-rank procedure compares the observed pattern of events over follow-up between treatment groups rather than comparing only the proportion of participants who have progressed at one fixed time point.

Hazard ratio as the effect measure

The registry reports the treatment effect as a hazard ratio (HR). The primary estimate was 0.553 with a two-sided 95% confidence interval of 0.449 to 0.681.

Interpretation of the hazard ratio
HR = estimated hazard in abemaciclib + fulvestrant ÷ estimated hazard in placebo + fulvestrant

An HR below 1 indicates a lower estimated instantaneous event rate in the abemaciclib group under the time-to-event model. It is not an absolute risk difference, a probability of benefit for an individual participant, or a statement that every participant experiences the same relative change.

Analysis population and censoring

The posted analysis population was all randomized participants. The analysis notes identify 224 censored participants in the abemaciclib group. Censoring is a fundamental feature of time-to-event analysis: a participant can contribute information up to the point at which their event status is no longer observed according to the analysis definition.

The important distinction is that censoring does not mean that a participant is treated as having experienced progression at the censoring time. Instead, the participant's observed follow-up contributes information to the survival analysis up to that point.

Planned event information and statistical power

The final analysis was planned at 378 PFS events. The registry states that this would provide approximately 90% power assuming a hazard ratio of 0.703 at a one-sided α of 0.025.

Planning assumption

The HR of 0.703 was a design assumption used for planning the event-driven analysis; it is not the observed treatment-effect estimate.

Observed estimate

The reported primary analysis estimate was HR 0.553 with a two-sided 95% CI of 0.449–0.681.

6. Results

Primary Endpoint: Progression-Free Survival

Hazard ratio for progression or death

0.553

95% CI: 0.449–0.681   ·   P < 0.0000001

Analysis: stratified log-rank test   ·   Population: all randomized participants

EndpointComparisonEffect estimate95% CIP-value
Progression-Free Survival Abemaciclib + Fulvestrant vs Placebo + Fulvestrant HR 0.553 0.449–0.681 <0.0000001

The reported estimate of 0.553 means that the estimated instantaneous rate of the PFS event—disease progression or death—was 55.3% of the corresponding rate in the placebo-plus-fulvestrant group under the reported analysis framework. Expressed as a simple relative-hazard interpretation, 1 − 0.553 = 0.447, so the estimate corresponds to an approximately 44.7% lower estimated hazard.

Clinical Biostats interpretation

What the estimate means: The HR of 0.553 is a relative time-to-event measure. It summarizes the estimated difference in the instantaneous rate of progression or death between the randomized groups over the analyzed follow-up.

What it does not mean: It does not mean that 44.7% of participants avoided progression, that an individual participant has exactly a 44.7% lower probability of progression, or that median PFS was reduced or increased by 44.7%. No median PFS is reported in the ClinicalTrials.gov record used for this page.

Confidence interval: The two-sided 95% CI of 0.449–0.681 describes uncertainty around the estimated hazard ratio under the statistical framework. It is an interval for the treatment-effect parameter, not a range containing 95% of individual patient outcomes.

P-value: The P-value of <0.0000001 quantifies the incompatibility of the observed data with the null hypothesis under the specified testing framework. It does not measure the magnitude of the treatment effect. The magnitude is described by the HR and its confidence interval.

Important cautions: The analysis is time-to-event based and includes censoring. The registry identifies a stratified analysis and a one-sided α of 0.025 for the event-driven design. The reported confidence interval is two-sided at 95%, so its interpretation should not be confused with the one-sided design alpha. As with hazard-ratio methods generally, interpretation of a single HR is most direct when the relative hazards are reasonably represented by the model over time.

Educational note: the ClinicalTrials.gov record contains the hazard ratio, confidence interval, P-value, analysis population, censoring information, and analysis method, but do not provide the underlying event-time and censoring records needed to reconstruct a Kaplan-Meier curve independently.

7. Understanding the PFS Result

The PFS analysis combines three pieces of information that should be read together: the effect estimate, the confidence interval, and the hypothesis-test result.

Effect size

HR 0.553 describes the relative treatment effect on the hazard scale. Its distance from 1 is more informative about effect magnitude than the P-value alone.

Precision

The 95% CI of 0.449–0.681 provides a range of values reflecting statistical uncertainty around the estimated HR.

Evidence against the null

The reported P-value of <0.0000001 indicates very strong statistical evidence against the null hypothesis used for the reported test.

Time-to-event context

PFS incorporates when events occur and also uses information from participants who are censored rather than simply classifying everyone as event or no event.

A useful statistical distinction is that effect size and statistical evidence are not interchangeable. With a sufficiently large amount of event information, even a relatively modest effect can produce a small P-value. Conversely, an important estimated effect can have a wide confidence interval when information is limited. The MONARCH 2 registry result supplies all three quantities, allowing them to be interpreted together.

8. Statistical Methods Explained

Why was a log-rank test used?

PFS is a time-to-event endpoint. Participants can experience progression or death at different times, and some participants can be censored. A log-rank test is designed to compare survival or event-time distributions between groups while incorporating the timing of observed events rather than reducing follow-up to a single binary outcome.

What does an HR of 0.553 mean?

An HR of 0.553 means that the estimated instantaneous event rate in the abemaciclib-plus-fulvestrant group was 0.553 times that in the placebo-plus-fulvestrant group under the reported analysis. The corresponding arithmetic expression, 1 − 0.553, is 0.447, or 44.7%. That does not mean that exactly 44.7% of participants benefited.

Why is the confidence interval important?

The point estimate is only one estimate of the underlying treatment effect. The 95% CI of 0.449–0.681 communicates statistical precision around the HR. A narrower interval generally indicates greater precision than a wider interval, all else equal. The interval should not be interpreted as a distribution of individual treatment effects.

Why doesn't the P-value measure effect size?

The P-value depends on both the observed treatment contrast and the amount of statistical information available. It addresses compatibility with the null hypothesis under the specified testing framework. The HR describes the estimated relative treatment effect, while the confidence interval describes uncertainty around that estimate.

Why was the analysis stratified?

The registry analysis notes specify stratification by endocrine sensitivity and nature of disease as determined by IWRS. Stratification allows the time-to-event comparison to account for prespecified groups rather than treating all participants as though they came from a single unstructured population.

Why does the one-sided alpha differ from the two-sided confidence interval?

The trial planning note specifies a one-sided α of 0.025 for the event-driven final analysis, whereas the posted effect estimate is accompanied by a two-sided 95% confidence interval. These are related but distinct inferential quantities. A one-sided testing threshold concerns the direction and error rate of a hypothesis test; a two-sided confidence interval expresses uncertainty around the parameter in both directions.

9. Randomization and Blinding

MONARCH 2 is described in the ClinicalTrials.gov record as randomized and double-blind. These are design features rather than statistical results, but they are central to interpretation.

Randomization
Random assignment creates the framework for comparing outcomes between treatment groups without assigning participants according to observed prognosis.
Double-blinding
Blinding is intended to reduce the influence of treatment knowledge on trial conduct and outcome assessment.
Parallel design
Participants remain in their randomized treatment pathways rather than receiving treatments sequentially in a crossover structure.
Treatment purpose
The registry identifies the primary purpose of the trial as treatment.

These features do not themselves establish the magnitude of the treatment effect. Instead, they define the structure within which the primary PFS comparison was made.

10. Planned Analysis and Event-Driven Design

The registry states that the final analysis was planned at 378 PFS events. The event-driven framework is important because the amount of information in a time-to-event trial depends substantially on the number of observed events, not simply the number of enrolled participants.

Design assumptions
378 PFS events  →  approximately 90% power   assuming HR = 0.703 and one-sided α = 0.025

These quantities describe the statistical design and planning assumptions. The observed HR of 0.553 is a result from the posted analysis and should not be substituted into the original power calculation.

This distinction between design assumptions and observed results is fundamental. A trial can be designed around one anticipated treatment effect and ultimately observe a different effect. The planned HR of 0.703 is therefore not a claim about what the trial ultimately demonstrated.

11. Stratified Analysis

The primary analysis was not described simply as an unstratified log-rank test. The registry specifically states that the log-rank test was stratified by:

FeatureRegistry-supported detail
AnalysisLog Rank
StratificationEndocrine sensitivity and nature of disease by IWRS
Effect measureHazard ratio
Analysis populationAll randomized participants

Stratification is useful when clinically relevant or design-defined characteristics are expected to influence the event process. Rather than discarding those factors, the analysis compares treatment groups while respecting the prespecified strata.

12. Censoring and the Analysis Population

The primary PFS analysis used all randomized participants. The registry specifically reports 224 censored participants in the abemaciclib group.

Why randomization matters

Analyzing participants according to randomized assignment preserves the treatment-comparison framework created at randomization.

Why censoring matters

A censored participant contributes follow-up information until the censoring time but is not counted as having experienced the PFS event at that time.

Time-to-event methods are particularly useful for clinical trials because follow-up is rarely identical for every participant. Some participants experience progression or death during observation, while others remain event-free when their usable follow-up ends.

Interpretation caution: censoring requires assumptions about the relationship between censoring and the event process. The ClinicalTrials.gov record identifies the censored population but do not provide enough information to evaluate the detailed censoring mechanism independently.

13. Safety

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by the corresponding at-risk count.

Safety measureAffected / at risk
Abemaciclib + Fulvestrant129 / 441
Placebo + Fulvestrant33 / 223

These figures should be read as serious adverse events by arm using the affected/at-risk format reported in the ClinicalTrials.gov record. They are not the same statistical endpoint as PFS and should not be combined with the efficacy estimate into a single numerical measure.

The denominators also should not be silently replaced by the overall enrollment of 669. The registry specifically supplies 441 and 223 as the at-risk denominators for these serious-adverse-event figures, so those values are retained here exactly as reported.

14. What the Hazard Ratio Does — and Does Not — Mean

Effect estimate

The PFS HR of 0.553 indicates a lower estimated instantaneous rate of progression or death in the abemaciclib-plus-fulvestrant group relative to the placebo-plus-fulvestrant group under the reported time-to-event analysis.

The arithmetic complement, 44.7%, is a convenient way to describe the estimated relative reduction in hazard. It is not a statement that 44.7% of patients avoided progression and does not describe an individual's probability of benefit.

Confidence interval

The 95% CI of 0.449–0.681 communicates uncertainty around the estimated HR. It provides a statistical interval for the treatment-effect parameter rather than a range of outcomes that individual participants can be expected to experience.

P-value

The reported P < 0.0000001 is evidence against the null hypothesis under the reported testing framework. It should not be interpreted as the probability that the null hypothesis is true, nor as a measure of how large the treatment effect is.

Absolute versus relative information

The ClinicalTrials.gov record does not report median PFS or fixed-time PFS rates. Consequently, this page does not create those quantities from the HR. The available result is a relative time-to-event effect and should be interpreted as such.

15. Important Limitations and Interpretation Issues

16. Why This Trial Matters Statistically

MONARCH 2 provides a compact teaching example of how a randomized phase 3 trial can use a time-to-event endpoint, stratified hypothesis testing, hazard ratios, confidence intervals, event-driven planning, and censoring within one coherent analysis framework.

ConceptHow it appears in MONARCH 2
RandomizationThe trial uses randomized allocation with two parallel arms.
BlindingThe registry describes the trial as double-blind.
Time-to-event endpointPFS is defined from randomization until disease progression or death.
Log-rank testThe registry reports the log-rank test as the primary statistical method.
Stratified analysisThe log-rank test was stratified by endocrine sensitivity and nature of disease by IWRS.
Hazard ratioPFS treatment effect is reported as HR 0.553.
Confidence intervalThe HR has a two-sided 95% CI of 0.449–0.681.
P-valueThe reported P-value is <0.0000001.
CensoringThe analysis identifies 224 censored participants in the abemaciclib group.
Event-driven planningThe final analysis was planned at 378 PFS events.
Statistical powerThe registry states approximately 90% power under an assumed HR of 0.703 and one-sided α of 0.025.

The trial is especially useful for understanding why clinical-trial survival analysis is not simply a comparison of percentages. The timing of events, censoring, prespecified strata, and the number of observed events all influence how the treatment comparison is estimated and tested.

17. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

18. Related Statistical Calculators

19. Sources

Continue with the statistical methods behind MONARCH 2

Explore the broader Clinical Biostats tutorials and calculators for survival analysis, hazard ratios, confidence intervals, log-rank testing, and clinical-trial methodology.

20. Record Summary

MONARCH 2 provides a focused example of randomized time-to-event analysis. The trial enrolled 669 participants in a randomized, double-blind, parallel phase 3 design comparing abemaciclib plus fulvestrant with placebo plus fulvestrant. Its registered primary endpoint was PFS, defined from randomization until disease progression or death due to any cause, with a time frame of up to 31 months.

The posted primary analysis used a stratified log-rank test in all randomized participants. The reported PFS hazard ratio was 0.553, with a two-sided 95% CI of 0.449–0.681 and a P-value of <0.0000001. The registry specifies stratification by endocrine sensitivity and nature of disease by IWRS. The final analysis was planned at 378 PFS events, with approximately 90% power under an assumed HR of 0.703 and one-sided α of 0.025.

Statistically, the most important lesson is that the result should be read as a complete set of quantities rather than as a P-value alone: the hazard ratio describes the estimated relative event rate, the confidence interval describes statistical uncertainty around that estimate, and the P-value addresses the evidence against the null hypothesis under the specified testing framework. Censoring, stratification, randomization, and the event-driven design are integral parts of that interpretation.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical evidence from educational interpretation. Where the ClinicalTrials.gov record does not provide a median, fixed-time survival estimate, subgroup analysis, or additional formal comparison, those quantities are not reconstructed or inferred from the reported hazard ratio.