← Clinical Trials
Breast Cancer Phase 3 Time-to-Event NCT02246621

MONARCH 3: Complete Statistical Analysis of Abemaciclib in Breast Cancer

An independent statistical analysis of the randomized phase 3 MONARCH 3 trial evaluating abemaciclib plus a nonsteroidal aromatase inhibitor versus placebo plus a nonsteroidal aromatase inhibitor in postmenopausal women with breast cancer.

MONARCH 3  ·  Phase 3  ·  Enrollment 493  ·  Primary endpoint: Progression Free Survival
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

MONARCH 3 was a randomized, double-blind, parallel-group phase 3 trial evaluating abemaciclib combined with a nonsteroidal aromatase inhibitor (NSAI) versus placebo combined with an NSAI in postmenopausal women with breast cancer. The registry reports 493 participants, two arms, and a primary time-to-event endpoint of progression-free survival.

493
Enrollment
Randomized phase 3 trial
2
Arms
Parallel-group design
0.54
PFS HR
95% CI 0.418–0.698
0.000002
PFS P-value
Two-sided log-rank test
FeatureMONARCH 3
PhasePhase 3
ConditionBreast Cancer
Brief titleA Study of Nonsteroidal Aromatase Inhibitors Plus Abemaciclib (LY2835219) in Postmenopausal Women With Breast Cancer
DesignRandomized, double-blind, parallel-group
AllocationRandomized
Primary purposeTreatment
Enrollment493
Primary endpointProgression Free Survival (PFS)
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Results postedYes
Statistical analyses posted6
ClinicalTrials.govNCT02246621

2. Clinical Question

The central statistical question was whether adding abemaciclib to a nonsteroidal aromatase inhibitor improved progression-free survival compared with an NSAI regimen without abemaciclib in postmenopausal women with breast cancer.

Population

Postmenopausal women with breast cancer, as specified by the trial's registered brief title.

Intervention

Abemaciclib combined with a nonsteroidal aromatase inhibitor. The registered interventions include abemaciclib, anastrozole, and letrozole.

Comparator

Placebo combined with a nonsteroidal aromatase inhibitor. Placebo is listed as a registered intervention.

Primary question

Does abemaciclib plus an NSAI improve progression-free survival relative to placebo plus an NSAI?

3. Trial Design

01
Randomize493 participants
02
Two armsAbemaciclib + NSAI or placebo + NSAI
03
Double-blindBlinded treatment comparison
04
FollowProgression, death, response, health status
05
AnalyzeTime-to-event and secondary outcomes
ARM A

Abemaciclib + NSAI

  • Abemaciclib
  • Nonsteroidal aromatase inhibitor
  • Registered NSAI interventions include anastrozole and letrozole
ARM B

Placebo + NSAI

  • Placebo
  • Nonsteroidal aromatase inhibitor
  • Registered NSAI interventions include anastrozole and letrozole
Allocation
Randomized
Masking
Double
Design model
Parallel
Primary purpose
Treatment

The registry lists the study start as 2014-11-06 and the primary completion date as 2017-01-31. The current registry status in the ClinicalTrials.gov record is ACTIVE_NOT_RECRUITING.

4. Endpoints

The registry identifies one registered primary endpoint. Results were posted for that endpoint, and the registry contains six statistical analyses overall: one primary-endpoint analysis and five secondary analyses.

EndpointTime frameTypeAnalysis reported
Progression Free Survival (PFS) Randomization to Progressive Disease or Death Due to Any Cause (Up to 32 Months) Time-to-event Log-rank test; hazard ratio
Change From Baseline to End of Study in Health Status on the EuroQuol 5-Dimension 5 Level (EuroQol-5D 5L) Index Value Baseline, End of Study (Up to 32 Months) Continuous Mixed-effects model; mean difference
Change From Baseline to End of Study in Health Status on the EuroQol-5D 5L Visual Analog Scale (VAS) Scores Scale Baseline, End of Study (Up to 32 Months) Continuous Mixed-effects model; mean difference
Percentage of Participants With Complete Response (CR) or Partial Response (PR) (Objective Response Rate [ORR]) Randomization to Progressive Disease or Death Due to Any Cause (Up to 32 Months) Binary Cochran-Mantel-Haenszel test
Percentage of Participants With CR, PR or Stable Disease (SD) (Disease Control Rate [DCR]) Randomization to Progressive Disease or Death Due to Any Cause (Up to 32 Months) Binary Cochran-Mantel-Haenszel test
Percentage of Participants With Tumor Response of SD for at Least 6 Months, PR, or CR (Clinical Benefit Rate [CBR]) Randomization to Progressive Disease or Death Due to Any Cause (Up to 32 Months) Binary Cochran-Mantel-Haenszel test

Registered PFS definition

The registry defines PFS as the time from the first day of therapy to the first evidence of disease progression as defined by RECIST v1.1 or death from any cause. Progressive Disease (PD) was defined as at least a 20% increase in the sum of the diameters of target lesions, with reference being the smallest sum on study and an absolute increase of at least 5 mm, or unequivocal progression of non-target lesions, or the additional registry wording in the ClinicalTrials.gov record.

Endpoint distinction: PFS is a time-to-event endpoint. ORR, DCR, and CBR are binary response endpoints, while the two EuroQol outcomes are continuous changes from baseline. These endpoint types naturally lead to different statistical methods.

5. Statistical Methodology

Log-rank test for progression-free survival

The registered primary analysis uses the log-rank test to compare PFS between the abemaciclib + NSAI and placebo + NSAI groups. The analysis population is described as all randomized participants who had evaluable data.

Primary analysis framework
H0: no difference in the time-to-event distributions between treatment groups

The log-rank procedure compares the observed pattern of events between randomized groups over follow-up while accounting for the timing of events and censoring. The registry separately reports a hazard ratio as the effect measure.

Hazard ratio

The PFS effect measure is a hazard ratio. The posted estimate is 0.54, with a two-sided 95% confidence interval of 0.418 to 0.698.

Interpretation of the primary effect measure
HR = 0.54  →  estimated hazard is 54% of the comparator hazard

Equivalently, 1 − 0.54 = 0.46, so the estimated hazard is 46% lower in the abemaciclib + NSAI group under the hazard-ratio interpretation. This is a relative time-to-event measure, not a statement that 46% of participants avoided progression or death.

Mixed-effects model for health-status outcomes

The two EuroQol outcomes were analyzed using a mixed-effects model. This is appropriate conceptually for repeated or longitudinal measurements because observations from the same participant are correlated rather than statistically independent.

Longitudinal data principle
Observed outcome = fixed effects + participant-level variation + residual variation

A mixed-effects framework can represent both population-level effects and within-participant correlation. In MONARCH 3, the registry reports the resulting effect as a mean difference (net).

Cochran-Mantel-Haenszel test for binary outcomes

ORR, DCR, and CBR were analyzed using the Cochran-Mantel-Haenszel test. This is a family of methods for comparing categorical outcomes while allowing analysis across strata when the analysis specifies them. The ClinicalTrials.gov record does not identify the particular stratification variables used for these analyses.

Effect measures used in the registry analyses

Endpoint typeReported methodReported effect measure
Time-to-eventLog-rank testHazard ratio
Continuous health-status outcomeMixed-effects modelMean difference (net)
Binary response outcomeCochran-Mantel-Haenszel testRegistry reports P-value; no estimate is reported in the posted analysis

6. Primary Result: Progression-Free Survival

The registry reports a formal statistical analysis for the primary endpoint of PFS. The comparison was between abemaciclib + NSAI and placebo + NSAI, using a log-rank test and a hazard ratio as the effect measure.

Progression-free survival

HR 0.54

95% CI: 0.418–0.698   ·   P = 0.000002

Two-sided confidence interval and superiority hypothesis.

Primary PFS analysis featureReported result
EndpointProgression Free Survival (PFS)
Time frameRandomization to Progressive Disease or Death Due to Any Cause (Up to 32 Months)
Analysis populationAll randomized participants who had evaluable data
Censored participantsAbemaciclib + NSAI = 190; Placebo + NSAI = 57
ComparisonAbemaciclib + NSAI vs Placebo + NSAI
MethodLog-rank test
Effect measureHazard ratio
Estimate0.54
95% CI0.418–0.698
P-value0.000002
HypothesisSuperiority
Clinical Biostats interpretation

The hazard ratio of 0.54 means that the estimated instantaneous rate of progression or death was 54% of the corresponding rate in the placebo + NSAI group, under the hazard-ratio interpretation. Expressed as a relative reduction, 1 − 0.54 = 0.46, corresponding to a 46% lower estimated hazard.

The HR does not mean that 46% of participants were protected from progression, that every participant experienced the same reduction, or that the median PFS was reduced or increased by 46%. It is a relative time-to-event measure.

The 95% confidence interval of 0.418 to 0.698 describes uncertainty around the estimated hazard ratio under the statistical framework. Because the interval lies below 1, the posted interval is consistent with a lower estimated hazard in the abemaciclib + NSAI group.

The P = 0.000002 value addresses evidence against the null hypothesis used for the log-rank comparison. It does not measure the size of the treatment effect, does not give the probability that the treatment is effective, and does not describe the probability that the observed hazard ratio will recur in another study.

The analysis also includes censoring: 190 participants in the abemaciclib + NSAI group and 57 in the placebo + NSAI group were reported as censored. Consequently, PFS analysis is not equivalent to simply counting how many participants progressed. Interpretation depends on the timing of events and censoring.

A Kaplan-Meier analysis is the natural descriptive companion to a time-to-event endpoint such as PFS, but the registry statistical analysis specifically identifies the log-rank test and hazard ratio. No median PFS or time-specific survival estimate is provided in the ClinicalTrials.gov record, so none is reported here.

7. Secondary Results

The registry contains five secondary statistical analyses. These analyses address health status, objective response, disease control, and clinical benefit. Unlike the primary PFS analysis, the ClinicalTrials.gov record does not provide effect estimates and confidence intervals for the three binary response endpoints.

Health status: EuroQol-5D 5L Index Value

Net mean difference

-0.01

Two-sided P = 0.688

Mixed-effects model; outcome measured in units on a scale.

The endpoint was change from baseline to end of study in health status on the EuroQol-5D 5L Index Value, with assessment at baseline and end of study up to 32 months. The analysis population was all randomized participants.

Interpretation

The reported net mean difference was -0.01. The negative sign indicates the direction of the estimated between-group difference as defined by the registry's analysis, but the ClinicalTrials.gov record does not provide the underlying group-specific means needed to interpret the absolute change in each arm.

The P = 0.688 value is not a measure of the magnitude or clinical importance of the difference. It is a test statistic result under the specified superiority framework. The ClinicalTrials.gov record does not provide a confidence interval for this analysis.

Health status: EuroQol-5D 5L VAS

Net mean difference

-1.01 mm

Two-sided P = 0.466

Mixed-effects model; outcome measured in millimeters.

The endpoint measured change from baseline to end of study in the EuroQol-5D 5L Visual Analog Scale (VAS) Scores Scale, at baseline and end of study up to 32 months. All randomized participants were included in the stated analysis population.

Interpretation

The reported net mean difference was -1.01 mm. The ClinicalTrials.gov record identifies the direction and size of the estimated between-group difference but does not provide the arm-specific means or a confidence interval.

The P = 0.466 result should not be interpreted as evidence that the two treatments are identical. A nonsignificant P-value means that the analysis did not provide strong statistical evidence against its null hypothesis; it does not establish equivalence.

Objective Response Rate

Cochran-Mantel-Haenszel analysis

P = 0.005

Superiority hypothesis; all randomized participants.

ORR was defined as the percentage of participants with complete response or partial response. The time frame was randomization to progressive disease or death due to any cause, up to 32 months. The reported method was the Cochran-Mantel-Haenszel test.

Important reporting distinction: the registry analysis reports the P-value but does not provide an ORR percentage for either arm, an effect estimate, or a confidence interval. Therefore, the statistical evidence can be described as a reported P = 0.005 without reconstructing an unreported response-rate difference.

Disease Control Rate

Cochran-Mantel-Haenszel analysis

P = 0.501

Superiority hypothesis; all randomized participants.

DCR was defined as the percentage of participants with complete response, partial response, or stable disease. The analysis covered randomization to progressive disease or death due to any cause, up to 32 months. The registry reports a Cochran-Mantel-Haenszel analysis with P = 0.501.

The ClinicalTrials.gov record does not report the arm-specific DCR percentages or an effect estimate and confidence interval. The P-value therefore should not be converted into an estimated treatment difference.

Clinical Benefit Rate

Cochran-Mantel-Haenszel analysis

P = 0.101

Superiority hypothesis; all randomized participants.

CBR was defined as the percentage of participants with tumor response of stable disease for at least 6 months, partial response, or complete response. The time frame was randomization to progressive disease or death due to any cause, up to 32 months.

The reported Cochran-Mantel-Haenszel P-value was 0.101. No arm-specific percentage, effect estimate, or confidence interval is reported in the ClinicalTrials.gov record.

8. Safety Results

The ClinicalTrials.gov record provides serious adverse events by randomized treatment arm as affected participants divided by participants at risk. These figures are reported separately from the efficacy analyses.

Safety measureAbemaciclib + NSAIPlacebo + NSAI
Serious adverse events, affected / at risk102 / 32727 / 161
Serious adverse events: affected participants among those at risk
Abemaciclib + NSAI
102 / 327
Placebo + NSAI
27 / 161

The affected/at-risk figures correspond to approximately 31.2% and 16.8%, respectively, when calculated directly from the reported counts. These percentages are displayed only as a simple arithmetic representation of the registry-reported fractions; the underlying registry values remain 102/327 and 27/161.

Safety interpretation: the serious-adverse-event denominators are not the same as the overall enrollment of 493. They are the at-risk populations reported by the registry for this safety measure. The figures therefore should not be silently treated as percentages of all randomized participants.

9. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint because both the timing of progression or death and the presence of censoring matter. A log-rank test compares the event-time experience of two groups across follow-up rather than reducing the outcome to a single binary status at one arbitrary date.

In MONARCH 3, the registry specifically reports the log-rank test for the primary PFS comparison. The corresponding effect measure is the hazard ratio.

What does a hazard ratio of 0.54 mean?

A hazard ratio of 0.54 means that the estimated instantaneous event rate in the abemaciclib + NSAI group was 0.54 times the corresponding rate in the placebo + NSAI group under the fitted time-to-event interpretation. This corresponds to an estimated 46% lower hazard.

It does not mean that 46% of participants avoided progression or death. It also does not mean that each individual experienced the same 46% reduction.

Why does the confidence interval matter?

The point estimate alone does not communicate statistical uncertainty. The 95% confidence interval of 0.418 to 0.698 shows the range of hazard-ratio values compatible with the specified confidence framework and observed data, subject to the assumptions of the analysis.

Because the entire reported interval is below 1, the interval is consistent with a lower estimated hazard for the abemaciclib + NSAI group.

Why does the P-value not measure effect size?

The P-value answers a hypothesis-testing question: how compatible are the observed data with the null hypothesis under the specified test? It does not quantify how large the treatment effect is. For MONARCH 3, the effect size is communicated by the hazard ratio and its confidence interval, while the P-value provides evidence concerning the statistical test.

Why use a mixed-effects model for the EuroQol outcomes?

The EuroQol outcomes are continuous measurements obtained at baseline and end of study. Measurements from the same participant are related, so treating every observation as independent can misrepresent uncertainty. A mixed-effects model provides a framework for modeling population-level treatment effects while accounting for participant-level variation and within-participant dependence.

What does the Cochran-Mantel-Haenszel test add to response analyses?

The Cochran-Mantel-Haenszel framework is designed for categorical comparisons and can account for stratification when strata are part of the analysis. In MONARCH 3, the registry identifies this method for ORR, DCR, and CBR. The ClinicalTrials.gov record does not specify the strata used for those analyses, so the page does not infer them.

Why should the secondary P-values not be treated as effect estimates?

A P-value such as 0.005 for ORR indicates the result of the reported statistical test, but it does not reveal the difference in response percentages. Without the arm-specific percentages or an effect estimate and confidence interval, the magnitude of the ORR difference cannot be reconstructed from the P-value alone.

10. Understanding Censoring in the PFS Analysis

The registry identifies 190 censored participants in the abemaciclib + NSAI group and 57 censored participants in the placebo + NSAI group within the primary analysis population.

What censoring means

A censored participant contributes information to the analysis up to the point at which their event status is no longer observed within the relevant follow-up framework.

What censoring does not mean

Censoring does not mean that a participant experienced progression at the censoring date. It means that the event time was not observed in the way required for an uncensored event.

Because PFS incorporates censoring, the analysis cannot be reproduced by simply dividing the number of participants who progressed by the total number randomized. The timing of events and follow-up contributes directly to the time-to-event comparison.

Conceptual survival-analysis quantity
S(t) = P(T > t)

A time-to-event analysis asks about the distribution of the event time T. Kaplan-Meier estimation is commonly used to describe this distribution, while the log-rank test compares groups and a hazard ratio summarizes relative event rates under its modeling interpretation.

11. Primary Analysis Population

The posted PFS analysis population is described as all randomized participants who had evaluable data. The registry separately reports the number of censored participants in each treatment group.

AnalysisPopulation reported in the trial data
Primary PFSAll randomized participants who had evaluable data
EuroQol Index ValueAll randomized participants
EuroQol VASAll randomized participants
ORRAll randomized participants
DCRAll randomized participants
CBRAll randomized participants

The consistency of the stated randomized analysis populations for the secondary endpoints is important because randomized comparisons are intended to preserve the balance created by randomization. The ClinicalTrials.gov record does not provide additional analysis-population definitions beyond those listed above.

12. Statistical Interpretation of the Secondary Endpoints

EndpointEstimateP-valueInterpretive point
EuroQol-5D 5L Index Value -0.01 mean difference 0.688 Direction and magnitude of the reported net mean difference; no CI reported.
EuroQol-5D 5L VAS -1.01 mm mean difference 0.466 Reported net mean difference; no CI reported.
ORR Not reported 0.005 CMH test result; arm-specific response percentages are not reported.
DCR Not reported 0.501 CMH test result; arm-specific disease-control percentages are not reported.
CBR Not reported 0.101 CMH test result; arm-specific clinical-benefit percentages are not reported.

This distinction illustrates an important statistical reporting principle: an analysis result should not be made more precise than the source data permit. The P-values can be reported exactly as reported in the registry, but they cannot be used to derive missing response rates or confidence intervals.

13. What the PFS Hazard Ratio Does — and Does Not — Mean

Effect size

The PFS hazard ratio of 0.54 indicates a lower estimated hazard of progression or death in the abemaciclib + NSAI group relative to placebo + NSAI. On the hazard-ratio scale, the estimate corresponds to a 46% lower estimated hazard.

Not an absolute risk difference

A hazard ratio is not the same as an absolute reduction in the probability of progression or death. Absolute effects depend on the underlying event risk and follow-up time. The ClinicalTrials.gov record does not provide time-specific PFS probabilities or median PFS, so those quantities are not reported here.

Not a probability of benefit

The P-value of 0.000002 is not the probability that abemaciclib works, and the hazard ratio is not the proportion of participants who benefit. These quantities answer different statistical questions.

Precision

The 95% confidence interval of 0.418 to 0.698 provides the principal uncertainty statement posted on ClinicalTrials.gov for the primary PFS effect estimate. It is narrower than the full range of possible hazard ratios below and above 1 would have been, because the observed data provide information about the treatment comparison.

14. Limitations

15. Why This Trial Matters Statistically

MONARCH 3 is a useful teaching case because its registered analyses illustrate several distinct statistical problems within the same randomized trial: time-to-event analysis for PFS, longitudinal modeling for health-status outcomes, and categorical testing for response endpoints.

ConceptHow it appears in MONARCH 3
RandomizationParticipants were randomized in a two-arm parallel design.
BlindingThe registry describes the study as double-blind.
Time-to-event endpointPFS is defined from randomization to progressive disease or death due to any cause, up to 32 months.
Log-rank testUsed for the posted primary PFS comparison.
Hazard ratioUsed to quantify the primary PFS treatment effect.
Confidence intervalThe PFS HR is accompanied by a two-sided 95% CI of 0.418–0.698.
Censoring190 and 57 censored participants are reported in the two PFS analysis groups.
Mixed-effects modelUsed for changes in EuroQol-5D 5L Index Value and VAS.
Cochran-Mantel-Haenszel testUsed for ORR, DCR, and CBR.
P-valuesReported for all six posted statistical analyses.
Safety analysisSerious adverse events are reported by arm using affected/at-risk counts.

16. A Practical Statistical Reading of MONARCH 3

A useful way to read the trial is to move from the endpoint definition to the analysis method and only then to the numerical result.

Step 1: Identify the endpoint

PFS is a time-to-event endpoint, so the timing of progression, death, and censoring matters.

Step 2: Identify the test

The registry reports a log-rank test for the primary PFS comparison.

Step 3: Identify the effect measure

The effect is summarized by a hazard ratio rather than a simple difference in percentages.

Step 4: Read uncertainty

The HR of 0.54 is accompanied by a two-sided 95% CI from 0.418 to 0.698.

Step 5: Read the P-value separately

The P-value of 0.000002 describes evidence under the statistical test; it is not an effect-size measure.

Step 6: Check the analysis population

The primary analysis includes randomized participants with evaluable data, with censoring reported separately by group.

This sequence prevents a common interpretive error: treating a small P-value as if it were itself evidence about the magnitude or clinical importance of an effect. The numerical effect, its uncertainty, the analysis population, and the endpoint definition all contribute to the statistical interpretation.

17. Trial Timeline

2014-11-06

Study start

The registry lists 2014-11-06 as the study start date.

2017-01-31

Primary completion

The registry lists 2017-01-31 as the primary completion date.

Results posted

Statistical results available

The ClinicalTrials.gov record indicates that results are posted, with 13 outcome measures and 6 statistical analyses.

Current registry status

Active, not recruiting

The ClinicalTrials.gov status is ACTIVE_NOT_RECRUITING.

18. Related Tutorials

Learn more about the methods used in this trial:

19. Related Calculators

20. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the time-to-event, categorical, longitudinal, and inference methods represented in MONARCH 3.

21. Record Summary

MONARCH 3 provides a compact example of how different clinical-trial endpoints require different statistical tools. The primary endpoint, PFS, is a time-to-event outcome analyzed with a log-rank test and summarized using a hazard ratio of 0.54 with a two-sided 95% confidence interval of 0.418–0.698 and P = 0.000002. The same trial also uses mixed-effects models for longitudinal health-status outcomes and Cochran-Mantel-Haenszel tests for binary response outcomes.

The most important statistical lesson is that these quantities should be interpreted according to what they actually measure. The hazard ratio describes a relative time-to-event effect; the confidence interval describes uncertainty around that estimate; the P-value describes evidence under a hypothesis test; and the secondary mean differences and categorical P-values address different outcome structures. Treating these quantities as interchangeable would obscure the statistical story of the trial.

Clinical Biostats methodology: A trial-results page should separate reported evidence from statistical interpretation. Where the ClinicalTrials.gov record contains an estimate and confidence interval, both are reported exactly. Where only a P-value is reported, the page does not manufacture an effect estimate or confidence interval.