← Clinical Trials
Early Alzheimer's Disease Phase 3 MMRM Analysis NCT02477800

ENGAGE: Complete Statistical Analysis of Aducanumab in Early Alzheimer's Disease

An independent statistical review of the randomized phase 3 ENGAGE trial of aducanumab (BIIB037) in early Alzheimer's disease, focusing on the Clinical Dementia Rating Sum of Boxes endpoint, longitudinal MMRM analysis, cognitive and functional secondary endpoints, and reported safety data.

Trial start: 2015-08-13  ·  Primary completion: 2019-08-08  ·  Status: Terminated
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ENGAGE was a randomized, parallel, quadruple-masked phase 3 study of aducanumab (BIIB037) in early Alzheimer's disease. The registry reports an enrollment of 1653 participants, two arms, a superiority hypothesis, and a primary endpoint based on change from baseline in the Clinical Dementia Rating Sum of Boxes (CDR-SB) score at Week 78.

1653
Enrollment
Phase 3
2
Arms
Parallel randomized design
78
Primary time point
Week 78
MMRM
Primary method
Longitudinal mixed model
FeatureENGAGE
Trial nameENGAGE
Brief title221AD301 Phase 3 Study of Aducanumab (BIIB037) in Early Alzheimer's Disease
PhasePhase 3
Therapeutic areaNeurology
ConditionAlzheimer's Disease
DesignRandomized, parallel
MaskingQuadruple
Primary purposeTreatment
Enrollment1653
Hypothesis typeSuperiority
Lead sponsorBiogen
Sponsor typeIndustry
StatusTerminated
ClinicalTrials.govNCT02477800

2. Clinical Question

The primary statistical question was whether the randomized treatment comparison was associated with a difference in change from baseline in CDR-SB score at Week 78. The registry classifies the hypothesis as superiority.

Population

Participants enrolled in a phase 3 study of aducanumab (BIIB037) for early Alzheimer's disease.

Intervention

Aducanumab (BIIB037), represented in the registry as low-dose and high-dose treatment groups in the posted statistical analyses.

Comparator

Placebo.

Primary question

Does randomized treatment assignment produce a difference in change from baseline in CDR-SB at Week 78?

3. Trial Design

01
Randomize 1653 participants
02
Mask Quadruple masking
03
Treat Aducanumab or placebo
04
Assess Repeated outcome measurements
05
Week 78 Primary endpoint
TREATMENT · BIIB037

Aducanumab (BIIB037)

  • The registry identifies aducanumab (BIIB037) as the study drug.
  • Posted statistical analyses distinguish low-dose and high-dose BIIB037 groups.
  • The primary efficacy comparisons are reported separately for low dose versus placebo and high dose versus placebo.
CONTROL · PLACEBO

Placebo

  • Placebo is the comparator in the posted primary and secondary statistical analyses.
  • The registry reports placebo participants during the primary-completion period and later treatment-period safety categories.
Registry design detail: the ClinicalTrials.gov record reports two arms and a parallel randomized design, while the posted statistical analyses distinguish BIIB037 low-dose and high-dose comparisons with placebo. This page preserves the registry's reported structure rather than inferring an alternative arm-count description.

4. Randomization and Analysis Population

The registry identifies the primary analysis population as an intention-to-treat (ITT) population. ITT was defined as all randomized participants who had received at least one dose of study treatment, either aducanumab or placebo.

Analysis populationRegistry definition / role
ITTAll randomized participants who had received at least one dose of study treatment (Aducanumab or Placebo).
Number analyzedThe registry notes that the reported number of participants analyzed signifies the number of participants analyzed for the specified statistical analysis.

The use of an ITT framework is important because treatment comparisons remain anchored to randomized assignment rather than being redefined according to later treatment exposure or observed response.

5. Primary Endpoint

EndpointRegistered definitionTime frameType
Change From Baseline in Clinical Dementia Rating Sum of Boxes (CDR-SB) Score at Week 78 CDR-SB integrates assessments from 3 domains of cognition (memory, orientation, judgment/problem-solving) and 3 domains of function (community affairs, home/hobbies, personal care). Following caregiver interview and systematic patient examination, the rater assigns a score describing the participant's current performance level in each of these domains of life functioning. Baseline, Week 78 Continuous; score on a scale

The registry identifies one registered primary endpoint. Its posted statistical analyses evaluate the endpoint separately for BIIB037 low dose versus placebo and BIIB037 high dose versus placebo.

6. Statistical Methodology

Mixed model for repeated measures

The primary endpoint was analyzed using an MMRM (mixed model for repeated measures). This is a longitudinal modeling approach designed for outcomes measured repeatedly over multiple visits. Rather than reducing each participant's longitudinal record to a single observed value, the model uses the repeated measurements within the statistical framework.

Registry-reported model structure
Change from baseline CDR-SB = treatment group + categorical visit + treatment × visit + baseline CDR-SB + baseline CDR-SB × visit + additional prespecified model terms

The registry analysis text states that adjusted means, treatment differences, 95% confidence intervals, and p-values at each time point were based on this MMRM framework.

Why the visit term matters

A categorical visit effect allows the expected mean change to vary across visits rather than imposing a simple linear time trend. The treatment-by-visit interaction allows the treatment difference to vary over time. That is especially relevant for a longitudinal endpoint because the treatment contrast at one visit need not be identical to the contrast at another.

Baseline adjustment

The reported model includes baseline CDR-SB and its interaction with visit. Baseline adjustment can improve precision by accounting for the participant's starting level of the outcome while allowing the relationship between baseline and subsequent change to vary by visit.

Superiority testing

The registry classifies the hypothesis as superiority. Accordingly, the statistical question is whether the observed treatment difference provides evidence of a difference rather than whether a treatment can be shown to remain within a prespecified non-inferiority margin. No non-inferiority margin is reported in the ClinicalTrials.gov record.

Two-sided confidence intervals

The posted primary analyses use 95% two-sided confidence intervals. A confidence interval communicates the statistical precision of the estimated treatment difference; it should be read alongside, rather than replaced by, the p-value.

7. Primary Results: CDR-SB at Week 78

BIIB037 Low Dose vs Placebo

Difference from placebo

0.110

95% CI: -0.469 to 0.403   ·   P = 0.2250

MMRM analysis of change from baseline in CDR-SB at Week 78.

Primary endpointBIIB037 Low Dose vs Placebo
OutcomeChange From Baseline in CDR-SB Score at Week 78
AnalysisMMRM Model
Effect measureDifference from Placebo
Estimate0.110
95% two-sided CI-0.469 to 0.403
P-value0.2250
HypothesisSuperiority
Clinical Biostats interpretation

The estimated difference from placebo was 0.110 points on the CDR-SB outcome scale under the reported MMRM analysis. The positive sign identifies the direction of the reported difference according to the registry's treatment-comparison convention; it should not be translated into a clinical benefit or harm without specifying how higher CDR-SB values correspond to clinical status.

The 95% confidence interval of -0.469 to 0.403 spans both negative and positive values. This indicates substantial statistical uncertainty around the estimated treatment difference and means that the interval includes values on both sides of zero.

The P = 0.2250 value is evidence from the specified statistical test against its null hypothesis; it is not a measure of the size of the treatment effect. A p-value does not tell us the probability that the treatment is effective, nor does it quantify clinical importance.

The analysis is based on the reported ITT definition and an MMRM. Interpretation therefore depends on the model specification and the assumptions underlying the longitudinal analysis, including how the observed repeated measurements provide information about the unobserved measurements. The ClinicalTrials.gov record does not provide a separate assessment of those assumptions.

BIIB037 High Dose vs Placebo

Difference from placebo

0.03

95% CI: -0.262 to 0.326   ·   P = 0.8330

MMRM analysis of change from baseline in CDR-SB at Week 78.

Primary endpointBIIB037 High Dose vs Placebo
OutcomeChange From Baseline in CDR-SB Score at Week 78
AnalysisMMRM Model
Effect measureDifference from late start group
Estimate0.03
95% two-sided CI-0.262 to 0.326
P-value0.8330
HypothesisSuperiority
Clinical Biostats interpretation

The estimated treatment difference was 0.03 points on the CDR-SB scale. The estimate itself is close to zero, but the estimate should not be interpreted in isolation from its uncertainty interval and the model used to obtain it.

The 95% confidence interval of -0.262 to 0.326 includes zero and extends in both directions. It therefore does not identify a narrow range of treatment differences around the point estimate.

The P = 0.8330 value describes the evidence against the specified null hypothesis under the reported analysis. It does not measure the magnitude of the estimated treatment difference and should not be interpreted as the probability that there is no treatment effect.

As with the low-dose comparison, this result is an MMRM estimate in the reported ITT population. The validity of the inference depends on the longitudinal model and its assumptions, as well as the prespecified analysis framework.

Important comparison point: the two primary analyses are separate comparisons of BIIB037 dose groups with placebo. The ClinicalTrials.gov record does not provide a formal statistical comparison of the low-dose and high-dose treatment effects with each other. The two estimates therefore should not be treated as a direct dose-versus-dose hypothesis test.

8. Secondary Cognitive Results

The registry also posts MMRM analyses for change from baseline in the Mini-Mental State Examination (MMSE) score at Week 78. These are secondary endpoint analyses and use the same general longitudinal modeling framework.

Secondary endpointComparisonEstimate95% CIP-value
Change From Baseline in MMSE Score at Week 78 Placebo vs BIIB037 Low Dose 0.2 -0.35 to 0.74 0.4795
Change From Baseline in MMSE Score at Week 78 Placebo vs BIIB037 High Dose -0.1 -0.62 to 0.49 0.8106

The low-dose comparison has an estimated difference of 0.2, with a 95% confidence interval from -0.35 to 0.74 and P = 0.4795. The high-dose comparison has an estimated difference of -0.1, with a 95% confidence interval from -0.62 to 0.49 and P = 0.8106.

For both comparisons, the confidence intervals include zero. These results should be interpreted as secondary analyses rather than as replacements for the registered primary endpoint.

9. Secondary ADAS-Cog 13 Results

The Alzheimer's Disease Assessment Scale-Cognitive Subscale (13 Items), or ADAS-Cog 13, was also analyzed as a secondary endpoint using MMRM.

Secondary endpointComparisonEstimate95% CIP-value
Change From Baseline in ADAS-Cog 13 Score at Week 78 Placebo vs BIIB037 Low Dose -0.583 -1.5835 to 0.4181 0.2536
Change From Baseline in ADAS-Cog 13 Score at Week 78 Placebo vs BIIB037 High Dose -0.588 -1.6067 to 0.4309 0.2578

The estimated differences are -0.583 for low dose and -0.588 for high dose. Their corresponding 95% confidence intervals both cross zero: -1.5835 to 0.4181 and -1.6067 to 0.4309, respectively.

These results illustrate why a point estimate alone is insufficient. The two estimates are similar in magnitude, yet their confidence intervals extend across zero, indicating uncertainty around the direction and magnitude of the treatment difference under the reported model.

10. Secondary Functional Results

The registry also reports change from baseline in the Alzheimer's Disease Cooperative Study-Activities of Daily Living Inventory (Mild Cognitive Impairment Version), or ADCS-ADL-MCI, at Week 78.

Secondary endpointComparisonEstimate95% CIP-value
Change From Baseline in ADCS-ADL-MCI Score at Week 78 Placebo vs BIIB037 Low Dose 0.7 -0.19 to 1.64 0.1225
Change From Baseline in ADCS-ADL-MCI Score at Week 78 Placebo vs BIIB037 High Dose 0.7 -0.25 to 1.61 0.1506

Both dose comparisons have an estimated difference of 0.7. The low-dose 95% confidence interval is -0.19 to 1.64, while the high-dose interval is -0.25 to 1.61. The corresponding p-values are 0.1225 and 0.1506.

Reading a repeated-measures result: all of these secondary estimates are model-based differences in change from baseline at Week 78. They are not simple differences between the last observed measurements in the two groups. The MMRM uses the longitudinal data and the prespecified model structure to estimate the treatment contrast.

11. Statistical Methods Explained

Why was an MMRM used?

The outcomes in ENGAGE were measured longitudinally, with a baseline assessment and subsequent visits. MMRM is designed for this structure because it models repeated outcome measurements rather than treating each time point as an unrelated analysis. The reported model includes treatment, categorical visit, treatment-by-visit interaction, baseline outcome, and baseline-by-visit interaction.

What does the treatment difference represent?

The reported effect measure is a difference between treatment groups in the modeled change from baseline. For example, the low-dose CDR-SB estimate is 0.110. It is not a ratio, percentage change, hazard ratio, or probability. Its interpretation is tied to the CDR-SB score scale and the direction in which that scale is defined.

Why does the confidence interval matter?

The 95% confidence interval shows the statistical uncertainty around the estimated treatment difference under the specified analysis framework. For the low-dose primary comparison, the interval is -0.469 to 0.403. For the high-dose comparison, it is -0.262 to 0.326. In both cases, zero lies inside the interval.

What does the p-value tell us?

A p-value measures the compatibility of the observed result with the null hypothesis under the statistical model and testing framework. It does not measure effect size, clinical importance, or the probability that the null hypothesis is true. The p-values of 0.2250 and 0.8330 therefore should not be described as percentages of certainty.

Why use an ITT population?

Analyzing randomized participants according to the trial's defined ITT framework helps preserve the treatment comparison established by randomization. In ENGAGE, the registry defines ITT as all randomized participants who had received at least one dose of aducanumab or placebo.

What does quadruple masking add?

Quadruple masking means that the registry classifies the trial as having four categories of participants or trial personnel masked to treatment assignment. Masking can reduce the potential for knowledge of assignment to influence behavior, assessments, treatment administration, or other aspects of the trial. The ClinicalTrials.gov record does not specify the four masked categories.

12. Safety Results

The registry provides serious adverse event counts for the primary-completion period and later LTE-period categories. The ClinicalTrials.gov record contains complete affected/at-risk counts for three primary-completion-period groups.

Period / groupSerious adverse eventsAffected / at risk
PC Period · PlaceboSerious adverse events70/540
PC Period · BIIB037 Low DoseSerious adverse events76/549
PC Period · BIIB037 High DoseSerious adverse events79/558

The registry also lists LTE-period serious adverse event counts for late-start and early-start groups. The registry field is incomplete for some of those entries, so this page reports only the complete affected/at-risk pairs rather than constructing denominators that are not provided.

LTE period groupSerious adverse events
BIIB037 Late Start: Low Dose19/150
BIIB037 Late Start: High Dose14/152
BIIB037 Early Start: Low Dose35/299
Safety interpretation: serious adverse-event counts and efficacy estimates answer different questions. The safety figures describe affected participants among those at risk in the specified treatment period, whereas the primary CDR-SB results are model-based differences in longitudinal outcome change. They should not be combined into a single efficacy-safety statistic.

13. Premature Termination and Futility

The registry states that the study was halted prematurely based on a prespecified futility analysis and not based on safety concerns.

Prespecified futility

The registry identifies the futility analysis as the basis for the premature halt.

Not safety-driven

The registry explicitly states that the study was not halted based on safety concerns.

Interpretation

A futility decision is a design-stage decision about whether continuing the trial is unlikely to achieve the prespecified objective under the monitoring framework.

Participant flow

The registry states that participants who discontinued because of study termination are included in the "Reason not Specified" category in participant-flow tables.

Futility monitoring is statistically different from stopping because of demonstrated benefit. A futility decision can affect how much information ultimately accumulates and therefore becomes an important part of the interpretation of a prematurely terminated study.

14. Multiplicity and Multiple Treatment Comparisons

The ClinicalTrials.gov record contains two primary treatment comparisons for the same registered CDR-SB endpoint: BIIB037 low dose versus placebo and BIIB037 high dose versus placebo. There are also multiple secondary endpoints, each analyzed for both dose comparisons.

FeatureReported structure
Registered primary endpointOne CDR-SB endpoint at Week 78
Primary comparisons postedBIIB037 Low Dose vs Placebo; BIIB037 High Dose vs Placebo
Secondary endpoints postedMMSE, ADAS-Cog 13, and ADCS-ADL-MCI at Week 78
Statistical analyses posted8
Hypothesis typeSuperiority

When several hypotheses are evaluated, the interpretation of individual p-values depends on the prespecified multiplicity strategy. The ClinicalTrials.gov record identifies the multiple comparisons but do not provide an alpha-allocation or multiplicity-adjustment procedure. Accordingly, the individual p-values should be read as the reported results of their respective analyses rather than assuming an unreported multiplicity procedure.

15. Missing Data and Longitudinal Interpretation

MMRM is commonly used in longitudinal clinical trials because it can incorporate repeated observations without requiring every participant to have an observed value at every scheduled visit. The statistical advantage is that the model can use available repeated measurements rather than automatically discarding a participant after the first missing visit.

However, an MMRM is not automatically immune to missing-data bias. Its validity depends on the assumptions connecting the observed and unobserved outcomes and on the correctness of the model specification. The ClinicalTrials.gov record identifies the MMRM and its major fixed effects but do not provide a complete missing-data sensitivity-analysis strategy.

Participant-flow caveat: because the registry states that the study was terminated prematurely and that participants who discontinued because of study termination were included in the "Reason not Specified" category, participant discontinuation is particularly relevant when interpreting a longitudinal Week 78 analysis.

16. What the CDR-SB Estimate Does — and Does Not — Mean

Treatment difference

An estimated difference such as 0.110 is a model-based difference in change from baseline at the specified Week 78 time point. It does not mean that every participant experienced a change of 0.110 points, nor does it describe the probability that an individual participant benefits.

Confidence interval

The low-dose 95% confidence interval of -0.469 to 0.403 describes uncertainty around the estimated treatment difference under the specified model and sampling framework. It is not a range containing 95% of individual patient outcomes.

P-value

The p-value of 0.2250 is a measure associated with the statistical test under the specified model. It is not an effect-size measure and should not be converted into a statement such as "77.5% probability of no effect."

High-dose comparison

The high-dose estimate of 0.03 has a 95% confidence interval of -0.262 to 0.326. The interval's width illustrates why a small point estimate does not, by itself, establish that the true treatment difference is exactly zero.

17. Primary and Secondary Results in Context

OutcomeLow Dose vs PlaceboHigh Dose vs Placebo
CDR-SB at Week 78 0.110; 95% CI -0.469 to 0.403; P = 0.2250 0.03; 95% CI -0.262 to 0.326; P = 0.8330
MMSE at Week 78 0.2; 95% CI -0.35 to 0.74; P = 0.4795 -0.1; 95% CI -0.62 to 0.49; P = 0.8106
ADAS-Cog 13 at Week 78 -0.583; 95% CI -1.5835 to 0.4181; P = 0.2536 -0.588; 95% CI -1.6067 to 0.4309; P = 0.2578
ADCS-ADL-MCI at Week 78 0.7; 95% CI -0.19 to 1.64; P = 0.1225 0.7; 95% CI -0.25 to 1.61; P = 0.1506

The table illustrates a consistent statistical feature of the ClinicalTrials.gov record: all eight posted analyses use the same broad MMRM approach, and every listed 95% confidence interval spans zero. That pattern is descriptive of the reported estimates and uncertainty intervals; it does not by itself answer questions that were not tested in the registry analyses, such as direct dose-to-dose comparisons or unreported subgroup effects.

18. Important Limitations and Interpretation Issues

19. Why This Trial Matters Statistically

ENGAGE is a useful statistical teaching case because the ClinicalTrials.gov record bring together randomized treatment assignment, quadruple masking, a longitudinal continuous endpoint, repeated-measures modeling, baseline adjustment, multiple treatment comparisons, and prespecified futility monitoring.

ConceptHow it appears in ENGAGE
RandomizationThe registry identifies a randomized parallel phase 3 design.
BlindingThe study is classified as quadruple masked.
ITT analysisThe reported efficacy population includes randomized participants who received at least one dose.
Repeated measuresCDR-SB, MMSE, ADAS-Cog 13, and ADCS-ADL-MCI were analyzed longitudinally at Week 78.
MMRMThe posted statistical analyses use a mixed model for repeated measures.
Confidence intervalsPrimary and secondary analyses report 95% two-sided confidence intervals.
P-valuesEach posted analysis includes a p-value for its specified treatment comparison.
Superiority testingThe registry identifies superiority as the hypothesis type.
Futility monitoringThe study was halted prematurely based on a prespecified futility analysis.
MultiplicityBoth low-dose and high-dose comparisons were evaluated against placebo, alongside multiple secondary endpoints.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Explore the underlying clinical-trial and longitudinal-data concepts through focused tutorials and statistical calculators.

23. Record Summary

ENGAGE provides a detailed example of how a randomized phase 3 trial can analyze a longitudinal clinical outcome. The registry reports a single primary endpoint—change from baseline in CDR-SB at Week 78—with separate BIIB037 low-dose and high-dose comparisons against placebo. Both analyses use an MMRM and report model-based treatment differences, 95% two-sided confidence intervals, and p-values. Secondary MMSE, ADAS-Cog 13, and ADCS-ADL-MCI analyses use the same broad framework.

The statistical interpretation is inseparable from the trial's design. ENGAGE was randomized, parallel, quadruple masked, and classified under a superiority hypothesis. The study was subsequently halted prematurely following a prespecified futility analysis rather than because of safety concerns. The registry also identifies termination-related participant discontinuation in the "Reason not Specified" category of participant-flow tables.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical analysis from its statistical interpretation. For ENGAGE, that means reading the CDR-SB estimate together with its confidence interval and p-value, understanding why an MMRM was used, recognizing the role of the ITT population and longitudinal measurements, and accounting for premature termination and multiple treatment comparisons when interpreting the evidence.