This page separates reported trial results from statistical interpretation. The numerical results presented here are limited to the ClinicalTrials.gov record and its posted statistical analyses.
1. Trial at a Glance
TRAILBLAZER-ALZ 2 is a randomized, double-masked, parallel phase 3 study of donanemab versus placebo in participants with early Alzheimer's disease. The registry reports 1736 participants and two study arms, with 15 posted outcome measures and 15 posted statistical analyses.
| Feature | TRAILBLAZER-ALZ 2 |
|---|---|
| Phase | Phase 3 |
| Condition | Alzheimer Disease |
| Brief title | A Study of Donanemab (LY3002813) in Participants With Early Alzheimer's Disease (TRAILBLAZER-ALZ 2) |
| Design | Randomized, double-masked, parallel |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Interventions | Donanemab; placebo |
| Enrollment | 1736 |
| Status | Active, not recruiting |
| Lead sponsor | Eli Lilly and Company |
| Results | Posted |
| ClinicalTrials.gov | NCT04437511 |
2. Clinical Question
The statistical question is whether participants randomized to donanemab differ from participants randomized to placebo in change from baseline on the Integrated Alzheimer's Disease Rating Scale (iADRS) at Week 76.
Population
Participants with early Alzheimer's disease enrolled in the phase 3 TRAILBLAZER-ALZ 2 trial.
Intervention
Donanemab.
Comparator
Placebo.
Primary question
Does donanemab produce a different change from baseline in iADRS at Week 76 compared with placebo?
3. Trial Design
Donanemab
- Intervention: Donanemab
- Drug intervention
- Compared with placebo
Placebo
- Comparator: Placebo
- Drug intervention category in the registry
- Compared with donanemab
The study is described in the registry as randomized, parallel, and double-masked. Those design features are important statistically: randomization establishes the basis for the between-group comparison, the parallel structure means the randomized groups are followed as separate treatment groups, and masking is intended to reduce the influence of treatment knowledge on trial conduct and assessment.
4. Endpoints
The registry identifies two primary endpoints. Both use change from baseline to Week 76 on the Integrated Alzheimer's Disease Rating Scale (iADRS), with one analysis in the overall population and one in the intermediate (low-medium) tau population.
| Primary endpoint | Time frame | Endpoint type |
|---|---|---|
| Change From Baseline on the Integrated Alzheimer's Disease Rating Scale (iADRS) (Overall Population) | Baseline, Week 76 | Continuous |
| Change From Baseline on the Integrated Alzheimer's Disease Rating Scale (iADRS) (Intermediate (Low-medium) Tau Population) | Baseline, Week 76 | Continuous |
The registry describes iADRS as an integrated assessment of cognition and daily function comprised of items from the ADAS-Cog13 and the Alzheimer's disease cooperative study-instrumental activities of daily living scale (ADCS-iADL). The registry definition states that the scale ranges from 0 to 144 and is used to assess whether donanemab slows clinical decline associated with Alzheimer's disease compared with placebo.
5. Primary Results
Overall Population: iADRS
LS mean change difference at Week 76
95% CI: 1.508 to 4.331 · P < 0.001
Mixed models analysis; two-sided 95% confidence interval; superiority hypothesis.
| Feature | Reported result |
|---|---|
| Outcome | Change From Baseline on the Integrated Alzheimer's Disease Rating Scale (iADRS) (Overall Population) |
| Time frame | Baseline, Week 76 |
| Groups compared | Donanemab vs Placebo |
| Analysis population | All randomized participants with a baseline and at least one postbaseline iADRS data point. |
| Method | Mixed Models Analysis |
| Effect measure | LS Mean change difference (Final Values) |
| Estimate | 2.92 |
| 95% CI | 1.508 to 4.331 |
| P-value | <0.001 |
| Hypothesis | Superiority |
The reported estimate of 2.92 is the model-based difference in least-squares mean change from baseline between donanemab and placebo at the reported final-value analysis. The positive direction is consistent with the registry's stated purpose of assessing whether donanemab slows clinical decline relative to placebo.
The estimate does not mean that every participant experienced a 2.92-point difference, nor does it describe an individual patient's expected change. It is a population-level model estimate for the specified analysis population.
The 95% confidence interval, 1.508 to 4.331, describes statistical uncertainty around the estimated treatment-group difference under the analysis framework. It is not a range containing 95% of individual treatment effects or individual patient outcomes.
The P < 0.001 value addresses the statistical evidence against the null hypothesis under the specified test. It is not a measure of effect size, clinical importance, or the probability that the treatment hypothesis is true.
Because the analysis is longitudinal and model-based, interpretation depends on the mixed-effects model specification and the handling of repeated observations. The ClinicalTrials.gov record does not provide the complete model specification or a separate description of missing-data assumptions, so those features should not be inferred from the reported estimate alone.
Intermediate (Low-medium) Tau Population: iADRS
LS mean change difference at Week 76
95% CI: 1.883 to 4.618 · P < 0.001
Mixed models analysis; two-sided 95% confidence interval; superiority hypothesis.
| Feature | Reported result |
|---|---|
| Outcome | Change From Baseline on the Integrated Alzheimer's Disease Rating Scale (iADRS) (Intermediate (Low-medium) Tau Population) |
| Time frame | Baseline, Week 76 |
| Groups compared | Donanemab vs Placebo |
| Analysis population | All randomized participants with baseline Intermediate Tau level and with baseline and at least one postbaseline iADRS data point. |
| Method | Mixed Models Analysis |
| Effect measure | LS Mean change difference (Final Values) |
| Estimate | 3.25 |
| 95% CI | 1.883 to 4.618 |
| P-value | <0.001 |
| Hypothesis | Superiority |
The reported 3.25 represents the least-squares mean difference in change from baseline between the randomized groups within the prespecified intermediate (low-medium) tau population analyzed by the registry.
The confidence interval of 1.883 to 4.618 provides a measure of precision for this model-based estimate. It does not describe the variability of individual participants' responses.
The P < 0.001 result indicates strong statistical evidence against the null hypothesis used for this superiority comparison. It should not be converted into a percentage probability of benefit and should not be interpreted as a measure of the magnitude of the treatment effect.
The subgroup-specific estimate also should not automatically be interpreted as proof that the treatment effect differs from the overall population estimate. Establishing treatment-effect heterogeneity requires a formal interaction or other prespecified comparison; the ClinicalTrials.gov record does not report such an analysis.
6. Secondary Endpoint Results
The registry reports 13 secondary statistical analyses in addition to the two primary analyses. Most use mixed models analysis for continuous change-from-baseline outcomes; brain tau deposition is analyzed with ANCOVA.
Mini Mental State Examination
| Population | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Overall Population | 0.47 | 0.104 to 0.841 | 0.012 | Mixed Models Analysis |
| Intermediate (Low-medium) Tau Population | 0.48 | 0.089 to 0.868 | 0.016 | Mixed Models Analysis |
Both analyses use the outcome "Change From Baseline on the Mini Mental State Examination (MMSE) Score" at Baseline and Week 76. The overall-population analysis includes all randomized participants with a baseline and at least one postbaseline MMSE data point. The intermediate-tau analysis includes all randomized participants with baseline Intermediate Tau level and with baseline and at least one postbaseline MMSE data point.
Alzheimer's Disease Assessment Scale — Cognitive Subscale
| Population | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Overall Population | -1.33 | -2.086 to -0.565 | 0.0006 | Mixed Models Analysis |
| Intermediate (Low-medium) Tau Population | -1.52 | -2.250 to -0.794 | <0.001 | Mixed Models Analysis |
These analyses evaluate "Change From Baseline on the Alzheimer's Disease Assessment Scale - Cognitive Subscale (ADAS-Cog13)" at Baseline and Week 76. Both are superiority analyses using the LS Mean change difference (Final Values).
Clinical Dementia Rating Scale — Sum of Boxes
| Population | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Overall Population | -0.70 | -0.95 to -0.45 | <0.001 | Mixed Models Analysis |
| Intermediate (Low-medium) Tau Population | -0.67 | -0.95 to -0.40 | <0.001 | Mixed Models Analysis |
The registered outcome is "Change From Baseline on the Clinical Dementia Rating Scale-Sum of Boxes (CDR-SB)" with the time frame Baseline, Week 76. The reported analysis population for the overall analysis consists of all randomized participants with a baseline and at least one postbaseline CDR-SB data point. The intermediate-tau analysis additionally requires baseline Intermediate Tau level.
Alzheimer's Disease Cooperative Study — Instrumental Activities of Daily Living
| Population | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Overall Population | 1.70 | 0.840 to 2.566 | 0.0001 | Mixed Models Analysis |
| Intermediate (Low-medium) Tau Population | 1.83 | 0.913 to 2.748 | <0.001 | Mixed Models Analysis |
The outcome is "Change From Baseline on the Alzheimer's Disease Cooperative Study - Instrumental Activities of Daily Living (ADCS-iADL) Score," assessed at Baseline and Week 76. The overall analysis includes randomized participants with baseline and at least one postbaseline ADCS-iADL data point; the intermediate-tau analysis additionally requires baseline Intermediate Tau level.
Brain Amyloid Plaque Deposition
Change in amyloid PET measurement
95% CI: -88.87 to -83.87 · P < 0.0001
Outcome unit: centiloids. Analysis: mixed models.
The registered outcome is "Change From Baseline in Brain Amyloid Plaque Deposition as Measured by Amyloid Positron Emission Tomography (PET) Scan," with time frame Baseline, Week 76. The analysis population consists of all randomized participants with a baseline and at least one postbaseline amyloid PET scan data point.
Brain Tau Deposition
Change in tau PET measurement
95% CI: -0.0148 to 0.0066 · P = 0.4522
Outcome unit: standardized uptake value ratio (SUVR). Analysis: ANCOVA.
The registered outcome is "Change From Baseline in Brain Tau Deposition as Measured by Flortaucipir F18 PET Scan," assessed from Baseline to Week 76. The analysis population includes all randomized participants with a baseline and post-baseline tau PET scan. Unlike the other reported continuous outcomes, this endpoint uses ANCOVA.
The point estimate of -0.0041 is accompanied by a 95% confidence interval extending from -0.0148 to 0.0066. Because the interval includes zero, the reported estimate is compatible with both a negative and positive difference under the specified statistical framework.
The P = 0.4522 value does not measure the size of the observed difference. It describes the evidence against the null hypothesis used for this comparison. A nonsignificant P-value also does not establish that the two treatments are equivalent or that their effects are exactly the same.
Brain Volume: Bilateral Hippocampus
Change in bilateral hippocampal volume
95% CI: 0.01 to 0.04 cm³ · P = 0.002
Outcome unit: cubic centimeter (cm3). Analysis: mixed models.
The registry identifies the outcome as "Change From Baseline in Brain Volume as Measured by Volumetric Magnetic Resonance Imaging (vMRI)" and specifies the analysis note "Bilateral Hippocampus." The time frame is Baseline, Week 76.
Brain Volume: Bilateral Whole Brain
Change in bilateral whole-brain volume
95% CI: -7.76 to -5.56 cm³ · P < 0.001
Outcome unit: cubic centimeter (cm3). Analysis: mixed models.
The registry specifies "Bilateral Whole Brain" as the analysis note for this vMRI outcome. The analysis population includes all randomized participants with a baseline and at least one postbaseline vMRI data point.
Brain Volume: Bilateral Ventricles
Change in bilateral ventricular volume
95% CI: 2.52 to 3.52 cm³ · P < 0.001
Outcome unit: cubic centimeter (cm3). Analysis: mixed models.
The registry specifies "Bilateral Ventricles" as the analysis note for this vMRI outcome. The time frame is Baseline, Week 76, and the analysis population is all randomized participants with a baseline and at least one postbaseline vMRI data point.
7. Secondary Results: Consolidated View
| Outcome | Population | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| MMSE | Overall | 0.47 | 0.104 to 0.841 | 0.012 |
| MMSE | Intermediate Tau | 0.48 | 0.089 to 0.868 | 0.016 |
| ADAS-Cog13 | Overall | -1.33 | -2.086 to -0.565 | 0.0006 |
| ADAS-Cog13 | Intermediate Tau | -1.52 | -2.250 to -0.794 | <0.001 |
| CDR-SB | Overall | -0.70 | -0.95 to -0.45 | <0.001 |
| CDR-SB | Intermediate Tau | -0.67 | -0.95 to -0.40 | <0.001 |
| ADCS-iADL | Overall | 1.70 | 0.840 to 2.566 | 0.0001 |
| ADCS-iADL | Intermediate Tau | 1.83 | 0.913 to 2.748 | <0.001 |
| Amyloid PET | Overall | -86.37 | -88.87 to -83.87 | <0.0001 |
| Tau PET | Overall | -0.0041 | -0.0148 to 0.0066 | 0.4522 |
| Bilateral hippocampus vMRI | Overall | 0.02 | 0.01 to 0.04 | 0.002 |
| Bilateral whole-brain vMRI | Overall | -6.66 | -7.76 to -5.56 | <0.001 |
| Bilateral ventricles vMRI | Overall | 3.02 | 2.52 to 3.52 | <0.001 |
This table is useful for seeing the statistical pattern without treating the different measurement scales as interchangeable. A change of 1 unit on one instrument cannot be interpreted as numerically equivalent to a change of 1 unit on another instrument. The sign of an estimate also depends on the direction in which the outcome is coded and on whether an increase or decrease represents improvement for that particular measure.
8. Statistical Methodology
Mixed-effects models
The registry reports "Mixed Models Analysis" for both primary iADRS endpoints and for most secondary continuous outcomes. Mixed-effects models are designed for repeated or longitudinal observations in which measurements from the same participant are correlated.
The fixed-effects component represents systematic treatment and time-related differences, while random effects can represent participant-level variation and account for correlation among repeated measurements.
For this trial, the registry reports an effect measure of LS Mean change difference (Final Values). Least-squares means are model-based means adjusted according to the covariates and structure specified in the fitted model. They are not necessarily the same as simple arithmetic means calculated directly from observed final scores.
Why use a longitudinal model?
A longitudinal model can use repeated measurements rather than reducing each participant's follow-up to a single observed value. This can be statistically efficient when the model is appropriately specified and when its assumptions are reasonable.
The ClinicalTrials.gov record does not provide the full covariance structure, fixed-effect terms, random-effect terms, visit structure, or missing-data assumptions. Those details matter when reproducing an analysis and should not be reconstructed from the final estimate alone.
ANCOVA
The registry reports ANCOVA for the change from baseline in brain tau deposition measured by Flortaucipir F18 PET scan. ANCOVA is a linear-model approach that can compare treatment groups while accounting for prespecified covariates, such as a baseline measurement, when included in the model.
The treatment coefficient represents the adjusted between-group difference under the fitted model. The actual ClinicalTrials.gov record does not list the complete ANCOVA covariate specification.
Mean difference
The effect measure for all 15 posted analyses is a mean difference. For the mixed-model analyses, the reported form is specifically the LS Mean change difference (Final Values). The units therefore remain those of the underlying endpoint: score units for clinical scales, centiloids for amyloid PET, SUVR for tau PET, and cm3 for vMRI.
Superiority testing
The registered analyses use a superiority hypothesis. In a superiority framework, the statistical question is whether the treatment groups differ, rather than whether one treatment is no worse than another by a prespecified non-inferiority margin.
9. Confidence Intervals and P-values
Every statistical analysis reported in the registry reports a two-sided 95% confidence interval. The confidence interval and P-value answer related but different questions.
Confidence interval
The 95% CI describes uncertainty around the estimated treatment difference under the model and sampling framework. Narrower intervals indicate greater statistical precision than wider intervals, all else equal.
P-value
The P-value quantifies how compatible the observed result is with the null hypothesis under the specified test. It does not measure the magnitude of the treatment effect.
Effect estimate
The estimate describes the modeled between-group difference. Its practical meaning depends on the endpoint's measurement scale and direction.
Statistical significance
A small P-value does not by itself establish clinical importance, and a larger P-value does not establish equivalence.
10. Statistical Methods Explained
Why was a mixed-effects model used for the iADRS endpoint?
The registry identifies the primary iADRS analysis as a Mixed Models Analysis. Because the endpoint is assessed longitudinally from baseline through Week 76, a mixed-effects framework can account for the correlation among repeated measurements within participants while estimating treatment-related differences in change over time.
What does an LS Mean change difference of 2.92 mean?
It is the reported model-based difference between the treatment groups in least-squares mean change from baseline at the final reported analysis. It is not the average raw difference for every participant and does not mean that each participant's outcome differed by exactly 2.92 points.
Why does the confidence interval matter?
The estimate alone provides only one point on the statistical scale. The 95% CI of 1.508 to 4.331 for the overall iADRS analysis shows the range of values compatible with the model and sampling uncertainty represented by that interval. It also provides more information about precision than the P-value alone.
Why is P < 0.001 not an effect-size measure?
A P-value depends on both the size of the observed difference and the amount of statistical information available. A relatively small effect can have a small P-value in a large study, while a potentially meaningful effect can have a larger P-value when uncertainty is substantial. Effect estimates and confidence intervals therefore need to be considered alongside P-values.
Why was ANCOVA used for the tau PET endpoint?
The registry specifically identifies ANCOVA as the method for change from baseline in brain tau deposition measured by Flortaucipir F18 PET. ANCOVA is a linear modeling approach suited to continuous outcomes and can incorporate baseline information or other prespecified covariates. The ClinicalTrials.gov record does not specify the complete covariate list, so the exact fitted model should not be inferred.
Does a nonsignificant P-value prove that two groups are equivalent?
No. The tau PET result has an estimate of -0.0041, a 95% CI of -0.0148 to 0.0066, and P = 0.4522. These results do not establish equivalence. Equivalence requires an equivalence design and prespecified equivalence margins; the ClinicalTrials.gov record identifies the hypothesis type as superiority.
11. Analysis Populations
The registry's posted analyses explicitly define their analysis populations. These definitions are important because the denominator for an analysis can differ from the total enrollment when an outcome requires baseline and postbaseline measurements.
| Analysis | Analysis population |
|---|---|
| Overall iADRS | All randomized participants with a baseline and at least one postbaseline iADRS data point. |
| Intermediate-tau iADRS | All randomized participants with baseline Intermediate Tau level and with baseline and at least one postbaseline iADRS data point. |
| Overall MMSE | All randomized participants with a baseline and at least one postbaseline MMSE data point. |
| Intermediate-tau MMSE | All randomized participants with baseline Intermediate Tau level and with baseline and at least one postbaseline MMSE data point. |
| Overall ADAS-Cog13 | All randomized participants with a baseline and at least one postbaseline ADAS-Cog13 data point. |
| Intermediate-tau ADAS-Cog13 | All randomized participants with baseline Intermediate Tau level and with baseline and at least one postbaseline ADAS-Cog13 data point. |
| Overall CDR-SB | All randomized participants with a baseline and at least one postbaseline CDR-SB data point. |
| Intermediate-tau CDR-SB | All randomized participants with baseline Intermediate Tau level and with baseline and at least one postbaseline CDR-SB data point. |
| Overall ADCS-iADL | All randomized participants with a baseline and at least one postbaseline ADCS-iADL data point. |
| Intermediate-tau ADCS-iADL | All randomized participants with a baseline Intermediate Tau level and with baseline and at least one postbaseline ADCS-iADL data point. |
| Amyloid PET | All randomized participants with a baseline and at least one postbaseline amyloid PET scan data point. |
| Tau PET | All randomized participants with a baseline and post-baseline tau PET scan. |
| vMRI | All randomized participants with a baseline and at least one postbaseline vMRI data point. |
These definitions illustrate an important statistical distinction: randomized enrollment and the analysis population for a particular outcome are not necessarily identical. An outcome requiring at least one postbaseline measurement cannot include a participant for whom no qualifying postbaseline measurement exists.
12. Safety
The registry reports serious adverse events by randomized treatment arm as follows:
| Treatment arm | Serious adverse events affected / at risk |
|---|---|
| Donanemab | 148 / 853 |
| Placebo | 138 / 874 |
The reported denominators for this safety measure are different from the overall enrollment of 1736. The registry-the ClinicalTrials.gov record do not provide additional information here about the timing, individual adverse-event terms, severity distribution, causality, or formal between-group hypothesis test for serious adverse events.
13. What the Primary Estimates Do — and Do Not — Mean
The reported estimate of 2.92 is a model-based difference in change from baseline between donanemab and placebo at Week 76. The confidence interval of 1.508 to 4.331 communicates uncertainty around that estimate, while P < 0.001 addresses the statistical test of the superiority hypothesis.
It does not mean that every participant experienced a 2.92-point treatment difference, nor does it quantify an individual patient's probability of improvement.
The estimate of 3.25 represents the reported LS Mean change difference in the specified intermediate-tau population. Its 95% CI is 1.883 to 4.618, with P < 0.001.
This subgroup-specific estimate should be read as an estimate within the defined analysis population. It does not by itself demonstrate that the effect differs from the effect in the overall population.
The trial reports multiple continuous endpoints with different units and measurement scales. A positive estimate on one instrument and a negative estimate on another cannot be compared numerically without first considering how each outcome is defined and coded.
14. Interpreting the Pattern Across Endpoints
The posted results include clinical rating scales, a PET measure of amyloid plaque deposition, a PET measure of tau deposition, and three vMRI measures. These endpoints operate on different statistical and scientific scales.
| Endpoint family | Measurement | Reported method | Interpretive issue |
|---|---|---|---|
| Clinical function/cognition | iADRS, MMSE, ADAS-Cog13, CDR-SB, ADCS-iADL | Mixed models | Direction and magnitude depend on each scale's definition. |
| Amyloid pathology | Centiloids | Mixed models | Estimate is expressed in centiloids, not clinical score units. |
| Tau pathology | SUVR | ANCOVA | CI crosses zero and P = 0.4522 in the posted analysis. |
| Brain structure | cm3 | Mixed models | Separate estimates are reported for hippocampus, whole brain, and ventricles. |
A statistical analysis should therefore resist the temptation to combine these outcomes into a single numerical "effect." They measure different constructs, use different units, and are analyzed according to the methods specified for each endpoint.
15. Multiplicity and Multiple Endpoints
The registry reports 2 primary endpoints and a total of 15 statistical analyses. The primary endpoints are both versions of the iADRS endpoint: one for the overall population and one for the intermediate (low-medium) tau population.
| Analysis group | Number reported | Role |
|---|---|---|
| Primary analyses | 2 | Overall iADRS and intermediate-tau iADRS |
| Secondary analyses | 13 | Clinical, imaging, and structural outcomes |
| Total statistical analyses | 15 | Posted on ClinicalTrials.gov |
Because multiple endpoints are evaluated, interpretation of individual P-values should be tied to the trial's prespecified multiplicity strategy when one is available. The ClinicalTrials.gov record identifies superiority hypotheses and provide the individual analyses, but they do not provide an alpha-allocation or multiplicity-adjustment procedure. It would therefore be inappropriate to invent a hierarchy or claim that every secondary P-value was independently confirmatory.
16. Randomization and Blinding
The registry classifies the allocation as randomized and the masking as double. Randomization is central to causal comparison because treatment assignment is determined independently of participants' subsequent outcomes, subject to the trial's actual implementation. Double masking is intended to reduce the possibility that knowledge of assignment influences participant behavior, investigator decisions, outcome assessment, or other trial processes.
Randomization
Creates the design basis for comparing donanemab with placebo while reducing systematic differences in treatment assignment.
Double masking
Reduces the opportunity for treatment knowledge to influence trial conduct and assessment.
The ClinicalTrials.gov record does not provide a randomization ratio or a list of stratification factors, so neither should be inferred from the two-arm design alone.
17. Trial Timeline
Study start
The registry lists June 19, 2020 as the trial start date.
Randomized parallel study
The study is classified as a randomized, double-masked, parallel phase 3 trial with donanemab and placebo arms.
Primary endpoint assessment
Both registered primary endpoints evaluate change from baseline on iADRS at Week 76.
Primary completion
The registry lists April 14, 2023 as the primary completion date.
Active, not recruiting
The ClinicalTrials.gov record classifies the study as ACTIVE_NOT_RECRUITING.
18. Important Limitations and Interpretation Issues
- Registry-level reporting: this analysis is restricted to the numerical results and methodological fields from the ClinicalTrials.gov record.
- Incomplete model specification: the registry identifies mixed models and ANCOVA but does not provide the complete model equations, covariance structure, covariate list, or all estimation details in the ClinicalTrials.gov record.
- Missing-data assumptions: the analysis populations require baseline and postbaseline observations, but the ClinicalTrials.gov record does not describe the complete missing-data or imputation strategy.
- Multiple endpoints: 15 statistical analyses are posted. Without a registry-reported multiplicity procedure, individual P-values should not automatically be treated as independently confirmatory.
- Subgroup interpretation: the intermediate (low-medium) tau population is a defined analysis population, but a difference between its estimate and another population's estimate would require formal interaction testing to establish effect modification.
- Different measurement scales: iADRS, MMSE, ADAS-Cog13, CDR-SB, ADCS-iADL, centiloids, SUVR, and cm3 cannot be placed on a common numerical scale.
- Safety denominators: serious adverse events are reported with denominators of 853 for donanemab and 874 for placebo, which differ from the overall enrollment of 1736.
- Statistical versus clinical importance: P-values and confidence intervals describe statistical uncertainty; they do not independently determine the clinical importance of a given score difference.
- Superiority design: a nonsignificant superiority test does not establish equivalence or non-inferiority.
19. Why This Trial Matters Statistically
TRAILBLAZER-ALZ 2 is a useful statistical teaching case because it combines randomized treatment comparison with repeated clinical measurements and several distinct continuous outcome types.
| Concept | How it appears in TRAILBLAZER-ALZ 2 |
|---|---|
| Randomization | The registry classifies the allocation as randomized. |
| Blinding | The study is double-masked. |
| Parallel design | Two treatment arms are evaluated in parallel. |
| Longitudinal analysis | Primary and most secondary outcomes are evaluated from baseline to Week 76 using mixed models. |
| Mixed-effects model | Used for the two primary iADRS analyses and most secondary continuous outcomes. |
| ANCOVA | Used for change from baseline in brain tau deposition measured by Flortaucipir F18 PET. |
| Least-squares means | The reported effect measure is LS Mean change difference (Final Values). |
| Confidence intervals | All registry-reported analyses report two-sided 95% CIs. |
| P-values | Each posted statistical analysis includes a P-value. |
| Subpopulation analysis | The second primary endpoint and several secondary outcomes are reported for the intermediate (low-medium) tau population. |
| Multiple endpoints | Two primary analyses and 13 secondary analyses are posted. |
| Safety analysis | Serious adverse events are reported separately by treatment arm. |
20. A Practical Framework for Reading the Results
A useful way to read this trial is to proceed in layers rather than beginning with the smallest P-value.
For example, the overall iADRS analysis should first be identified as a Week 76 change-from-baseline endpoint analyzed in randomized participants with baseline and postbaseline data. The estimate is then read as a model-based LS mean change difference of 2.92, followed by its 95% CI of 1.508 to 4.331 and P < 0.001. Only after those quantities are understood should the result be interpreted in the context of the randomized, double-masked trial design and the limitations of the ClinicalTrials.gov record.
The same discipline is particularly important for the tau PET analysis. Its estimate of -0.0041, 95% CI of -0.0148 to 0.0066, and P = 0.4522 cannot be interpreted simply by comparing the P-value with those of other endpoints. The measurement scale, ANCOVA method, analysis population, confidence interval, and superiority hypothesis all form part of the statistical context.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Calculators
23. Sources
- ClinicalTrials.gov: TRAILBLAZER-ALZ 2, NCT04437511.
- PubMed: PMID 42644303.
- PubMed: PMID 42552749.
- PubMed: PMID 42273802.
- PubMed: PMID 42128444.
- PubMed: PMID 41804324.
Continue through the Clinical Biostats statistical pathway
Connect the trial's clinical endpoints to tutorials on study design, inference, and longitudinal analysis, then explore statistical calculators for the underlying methods.
24. Record Summary
TRAILBLAZER-ALZ 2 provides a detailed example of how randomized clinical-trial evidence can be analyzed across repeated clinical and imaging outcomes. The two primary analyses evaluate change from baseline in iADRS at Week 76, with reported LS Mean change differences of 2.92 in the overall population and 3.25 in the intermediate (low-medium) tau population. Both use mixed models analysis, report two-sided 95% confidence intervals, and have P-values below 0.001.
The secondary analyses extend the same framework across MMSE, ADAS-Cog13, CDR-SB, ADCS-iADL, amyloid PET, tau PET, and vMRI outcomes. Most use mixed models, while tau PET uses ANCOVA. The statistical results should be interpreted endpoint by endpoint because the outcomes have different measurement scales, analysis populations, and model specifications.
The trial also illustrates why an analysis cannot be reduced to its P-values. Estimates describe the magnitude and direction of the modeled differences; confidence intervals describe their statistical precision; P-values address compatibility with the null hypothesis; and the design and analysis population determine the context in which those quantities should be understood.