This page separates reported trial results from statistical interpretation. The numerical results presented here are limited to the ClinicalTrials.gov record for NCT02614196. ClinicalTrials.gov provides the official trial registry record. View the ClinicalTrials.gov record.
1. Trial at a Glance
EVOLVE-2 was a randomized, double-blind, parallel phase 3 trial evaluating galcanezumab versus placebo in migraine. The registry reports 986 enrolled participants, six arms, 10 posted outcome measures, and 20 posted statistical analyses.
| Feature | EVOLVE-2 |
|---|---|
| Trial name | EVOLVE-2 |
| Brief title | Evaluation of Efficacy & Safety of Galcanezumab in the Prevention of Episodic Migraine- the EVOLVE-2 Study |
| Phase | Phase 3 |
| Condition | Migraine |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 986 |
| Arms | 6 |
| Interventions | Galcanezumab; Placebo |
| Trial status | Completed |
| Lead sponsor | Eli Lilly and Company |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT02614196 |
The registry identifies the primary hypothesis type as superiority. The posted statistical analyses use two principal method categories: mixed-effects models for continuous longitudinal outcomes and Fisher exact tests for the reported categorical anti-drug-antibody analysis.
2. Clinical Question
The primary statistical question was whether treatment with galcanezumab was associated with a different overall mean change from baseline in the number of monthly migraine headache days compared with placebo over the registered period from baseline through Month 6.
Population
Participants enrolled in the phase 3 EVOLVE-2 trial for the condition of migraine.
Intervention
Galcanezumab, with the posted primary comparisons separately evaluating 120 mg and 240 mg.
Comparator
Placebo.
Primary question
Does galcanezumab produce a different overall mean change from baseline in monthly migraine headache days than placebo over Baseline, Month 1 through Month 6?
3. Trial Design
Active-treatment comparison
- Compared with placebo for the primary endpoint.
- Primary efficacy analysis used a mixed models analysis.
- Posted primary estimate: LSMean Difference -2.02.
Active-treatment comparison
- Compared with placebo for the primary endpoint.
- Primary efficacy analysis used a mixed models analysis.
- Posted primary estimate: LSMean Difference -1.90.
Control treatment
- Served as the comparator for both posted primary efficacy analyses.
- Serious adverse events during the treatment phase: 5/461.
- Serious adverse events during the post-treatment phase: 3/410.
Six-arm parallel trial
- Allocation: randomized.
- Masking: double.
- Primary purpose: treatment.
The ClinicalTrials.gov record identifies six arms but do not provide a complete arm-by-arm description of all six arms. Accordingly, this page does not infer the remaining arm assignments or sample sizes from the total enrollment.
4. Endpoints
| Endpoint | Time frame | Type | Posted statistical method |
|---|---|---|---|
| Overall Mean Change From Baseline in the Number of Monthly Migraine Headache Days | Baseline, Month 1 through Month 6 | Continuous | Mixed Models Analysis |
| Mean Percentage of Participants With Reduction From Baseline ≥50%, ≥75%, and 100% in Monthly Migraine Headache Days | Baseline, Month 1 through Month 6 | Binary | Not reported |
| Mean Change From Baseline in the Migraine-Specific Quality of Life Questionnaire (MSQ) Version 2.1 Role Function Restrictive Domain | Baseline, Month 4 through Month 6 | Continuous | Mixed Models Analysis |
| Overall Mean Change From Baseline in the Number of Monthly Migraine Headache Days Requiring Medication for the Acute Treatment of Migraine or Headache | Baseline, Month 1 through Month 6 | Continuous | Mixed Models Analysis |
| Mean Change From Baseline in Patient Global Impression of Severity (PGI-S) Rating | Baseline, Month 4 through Month 6 | Continuous | Mixed Models Analysis |
| Overall Mean Change From Baseline in Headache Hours | Baseline, Month 1 through Month 6 | Continuous | Mixed Models Analysis |
| Mean Change From Baseline in the Migraine Disability Assessment Test (MIDAS) Total Score | Baseline, Month 6 | Continuous | Mixed Models Analysis |
| Percentage of Participants Developing Anti-drug Antibodies (ADA) to Galcanezumab | Month 1 through Month 6 | Binary | Fisher Exact |
The registry identifies one primary endpoint type as “Other / unclear.” Its registered name is Overall Mean Change From Baseline in the Number of Monthly Migraine Headache Days, with the time frame Baseline, Month 1 through Month 6.
Registered primary endpoint definition
5. Analysis Populations
| Analysis | Population / information reported |
|---|---|
| Primary monthly migraine headache days | All randomized participants who received at least one dose of study drug and had baseline and at least one post baseline value. |
| Responder endpoint | All randomized participants who received at least one dose of study drug and had baseline and at least one post baseline value. |
| MSQ Role Function Restrictive Domain | All randomized participants who received at least one dose of study drug and had baseline and at least one post baseline value. |
| PGI-S | All randomized participants who received at least one dose of study drug and had baseline and post baseline value. |
| ADA | All randomized participants who received at least one dose of study drug and had at least one non-missing test result for ADA for each of the baseline period and the post-baseline period. |
6. Primary Results
Two formal primary analyses were posted for the registered endpoint, one comparing placebo with galcanezumab 120 mg and one comparing placebo with galcanezumab 240 mg. Both used a mixed models analysis and reported least-squares mean differences with two-sided 95% confidence intervals.
Galcanezumab 120 mg vs Placebo
Overall mean change in monthly migraine headache days
LSMean Difference · 95% CI: -2.55 to -1.48 · P < .001
Time frame: Baseline, Month 1 through Month 6
The reported LSMean Difference of -2.02 is the model-based difference in adjusted mean change from baseline between the two randomized treatment comparisons as reported by the registry. Because the comparison is expressed as placebo versus galcanezumab 120 mg, the negative value indicates a more negative estimated change in the galcanezumab comparison under the registry's reported direction.
The 95% confidence interval of -2.55 to -1.48 describes statistical uncertainty around the estimated difference. It does not mean that 95% of individual participants experienced a change within that interval, nor does it describe the range of individual treatment responses.
The reported P < .001 addresses evidence against the null hypothesis under the specified statistical test; it does not measure the magnitude or clinical importance of the effect. Effect size and uncertainty are better conveyed by considering the LSMean Difference together with its confidence interval.
Because this is a longitudinal mixed-model result, interpretation depends on the model specification and assumptions underlying the repeated observations. The ClinicalTrials.gov record identifies the method as a mixed models analysis but do not provide enough detail to independently assess covariance structure, fixed-effect specification, or missing-data assumptions.
Galcanezumab 240 mg vs Placebo
Overall mean change in monthly migraine headache days
LSMean Difference · 95% CI: -2.44 to -1.36 · P < .001
Time frame: Baseline, Month 1 through Month 6
The reported LSMean Difference of -1.90 is the model-based difference in mean change from baseline for the placebo versus galcanezumab 240 mg comparison. As with the 120 mg analysis, the negative estimate indicates a more negative estimated change in the galcanezumab comparison under the reported direction of the effect measure.
The 95% confidence interval of -2.44 to -1.36 quantifies uncertainty around the estimated mean difference. Its width is relevant because a point estimate alone does not communicate statistical precision.
The reported P < .001 is evidence against the null hypothesis for this formal comparison under the posted analysis. It is not a measure of effect size, probability that the treatment is effective, or probability that the null hypothesis is true.
The result comes from a mixed-effects longitudinal analysis. A mixed model is particularly useful when an endpoint is observed repeatedly over time because it can model the pattern of observations jointly rather than treating every time point as an unrelated analysis. The registry excerpt does not specify all model details, so more specific claims about covariance structure or imputation would go beyond the ClinicalTrials.gov record.
Primary-result comparison
| Comparison | Method | LSMean Difference | 95% CI | P-value |
|---|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | Mixed Models Analysis | -2.02 | -2.55 to -1.48 | <.001 |
| Placebo vs Galcanezumab 240 mg | Mixed Models Analysis | -1.90 | -2.44 to -1.36 | <.001 |
The two posted primary estimates are close in magnitude, but a numerical difference between point estimates does not establish that the two galcanezumab doses have different treatment effects. A formal dose-comparison or interaction analysis would be needed for that question, and none is reported in the ClinicalTrials.gov record.
7. Secondary Results: Migraine Headache Day Responders
The registry reports a secondary endpoint defined as the Mean Percentage of Participants With Reduction From Baseline ≥50%, ≥75%, and 100% in Monthly Migraine Headache Days. Six statistical analyses are posted for this endpoint: three response thresholds for each of the two galcanezumab dose comparisons. The analyses posted on ClinicalTrials.gov report odds ratios and 95% confidence intervals, but no P-values.
| Comparison | Reduction threshold | Odds ratio | 95% CI |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | ≥50% | 2.60 | 2.03 to 3.32 |
| Placebo vs Galcanezumab 240 mg | ≥50% | 2.31 | 1.81 to 2.96 |
| Placebo vs Galcanezumab 120 mg | ≥75% | 2.34 | 1.78 to 3.06 |
| Placebo vs Galcanezumab 240 mg | ≥75% | 2.42 | 1.84 to 3.17 |
| Placebo vs Galcanezumab 120 mg | ≥100% | 2.16 | 1.50 to 3.12 |
| Placebo vs Galcanezumab 240 mg | ≥100% | 2.67 | 1.87 to 3.81 |
The registry does not report a formal statistical method for these six odds-ratio analyses in the ClinicalTrials.gov record. Therefore, this page reports the posted odds ratios and confidence intervals without assigning an unreported regression model or hypothesis-testing procedure to them.
How to interpret an odds ratio
An odds ratio of 2.60 means that the estimated odds of meeting the specified response threshold are 2.60 times the odds in the comparison group, according to the analysis as reported. It is not equivalent to saying that the probability of response is 2.60 times as large.
This distinction matters especially when the outcome is common. Odds and probabilities are mathematically related but are not interchangeable. The actual responder percentages are not included in the ClinicalTrials.gov record, so they cannot be reconstructed from the odds ratios alone.
8. Secondary Results: Quality of Life
MSQ Version 2.1 Role Function Restrictive Domain
| Comparison | LSMean Difference | 95% CI | P-value |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | 8.82 | 6.33 to 11.31 | <.001 |
| Placebo vs Galcanezumab 240 mg | 7.39 | 4.88 to 9.90 | <.001 |
The registered time frame was Baseline, Month 4 through Month 6. Both analyses used a mixed models analysis and reported LSMean Difference as the effect measure.
The estimates are expressed in units on a scale. A positive difference therefore represents a higher modeled change in the galcanezumab comparison relative to placebo under the reported direction of the effect measure. The registry data do not supply a minimally important difference for this questionnaire domain, so statistical evidence should not be converted into a claim about clinical importance without an external prespecified threshold.
9. Secondary Results: Acute-Treatment Medication Days
| Comparison | LSMean Difference | 95% CI | P-value |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | -1.82 | -2.29 to -1.36 | <.001 |
| Placebo vs Galcanezumab 240 mg | -1.78 | -2.25 to -1.31 | <.001 |
The endpoint was Overall Mean Change From Baseline in the Number of Monthly Migraine Headache Days Requiring Medication for the Acute Treatment of Migraine or Headache, with the time frame Baseline, Month 1 through Month 6. Both analyses used mixed models analysis.
These results are a useful example of why the endpoint definition matters. The outcome is not simply all migraine headache days; it specifically concerns monthly migraine headache days requiring medication for acute treatment. The negative estimates indicate lower modeled change in the galcanezumab comparison under the registry's reported direction.
10. Secondary Results: Patient Global Impression of Severity
| Comparison | LSMean Difference | 95% CI | P-value |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | -0.29 | -0.47 to -0.11 | .002 |
| Placebo vs Galcanezumab 240 mg | -0.23 | -0.41 to -0.05 | .012 |
The registered endpoint was Mean Change From Baseline in Patient Global Impression of Severity (PGI-S) Rating, assessed from Baseline, Month 4 through Month 6. The registry classifies this as a continuous endpoint measured in units on a scale and reports mixed models analysis.
Both confidence intervals lie below zero, consistent with the direction of the posted estimates. The P-values, .002 and .012, indicate evidence against the respective null hypotheses under the reported analyses; they do not themselves tell us whether the magnitude of the change is clinically meaningful.
11. Secondary Results: Headache Hours
| Comparison | LSMean Difference | 95% CI | P-value |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | -15.19 | -20.27 to -10.11 | <.001 |
| Placebo vs Galcanezumab 240 mg | -13.56 | -18.67 to -8.44 | <.001 |
The endpoint was Overall Mean Change From Baseline in Headache Hours, measured in headache hours per month from Baseline, Month 1 through Month 6. Both comparisons used mixed models analysis.
Here the effect measure is expressed directly in hours per month rather than as an odds ratio or standardized effect. This makes the unit of the estimated difference immediately interpretable: the model estimates a difference of -15.19 headache hours per month for the 120 mg comparison and -13.56 for the 240 mg comparison, according to the registry's reported direction.
12. Secondary Results: MIDAS Total Score
| Comparison | LSMean Difference | 95% CI | P-value |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | -9.15 | -12.61 to -5.69 | <.001 |
| Placebo vs Galcanezumab 240 mg | -8.22 | -11.71 to -4.72 | <.001 |
The registered endpoint was Mean Change From Baseline in the Migraine Disability Assessment Test (MIDAS) Total Score, assessed at Baseline, Month 6. Both comparisons used mixed models analysis.
The negative LSMean Differences indicate a more negative modeled change in the galcanezumab comparisons under the reported effect direction. The confidence intervals provide substantially more information than the point estimates alone because they show the range of values compatible with the statistical model and observed data under the stated confidence level.
13. Secondary Results: Anti-drug Antibodies
| Comparison | Method | Analysis note | P-value |
|---|---|---|---|
| Placebo vs Galcanezumab 120 mg | Fisher Exact | TE ADA Positive | <.001 |
| Placebo vs Galcanezumab 240 mg | Fisher Exact | TE ADA Positive | <.001 |
The registered endpoint was Percentage of Participants Developing Anti-drug Antibodies (ADA) to Galcanezumab, with the time frame Month 1 through Month 6. The analysis population required at least one non-missing ADA test result for each of the baseline and post-baseline periods. The registry reports Fisher Exact as the statistical method and “TE ADA Positive” as the analysis note.
No effect estimate or confidence interval is posted on ClinicalTrials.gov for these analyses in the ClinicalTrials.gov record. Accordingly, the result is reported as a Fisher exact test with the posted P-value rather than converting the P-value into an unreported effect measure.
14. Statistical Methodology
Mixed-effects model
The dominant method in the posted EVOLVE-2 analyses is the mixed-effects model. It was used for the primary monthly migraine headache-day endpoint and for several continuous secondary outcomes, including MSQ Role Function Restrictive Domain, acute-treatment medication days, PGI-S, headache hours, and MIDAS.
The important statistical idea is that repeated observations from the same participant are related. A mixed model can represent population-level treatment effects while accounting for within-participant correlation through random effects and/or the model's covariance structure.
The registry specifically calls the primary method “Mixed Models Analysis” and normalizes it to “mixed-effects model.” It does not supply the complete model specification in the ClinicalTrials.gov record. Therefore, this analysis does not claim a particular covariance structure, link function, baseline adjustment, interaction structure, or imputation rule.
Least-squares mean difference
The continuous outcomes report LSMean Difference. A least-squares mean is a model-based adjusted mean rather than simply the arithmetic average observed in each group. The difference compares the model-estimated means under the specified analysis framework.
This distinction is important when data are collected repeatedly. The reported LSMean Difference summarizes the treatment contrast generated by the longitudinal model, rather than being a simple subtraction of two raw endpoint averages.
Fisher exact test
The anti-drug-antibody analyses use the Fisher exact test, a method for testing association between categorical outcomes in a two-group comparison. It is particularly useful when sample sizes or cell counts are small enough that large-sample approximations may be questionable.
The ClinicalTrials.gov record identifies the endpoint as binary and the method as Fisher Exact. It does not provide the underlying 2×2 cell counts or an effect estimate, so those quantities cannot be reconstructed here.
Confidence intervals
The posted continuous and responder analyses use two-sided 95% confidence intervals. A confidence interval communicates statistical precision around an estimate. For example, the primary 120 mg estimate of -2.02 is accompanied by an interval from -2.55 to -1.48.
A confidence interval should not be interpreted as a probability statement about the fixed true effect. Rather, it is a procedure that, under its statistical assumptions, produces intervals with the stated long-run coverage property.
P-values
The registry reports P-values for the primary analyses and several secondary continuous outcomes. A P-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model. It does not measure the size of an effect and does not provide the probability that the null hypothesis is true.
15. Statistical Methods Explained
Why was a mixed-effects model used?
The primary endpoint is measured repeatedly from baseline through Month 6. Repeated measurements from the same participant are not independent observations. A mixed-effects model is designed for this longitudinal setting because it can account for participant-level correlation while estimating treatment-related differences over the assessment period.
What does an LSMean Difference of -2.02 mean?
It is the reported model-based difference in mean change from baseline for the placebo versus galcanezumab 120 mg comparison. The negative direction indicates a more negative modeled change in the galcanezumab comparison according to the registry's effect-direction convention. It does not mean that every participant experienced exactly 2.02 fewer migraine headache days.
Why does the confidence interval matter?
The 95% CI of -2.55 to -1.48 communicates how precisely the model estimated the primary difference. A point estimate without an interval can obscure uncertainty. The interval also gives context for whether values close to the null are compatible with the analysis.
Why does P < .001 not measure the size of the effect?
A P-value and an effect estimate answer different questions. The P-value describes evidence against a null hypothesis under the specified test, whereas the LSMean Difference describes the magnitude and direction of the modeled treatment contrast. The confidence interval connects these ideas by showing the estimated effect together with uncertainty.
Why is Fisher's exact test different from the mixed model?
The two methods address different outcome structures. The mixed model is used for continuous longitudinal outcomes such as migraine headache days and questionnaire scores. Fisher's exact test is used for the binary ADA analysis, where the registry compares categorical treatment-group information.
Does a larger odds ratio mean a larger probability difference?
Not necessarily. An odds ratio compares odds rather than probabilities. The same odds ratio can correspond to different absolute probability differences depending on the baseline probability. Because the ClinicalTrials.gov record does not provide the underlying response percentages for the responder endpoint, the absolute probability differences cannot be calculated from the posted odds ratios alone.
16. Multiplicity and Multiple Comparisons
The ClinicalTrials.gov record contains multiple formal statistical analyses: two primary endpoint analyses and numerous secondary analyses. The registry identifies the hypothesis type for these analyses as superiority.
The presence of multiple analyses is statistically important because repeated hypothesis testing can increase the chance of observing at least one small P-value by chance if no multiplicity strategy is applied. However, the ClinicalTrials.gov record does not specify an alpha-allocation or multiplicity-adjustment procedure for the full set of EVOLVE-2 analyses.
17. Missing Data and Model Assumptions
The primary analysis population requires baseline and at least one post-baseline value. This indicates that availability of longitudinal observations is part of the registry-defined analysis population. The ClinicalTrials.gov record does not provide a complete description of how intermittent missing observations, treatment discontinuation, or other missing data were handled within the mixed model.
That distinction matters because a longitudinal model does not automatically make missing data irrelevant. The validity of inference depends on the assumptions underlying the missing-data mechanism and the model specification. Without the full statistical analysis plan, it is not possible to determine from the ClinicalTrials.gov record whether sensitivity analyses, pattern-mixture models, multiple imputation, or other missing-data approaches were used.
18. Crossover, Interim Analysis, and Bayesian Methods
Crossover
The ClinicalTrials.gov record does not report a crossover design or a crossover analysis. No crossover effect should therefore be inferred.
Interim analysis
The ClinicalTrials.gov record does not identify an interim efficacy analysis, stopping boundary, or alpha-spending method.
Bayesian methods
No Bayesian statistical method is identified in the registry-reported normalized methods. The posted methods are Fisher exact test and mixed-effects model.
Non-inferiority
The registry identifies the hypothesis type as superiority. No non-inferiority margin is reported or applicable to the reported primary comparisons in this dataset.
19. Safety Results
The ClinicalTrials.gov record provides serious adverse event counts by treatment phase for several arms. Because the source string becomes incomplete after the beginning of an additional arm entry, this page reports only the complete serious-adverse-event values explicitly reported.
| Arm / phase | Serious adverse events | Affected / at risk |
|---|---|---|
| Placebo — Treatment Phase | Serious adverse events | 5/461 |
| Galcanezumab 120 mg — Treatment Phase | Serious adverse events | 5/226 |
| Galcanezumab 240 mg — Treatment Phase | Serious adverse events | 7/228 |
| Placebo — Post-treatment Phase | Serious adverse events | 3/410 |
| Galcanezumab 120 mg — Post-treatment Phase | Serious adverse events | 1/212 |
| Galcanezumab 240 mg — Post-treatment Phase | Serious adverse events | 3/208 |
These are counts of affected participants over the stated at-risk denominators. They should not be treated as efficacy outcomes or combined with the longitudinal efficacy estimates. Safety and efficacy answer different statistical questions and generally use different analysis populations.
20. What the Primary Effect Estimates Do — and Do Not — Mean
The LSMean Difference of -2.02 describes a model-based difference in mean change from baseline in monthly migraine headache days for the placebo versus galcanezumab 120 mg comparison. It is not a claim that every participant had exactly two fewer migraine headache days, and it is not an individual-level treatment effect.
The LSMean Difference of -1.90 describes the corresponding model-based contrast for placebo versus galcanezumab 240 mg. It summarizes the estimated group-level treatment difference under the longitudinal model rather than describing the response of any particular participant.
The 95% CIs, -2.55 to -1.48 for 120 mg and -2.44 to -1.36 for 240 mg, communicate uncertainty around the corresponding estimates. They do not describe the distribution of individual responses or guarantee that future studies will reproduce exactly those ranges.
The two primary P-values are reported as <.001. They indicate strong evidence against the respective null hypotheses under the posted analyses, but they do not quantify the magnitude of the treatment effect, the probability of benefit for an individual participant, or the probability that the observed estimate will recur in another study.
21. Understanding the Longitudinal Analysis
The primary endpoint spans multiple post-baseline months rather than a single endpoint visit. That design creates both an analytical opportunity and an analytical challenge.
| Statistical feature | Why it matters in EVOLVE-2 |
|---|---|
| Repeated measurements | Monthly migraine headache days are assessed across a longitudinal period rather than at only one time point. |
| Within-participant correlation | Repeated observations from the same participant are related, motivating a longitudinal modeling framework. |
| LSMean Difference | Provides a model-based treatment contrast rather than a simple raw difference in arithmetic averages. |
| Confidence interval | Quantifies uncertainty around the model-based treatment contrast. |
| Missing observations | The analysis population requires baseline and post-baseline information, while the ClinicalTrials.gov record does not specify the complete missing-data strategy. |
| Multiple comparisons | Two primary dose comparisons and numerous secondary analyses create a broader testing context than a single hypothesis. |
This is one reason clinical-trial statistics cannot be reduced to a single P-value. The treatment estimate, its confidence interval, the outcome definition, the repeated-measures structure, the analysis population, and the broader multiplicity context all contribute to interpretation.
22. Statistical Interpretation of the Secondary Outcomes
Responder outcomes
The odds ratios range from 2.16 to 2.67 across the posted ≥50%, ≥75%, and ≥100% response analyses. These are relative odds measures, not percentages of participants and not risk ratios.
Quality of life
The MSQ Role Function Restrictive Domain analyses report positive LSMean Differences of 8.82 and 7.39 with two-sided 95% CIs that remain above zero.
Headache burden
The headache-hours estimates are -15.19 and -13.56, illustrating how a continuous outcome can retain its natural unit rather than being converted to a relative measure.
Patient-reported severity
PGI-S estimates are -0.29 and -0.23. Statistical significance does not by itself establish the magnitude of clinically meaningful change on the underlying scale.
A consistent statistical pattern across several endpoints can provide a broader description of the trial's reported findings, but each endpoint still has its own definition, scale, analysis population, and inferential context. Secondary endpoints should not automatically be interpreted as independently powered confirmatory conclusions.
23. Important Limitations and Interpretation Issues
- Registry-level detail: the ClinicalTrials.gov record provides the reported estimates and methods but do not include a complete statistical analysis plan.
- Incomplete model specification: the registry identifies mixed models analysis but does not provide all fixed effects, covariance assumptions, random-effects structure, or other implementation details in the ClinicalTrials.gov record.
- Analysis population: the primary population requires treatment exposure and baseline/post-baseline measurements, so it should not simply be described as every randomized participant.
- Missing-data handling: the ClinicalTrials.gov record does not establish the complete missing-data or imputation strategy.
- Multiple comparisons: the trial has two posted primary dose comparisons and numerous secondary analyses. The ClinicalTrials.gov record does not specify the full multiplicity-control strategy.
- Odds ratios versus probabilities: responder odds ratios cannot be translated into absolute response probabilities without the underlying event proportions or counts.
- No dose-comparison test: the difference between the 120 mg and 240 mg point estimates does not establish that the doses differ statistically.
- Proportional-hazards assumptions: not relevant to the posted primary method because the registry-reported primary analyses use mixed models rather than Cox proportional-hazards models.
24. Why This Trial Matters Statistically
EVOLVE-2 is a useful teaching case because its registry results bring together several core principles of clinical-trial statistics without relying on a time-to-event primary endpoint.
| Concept | How it appears in EVOLVE-2 |
|---|---|
| Randomization | The registry identifies randomized allocation. |
| Blinding | The trial is double masked. |
| Parallel design | The registry identifies a parallel design model. |
| Repeated measures | The primary endpoint is assessed from baseline through Month 6. |
| Mixed-effects model | Used for the primary endpoint and multiple continuous secondary endpoints. |
| Least-squares mean | The primary and continuous secondary analyses report LSMean Difference. |
| Confidence intervals | Reported as two-sided 95% intervals for the posted estimates. |
| Odds ratio | Used for the responder analyses. |
| Fisher exact test | Used for the posted anti-drug-antibody analyses. |
| Superiority testing | The registry identifies the hypothesis type as superiority. |
| Multiplicity | Two primary comparisons and numerous secondary analyses require careful interpretation of the inferential hierarchy. |
25. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: NCT02614196 — EVOLVE-2.
- Linked publication: PubMed 35076090.
- Linked publication: PubMed 34264500.
- Linked publication: PubMed 34049484.
- Linked publication: PubMed 33950375.
- Linked publication: PubMed 33549036.
Continue through the Clinical Biostats statistical tutorials
Explore the statistical concepts that recur across randomized trials, longitudinal outcomes, categorical endpoints, confidence intervals, and hypothesis testing.
28. Record Summary
EVOLVE-2 provides a useful example of longitudinal clinical-trial analysis. The registry identifies a randomized, double-blind, parallel phase 3 design with 986 enrolled participants and six arms. Its primary endpoint is overall mean change from baseline in monthly migraine headache days from Baseline through Month 6, with two formal mixed-model comparisons: placebo versus galcanezumab 120 mg and placebo versus galcanezumab 240 mg. The posted LSMean Differences are -2.02 and -1.90, respectively, with two-sided 95% confidence intervals and P-values of <.001.
The secondary analyses extend the statistical picture across responder thresholds, quality of life, acute-treatment medication days, PGI-S, headache hours, MIDAS, and anti-drug antibodies. Several continuous endpoints use mixed models, while the ADA endpoint uses Fisher's exact test. The responder endpoint is reported using odds ratios, illustrating the importance of distinguishing odds from probabilities. The available safety information includes serious-adverse-event counts for several treatment and post-treatment phases.
The central statistical lesson is that a clinical-trial result is more than a P-value. Correct interpretation requires attention to the endpoint definition, analysis population, longitudinal structure, effect measure, confidence interval, hypothesis type, missing-data assumptions, multiplicity, and the limits of what the registry actually reports.