← Clinical Trials
COPD Phase 3 Exacerbation Rate NCT02138916

GALATHEA: Complete Statistical Analysis of Benralizumab in COPD

An independent statistical review of the randomized, triple-masked, placebo-controlled phase 3 GALATHEA trial, which compared two doses of benralizumab (30 mg and 100 mg) with placebo in patients with moderate to very severe chronic obstructive pulmonary disease and a history of exacerbations, with the primary comparison made in patients with baseline blood eosinophils (EOS) of at least 220/µL.

Sponsor: AstraZeneca  ·  Study start: 2014-06-13  ·  Primary completion: 2018-04-10  ·  Status: Completed
About this analysis

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Reported estimates are kept separate from the interpretation that follows them, and all numerical results are those posted to the registry.

1. Trial at a Glance

GALATHEA asked whether benralizumab, given at 30 mg or 100 mg, reduced the annual rate of COPD exacerbations compared with placebo over 56 weeks in patients with a baseline blood eosinophil count of at least 220/µL. Neither dose produced a statistically significant reduction in the primary endpoint.

1656
Enrolled
Three parallel arms
56 wk
Primary time frame
From first IP to week 56
0.96
Rate ratio, 30 mg
95% CI 0.8–1.15 · P = 0.6490
0.83
Ratio, 100 mg
95% CI 0.69–1.00 · P = 0.0525
FeatureGALATHEA
PhasePhase 3
ConditionModerate to very severe chronic obstructive pulmonary disease, with exacerbation history
DesignRandomized, parallel-group, triple-masked, placebo-controlled
Arms3: benralizumab 30 mg (Arm A), benralizumab 100 mg (Arm B), placebo
Enrollment1656
Primary endpointAnnual COPD exacerbation rate over 56 weeks in patients with baseline EOS ≥220/µL
Primary analysisNegative binomial model; rate ratio versus placebo for each dose
Primary purposeTreatment
ClinicalTrials.govNCT02138916
Lead sponsorAstraZeneca (industry)

2. Clinical Question

The central question was whether benralizumab could reduce the frequency of COPD exacerbations in patients whose blood eosinophil count suggested an eosinophilic component to their disease. Because two doses were tested against a common placebo group, the trial answers two related superiority questions rather than one.

Population

Patients with moderate to very severe COPD and a history of exacerbations. The primary analysis population was the full analysis set with baseline EOS ≥220/µL; patients with EOS <220/µL formed a separate cohort analysed as a secondary endpoint.

Intervention

Benralizumab 30 mg (Arm A) or benralizumab 100 mg (Arm B), administered as investigational product (IP) through the 56-week treatment period.

Comparator

Placebo, under triple masking, so that participants, care providers and investigators were unaware of assignment.

Primary question

In patients with baseline EOS ≥220/µL, does either dose of benralizumab lower the annual COPD exacerbation rate relative to placebo?

3. Trial Design

01
Enrol1656 patients
02
Randomize30 mg / 100 mg / placebo
03
TreatFirst IP to week 56
04
PrimaryExacerbation rate, EOS ≥220/µL
05
SecondaryEOS <220/µL, lung function, symptoms
ARM A · 554 at risk (safety)

Benralizumab 30 mg

  • Lower of the two benralizumab doses
  • Compared with the shared placebo group
  • Treatment period from first IP to week 56
ARM B · 552 at risk (safety)

Benralizumab 100 mg

  • Higher of the two benralizumab doses
  • Compared with the shared placebo group
  • Treatment period from first IP to week 56
CONTROL · 550 at risk (safety)

Placebo

  • Common comparator for both doses
  • Matched under triple masking
  • Patients continued background COPD therapy, which was included as a model covariate
DESIGN FEATURES

Shared framework

  • Randomized, parallel-group allocation
  • Triple masking
  • Eosinophil cohort (≥220 vs <220/µL) defines the primary population
  • All tests framed as superiority, two-sided 95% CIs
Two doses, one placebo group. Each benralizumab dose is compared with the same placebo patients. The two comparisons are therefore correlated: an unusually high or low exacerbation rate in the placebo group would move both rate ratios in the same direction. This is a structural feature of multi-arm trials and is one reason the two primary results should be read together rather than as independent experiments.

4. Eosinophil Cohorts and Analysis Populations

The trial was built around a biomarker hypothesis: that patients with higher blood eosinophil counts would be most likely to benefit from an eosinophil-targeted therapy. The registry defines analysis populations by baseline EOS, and the eosinophil cohort also appears as a covariate in most models.

Analysis populationRole
Full analysis set, baseline EOS ≥220/µLPrimary endpoint and most secondary efficacy endpoints
Full analysis set, baseline EOS <220/µLSecondary analysis of the annual exacerbation rate in the lower-eosinophil cohort
Participants at risk for adverse eventsSerious adverse events by arm: 554 (30 mg), 552 (100 mg), 550 (placebo)

A full analysis set generally follows the intention-to-treat principle, analysing randomized patients according to their assigned group. Restricting the primary analysis to a biomarker-defined subset is not the same as a post hoc subgroup analysis: here, the EOS ≥220/µL population was the prespecified target of the primary endpoint, so its comparison carries confirmatory status.

5. Endpoints

Primary endpoint

EndpointRegistry definitionTime frame
Annual COPD Exacerbation Rate Over 56 Weeks, baseline EOS ≥220/µLA COPD exacerbation is defined by symptomatic worsening of COPD requiring use of systemic corticosteroids for at least 3 days (a single depot injectable dose of corticosteroids is considered equivalent to a 3-day course), and/or use of antibiotics, and/or an inpatient hospitalization or death due to COPD. The annual exacerbation rate is the number of exacerbations per year; its raw rate is the number of exacerbations divided by the treatment period, normalized to an annual rate, and is estimated by a negative binomial model. The rate ratio between two treatment groups is also estimated through this model.From first IP to week 56

Secondary endpoints with posted analyses

Endpoint (shortened)PopulationUnitTime frame
Annual COPD exacerbation rate over 56 weeksEOS <220/µLExacerbations per yearFrom first IP to week 56
Mean change from baseline to week 56 in pre-bronchodilator FEV1 (L)EOS ≥220/µLLiterFirst IP up to end of treatment week 56
Mean change from baseline in SGRQ total scoreEOS ≥220/µLPercentageFirst IP up to week 56
Mean change from baseline in CAT total scoreEOS ≥220/µLScore on a scaleFirst IP up to week 56
Mean change from baseline in E-RS: COPD total scoreEOS ≥220/µLScore on a scaleFirst IP up to week 56
Mean change from baseline in total rescue medication useEOS ≥220/µLPuffs/dayFirst IP up to week 56
Mean change from baseline in proportion of nights with awakenings due to respiratory symptomsEOS ≥220/µLProportion of nightsFirst IP up to week 56
Annual EXACT-PRO exacerbation rate over 56 weeksEOS ≥220/µLExacerbations per yearImmediately following first IP up to week 56
Number of participants having at least 1 COPD exacerbationEOS ≥220/µLParticipantsImmediately following first IP up to week 56
Annual COPD exacerbation rate associated with ER or hospitalization over 56 weeksEOS ≥220/µLExacerbations per yearImmediately following first IP up to week 56

The primary endpoint is a count outcome with variable exposure time: patients contribute different lengths of follow-up, and some experience several exacerbations while many experience none. That structure drives the choice of the negative binomial model described below.

6. Primary Endpoint Results

Both primary comparisons used a negative binomial model in the full analysis set with baseline EOS ≥220/µL. The model included treatment group, EOS cohort, region, background therapy and the number of exacerbations in the previous year. Both were superiority hypotheses with two-sided 95% confidence intervals.

Benralizumab 30 mg vs placebo

Rate ratio for annual COPD exacerbations

0.96

95% CI: 0.8–1.15   ·   P = 0.6490

Full analysis set, baseline EOS ≥220/µL  ·  Negative binomial model

Clinical Biostats interpretation

What the estimate means. A rate ratio of 0.96 means that, after adjustment for the model covariates, the estimated annual exacerbation rate in the 30 mg group was 4% lower than in the placebo group. This is a ratio of model-estimated rates, not a difference in the number of exacerbations and not the proportion of patients who avoided an exacerbation.

What it does not mean. It does not show that benralizumab 30 mg reduces exacerbations. The point estimate is close to 1, and the result is fully compatible with no effect.

Precision. The 95% CI of 0.8 to 1.15 spans values from a 20% lower rate to a 15% higher rate. The data are consistent with modest benefit, no difference, or modest harm, so the comparison does not pin down the direction of any effect.

The p-value. P = 0.6490 describes how compatible the observed data are with a true rate ratio of 1 under the model. It does not measure the size of the effect, and a large p-value is not evidence that the effect is exactly zero; it simply indicates the trial did not distinguish this dose from placebo.

Cautions. Negative binomial rate ratios assume the covariates act multiplicatively on the rate and that overdispersion is adequately captured by a single dispersion parameter. Because the primary population is defined by baseline EOS, conclusions apply to that biomarker-selected group.

Benralizumab 100 mg vs placebo

Ratio for annual COPD exacerbations

0.83

95% CI: 0.69–1.00   ·   P = 0.0525

Full analysis set, baseline EOS ≥220/µL  ·  Negative binomial model

ComparisonEffect measure (registry label)Estimate95% CI (two-sided)P-value
Benralizumab 30 mg vs placeboRate ratio0.960.8 to 1.150.6490
Benralizumab 100 mg vs placeboRisk Ratio (RR)0.830.69 to 1.000.0525
A note on labelling. The registry labels the 100 mg comparison a "Risk Ratio (RR)", whereas the 30 mg comparison is labelled a "Rate ratio". Both come from the same negative binomial model of exacerbations per year, and the endpoint definition states that the model estimates a rate ratio between treatment groups. The 100 mg estimate is therefore best read as a ratio of annual exacerbation rates, not as a ratio of the probability of having an exacerbation.
Clinical Biostats interpretation

What the estimate means. A ratio of 0.83 means the model-estimated annual exacerbation rate with benralizumab 100 mg was 17% lower than with placebo in the EOS ≥220/µL population.

What it does not mean. It does not establish a treatment effect. With P = 0.0525 and an upper confidence limit of 1.00, the comparison did not meet the conventional two-sided 0.05 threshold. Describing it as a "trend" or "near-significant" result invites over-interpretation; the prespecified test was not met.

Precision. The 95% CI runs from 0.69 to 1.00, ranging from a 31% lower rate to no difference. The interval sits almost entirely below 1, so the data lean toward a reduction, but they cannot exclude a rate ratio of 1.

The p-value. A p-value just above or just below 0.05 carries nearly the same evidential weight; the 0.05 line is a decision rule, not a boundary in nature. At the same time, a p-value just above 0.05 does not measure how large the effect is, and it cannot be rescued by pointing to the point estimate.

Cautions. With two doses tested against one placebo group, some form of multiplicity control is normally needed to keep the familywise type I error at its nominal level. The registry does not describe the procedure, but any adjustment would make the threshold for each comparison stricter, not looser. The consistency of this result with the 30 mg arm should also be weighed: a clearer effect at 100 mg than at 30 mg is compatible with a dose-response, but also with chance variation around a small or null effect.

7. Secondary Exacerbation Endpoints

Several secondary endpoints examined exacerbations from different angles: in the lower-eosinophil cohort, using a patient-reported definition (EXACT-PRO), as the proportion of patients with any exacerbation, and restricted to more severe events leading to an emergency room (ER) visit or hospitalization.

EndpointComparisonMethodEstimate (registry label)95% CIP-value
Annual exacerbation rate, EOS <220/µL30 mg vs placeboNegative binomial1.07 (Risk Ratio)0.86 to 1.340.5236
Annual exacerbation rate, EOS <220/µL100 mg vs placeboNegative binomial1.02 (Risk Ratio)0.82 to 1.270.8812
Annual EXACT-PRO exacerbation rate30 mg vs placeboNegative binomial1.09 (Rate ratio)0.89 to 1.340.4080
Annual EXACT-PRO exacerbation rate100 mg vs placeboNegative binomial0.98 (Rate ratio)0.80 to 1.210.8688
At least 1 COPD exacerbation30 mg vs placeboCochran-Mantel-Haenszel0.90 (Odds Ratio)0.66 to 1.220.4850
At least 1 COPD exacerbation100 mg vs placeboCochran-Mantel-Haenszel0.89 (Odds Ratio)0.65 to 1.210.4489
Exacerbations with ER visit or hospitalization30 mg vs placeboNegative binomial1.06 (Rate ratio)0.73 to 1.530.7733
Exacerbations with ER visit or hospitalization100 mg vs placeboNegative binomial0.58 (Rate ratio)0.39 to 0.890.0114

All populations are the full analysis set with baseline EOS ≥220/µL except the first two rows (EOS <220/µL). The EOS <220/µL models included treatment group, region and number of exacerbations in the previous year (the 100 mg model also included EOS cohort). The ER/hospitalization models replaced the prior-exacerbation count with an indicator for previous-year exacerbations associated with hospitalization.

Reading the secondary exacerbation results

8. Secondary Lung-Function and Patient-Reported Endpoints

Continuous endpoints were analysed with mixed models for repeated measures. Each model included treatment group, the baseline value of the outcome, EOS cohort, region, background therapy, visit and a treatment-by-visit interaction. Estimates are mean differences versus placebo; negative values favour benralizumab for the symptom scores, rescue medication and night awakenings, while positive values favour benralizumab for FEV1.

Endpoint (EOS ≥220/µL)ComparisonMean difference95% CIP-value
Pre-bronchodilator FEV1 (L)30 mg vs placebo0.007-0.035 to 0.0480.7550
Pre-bronchodilator FEV1 (L)100 mg vs placebo0.021-0.021 to 0.0620.3285
SGRQ total score30 mg vs placebo-1.011-2.887 to 0.8650.2906
SGRQ total score100 mg vs placebo-2.136-4.020 to -0.2510.0264
CAT total score30 mg vs placebo-0.19-1.08 to 0.700.6782
CAT total score100 mg vs placebo-0.81-1.70 to 0.080.0753
E-RS: COPD total score30 mg vs placebo-0.585-1.260 to 0.0890.0889
E-RS: COPD total score100 mg vs placebo-0.703-1.378 to -0.0280.0413
Rescue medication use (puffs/day)30 mg vs placebo-0.348-0.728 to 0.0320.0728
Rescue medication use (puffs/day)100 mg vs placebo-0.487-0.868 to -0.1070.0121
Proportion of nights with awakeningsBenralizumab 30 mg-0.041-0.077 to -0.0060.0235
Proportion of nights with awakeningsBenralizumab 100 mg-0.044-0.080 to -0.0080.0158

For the night-awakening endpoint the registry names only the benralizumab group in the comparison field; as with the other mean-difference analyses, the estimate is a difference from placebo.

What the pattern shows

FEV1 differences of 0.007 L and 0.021 L are small, with intervals centred near zero, so the trial provides no evidence of a lung-function effect. Among patient-reported and diary outcomes, the point estimates are consistently in the favourable direction, and several 100 mg comparisons (SGRQ, E-RS: COPD, rescue medication) and both night-awakening comparisons have intervals excluding zero. The 100 mg SGRQ difference of -2.136 units has a 95% CI of -4.020 to -0.251, meaning the data are compatible with anything from a very small to a moderate improvement relative to placebo.

These nominally significant results must be read in context. They are secondary endpoints, several are correlated measures of the same underlying symptom burden, and the primary endpoint was not met. A consistent direction across correlated outcomes is less persuasive than it looks, because the outcomes are not independent pieces of evidence.

9. Safety: Serious Adverse Events

The registry reports the number of participants with at least one serious adverse event in each arm.

ArmParticipants with serious adverse eventsParticipants at risk
Benralizumab 30 mg151554
Benralizumab 100 mg177552
Placebo176550

The counts in the 100 mg and placebo arms are nearly identical, and the 30 mg arm has fewer participants with serious events. No formal statistical comparison of serious adverse events is posted, and safety counts of this kind are descriptive: they are not adjusted for exposure time and group all serious events together, regardless of cause or relatedness. A risk ratio with an exact or Wald confidence interval is the usual summary if a formal comparison is wanted, but individual event types, not the overall serious-event count, are generally more informative about specific harms.

10. Statistical Methodology

Negative binomial regression for exacerbation rates

Exacerbations are counts, recorded over treatment periods that vary between patients because of discontinuation or withdrawal. A Poisson model assumes the variance equals the mean, but exacerbation counts are typically overdispersed: a minority of patients exacerbate repeatedly while many have none. The negative binomial model adds a dispersion parameter so that between-patient heterogeneity widens the standard errors appropriately, and it uses the log of the treatment period as an offset so that rates are compared per unit of time.

Conceptual form
log E[Yi] = log(ti) + β0 + βtrt·Trti + β·Xi    Var(Yi) = μi + k·μi2

Here Yi is the exacerbation count, ti the treatment period, Xi the covariates (EOS cohort, region, background therapy, prior-year exacerbations) and k the dispersion parameter. The rate ratio is exp(βtrt).

Mixed models for repeated measures

FEV1, SGRQ, CAT, E-RS: COPD, rescue medication and night awakenings were measured repeatedly. The mixed models included visit, treatment and a treatment-by-visit interaction, allowing a separate treatment difference at each visit; the week-56 (or overall) contrast is then extracted from the model. Adjusting for the baseline value of each outcome reduces residual variance and improves precision. Mixed models use all available post-baseline measurements and give valid estimates if missing data are missing at random given the observed data, an assumption that cannot be verified from the data alone.

Cochran-Mantel-Haenszel test for the binary exacerbation endpoint

The proportion of patients with at least one exacerbation was compared with a CMH test controlling for EOS cohort, region and background therapy. The CMH approach combines 2×2 tables across strata into a common odds ratio, which protects against confounding by the stratifying variables and assumes the odds ratio is broadly similar across strata.

Covariate adjustment

Every posted model adjusts for prognostic variables, most importantly the number of exacerbations in the previous year, which is one of the strongest predictors of future exacerbations. In a randomized trial, covariate adjustment is not needed to remove bias; its role is to increase precision and to align the analysis with factors that may have been used to balance the design.

Hypothesis framework

All posted analyses are superiority tests with two-sided 95% confidence intervals. For ratio measures the null value is 1; for mean differences it is 0. Superiority is concluded when the interval excludes the null value and the p-value falls below the prespecified significance level.

11. Multiple Comparisons and Nominal P-values

GALATHEA produced two primary comparisons and twenty posted secondary comparisons. The registry does not describe the multiplicity procedure, so the table below classifies results by their role rather than by any formal testing hierarchy.

AnalysisRoleInterpretation
Exacerbation rate, EOS ≥220/µL, 30 mg and 100 mgPrimaryConfirmatory; neither comparison reached P < 0.05
Exacerbation rate, EOS <220/µLSecondarySupports understanding of the biomarker hypothesis; not a test of interaction
EXACT-PRO, any exacerbation, ER/hospitalizationSecondarySupportive; nominal p-values only
FEV1, SGRQ, CAT, E-RS: COPD, rescue medication, night awakeningsSecondarySupportive; correlated outcomes; nominal p-values only
Why nominal significance is not enough: with many secondary comparisons, some p-values below 0.05 are expected even without any true effect. In a trial where the primary endpoint was not met, a conventional fixed-sequence hierarchy would stop formal testing at the primary endpoint, and all subsequent p-values would be descriptive. Nominally significant secondary results can motivate further research, but they do not stand in for the failed primary comparison.

12. Statistical Methods Explained

Why was a negative binomial model used instead of a Poisson model or a t-test?

Exacerbation counts are non-negative integers, heavily skewed, and recorded over unequal follow-up. A t-test on raw counts ignores both the count nature of the data and unequal exposure. A Poisson model handles counts and exposure but assumes variance equals the mean; when some patients exacerbate much more often than others, Poisson standard errors are too small and p-values too optimistic. The negative binomial model relaxes that assumption through its dispersion parameter.

What does a rate ratio of 0.83 with a 95% CI of 0.69 to 1.00 tell us?

It says the best estimate is a 17% lower annual exacerbation rate with 100 mg than with placebo, but the plausible range extends to no difference at all. The upper limit touching 1.00 matches the P value of 0.0525: the interval and the test carry the same information, and neither supports a claim of superiority at the two-sided 0.05 level.

Why is the 100 mg primary result labelled a "risk ratio" when it is a rate ratio?

The label is a registry entry choice. The analysis is a negative binomial model of exacerbations per year, which estimates a ratio of rates. A true risk ratio compares the probability of an event over a fixed period and would be estimated from binary data. Confusing the two can change the clinical reading: a 17% lower rate of events is not the same as a 17% lower chance of having any event.

Why was the proportion with at least one exacerbation analysed with a CMH odds ratio?

Dichotomising the outcome converts each patient's count into "any exacerbation: yes or no". A CMH test compares these proportions while stratifying on EOS cohort, region and background therapy, producing a common odds ratio across strata. The odds ratio is not the same as a risk ratio; when the outcome is common, an odds ratio lies further from 1 than the corresponding risk ratio would.

Does the positive ER/hospitalization result at 100 mg rescue the trial?

No. A rate ratio of 0.58 with a 95% CI of 0.39 to 0.89 is a notable estimate, but it is one of many secondary comparisons after a primary endpoint that was not met, and the 30 mg estimate for the same endpoint was 1.06. Its role is to generate hypotheses about whether severe events respond differently from overall exacerbations, not to confirm efficacy.

Why include a treatment-by-visit interaction in the mixed models?

Without it, the model forces the treatment effect to be identical at every visit. With it, the difference versus placebo can evolve over the 56 weeks, which is realistic for outcomes such as symptom scores. The reported mean difference is then a specific contrast from that flexible model, while the repeated-measures structure accounts for the correlation of measurements within the same patient.

13. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

Neither prespecified primary comparison excluded a rate ratio of 1. The 30 mg estimate was close to 1; the 100 mg estimate favoured benralizumab but its interval reached 1.00. Secondary findings are nominal and should be viewed as exploratory.

Clinical interpretation

In this population the trial did not demonstrate that benralizumab lowers the overall COPD exacerbation rate. Signals in severe exacerbations and some patient-reported outcomes at 100 mg describe where further study might focus, not established benefit.

14. Limitations

15. Why This Trial Matters Statistically

GALATHEA is a useful teaching case precisely because its primary result was not positive. It shows how count outcomes are modelled, how a near-threshold p-value should be read, and why secondary signals after a failed primary endpoint need restraint.

ConceptHow it appears in GALATHEA
Negative binomial regressionAnnual exacerbation rates with overdispersion and variable exposure time
Rate ratioPrimary effect measure: 0.96 (30 mg) and 0.83 (100 mg) versus placebo
Rate ratio vs risk ratioRegistry labels differ for estimates from the same model
Confidence intervalsAn upper limit of 1.00 aligned with P = 0.0525
P-valuesNear-threshold primary result; many nominal secondary p-values
Mixed-effects modelsRepeated measures for FEV1, SGRQ, CAT, E-RS: COPD, rescue medication, night awakenings
Cochran-Mantel-Haenszel testStratified odds ratio for at least one exacerbation
Odds ratioBinary exacerbation endpoint summary
Multi-arm designTwo doses sharing a placebo group
Biomarker-defined populationPrimary analysis restricted to baseline EOS ≥220/µL
MultiplicitySecondary findings after a primary endpoint that was not met

16. Related Tutorials

Learn more about the methods used in this trial:

17. Related Calculators

18. Sources

Continue learning with Clinical Biostats

Explore the statistical methods behind this trial in more depth, or apply them to your own data with our calculators.

19. Record Summary

GALATHEA was a randomized, triple-masked, three-arm phase 3 trial enrolling 1656 patients with moderate to very severe COPD. In the prespecified EOS ≥220/µL population, the negative binomial rate ratios for annual exacerbations were 0.96 (95% CI 0.8–1.15; P = 0.6490) for 30 mg and 0.83 (95% CI 0.69–1.00; P = 0.0525) for 100 mg versus placebo, so neither comparison met its superiority test. The most useful reading combines the relative rate estimates, their confidence intervals, an honest treatment of the near-threshold p-value, and appropriate caution about nominally significant secondary endpoints in a multi-arm, multi-endpoint design.

Clinical Biostats methodology: A trial-results page should do more than repeat the registry entry. The aim is to explain the statistical story of the trial in a consistent format while keeping reported estimates clearly separate from educational interpretation.