← Clinical Trials
Severe Eosinophilic Asthma Phase 3 Non-Inferiority NCT04718389

NIMBLE: Complete Statistical Analysis of Depemokimab in Severe Eosinophilic Asthma

An independent statistical review of the randomized, double-blind phase 3 NIMBLE trial comparing depemokimab with mepolizumab or benralizumab, each added to standard of care, in participants with severe asthma with an eosinophilic phenotype. The primary question was one of non-inferiority on the annualized rate of clinically significant exacerbations over 52 weeks.

Sponsor: GlaxoSmithKline  ·  Start: 2021-01-26  ·  Primary completion: 2025-09-14  ·  Status: Completed
About this page

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

NIMBLE was a randomized, double-blind, parallel-group phase 3 trial asking whether depemokimab could be shown to be non-inferior to the established anti-IL-5 pathway biologics mepolizumab or benralizumab in preventing clinically significant asthma exacerbations over 52 weeks.

1717
Enrolled
2 randomized arms
1.16
Exacerbation rate ratio
95% CI 0.98–1.38
0.079
Primary P-value
As posted to the registry
52 wk
Primary time frame
Up to Week 52
FeatureNIMBLE
PhasePhase 3
ConditionAsthma (severe asthma with an eosinophilic phenotype)
DesignRandomized, parallel-group, double-masked, treatment purpose
Arms2: depemokimab 100 mg SC versus mepolizumab 100 mg / benralizumab 30 mg SC, each with standard of care
Enrollment1717
Primary endpointAnnualized rate of clinically significant exacerbations over 52 weeks
Primary hypothesisNon-inferiority
Primary analysis modelGeneralized linear model with a negative binomial distribution
ClinicalTrials.govNCT04718389
FundingGlaxoSmithKline (industry)

2. Clinical Question

The central question was whether depemokimab added to standard of care could preserve the exacerbation-reducing effect already achieved by an active anti-IL-5 pathway biologic. Because the comparator is itself an effective therapy, the relevant question is not "is depemokimab better?" but "is it not unacceptably worse?" That framing determines almost every statistical choice in the trial.

Population

Participants with severe asthma with an eosinophilic phenotype. The primary model adjusts for pre-study biologic therapy (mepolizumab or benralizumab), indicating that participants entered the trial from one of these existing treatments.

Intervention

Depemokimab (GSK3511294) 100 mg subcutaneously, plus standard of care.

Comparator

Mepolizumab 100 mg or benralizumab 30 mg subcutaneously, plus standard of care, analyzed as a single active-comparator group.

Primary question

Is depemokimab non-inferior to mepolizumab or benralizumab for the annualized rate of clinically significant exacerbations over 52 weeks?

3. Trial Design

01
Enroll1717 participants
02
Randomize2 parallel arms, double-blind
03
TreatBiologic + standard of care
04
RecordExacerbations in the eCRF
05
AnalyzeRate ratio at Week 52
ARM 1 · Experimental

Depemokimab 100mg SC

  • Depemokimab (GSK3511294), a biological
  • Subcutaneous administration
  • Standard of care (SoC) continued
  • Serious adverse events: 81/859 affected / at risk
ARM 2 · Active comparator

Mepolizumab 100 mg / Benralizumab 30 mg SC

  • Mepolizumab 100 mg or benralizumab 30 mg
  • Subcutaneous administration
  • Standard of care (SoC) continued
  • Serious adverse events: 76/855 affected / at risk
Blinding across different products. The registry lists placebo and pre-filled syringes among the interventions alongside the three active biologics. When active treatments differ in product and schedule, placebo injections are the usual way to keep participants and investigators unaware of assignment. The registry does not describe the blinding scheme in detail, but the double masking matters statistically: exacerbations requiring systemic corticosteroids involve clinical judgement, and unblinded assessment could bias event counts.
Allocation
Randomized, two parallel groups
Masking
Double
Comparator structure
One pooled active-comparator arm containing two different biologics
Hypothesis
Non-inferiority of depemokimab + SoC versus active comparator + SoC

4. Analysis Population and Covariate Adjustment

Efficacy analyses used the Full Analysis Set (FAS). The registry defines it as all randomized participants who received at least one dose of study intervention, excluding participants from sites with GCP violations, analyzed according to the intervention allocated at randomization.

ElementRegistry descriptionStatistical role
Full Analysis SetRandomized and dosed at least once; sites with GCP violations excludedPrimary and secondary efficacy population; a modified intention-to-treat set
Analysis by allocationParticipants analyzed according to the intervention allocated at randomizationPreserves the randomized comparison
Safety denominators859 (depemokimab) and 855 (comparator) at riskDenominators for serious adverse events

The primary model adjusted for the following covariates, which the registry lists explicitly:

Prior-year exacerbation count is typically the strongest predictor of future exacerbations, so adjusting for it removes a large share of between-participant variability and sharpens the treatment comparison. Adjusting for pre-study biologic is especially relevant here, because participants switching from mepolizumab and from benralizumab may differ in baseline risk.

5. Endpoints

Primary endpoint

EndpointRegistry definitionTime frame
Annualized Rate of Clinically Significant Exacerbations Over 52 WeeksWorsening of asthma requiring systemic corticosteroids (IM, IV or oral) and/or hospitalization and/or Emergency Department visit. For all participants, IV or oral steroids for at least 3 days or a single IM corticosteroid dose is required; for participants on maintenance systemic corticosteroids, at least double the existing maintenance dose for at least 3 days is required. Exacerbations recorded in the eCRF were considered verified clinically significant exacerbations and included in the primary analysis. Exacerbations separated by less than 7 days were treated as a continuation of the same exacerbation.Up to Week 52

Two features of this definition have direct statistical consequences. First, the 7-day rule defines what counts as a separate event, which directly shapes the count distribution being modelled. Second, the requirement of systemic corticosteroid use, hospitalization or ED visit anchors the endpoint to an objective treatment action rather than symptoms alone, which reduces measurement noise.

Secondary endpoints with posted analyses

EndpointUnitTime framePosted method
Weighted Mean (WM) Change From Baseline in St. George's Respiratory Questionnaire (SGRQ) Total ScoreScores on ScaleFrom Baseline (Day 1) up to Week 52Listed as negative binomial model
Weighted Mean Change From Baseline in Asthma Control Questionnaire-5 (ACQ-5) ScoreScores on ScaleFrom Baseline (Day 1) up to Week 52ANCOVA
Weighted Mean Change From Baseline in Pre-bronchodilator FEV1LiterFrom Baseline (Day 1) up to Week 52ANCOVA

6. Primary Result: Annualized Exacerbation Rate

The registry reports a single model-based comparison of depemokimab 100mg SC versus mepolizumab 100 mg / benralizumab 30 mg SC in the FAS, with the stated purpose of demonstrating the non-inferiority of depemokimab + SoC relative to the active comparator + SoC over the 52-week intervention period.

Rate ratio, depemokimab vs mepolizumab/benralizumab

1.16

95% CI (two-sided): 0.98–1.38   ·   P = 0.079

Negative binomial generalized linear model · Unit: exacerbations per participant per year

ItemPosted value
Effect measureRate ratio (depemokimab / comparator)
Estimate1.16
95% confidence interval0.98 to 1.38 (two-sided)
P-value0.079
Hypothesis typeNon-inferiority
ModelGeneralized linear model, negative binomial distribution, covariate-adjusted
Clinical Biostats interpretation

What the estimate means. A rate ratio of 1.16 means that, after covariate adjustment, the model-estimated annualized rate of clinically significant exacerbations was 16% higher in the depemokimab group than in the mepolizumab/benralizumab group. The ratio is written with depemokimab in the numerator, so values above 1 point in the direction unfavourable to depemokimab and values below 1 favour it.

What it does not mean. It does not mean that 16% more participants had an exacerbation, nor that each participant's own exacerbation rate rose by 16%. It is a ratio of adjusted population-average event rates, not a proportion of patients and not an individual-level effect. A relative increase in a rate also says nothing on its own about the absolute number of additional exacerbations per year, which depends on the baseline rate in the comparator group.

What the confidence interval says. The two-sided 95% CI of 0.98 to 1.38 is compatible with anything from a slightly lower rate (2% lower) to a materially higher rate (38% higher) with depemokimab. Because the interval includes 1, the data do not distinguish depemokimab from the comparator in either direction. The width of the interval, despite a large trial, reflects the overdispersed nature of exacerbation counts.

Non-inferiority logic. In a non-inferiority trial the decision is made by comparing the upper confidence limit (here 1.38) against a prespecified margin: non-inferiority is shown if the upper limit lies below the margin. The margin itself is not reported in the posted analysis on the ClinicalTrials.gov record, so the registry entry alone does not state whether the non-inferiority criterion was met. The conclusion hinges entirely on where 1.38 sits relative to that margin.

Why the P-value does not measure effect size. P = 0.079 is a statement about compatibility of the data with a particular null hypothesis, not about how large the difference is. The registry does not specify which null hypothesis this P-value tests. If it tests a rate ratio of 1 (no difference), a value above 0.05 simply mirrors the CI crossing 1 and is not evidence of equivalence or non-inferiority. If it tests the non-inferiority margin, its interpretation would be different. Without the margin and the null stated, the P-value should not be read as a verdict either way.

Cautions. The analysis population excludes participants from sites with GCP violations and those who never received study intervention; in non-inferiority trials, departures from strict intention-to-treat can bias results toward similarity, so sponsors usually examine consistency across populations. The comparator arm pools two different biologics, so the estimate is relative to that mixture rather than to either drug alone.

Group-level annualized exacerbation rates, their confidence intervals and the non-inferiority margin are not included in the posted statistical analysis on the ClinicalTrials.gov record; only the adjusted rate ratio, its 95% CI and the P-value are reported.

7. Secondary Efficacy Results

Three secondary endpoints have posted between-group comparisons. All are expressed as depemokimab minus comparator differences in weighted mean change from baseline, with two-sided 95% confidence intervals. No P-values are posted for these comparisons, and the registry lists their hypothesis type as "Other / not stated".

EndpointMethodDifference (depemokimab − comparator)95% CI
SGRQ total score, WM change from baselineNegative binomial model (as listed)0.17−0.85 to 1.20
ACQ-5 score, WM change from baselineANCOVA−0.03−0.09 to 0.02
Pre-bronchodilator FEV1, WM change from baseline (L)ANCOVA0.004−0.016 to 0.024

Reading the direction of each difference

SGRQ: 0.17

Lower SGRQ scores indicate better health status, so a positive difference points slightly against depemokimab. The interval from −0.85 to 1.20 spans zero and is narrow relative to the scale, suggesting the two groups had very similar changes in health status.

ACQ-5: −0.03

Lower ACQ-5 scores indicate better asthma control, so a negative difference points slightly toward depemokimab. The interval −0.09 to 0.02 includes zero and is tightly concentrated around it.

FEV1: 0.004 L

Higher FEV1 is better. The estimated difference of 0.004 L, with an interval from −0.016 to 0.024 L, is very close to zero on a scale measured in litres.

Weighted mean

The registry notes that the least squares mean represents the weighted mean change from baseline, summarizing change across repeated visits over the 52 weeks rather than at a single time point.

A note on the SGRQ method label. The registry lists "Negative binomial model" as the method for the SGRQ total score comparison, while the effect measure is a mean difference on a continuous score. A negative binomial model is designed for counts and produces ratios, not differences, so this label is best treated with caution; a continuous-score endpoint of this kind is ordinarily analysed with ANCOVA or a related linear model, as the ACQ-5 and FEV1 endpoints were.
Why "no difference" is not "shown equivalent"

Each secondary interval includes zero. That means the data are compatible with no difference, but the secondary endpoints were not posted with equivalence or non-inferiority margins, so they cannot formally establish that the treatments are equivalent. Their value lies in the precision of the intervals: narrow intervals around zero place informal bounds on how large any difference in symptoms, control or lung function could plausibly be.

8. Safety: Serious Adverse Events

The registry reports serious adverse events as the number of participants affected over the number at risk in each arm.

ArmParticipants with serious adverse eventsParticipants at risk
Depemokimab 100mg SC81859
Mepolizumab 100 mg / Benralizumab 30 mg76855

The counts and denominators are closely matched between arms. No formal statistical comparison of serious adverse events is posted. Safety counts of this kind are descriptive: trials are rarely powered to test differences in serious adverse events, and a lack of statistical difference in such counts should not be interpreted as proof of equal safety. Individual event types, severity and relatedness would also need to be examined separately.

9. Statistical Methodology

Negative binomial regression for exacerbation counts

Exacerbations are counts of events accumulated over a period of follow-up. A Poisson model assumes the variance equals the mean, but exacerbation counts are typically overdispersed: most participants have zero or one event while a minority have many. The negative binomial model adds a dispersion parameter that allows the variance to exceed the mean, producing more honest (wider) confidence intervals.

Conceptual form
log(E[Yi]) = log(ti) + β0 + β1·Treatmenti + β2·FEV1%predi + β3·PriorBiologici + β4·Regioni + β5·PriorExaci
Var(Yi) = μi + k·μi2   ·   Rate ratio = exp(β1)

Here Yi is the exacerbation count, ti is time on study (an offset in models of this type, which is what turns counts into annualized rates), and k is the dispersion parameter. The covariates are those listed in the registry; the offset form is the standard convention rather than a detail posted to the registry.

Confidence interval for a rate ratio

The model is fitted on the log scale, so the confidence interval is computed for β1 and then exponentiated. This is why the posted interval of 0.98 to 1.38 is not symmetric around 1.16: on the ratio scale, intervals stretch further above the estimate than below it.

Non-inferiority testing

For a ratio where higher values are worse for the experimental treatment, non-inferiority is declared if the upper bound of the confidence interval lies below a prespecified margin M > 1. Using the upper limit of a two-sided 95% interval corresponds to a one-sided test at the 2.5% level.

Non-inferiority hypotheses for a rate ratio
H0: RR ≥ M   (depemokimab unacceptably worse)
H1: RR < M   (depemokimab non-inferior)

Rejecting H0 requires the upper confidence limit of the rate ratio to fall below M. The P-value that matters for this decision is the one computed against M, not against a ratio of 1.

ANCOVA for continuous secondary endpoints

ACQ-5 and FEV1 changes were analysed by ANCOVA, which models the outcome as a linear function of treatment and baseline covariates. Adjusting for the baseline value removes variability that has nothing to do with treatment, typically narrowing confidence intervals relative to an unadjusted comparison. The reported least squares means are model-adjusted group means, and the posted estimates are differences between them.

Analysis by randomized allocation

Participants in the FAS were analysed according to the intervention allocated at randomization. This preserves the randomized comparison, although the requirement of at least one dose and the exclusion of sites with GCP violations mean the FAS is narrower than a pure all-randomized population.

10. Statistical Methods Explained

Why was a negative binomial model used instead of comparing proportions?

A proportion ("did the participant have any exacerbation?") throws away information: a participant with five exacerbations counts the same as one with one. Modelling the count over time uses all events and, via the dispersion parameter, accounts for the clustering of events within high-risk participants. The result is expressed as a rate ratio rather than a risk ratio or odds ratio.

What does a rate ratio of 1.16 mean here?

It means the adjusted annualized exacerbation rate was estimated to be 16% higher with depemokimab than with mepolizumab/benralizumab. Because depemokimab is in the numerator, a ratio above 1 is the unfavourable direction for the new drug. The confidence interval from 0.98 to 1.38 shows the data are also compatible with a small advantage or a larger disadvantage.

Why is non-inferiority judged against the margin rather than the P-value of 0.079?

A non-inferiority conclusion depends on excluding an unacceptably large disadvantage, which is defined by the margin. A P-value against "no difference" answers a superiority-style question and cannot establish non-inferiority. Failing to show a difference (P > 0.05) is never, by itself, evidence that two treatments are similar. The decisive quantity is the upper confidence limit of 1.38 compared with the prespecified margin.

Why adjust for exacerbations in the prior year?

Past exacerbation frequency is typically the strongest predictor of future exacerbations. Randomization balances it on average, but including it as a covariate explains a large part of the outcome variance, making the treatment estimate more precise. It also protects against chance imbalances between arms.

What does pooling two comparator biologics imply?

The comparator arm combines participants receiving mepolizumab and participants receiving benralizumab. The rate ratio therefore compares depemokimab with a mixture whose composition reflects the enrolled population. Including pre-study biologic as a covariate adjusts for differences in baseline risk between these subgroups, but the posted analysis does not provide separate estimates versus each drug.

Why are the ANCOVA differences reported with confidence intervals but no P-values?

The registry lists no hypothesis type for these secondary comparisons and posts no P-values. The intervals still convey the most useful information: the size and precision of the estimated difference. An interval such as −0.09 to 0.02 for ACQ-5 shows directly how large a difference is compatible with the data, which a P-value alone would not.

11. Limitations

12. Why This Trial Matters Statistically

NIMBLE is a useful teaching case because it combines an active-controlled non-inferiority question with count-data modelling, and because the posted result sits in the zone where interpretation depends entirely on the design rather than on the headline P-value.

ConceptHow it appears in NIMBLE
Non-inferiority designPrimary hypothesis framed as non-inferiority to active biologics
Negative binomial regressionOverdispersed exacerbation counts over 52 weeks
Rate ratioPrimary effect measure: 1.16 (95% CI 0.98–1.38)
Confidence intervalsAsymmetric on the ratio scale; the upper limit drives the non-inferiority decision
P-valuesP = 0.079 illustrates why a non-significant difference is not evidence of similarity
ANCOVABaseline-adjusted comparison of ACQ-5 and FEV1 change
Covariate adjustmentPrior exacerbations, prior biologic, region and baseline FEV1
Analysis populationsFAS excluding undosed participants and GCP-violation sites
Endpoint definition7-day rule for separating exacerbation events

Statistical interpretation

The adjusted rate ratio exceeded 1 with a confidence interval spanning 0.98 to 1.38. Whether this satisfies non-inferiority depends on the prespecified margin, which is not reported in the posted registry analysis.

Clinical interpretation

Secondary measures of health status, asthma control and lung function showed differences close to zero with narrow intervals, and serious adverse event counts were similar. These supportive data should be read alongside, not instead of, the primary non-inferiority assessment.

13. Related Tutorials

Learn more about the methods used in this trial:

14. Related Calculators

15. Sources

Explore the methods behind the results

Connect this trial's endpoints and analysis methods to in-depth statistical tutorials and calculators.

16. Record Summary

NIMBLE compared depemokimab with mepolizumab or benralizumab in 1717 participants with severe eosinophilic asthma using a double-blind, randomized, non-inferiority design. The primary covariate-adjusted negative binomial analysis gave a rate ratio of 1.16 (95% CI 0.98–1.38; P = 0.079) for the annualized rate of clinically significant exacerbations. Secondary ANCOVA-type comparisons of SGRQ, ACQ-5 and FEV1 showed differences close to zero, and serious adverse events were reported in 81/859 and 76/855 participants. The most useful reading of this trial combines the rate ratio, the upper confidence limit relative to the non-inferiority margin, the precision of the secondary estimates and the analysis population, rather than the P-value alone.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry entry. The goal is to reconstruct the statistical story of the trial in a standardized format while clearly separating reported evidence from educational interpretation.