This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
NOTUS was a quadruple-masked, randomized, parallel-group phase 3 trial asking whether dupilumab, added to inhaled maintenance therapy, reduces the rate of moderate or severe exacerbations in patients with moderate to severe COPD and evidence of type 2 inflammation. Its primary endpoint is a count outcome over 52 weeks, which shapes almost every statistical choice that follows.
| Feature | NOTUS |
|---|---|
| Registry title | Pivotal Study to Assess the Efficacy, Safety and Tolerability of Dupilumab in Patients With Moderate to Severe COPD With Type 2 Inflammation |
| Phase | Phase 3 |
| Condition | Chronic Obstructive Pulmonary Disease |
| Design | Randomized, parallel-group, quadruple-masked, placebo-controlled |
| Arms | 2: dupilumab 300 mg q2w; placebo |
| Enrollment | 935 |
| Primary endpoint | Annualized rate of moderate or severe COPD exacerbations over the 52-week treatment period |
| Primary analysis | Negative binomial model, ITT population, superiority |
| Status | Completed; results posted |
| Dates | Start 2020-07-06; primary completion 2024-02-28 |
| ClinicalTrials.gov | NCT04456673 |
| Lead sponsor | Sanofi (industry) |
2. Clinical Question
The central question was whether adding dupilumab to background inhaled maintenance therapy lowers the annualized rate of moderate or severe COPD exacerbations compared with placebo on the same background therapy, in a population selected for type 2 inflammation.
Population
Patients with moderate to severe chronic obstructive pulmonary disease with type 2 inflammation.
Intervention
Dupilumab (SAR231893) 300 mg every two weeks, alongside inhaled maintenance therapy (inhaled corticosteroid, long-acting beta agonist and long-acting muscarinic antagonist).
Comparator
Placebo, alongside the same inhaled maintenance therapy.
Primary question
Is the annualized rate of moderate or severe exacerbations over 52 weeks lower with dupilumab than with placebo?
Because both arms received background inhaled therapy, the comparison estimates the add-on effect of dupilumab. The trial does not compare dupilumab with any of the inhaled drugs, and it does not estimate the effect of dupilumab used alone.
3. Trial Design
Dupilumab 300 mg q2w
- Dupilumab (SAR231893) 300 mg every two weeks
- Inhaled corticosteroid
- Inhaled long-acting beta agonist
- Inhaled long-acting muscarinic antagonist
Placebo
- Placebo on the same schedule
- Inhaled corticosteroid
- Inhaled long-acting beta agonist
- Inhaled long-acting muscarinic antagonist
Analysis populations
| Population | Definition in the registry | Used for |
|---|---|---|
| Intent-to-treat (ITT) | All randomized participants analyzed according to the treatment group allocated by randomization | Primary exacerbation endpoint; Week 12 FEV1 (participants with data collected) |
| ITT with an opportunity to reach Week 52 | Participants who had an opportunity to reach Week 52 assessments | Continuous and proportion-type endpoints at Week 52 (SGRQ, SGRQ responders, Week 52 FEV1) |
4. Endpoints
Primary endpoint
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Annualized rate of moderate or severe COPD exacerbations over the 52-week treatment period | Moderate exacerbations: acute exacerbation of COPD (AECOPD) requiring systemic corticosteroids (such as intramuscular, intravenous or oral) and/or antibiotics. Severe exacerbations: AECOPD requiring hospitalization or observation for >24 hours in an emergency department/urgent care facility, or resulting in death. Both recorded by the investigator. Events were counted separately only if at least 14 days apart. The annualized rate is the total number of events during the 52-week treatment period divided by the total participant-years followed. | Baseline (Day 1) to Week 52 |
Secondary endpoints with posted statistical analyses
| Endpoint | Time frame | Unit | Endpoint type |
|---|---|---|---|
| Change from baseline in pre-bronchodilator FEV1 to Week 12 | Baseline (Day 1) and Week 12 | Liters | Continuous, repeated measures |
| Change from baseline in SGRQ total score to Week 52 | Baseline (Day 1) and Week 52 | Score on a scale | Continuous, repeated measures |
| Percentage of participants with SGRQ improvement ≥4 points at Week 52 | Baseline (Day 1) and Week 52 | Percentage of participants | Binary responder |
| Change from baseline in pre-bronchodilator FEV1 to Week 52 | Baseline (Day 1) and Week 52 | Liters | Continuous, repeated measures |
The endpoint family illustrates three distinct data types, each requiring its own model: an event count accumulated over variable follow-up (negative binomial regression), continuous changes measured at several visits (MMRM), and a binary responder outcome (logistic regression).
5. Primary Endpoint Results
The primary analysis compared the annualized rate of moderate or severe exacerbations between arms in the ITT population using a negative binomial model. The registry reports the treatment effect as a difference in annualized rates, in exacerbations per participant-year, obtained from the model using the delta method.
Difference in annualized exacerbation rate (dupilumab 300 mg q2w vs placebo)
Exacerbations per participant-year · 95% CI: −0.682 to −0.188 · P = 0.0002
Negative binomial model, ITT population, two-sided superiority test
| Item | Registry entry |
|---|---|
| Groups compared | Placebo vs Dupilumab 300 mg q2w |
| Statistical method | Negative binomial model |
| Effect measure | Labelled "Risk Difference (RD)"; in substance a difference in annualized event rates |
| Estimate (95% CI) | −0.435 (−0.682 to −0.188) |
| P-value | 0.0002 |
| Response variable | Total number of events during the 52-week treatment period |
| Covariates | Treatment group, region (pooled country), ICS dose, smoking status at screening, baseline disease severity, number of moderate or severe COPD exacerbation events within one year prior to the study |
| Offset | Log-transformed treatment duration |
| Derivation of difference | Delta method |
What the estimate means. The estimate of −0.435 is an absolute difference in the model-based annualized rate. Read in the direction consistent with the other endpoints (dupilumab minus placebo), it corresponds to roughly 0.435 fewer moderate or severe exacerbations per participant-year of follow-up in the dupilumab arm. Scaled to a group, that is about 43.5 fewer exacerbations per 100 participant-years.
What it does not mean. It is not a probability and not a proportion of patients: the registry's "risk difference" label should not be read as a difference in the percentage of patients who had an exacerbation. It also does not say that each patient has 0.435 fewer exacerbations; exacerbation counts are highly variable between patients, which is exactly why a negative binomial rather than a Poisson model was used. Nor is it a relative reduction: the posted statistical analysis presents the absolute rate difference, and a rate ratio would be a different (though related) summary of the same model.
Precision. The 95% confidence interval from −0.682 to −0.188 excludes zero, and all values in it favour dupilumab. Its width shows that the data are compatible with a modest absolute reduction near 0.19 exacerbations per participant-year as well as a larger one near 0.68. The interval is symmetric around the estimate, which is characteristic of a delta-method standard error applied on the rate-difference scale.
Why the p-value is not the effect size. P = 0.0002 says that a difference at least this large would be very unlikely if the true rates were equal. It reflects both effect size and the amount of information (935 participants followed for up to a year); it does not measure how large or clinically important the reduction is. The estimate and interval carry that information.
Cautions. The analysis follows ITT, so it estimates the effect of being assigned to dupilumab, including any treatment discontinuation. Participants followed for less than 52 weeks contribute through the log-duration offset rather than being dropped, which assumes that the event rate is reasonably constant over the time each participant is observed. Because the model adjusts for baseline covariates, the result is a covariate-adjusted estimate, and the delta-method interval relies on a large-sample normal approximation.
Yi is the number of exacerbations for participant i, ti the treatment duration (entering as a log offset so the model describes a rate), Xi the covariates, and k the overdispersion parameter. exp(β1) is a rate ratio; the registry's rate difference is obtained by transforming the model-based rates in each arm and computing a delta-method standard error.
6. Secondary Endpoint Results
Four secondary endpoints have posted statistical analyses. All compare dupilumab 300 mg q2w with placebo and all were tested for superiority.
| Endpoint | Method | Effect measure | Estimate (95% CI) | P-value |
|---|---|---|---|---|
| Change in pre-bronchodilator FEV1 to Week 12 (L) | MMRM | LS mean difference | 0.082 (0.040 to 0.124) | 0.0001 |
| Change in SGRQ total score to Week 52 | MMRM | LS mean difference | −3.371 (−5.811 to −0.931) | 0.0068 |
| SGRQ improvement ≥4 points at Week 52 | Logistic regression | Odds ratio | 1.164 (0.856 to 1.581) | 0.3329 |
| Change in pre-bronchodilator FEV1 to Week 52 (L) | MMRM | LS mean difference | 0.062 (0.011 to 0.113) | 0.0182 |
Pre-bronchodilator FEV1 at Week 12
LS mean difference in change from baseline
95% CI: 0.040 to 0.124 · P = 0.0001
MMRM; ITT population, participants with data collected
The MMRM used change from baseline in pre-bronchodilator FEV1 up to Week 12 as the response, with treatment group, age, sex, height, region (pooled country), ICS dose, smoking status at screening, visit, treatment-by-visit interaction, baseline pre-bronchodilator FEV1 and FEV1 baseline-by-visit interaction as covariates. The estimate of 0.082 L (82 mL) is a difference between arms in mean change from baseline, not the change within the dupilumab arm; both arms may have changed from baseline. The interval excludes zero and indicates a between-arm difference of roughly 40 to 124 mL.
SGRQ total score at Week 52
LS mean difference in change from baseline
95% CI: −5.811 to −0.931 · P = 0.0068
MMRM; ITT population with an opportunity to reach Week 52, participants with data collected
Lower SGRQ scores indicate better respiratory health status, so a negative difference favours dupilumab. The covariates were treatment group, region (pooled country), ICS dose, smoking status at screening, treatment-by-visit interaction, baseline SGRQ total score and SGRQ baseline-by-visit interaction. The mean difference of −3.371 points lies below the 4-point threshold used to define an individual responder, and the confidence interval spans values from under 1 point to almost 6 points.
SGRQ responders (improvement ≥4 points) at Week 52
Odds ratio for SGRQ improvement ≥4 points
95% CI: 0.856 to 1.581 · P = 0.3329
Logistic regression; ITT population with an opportunity to reach Week 52
The logistic model included treatment group, region (pooled country), ICS dose, smoking status at screening and baseline SGRQ total score. An odds ratio of 1.164 means the estimated odds of a ≥4-point improvement were 16.4% higher with dupilumab, but the interval runs from 0.856 (lower odds) to 1.581 (higher odds) and includes 1. The data are therefore compatible with no difference in the odds of response.
Why the continuous and responder SGRQ results differ
The mean-change analysis showed a between-arm difference with P = 0.0068, while the responder analysis of the same questionnaire did not. This is not a contradiction. Dichotomizing a continuous score at 4 points discards information: a participant who improves by 3.9 points and one who worsens by 10 points are counted identically as non-responders. Responder analyses therefore usually have less statistical power than analyses of the underlying continuous variable. A shift in the mean can also occur through many small changes without moving many participants across a fixed threshold. The odds ratio is, in addition, a different scale from the mean difference, and odds ratios are not the same as ratios of proportions when responses are common.
Pre-bronchodilator FEV1 at Week 52
LS mean difference in change from baseline
95% CI: 0.011 to 0.113 · P = 0.0182
MMRM; ITT population with an opportunity to reach Week 52, participants with data collected
The model specification matched the Week 12 FEV1 analysis. The Week 52 estimate (0.062 L) is smaller than the Week 12 estimate (0.082 L) and its interval is closer to zero. Whether this P-value can be read as confirmatory depends on its position in the testing hierarchy, discussed next.
7. Multiplicity and the Testing Hierarchy
For every posted analysis the registry states that "a hierarchical testing procedure was used to control type I error and handle primary and first 4 secondary endpoints (reported sequentially) analyses at a 2-sided significance level of 0.05."
Each hypothesis uses the full 0.05, but only if every hypothesis before it was rejected. Once one test fails, all later hypotheses are considered not formally tested, whatever their P-values.
| Order as reported | Endpoint | P-value | Reading if the hierarchy follows this order |
|---|---|---|---|
| 1 | Annualized moderate or severe exacerbation rate | 0.0002 | Significant at 0.05; proceed |
| 2 | Pre-bronchodilator FEV1, Week 12 | 0.0001 | Significant at 0.05; proceed |
| 3 | SGRQ total score change, Week 52 | 0.0068 | Significant at 0.05; proceed |
| 4 | SGRQ improvement ≥4 points, Week 52 | 0.3329 | Not significant; hierarchy stops |
| 5 | Pre-bronchodilator FEV1, Week 52 | 0.0182 | Nominal only; not formally tested under the hierarchy |
8. Safety: Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by arm as the number of participants affected over the number at risk.
| Arm | Participants with serious adverse events | Participants at risk |
|---|---|---|
| Dupilumab 300 mg q2w | 65 | 469 |
| Placebo | 79 | 464 |
Fewer participants in the dupilumab arm had a serious adverse event (65 of 469) than in the placebo arm (79 of 464). No formal statistical comparison of serious adverse events is posted, and safety tables of this kind are normally descriptive: they are not powered or multiplicity-controlled, and serious adverse events in COPD include exacerbation-related hospitalizations, which overlap with the efficacy endpoint. The numbers at risk (469 and 464, total 933) differ slightly from the 935 enrolled, reflecting the usual convention that safety is summarized among participants who were actually exposed to study treatment.
9. Statistical Methodology
Negative binomial regression for exacerbation counts
Exacerbations are discrete events that can recur within a patient, so the natural outcome is a count per unit of follow-up time. A Poisson model assumes the variance equals the mean; in COPD, a minority of patients typically have many exacerbations while most have few or none, producing variance larger than the mean (overdispersion). The negative binomial model adds a dispersion parameter to absorb this extra variability, giving standard errors that are not artificially small. The log-transformed treatment duration offset converts counts into rates, so participants with shorter follow-up contribute in proportion to their observation time.
Covariate adjustment
Every posted analysis adjusts for baseline covariates, including region (pooled country), ICS dose and smoking status at screening, and for the exacerbation endpoint, baseline disease severity and prior-year exacerbation count. In a randomized trial, adjustment is not needed to remove bias; its purpose is to account for strong prognostic factors, which typically increases precision. Prior exacerbation history is usually the strongest predictor of future exacerbations, making it an especially valuable covariate for this endpoint.
Mixed model for repeated measures (MMRM)
FEV1 and SGRQ were measured at several visits. The MMRM uses all post-baseline visits in one model, with visit, treatment-by-visit interaction and baseline-by-visit interaction terms, so that the treatment effect is estimated at a specific visit (Week 12 or Week 52) while borrowing information across visits through the within-patient correlation structure. The reported least-squares (LS) mean difference is the model-adjusted difference between arms at the target visit.
Missing data in the continuous endpoints
The registry notes that only participants with data collected are reported for the MMRM endpoints. MMRM does not impute missing values explicitly; it produces valid estimates under the assumption that data are missing at random, meaning that, given the observed data and covariates, the probability of dropout does not depend on unobserved values. If participants who left early would have done systematically worse, that assumption fails and the estimate may be biased. Sensitivity analyses are the usual way of examining this, and none are part of the posted statistical analyses.
Logistic regression for the responder endpoint
The proportion of participants with a ≥4-point SGRQ improvement was analyzed with logistic regression adjusted for region, ICS dose, smoking status and baseline SGRQ. The treatment coefficient, exponentiated, gives the adjusted odds ratio. Because adjusted odds ratios are non-collapsible, their value can differ from an unadjusted odds ratio even when covariates are perfectly balanced; this is a property of the scale, not a sign of confounding.
Intention-to-treat principle
The primary analysis used all randomized participants analyzed according to their assigned group. This preserves the comparability created by randomization and estimates the effect of a treatment policy, rather than the effect in participants who adhered perfectly.
10. Statistical Methods Explained
Why was a negative binomial model used rather than a t-test on exacerbation rates?
Individual exacerbation rates are skewed counts with many zeros and varying follow-up time. The negative binomial model handles the count nature of the data, the overdispersion between patients and unequal exposure (through the log-duration offset) in one framework, while also allowing covariate adjustment. A t-test on per-patient rates would weight a patient observed for one month the same as one observed for a full year.
What does a rate difference of −0.435 mean, and why is it labelled a "risk difference"?
It is the difference in model-estimated exacerbations per participant-year between arms, about 43.5 fewer events per 100 participant-years with dupilumab. The registry's effect-measure category is "Risk Difference (RD)", but the unit, exacerbations per participant-year, shows it is a difference in rates rather than in the proportion of patients with an event.
What is the delta method doing here?
The negative binomial model estimates effects on the log-rate scale. Converting to an absolute rate difference requires a non-linear transformation of the model parameters. The delta method approximates the standard error of that transformed quantity using a first-order Taylor expansion, giving the symmetric 95% interval of −0.682 to −0.188.
Why does the FEV1 analysis include height, age and sex, while the SGRQ analysis does not?
FEV1 depends strongly on body size, age and sex, so including them explains a large share of between-patient variability and sharpens the treatment comparison. SGRQ is a symptom and quality-of-life score whose main predictor is its own baseline value, which the SGRQ model includes together with a baseline-by-visit interaction.
Why was the SGRQ mean change significant but the responder analysis not?
Dichotomizing at 4 points throws away information and lowers power, and a mean difference of −3.371 can arise without many more participants crossing a 4-point threshold. The odds ratio of 1.164 (95% CI 0.856 to 1.581) shows a direction in favour of dupilumab but with an interval that includes no effect.
Why is the Week 52 FEV1 P-value of 0.0182 not automatically "significant"?
Under a hierarchical procedure, a hypothesis can be formally rejected only if all hypotheses before it in the sequence were rejected. If the Week 52 FEV1 endpoint came after the SGRQ responder endpoint, as the reported order suggests, the chain broke at P = 0.3329, and the 0.0182 is a nominal P-value outside the controlled family.
11. Limitations
- Absolute scale only in the posted analysis: the posted primary analysis gives a rate difference; its meaning depends on the underlying placebo rate, which varies between populations. The same model-based rate ratio would transport differently.
- Hierarchy break: the non-significant SGRQ responder result limits the confirmatory status of any endpoint tested after it.
- Missing data assumptions: MMRM results rely on data being missing at random; the Week 52 analyses also restrict to participants who had an opportunity to reach Week 52.
- Investigator-recorded events: exacerbation definitions depend on treatment decisions (systemic corticosteroids, antibiotics, hospitalization), which can vary by region and practice; masking reduces but does not remove this variability.
- Constant-rate assumption: the offset approach treats the exacerbation rate as constant over each participant's follow-up, whereas COPD exacerbations are seasonal and can cluster.
- Selected population: the trial enrolled patients with type 2 inflammation on inhaled maintenance therapy; the results do not describe the effect in COPD without type 2 inflammation.
- Descriptive safety: serious adverse events are reported as counts of affected participants without formal between-arm testing.
- Subgroups not reported: the ClinicalTrials.gov record does not report subgroup analyses, so consistency of the effect across patient characteristics cannot be assessed from the registry.
12. Why This Trial Matters Statistically
NOTUS is a useful teaching case because a single trial brings together three different outcome types, each matched to an appropriate model, under one multiplicity framework in which one link in the chain fails.
| Concept | How it appears in NOTUS |
|---|---|
| Count outcomes and overdispersion | Negative binomial model for annualized exacerbation rate |
| Exposure offsets | Log-transformed treatment duration converts counts to rates |
| Absolute vs relative effects | Rate difference derived by the delta method |
| Repeated measures | MMRM for FEV1 at Weeks 12 and 52 and SGRQ at Week 52 |
| Covariate adjustment | Prognostic baseline factors in every model |
| Continuous vs responder analysis | SGRQ mean change significant; ≥4-point responder odds ratio not |
| Odds ratio | 1.164 with a 95% CI including 1 |
| Multiplicity | Fixed-sequence hierarchy across the primary and first 4 secondary endpoints |
| Nominal vs confirmatory P-values | Week 52 FEV1 P = 0.0182 after the hierarchy break |
| Intention-to-treat | All randomized participants analyzed as allocated |
| Missing data | MMRM under missing at random; analysis restricted to participants with data |
Statistical interpretation
The primary exacerbation-rate difference and the first two secondary endpoints favoured dupilumab with confidence intervals excluding no effect; the SGRQ responder endpoint did not, which ends formal testing in the reported sequence.
Clinical interpretation
The registry results describe fewer moderate or severe exacerbations and better lung function over the trial period with dupilumab added to inhaled therapy in this selected population, with health-status gains that were smaller when judged by individual response thresholds.
13. Trial Timeline
Study start
Enrollment of participants with moderate to severe COPD with type 2 inflammation began.
Primary completion
Final data collection for the primary outcome; 935 participants enrolled.
Results posted
Status completed, with results and statistical analyses posted to ClinicalTrials.gov.
14. Related Tutorials
Learn more about the methods used in this trial:
15. Related Calculators
16. Sources
- ClinicalTrials.gov: NCT04456673, Pivotal Study to Assess the Efficacy, Safety and Tolerability of Dupilumab in Patients With Moderate to Severe COPD With Type 2 Inflammation (NOTUS).
- PubMed: PMID 39894389
- PubMed: PMID 38767614
- PubMed: PMID 41794122
- PubMed: PMID 42153327
- PubMed: PMID 42337097
Explore more trial statistics
Follow the methods used in NOTUS to step-by-step tutorials and calculators, or browse other trials analyzed in the same format.
17. Record Summary
NOTUS combines a randomized, quadruple-masked comparison with an event-rate primary endpoint analyzed by covariate-adjusted negative binomial regression, continuous secondary endpoints analyzed by MMRM, and a binary responder endpoint analyzed by logistic regression, all within a fixed-sequence hierarchy. The most informative reading keeps the absolute rate difference and its confidence interval in view, distinguishes confirmatory from nominal P-values along the testing sequence, and recognizes that mean-change and responder analyses of the same instrument answer different questions.