This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
DYNAGITO was a large, double-blind, randomized, parallel-group phase 3 trial asking whether adding the long-acting beta-agonist olodaterol to the long-acting muscarinic antagonist tiotropium, as a fixed-dose combination, reduces the rate of moderate to severe COPD exacerbations compared with tiotropium alone.
| Feature | DYNAGITO |
|---|---|
| Phase | Phase 3 |
| Condition | Pulmonary Disease, Chronic Obstructive (severe to very severe COPD) |
| Design | Randomized, parallel-group, double-masked |
| Arms | Tiotropium (5 µg) + olodaterol (5 µg) fixed-dose combination versus tiotropium 5 µg |
| Primary endpoint | Annualised rate of moderate to severe COPD exacerbations during the actual treatment period |
| Primary analysis | Negative binomial model with log treatment exposure as offset; treated set |
| Type I error control | Overall type I error protected at the 2-sided 0.01 level |
| Status | Completed; results posted |
| ClinicalTrials.gov | NCT02296138 |
| Sponsor | Boehringer Ingelheim (industry) |
2. Clinical Question
The central question was whether dual long-acting bronchodilation with tiotropium plus olodaterol lowers the frequency of moderate to severe exacerbations relative to tiotropium monotherapy in patients with severe to very severe COPD. Because exacerbations can occur repeatedly in the same patient, the primary endpoint was framed as an event rate per patient-year rather than as a single yes/no outcome.
Population
Patients with severe to very severe chronic obstructive pulmonary disease.
Intervention
Tiotropium + olodaterol (5/5 µg) fixed-dose combination.
Comparator
Tiotropium 5 µg alone.
Primary question
Does the combination reduce the annualised rate of moderate to severe COPD exacerbations, tested for superiority at a 2-sided 0.01 level?
3. Trial Design
Tiotropium + olodaterol
- Tiotropium 5 µg + olodaterol 5 µg
- Fixed-dose combination
- Long-acting muscarinic antagonist plus long-acting beta-agonist
Tiotropium monotherapy
- Tiotropium 5 µg
- Active comparator
- Long-acting muscarinic antagonist alone
The design is a straightforward two-arm, parallel-group superiority comparison between two active treatments. There was no placebo arm: both groups received tiotropium, so the comparison isolates the incremental contribution of olodaterol. In an active-controlled trial of this kind, the expected between-group difference is inherently smaller than it would be against placebo, which is one reason a very large sample was required.
4. Analysis Population
All posted efficacy analyses used the treated set (TS), defined in the registry as all randomised patients who were documented to have taken at least 1 dose of trial medication. This is sometimes described as a modified intention-to-treat population: patients are analysed by randomized arm, but those who never took study medication are excluded.
| Analysis set | Definition / role |
|---|---|
| Treated set (TS) | All randomised patients documented to have taken at least 1 dose of trial medication; used for the primary, sensitivity and secondary efficacy analyses posted to the registry. |
| Serious adverse event denominators | 3941 patients at risk in the tiotropium arm and 3939 at risk in the tiotropium + olodaterol arm. |
| Exclusion | 21 patients from one site closed for cause were excluded from the analysis sets. |
In a double-blind trial comparing two active inhalers, the gap between randomized and treated patients is usually small, and the treated-set definition is unlikely to introduce meaningful bias. It is still a departure from a pure intention-to-treat analysis, and readers should know which population a given estimate refers to.
5. Endpoints
Primary endpoint
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Annualised rate of moderate to severe COPD exacerbations during the actual treatment period | Calculated per treatment per patient-year. The actual treatment period was defined as the interval from first in-take of study medication until 1 day after last in-take of study medication. Least Squares Means are exponentiated. | From first in-take of study medication until 1 day after last in-take, up to 361 days |
Secondary endpoints with posted analyses
| Endpoint (registry wording, shortened) | Unit | Method |
|---|---|---|
| Number of patients with at least one moderate to severe COPD exacerbation during the actual treatment period | Number of patients | Cox proportional hazards model (HR); log-rank test (p-value) |
| Annualised rate of exacerbations leading to hospitalisation during the actual treatment period | Rate per patient-year | Negative binomial model with log exposure offset |
| Number of patients with at least one COPD exacerbation leading to hospitalisation during the actual treatment period | Number of patients | Cox proportional hazards model (HR); log-rank test (p-value) |
| Number of patients with all-cause mortality occurring during the actual treatment period | Number of patients | Cox proportional hazards model (HR); log-rank test (p-value) |
All secondary endpoints share the primary endpoint's time frame: from first in-take of study medication until 1 day after last in-take, up to 361 days. Although three of them are labelled as counts of patients, the posted hazard ratios and log-rank tests mean they were analysed as time to first event, not as simple proportions.
6. Primary Endpoint Results
The registry posts one primary analysis tied to the trial's type I error strategy, plus three covariate-adjusted models of the same endpoint (Section 7).
Rate ratio for moderate to severe exacerbations (combination vs tiotropium)
99% CI: 0.85–1.02 · P = 0.0498
Negative binomial model · treated set · superiority tested at 2-sided 0.01
| Item | Value |
|---|---|
| Effect measure | Ratio of rates vs. tiotropium 5 µg |
| Estimate | 0.93 |
| Confidence interval | 99%, two-sided: 0.85 to 1.02 |
| P-value | 0.0498 |
| Hypothesis | Superiority |
| Model | Negative binomial with fixed categorical effect of treatment and log of treatment exposure as offset |
| Error control | Hypothesis testing strategy protects overall type I error at 2-sided 0.01 |
What the estimate means. A rate ratio of 0.93 means that, under the fitted negative binomial model, the estimated annualised rate of moderate to severe exacerbations in the tiotropium + olodaterol group was 93% of the rate in the tiotropium group, a 7% lower estimated rate. It is a ratio of population-level event rates per patient-year of exposure.
What it does not mean. It does not mean that 7% of patients avoided an exacerbation, that each patient had 7% fewer exacerbations, or that the risk of having any exacerbation fell by 7%. Those are different quantities; the probability of a first exacerbation is addressed separately by the time-to-first-event secondary endpoint.
Precision. The 99% confidence interval of 0.85 to 1.02 is compatible with rate reductions of up to about 15% and also with a rate up to 2% higher in the combination arm. Because the interval includes 1.00, the data at this confidence level do not exclude no difference. The interval was reported at 99% rather than 95% because the trial's testing strategy spent alpha at the 2-sided 0.01 level; the 99% interval is the one that matches that decision rule.
The p-value is not the effect size. P = 0.0498 is the probability, under the null hypothesis of no difference, of a result at least this extreme. It says nothing on its own about whether a 7% rate reduction is clinically important. With nearly 4,000 patients per arm, even modest differences can generate small p-values.
Significance threshold. The registry states that the testing strategy protects the overall type I error at the 2-sided 0.01 level. A p-value of 0.0498 is below the conventional 0.05 threshold but above 0.01, so the primary endpoint did not meet the prespecified criterion for superiority. Describing this result as "significant at 0.05" would apply a threshold the trial did not use. The consistency between the p-value and the 99% interval (which crosses 1.00) is exactly what this rule implies.
Other cautions. The analysis is in the treated set, not all randomized patients; it excludes 21 patients from one closed site; and the negative binomial model assumes that between-patient variation in exacerbation rates follows a gamma-type pattern (overdispersion). The model included treatment only as a fixed effect, with no baseline covariates in this primary specification.
7. Covariate-Adjusted Models of the Primary Endpoint
The registry also posts three additional negative binomial analyses of the same endpoint, each using a covariate set modelled on another COPD trial programme. These can be read as sensitivity analyses showing how the estimate changes when prognostic baseline factors are added to the model.
| Model basis | Covariates (registry wording, shortened) | Rate ratio | 95% CI | P-value |
|---|---|---|---|---|
| SPARK/FLAME | Smoking status, baseline inhaled corticosteroid, GOLD stage, region, COPD Assessment Test score, exacerbations treated with antibiotics/steroids in previous year | 0.89 | 0.84–0.96 | 0.0010 |
| HERMES | Age, sex, smoking status, baseline LABA/ICS, region, percent predicted post-bronchodilator FEV1 | 0.91 | 0.85–0.98 | 0.0080 |
| TRINITY/TRILOGY | Treatment, region, severity of airflow limitation, smoking status, exacerbations treated with antibiotics/steroids in previous year | 0.89 | 0.84–0.96 | 0.0011 |
What the estimates mean. Rate ratios of 0.89 to 0.91 correspond to estimated exacerbation rates 9% to 11% lower in the combination arm after adjustment for baseline prognostic factors. All three adjusted estimates are slightly further from 1.00 than the unadjusted 0.93, and all three 95% intervals exclude 1.00.
Why adjustment can move the estimate. For a non-collapsible measure such as a rate ratio from a negative binomial model, adding strong prognostic covariates can shift the estimate away from 1 and tighten the interval even when randomization has balanced the arms. This is not evidence that the unadjusted analysis was biased; the adjusted and unadjusted estimates target slightly different (conditional versus marginal) quantities.
What they do NOT establish. These models were not the analysis tied to the 2-sided 0.01 testing strategy. Their smaller p-values (0.0010 to 0.0080) do not convert the primary endpoint into a positive confirmatory result: choosing among several model specifications after seeing the primary result would inflate the type I error that the prespecified strategy was designed to protect.
Confidence level. These results are reported with 95% intervals, whereas the primary analysis used a 99% interval. The widths are therefore not directly comparable; a 99% interval for the same data would be wider.
Direction wording. The registry's text for these three models describes the ratio as "Tiotropium (5 μg) versus Tiotropium (5 μg) + Olodaterol (5 μg)", whereas the primary analysis text describes the combination versus tiotropium. The values are of the same magnitude and direction as the primary estimate, but readers comparing the analyses should be aware of this inconsistency in the registry wording.
8. Secondary Endpoint Results
| Endpoint | Effect measure | Estimate | CI | P-value |
|---|---|---|---|---|
| At least one moderate to severe exacerbation (time to first) | Hazard ratio | 0.95 | 99%: 0.87–1.03 | 0.1188 |
| Annualised rate of exacerbations leading to hospitalisation | Ratio of events vs. tiotropium 5 µg | 0.89 | 95%: 0.76–1.03 | 0.1265 |
| At least one exacerbation leading to hospitalisation (time to first) | Hazard ratio | 0.93 | 95%: 0.82–1.06 | 0.2773 |
| All-cause mortality (time to death) | Hazard ratio | 1.09 | 95%: 0.67–1.75 | 0.7357 |
All comparisons are tiotropium + olodaterol versus tiotropium in the treated set, tested for superiority.
Time to first moderate to severe exacerbation
The hazard ratio of 0.95 indicates an estimated 5% lower hazard of a first moderate to severe exacerbation with the combination. The 99% interval (0.87 to 1.03) includes 1.00, and the log-rank p-value of 0.1188 is well above the 0.01 level. The registry notes that this endpoint falls under the same testing strategy that protects overall type I error at the 2-sided 0.01 level. Because the primary endpoint did not meet its criterion, a hierarchical strategy would typically not permit a confirmatory claim for this endpoint regardless of its p-value.
It is informative that the rate ratio (0.93) and the first-event hazard ratio (0.95) are both modest and in the same direction. The rate endpoint uses every exacerbation, including repeat events; the first-event endpoint uses only the first. A slightly stronger effect on the rate than on the first event could suggest some effect on recurrent events, but the difference between these two estimates is small and well within sampling variability.
Exacerbations leading to hospitalisation
For severe events requiring hospitalisation, the rate ratio was 0.89 (95% CI 0.76–1.03; P = 0.1265) and the first-event hazard ratio was 0.93 (95% CI 0.82–1.06; P = 0.2773). The point estimates favour the combination, but both intervals include 1.00. Hospitalised exacerbations are rarer than moderate ones, so these estimates rest on fewer events and are correspondingly less precise, as the wider intervals show.
All-cause mortality
The mortality hazard ratio of 1.09 means an estimated 9% higher hazard of death in the combination arm, but the 95% interval from 0.67 to 1.75 is very wide, spanning a substantial reduction and a substantial increase. With P = 0.7357, the data provide essentially no information to distinguish the arms on mortality over a treatment period of up to 361 days. A trial designed around exacerbation rates is not powered to detect mortality differences, and this estimate should not be read as either harm or safety with respect to death.
9. Safety: Serious Adverse Events
| Arm | Patients with serious adverse events | Patients at risk |
|---|---|---|
| Tiotropium + olodaterol (5/5 µg) | 810 | 3939 |
| Tiotropium 5 µg | 862 | 3941 |
Numerically fewer patients in the combination arm had at least one serious adverse event (810 of 3939) than in the tiotropium arm (862 of 3941), with near-identical denominators. The registry does not report a formal statistical comparison of serious adverse events. Serious adverse event counts in COPD populations include exacerbations requiring hospitalisation, so this safety summary partly overlaps with the efficacy endpoints and should not be read as an independent signal. Counts of patients with any event also do not reflect event severity, recurrence, or exposure time.
10. Statistical Methodology
Negative binomial regression for exacerbation rates
Exacerbations are count data: a patient can have zero, one or several during follow-up. A Poisson model assumes that the variance of the count equals its mean, but in COPD some patients are frequent exacerbators and others rarely exacerbate, so the observed variance is typically much larger than the mean (overdispersion). The negative binomial model accommodates this by allowing each patient's underlying rate to vary around the group mean.
Var(Yi) = μi + k·μi2 · Rate ratio = exp(β1)
Here Yi is the number of exacerbations for patient i, k is the overdispersion parameter, and log(Exposure) is the offset. The exponentiated treatment coefficient is the rate ratio, which is why the registry notes that least squares means "are actually exponentiated".
The exposure offset
Patients who discontinue early contribute less time at risk. Including the logarithm of treatment exposure as an offset converts the model from comparing raw counts to comparing rates per unit of exposure, so a patient treated for 120 days is not treated as equivalent to one treated for 361 days. The "actual treatment period" definition means events after treatment stops (beyond 1 day after last in-take) are not counted, which makes this an on-treatment estimand.
Cox proportional-hazards model and log-rank test
For the time-to-first-event secondary endpoints and mortality, hazard ratios and confidence intervals came from a Cox proportional hazards model, while p-values came from a log-rank test. The two methods are closely related: the log-rank test is the score test of an unadjusted Cox model with treatment as the only covariate. Both assume, for a single summary hazard ratio to be fully descriptive, that the ratio of hazards is roughly constant over follow-up.
A hazard ratio concerns the first event only. It is not a rate ratio for all events, not a relative risk at a fixed time, and not an absolute difference.
Covariate adjustment
The three additional primary-endpoint models added baseline covariates such as smoking status, region, airflow limitation severity, and prior exacerbation history. Adjusting for strongly prognostic variables typically improves precision in randomized trials; prior exacerbation history is the dominant predictor of future exacerbations in COPD, which explains why these models produced narrower intervals.
Type I error control at 2-sided 0.01
The registry states that the hypothesis testing strategy ensures the overall type I error is protected at the 2-sided 0.01 level. This is a stricter threshold than the conventional 0.05. A 0.01 threshold is often chosen when a single large trial is intended to provide evidence of a strength comparable to two independent trials each at 0.05. The primary and first secondary analyses were accordingly reported with 99% confidence intervals.
11. Multiplicity and Endpoint Hierarchy
| Analysis | Role | Interpretation |
|---|---|---|
| Annualised moderate to severe exacerbation rate (unadjusted) | Primary; under the 0.01 testing strategy | P = 0.0498, 99% CI crosses 1.00; prespecified superiority criterion not met |
| Covariate-adjusted rate models | Additional analyses of the primary endpoint | Supportive; 95% CIs; not a substitute for the prespecified test |
| Time to first moderate to severe exacerbation | Secondary; under the 0.01 testing strategy | P = 0.1188; not significant at 0.01 |
| Hospitalisation rate and time to first hospitalisation | Secondary | 95% CIs include 1.00; descriptive |
| All-cause mortality | Secondary | Very imprecise; descriptive |
When several endpoints are tested, the probability of at least one false-positive result rises unless the testing strategy controls it. The registry describes a strategy that protects the overall type I error; in such strategies, endpoints lower in the order are typically interpreted as confirmatory only if those above them succeed. The registry does not describe the full ordering or the method (for example, hierarchical or alpha-splitting) in detail.
12. Statistical Methods Explained
Why was a negative binomial model used instead of a Poisson model or a simple proportion?
The endpoint counts every moderate to severe exacerbation, and patients differ widely in how often they exacerbate. A Poisson model would understate the variance and produce confidence intervals that are too narrow. Analysing only the proportion of patients with any exacerbation would discard repeat events. The negative binomial model keeps all events and allows extra-Poisson variability.
What does a rate ratio of 0.93 with a 99% CI of 0.85 to 1.02 mean?
The estimated exacerbation rate with the combination was 93% of the rate with tiotropium, a 7% lower estimated rate. The 99% interval indicates the data are compatible, at that confidence level, with anything from roughly a 15% reduction to a 2% increase. Because the interval includes 1.00, the analysis cannot exclude no difference at the level the trial chose.
Why is P = 0.0498 not considered statistically significant here?
Statistical significance depends on the threshold fixed before the data are seen. The registry states that overall type I error was protected at the 2-sided 0.01 level. P = 0.0498 exceeds 0.01, so the prespecified superiority criterion was not met. The fact that it falls just under 0.05 is not relevant to the trial's decision rule, and treating it as a success would change the rules after the fact.
Why do the covariate-adjusted models give smaller p-values?
Adding baseline factors that strongly predict exacerbations, especially prior exacerbation history, explains part of the between-patient variation and can shift a non-collapsible rate ratio away from 1. The adjusted estimates (0.89 to 0.91) and their p-values (0.0010 to 0.0080) are therefore more favourable. But they were not the analysis tied to the error-controlled test, and selecting the most favourable of several specifications would inflate the false-positive rate.
Why use a log-rank test for the p-value but a Cox model for the hazard ratio?
The log-rank test is a nonparametric comparison of event-time distributions that is well suited to producing a p-value. It does not by itself produce an effect size, so a Cox proportional hazards model was used to estimate the hazard ratio and its confidence interval. In an unadjusted two-group comparison, the two approaches are closely aligned.
How should the mortality hazard ratio of 1.09 be read?
The point estimate is above 1, but the 95% interval from 0.67 to 1.75 is so wide that it is compatible with a large reduction or a large increase in the hazard of death. The trial was sized for exacerbation rates, not mortality, over a treatment period of up to 361 days. The appropriate conclusion is that the estimate is uninformative, not that the combination increases or decreases mortality.
13. Important Limitations and Interpretation Issues
- Primary criterion not met: the primary rate ratio did not reach the 2-sided 0.01 threshold, so the trial's confirmatory question was not answered in favour of superiority.
- Active comparator: both arms received tiotropium, so the estimates reflect the incremental effect of adding olodaterol, not the effect of either drug versus no treatment.
- On-treatment estimand: exacerbations were counted only during the actual treatment period, up to 1 day after last in-take. Events after discontinuation are not captured, so the estimate describes effects while on treatment rather than the effect of a treatment policy.
- Treated set rather than all randomized: the efficacy analyses excluded randomised patients without a documented dose, and 21 patients from a site closed for data irregularities were removed from all analysis sets.
- Model dependence: the rate ratio depends on the negative binomial specification, and the adjusted and unadjusted models give somewhat different estimates.
- Proportional hazards: the Cox-based hazard ratios summarise effects over up to 361 days as a single constant ratio; if the effect changed over time, one number is less descriptive.
- Multiplicity: secondary endpoints and additional models provide supporting information but do not carry confirmatory weight once the primary test is not significant.
- Safety summary: serious adverse events are reported as counts of affected patients by arm, without a formal comparison, and overlap with severe exacerbations.
14. Why This Trial Matters Statistically
DYNAGITO is a valuable teaching case precisely because its primary result sits in the ambiguous zone between conventional and prespecified thresholds. It shows how the choice of alpha, the confidence level, and the model specification shape what a trial can claim.
| Concept | How it appears in DYNAGITO |
|---|---|
| Recurrent-event endpoint | Annualised exacerbation rate counting all events per patient-year |
| Negative binomial regression | Overdispersed count model with log exposure offset |
| Rate ratio | 0.93 for moderate to severe exacerbations; 0.89 for hospitalised exacerbations |
| Prespecified alpha of 0.01 | P = 0.0498 below 0.05 but above 0.01; not significant under the trial's rule |
| 99% confidence interval | Interval (0.85–1.02) aligned with the 0.01 decision threshold |
| Covariate adjustment | Three adjusted models with smaller rate ratios and narrower intervals |
| Time-to-event endpoints | Time to first exacerbation, first hospitalisation and death |
| Hazard ratio and log-rank test | Cox model for estimates; log-rank for p-values |
| Multiplicity | Testing strategy protecting overall type I error across endpoints |
| Analysis population | Treated set with a data-integrity site exclusion |
Statistical interpretation
The primary rate ratio favoured the combination numerically but did not meet the prespecified 2-sided 0.01 superiority criterion; secondary endpoint estimates were in a similar direction for exacerbations but with intervals including 1.00.
Clinical interpretation
Any exacerbation benefit of adding olodaterol to tiotropium in this population appears modest in size. Whether such an effect is clinically meaningful is a separate judgement from its statistical significance.
15. Related Tutorials
Learn more about the methods used in this trial:
16. Related Calculators
17. Sources
- ClinicalTrials.gov: NCT02296138 — DYNAGITO registry record and posted results.
- PubMed: PMID 29605624.
- PubMed: PMID 30261995.
- PubMed: PMID 32776202.
- PubMed: PMID 32943047.
Explore more clinical trial statistics
Connect this trial's endpoints and methods to in-depth statistical tutorials and practical calculators.
18. Record Summary
DYNAGITO compared tiotropium + olodaterol with tiotropium alone in 7903 enrolled patients with severe to very severe COPD. Its primary negative binomial analysis estimated a rate ratio of 0.93 (99% CI 0.85–1.02; P = 0.0498) for moderate to severe exacerbations, which did not meet the trial's prespecified 2-sided 0.01 superiority threshold. Covariate-adjusted models gave estimates of 0.89 to 0.91 with 95% intervals excluding 1.00, and secondary exacerbation endpoints pointed in the same direction with intervals including 1.00. The most useful reading combines the size of the rate ratio, the confidence level matched to the testing strategy, the distinction between prespecified and supportive analyses, and the on-treatment, treated-set framing of the estimates.