← Clinical Trials
Severe to Very Severe COPD Phase 3 Recurrent-Event Endpoint NCT02296138

DYNAGITO: Complete Statistical Analysis of Tiotropium/Olodaterol in Severe COPD

A statistical review of the randomized, double-blind phase 3 DYNAGITO trial, which compared a tiotropium + olodaterol (5/5 µg) fixed-dose combination with tiotropium 5 µg for reducing moderate to severe exacerbations in patients with severe to very severe chronic obstructive pulmonary disease.

Study start: 2015-01-13  ·  Primary completion: 2017-03-08  ·  Sponsor: Boehringer Ingelheim
About this analysis

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

DYNAGITO was a large, double-blind, randomized, parallel-group phase 3 trial asking whether adding the long-acting beta-agonist olodaterol to the long-acting muscarinic antagonist tiotropium, as a fixed-dose combination, reduces the rate of moderate to severe COPD exacerbations compared with tiotropium alone.

7903
Enrolled
2 parallel arms
0.93
Primary rate ratio
99% CI 0.85–1.02
0.0498
Primary p-value
Tested at 2-sided 0.01
361
Days, maximum
Actual treatment period
FeatureDYNAGITO
PhasePhase 3
ConditionPulmonary Disease, Chronic Obstructive (severe to very severe COPD)
DesignRandomized, parallel-group, double-masked
ArmsTiotropium (5 µg) + olodaterol (5 µg) fixed-dose combination versus tiotropium 5 µg
Primary endpointAnnualised rate of moderate to severe COPD exacerbations during the actual treatment period
Primary analysisNegative binomial model with log treatment exposure as offset; treated set
Type I error controlOverall type I error protected at the 2-sided 0.01 level
StatusCompleted; results posted
ClinicalTrials.govNCT02296138
SponsorBoehringer Ingelheim (industry)

2. Clinical Question

The central question was whether dual long-acting bronchodilation with tiotropium plus olodaterol lowers the frequency of moderate to severe exacerbations relative to tiotropium monotherapy in patients with severe to very severe COPD. Because exacerbations can occur repeatedly in the same patient, the primary endpoint was framed as an event rate per patient-year rather than as a single yes/no outcome.

Population

Patients with severe to very severe chronic obstructive pulmonary disease.

Intervention

Tiotropium + olodaterol (5/5 µg) fixed-dose combination.

Comparator

Tiotropium 5 µg alone.

Primary question

Does the combination reduce the annualised rate of moderate to severe COPD exacerbations, tested for superiority at a 2-sided 0.01 level?

3. Trial Design

01
Enrol7903 patients
02
RandomizeDouble-masked, 2 arms
03
TreatUp to 361 days
04
Count eventsExacerbations per patient-year
05
AnalyseTreated set, negative binomial
ARM 1 · n = 3939 at risk (safety)

Tiotropium + olodaterol

  • Tiotropium 5 µg + olodaterol 5 µg
  • Fixed-dose combination
  • Long-acting muscarinic antagonist plus long-acting beta-agonist
ARM 2 · n = 3941 at risk (safety)

Tiotropium monotherapy

  • Tiotropium 5 µg
  • Active comparator
  • Long-acting muscarinic antagonist alone

The design is a straightforward two-arm, parallel-group superiority comparison between two active treatments. There was no placebo arm: both groups received tiotropium, so the comparison isolates the incremental contribution of olodaterol. In an active-controlled trial of this kind, the expected between-group difference is inherently smaller than it would be against placebo, which is one reason a very large sample was required.

Excluded site. The registry states that one site was closed for cause due to data irregularities, and that the data from the 21 patients at this site were excluded from the patient sets used for analysis. This is a pre-analysis data-integrity exclusion; with thousands of patients per arm, its numerical influence on the treatment comparison is likely to be small, but it does mean the analysed sets are not identical to the full enrolled population.

4. Analysis Population

All posted efficacy analyses used the treated set (TS), defined in the registry as all randomised patients who were documented to have taken at least 1 dose of trial medication. This is sometimes described as a modified intention-to-treat population: patients are analysed by randomized arm, but those who never took study medication are excluded.

Analysis setDefinition / role
Treated set (TS)All randomised patients documented to have taken at least 1 dose of trial medication; used for the primary, sensitivity and secondary efficacy analyses posted to the registry.
Serious adverse event denominators3941 patients at risk in the tiotropium arm and 3939 at risk in the tiotropium + olodaterol arm.
Exclusion21 patients from one site closed for cause were excluded from the analysis sets.

In a double-blind trial comparing two active inhalers, the gap between randomized and treated patients is usually small, and the treated-set definition is unlikely to introduce meaningful bias. It is still a departure from a pure intention-to-treat analysis, and readers should know which population a given estimate refers to.

5. Endpoints

Primary endpoint

EndpointRegistry definitionTime frame
Annualised rate of moderate to severe COPD exacerbations during the actual treatment periodCalculated per treatment per patient-year. The actual treatment period was defined as the interval from first in-take of study medication until 1 day after last in-take of study medication. Least Squares Means are exponentiated.From first in-take of study medication until 1 day after last in-take, up to 361 days

Secondary endpoints with posted analyses

Endpoint (registry wording, shortened)UnitMethod
Number of patients with at least one moderate to severe COPD exacerbation during the actual treatment periodNumber of patientsCox proportional hazards model (HR); log-rank test (p-value)
Annualised rate of exacerbations leading to hospitalisation during the actual treatment periodRate per patient-yearNegative binomial model with log exposure offset
Number of patients with at least one COPD exacerbation leading to hospitalisation during the actual treatment periodNumber of patientsCox proportional hazards model (HR); log-rank test (p-value)
Number of patients with all-cause mortality occurring during the actual treatment periodNumber of patientsCox proportional hazards model (HR); log-rank test (p-value)

All secondary endpoints share the primary endpoint's time frame: from first in-take of study medication until 1 day after last in-take, up to 361 days. Although three of them are labelled as counts of patients, the posted hazard ratios and log-rank tests mean they were analysed as time to first event, not as simple proportions.

6. Primary Endpoint Results

The registry posts one primary analysis tied to the trial's type I error strategy, plus three covariate-adjusted models of the same endpoint (Section 7).

Rate ratio for moderate to severe exacerbations (combination vs tiotropium)

0.93

99% CI: 0.85–1.02   ·   P = 0.0498

Negative binomial model · treated set · superiority tested at 2-sided 0.01

ItemValue
Effect measureRatio of rates vs. tiotropium 5 µg
Estimate0.93
Confidence interval99%, two-sided: 0.85 to 1.02
P-value0.0498
HypothesisSuperiority
ModelNegative binomial with fixed categorical effect of treatment and log of treatment exposure as offset
Error controlHypothesis testing strategy protects overall type I error at 2-sided 0.01
Clinical Biostats interpretation

What the estimate means. A rate ratio of 0.93 means that, under the fitted negative binomial model, the estimated annualised rate of moderate to severe exacerbations in the tiotropium + olodaterol group was 93% of the rate in the tiotropium group, a 7% lower estimated rate. It is a ratio of population-level event rates per patient-year of exposure.

What it does not mean. It does not mean that 7% of patients avoided an exacerbation, that each patient had 7% fewer exacerbations, or that the risk of having any exacerbation fell by 7%. Those are different quantities; the probability of a first exacerbation is addressed separately by the time-to-first-event secondary endpoint.

Precision. The 99% confidence interval of 0.85 to 1.02 is compatible with rate reductions of up to about 15% and also with a rate up to 2% higher in the combination arm. Because the interval includes 1.00, the data at this confidence level do not exclude no difference. The interval was reported at 99% rather than 95% because the trial's testing strategy spent alpha at the 2-sided 0.01 level; the 99% interval is the one that matches that decision rule.

The p-value is not the effect size. P = 0.0498 is the probability, under the null hypothesis of no difference, of a result at least this extreme. It says nothing on its own about whether a 7% rate reduction is clinically important. With nearly 4,000 patients per arm, even modest differences can generate small p-values.

Significance threshold. The registry states that the testing strategy protects the overall type I error at the 2-sided 0.01 level. A p-value of 0.0498 is below the conventional 0.05 threshold but above 0.01, so the primary endpoint did not meet the prespecified criterion for superiority. Describing this result as "significant at 0.05" would apply a threshold the trial did not use. The consistency between the p-value and the 99% interval (which crosses 1.00) is exactly what this rule implies.

Other cautions. The analysis is in the treated set, not all randomized patients; it excludes 21 patients from one closed site; and the negative binomial model assumes that between-patient variation in exacerbation rates follows a gamma-type pattern (overdispersion). The model included treatment only as a fixed effect, with no baseline covariates in this primary specification.

7. Covariate-Adjusted Models of the Primary Endpoint

The registry also posts three additional negative binomial analyses of the same endpoint, each using a covariate set modelled on another COPD trial programme. These can be read as sensitivity analyses showing how the estimate changes when prognostic baseline factors are added to the model.

Model basisCovariates (registry wording, shortened)Rate ratio95% CIP-value
SPARK/FLAMESmoking status, baseline inhaled corticosteroid, GOLD stage, region, COPD Assessment Test score, exacerbations treated with antibiotics/steroids in previous year0.890.84–0.960.0010
HERMESAge, sex, smoking status, baseline LABA/ICS, region, percent predicted post-bronchodilator FEV10.910.85–0.980.0080
TRINITY/TRILOGYTreatment, region, severity of airflow limitation, smoking status, exacerbations treated with antibiotics/steroids in previous year0.890.84–0.960.0011
Rate ratio point estimates (1.00 = no difference)
Primary (unadjusted)
0.93
SPARK/FLAME model
0.89
HERMES model
0.91
TRINITY/TRILOGY model
0.89
Clinical Biostats interpretation

What the estimates mean. Rate ratios of 0.89 to 0.91 correspond to estimated exacerbation rates 9% to 11% lower in the combination arm after adjustment for baseline prognostic factors. All three adjusted estimates are slightly further from 1.00 than the unadjusted 0.93, and all three 95% intervals exclude 1.00.

Why adjustment can move the estimate. For a non-collapsible measure such as a rate ratio from a negative binomial model, adding strong prognostic covariates can shift the estimate away from 1 and tighten the interval even when randomization has balanced the arms. This is not evidence that the unadjusted analysis was biased; the adjusted and unadjusted estimates target slightly different (conditional versus marginal) quantities.

What they do NOT establish. These models were not the analysis tied to the 2-sided 0.01 testing strategy. Their smaller p-values (0.0010 to 0.0080) do not convert the primary endpoint into a positive confirmatory result: choosing among several model specifications after seeing the primary result would inflate the type I error that the prespecified strategy was designed to protect.

Confidence level. These results are reported with 95% intervals, whereas the primary analysis used a 99% interval. The widths are therefore not directly comparable; a 99% interval for the same data would be wider.

Direction wording. The registry's text for these three models describes the ratio as "Tiotropium (5 μg) versus Tiotropium (5 μg) + Olodaterol (5 μg)", whereas the primary analysis text describes the combination versus tiotropium. The values are of the same magnitude and direction as the primary estimate, but readers comparing the analyses should be aware of this inconsistency in the registry wording.

8. Secondary Endpoint Results

EndpointEffect measureEstimateCIP-value
At least one moderate to severe exacerbation (time to first)Hazard ratio0.9599%: 0.87–1.030.1188
Annualised rate of exacerbations leading to hospitalisationRatio of events vs. tiotropium 5 µg0.8995%: 0.76–1.030.1265
At least one exacerbation leading to hospitalisation (time to first)Hazard ratio0.9395%: 0.82–1.060.2773
All-cause mortality (time to death)Hazard ratio1.0995%: 0.67–1.750.7357

All comparisons are tiotropium + olodaterol versus tiotropium in the treated set, tested for superiority.

Time to first moderate to severe exacerbation

The hazard ratio of 0.95 indicates an estimated 5% lower hazard of a first moderate to severe exacerbation with the combination. The 99% interval (0.87 to 1.03) includes 1.00, and the log-rank p-value of 0.1188 is well above the 0.01 level. The registry notes that this endpoint falls under the same testing strategy that protects overall type I error at the 2-sided 0.01 level. Because the primary endpoint did not meet its criterion, a hierarchical strategy would typically not permit a confirmatory claim for this endpoint regardless of its p-value.

It is informative that the rate ratio (0.93) and the first-event hazard ratio (0.95) are both modest and in the same direction. The rate endpoint uses every exacerbation, including repeat events; the first-event endpoint uses only the first. A slightly stronger effect on the rate than on the first event could suggest some effect on recurrent events, but the difference between these two estimates is small and well within sampling variability.

Exacerbations leading to hospitalisation

For severe events requiring hospitalisation, the rate ratio was 0.89 (95% CI 0.76–1.03; P = 0.1265) and the first-event hazard ratio was 0.93 (95% CI 0.82–1.06; P = 0.2773). The point estimates favour the combination, but both intervals include 1.00. Hospitalised exacerbations are rarer than moderate ones, so these estimates rest on fewer events and are correspondingly less precise, as the wider intervals show.

All-cause mortality

The mortality hazard ratio of 1.09 means an estimated 9% higher hazard of death in the combination arm, but the 95% interval from 0.67 to 1.75 is very wide, spanning a substantial reduction and a substantial increase. With P = 0.7357, the data provide essentially no information to distinguish the arms on mortality over a treatment period of up to 361 days. A trial designed around exacerbation rates is not powered to detect mortality differences, and this estimate should not be read as either harm or safety with respect to death.

9. Safety: Serious Adverse Events

ArmPatients with serious adverse eventsPatients at risk
Tiotropium + olodaterol (5/5 µg)8103939
Tiotropium 5 µg8623941

Numerically fewer patients in the combination arm had at least one serious adverse event (810 of 3939) than in the tiotropium arm (862 of 3941), with near-identical denominators. The registry does not report a formal statistical comparison of serious adverse events. Serious adverse event counts in COPD populations include exacerbations requiring hospitalisation, so this safety summary partly overlaps with the efficacy endpoints and should not be read as an independent signal. Counts of patients with any event also do not reflect event severity, recurrence, or exposure time.

10. Statistical Methodology

Negative binomial regression for exacerbation rates

Exacerbations are count data: a patient can have zero, one or several during follow-up. A Poisson model assumes that the variance of the count equals its mean, but in COPD some patients are frequent exacerbators and others rarely exacerbate, so the observed variance is typically much larger than the mean (overdispersion). The negative binomial model accommodates this by allowing each patient's underlying rate to vary around the group mean.

Conceptual form
log E[Yi] = β0 + β1·Treatmenti + log(Exposurei)
Var(Yi) = μi + k·μi2   ·   Rate ratio = exp(β1)

Here Yi is the number of exacerbations for patient i, k is the overdispersion parameter, and log(Exposure) is the offset. The exponentiated treatment coefficient is the rate ratio, which is why the registry notes that least squares means "are actually exponentiated".

The exposure offset

Patients who discontinue early contribute less time at risk. Including the logarithm of treatment exposure as an offset converts the model from comparing raw counts to comparing rates per unit of exposure, so a patient treated for 120 days is not treated as equivalent to one treated for 361 days. The "actual treatment period" definition means events after treatment stops (beyond 1 day after last in-take) are not counted, which makes this an on-treatment estimand.

Cox proportional-hazards model and log-rank test

For the time-to-first-event secondary endpoints and mortality, hazard ratios and confidence intervals came from a Cox proportional hazards model, while p-values came from a log-rank test. The two methods are closely related: the log-rank test is the score test of an unadjusted Cox model with treatment as the only covariate. Both assume, for a single summary hazard ratio to be fully descriptive, that the ratio of hazards is roughly constant over follow-up.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous rate of a first event in the combination group

A hazard ratio concerns the first event only. It is not a rate ratio for all events, not a relative risk at a fixed time, and not an absolute difference.

Covariate adjustment

The three additional primary-endpoint models added baseline covariates such as smoking status, region, airflow limitation severity, and prior exacerbation history. Adjusting for strongly prognostic variables typically improves precision in randomized trials; prior exacerbation history is the dominant predictor of future exacerbations in COPD, which explains why these models produced narrower intervals.

Type I error control at 2-sided 0.01

The registry states that the hypothesis testing strategy ensures the overall type I error is protected at the 2-sided 0.01 level. This is a stricter threshold than the conventional 0.05. A 0.01 threshold is often chosen when a single large trial is intended to provide evidence of a strength comparable to two independent trials each at 0.05. The primary and first secondary analyses were accordingly reported with 99% confidence intervals.

11. Multiplicity and Endpoint Hierarchy

AnalysisRoleInterpretation
Annualised moderate to severe exacerbation rate (unadjusted)Primary; under the 0.01 testing strategyP = 0.0498, 99% CI crosses 1.00; prespecified superiority criterion not met
Covariate-adjusted rate modelsAdditional analyses of the primary endpointSupportive; 95% CIs; not a substitute for the prespecified test
Time to first moderate to severe exacerbationSecondary; under the 0.01 testing strategyP = 0.1188; not significant at 0.01
Hospitalisation rate and time to first hospitalisationSecondary95% CIs include 1.00; descriptive
All-cause mortalitySecondaryVery imprecise; descriptive

When several endpoints are tested, the probability of at least one false-positive result rises unless the testing strategy controls it. The registry describes a strategy that protects the overall type I error; in such strategies, endpoints lower in the order are typically interpreted as confirmatory only if those above them succeed. The registry does not describe the full ordering or the method (for example, hierarchical or alpha-splitting) in detail.

12. Statistical Methods Explained

Why was a negative binomial model used instead of a Poisson model or a simple proportion?

The endpoint counts every moderate to severe exacerbation, and patients differ widely in how often they exacerbate. A Poisson model would understate the variance and produce confidence intervals that are too narrow. Analysing only the proportion of patients with any exacerbation would discard repeat events. The negative binomial model keeps all events and allows extra-Poisson variability.

What does a rate ratio of 0.93 with a 99% CI of 0.85 to 1.02 mean?

The estimated exacerbation rate with the combination was 93% of the rate with tiotropium, a 7% lower estimated rate. The 99% interval indicates the data are compatible, at that confidence level, with anything from roughly a 15% reduction to a 2% increase. Because the interval includes 1.00, the analysis cannot exclude no difference at the level the trial chose.

Why is P = 0.0498 not considered statistically significant here?

Statistical significance depends on the threshold fixed before the data are seen. The registry states that overall type I error was protected at the 2-sided 0.01 level. P = 0.0498 exceeds 0.01, so the prespecified superiority criterion was not met. The fact that it falls just under 0.05 is not relevant to the trial's decision rule, and treating it as a success would change the rules after the fact.

Why do the covariate-adjusted models give smaller p-values?

Adding baseline factors that strongly predict exacerbations, especially prior exacerbation history, explains part of the between-patient variation and can shift a non-collapsible rate ratio away from 1. The adjusted estimates (0.89 to 0.91) and their p-values (0.0010 to 0.0080) are therefore more favourable. But they were not the analysis tied to the error-controlled test, and selecting the most favourable of several specifications would inflate the false-positive rate.

Why use a log-rank test for the p-value but a Cox model for the hazard ratio?

The log-rank test is a nonparametric comparison of event-time distributions that is well suited to producing a p-value. It does not by itself produce an effect size, so a Cox proportional hazards model was used to estimate the hazard ratio and its confidence interval. In an unadjusted two-group comparison, the two approaches are closely aligned.

How should the mortality hazard ratio of 1.09 be read?

The point estimate is above 1, but the 95% interval from 0.67 to 1.75 is so wide that it is compatible with a large reduction or a large increase in the hazard of death. The trial was sized for exacerbation rates, not mortality, over a treatment period of up to 361 days. The appropriate conclusion is that the estimate is uninformative, not that the combination increases or decreases mortality.

13. Important Limitations and Interpretation Issues

14. Why This Trial Matters Statistically

DYNAGITO is a valuable teaching case precisely because its primary result sits in the ambiguous zone between conventional and prespecified thresholds. It shows how the choice of alpha, the confidence level, and the model specification shape what a trial can claim.

ConceptHow it appears in DYNAGITO
Recurrent-event endpointAnnualised exacerbation rate counting all events per patient-year
Negative binomial regressionOverdispersed count model with log exposure offset
Rate ratio0.93 for moderate to severe exacerbations; 0.89 for hospitalised exacerbations
Prespecified alpha of 0.01P = 0.0498 below 0.05 but above 0.01; not significant under the trial's rule
99% confidence intervalInterval (0.85–1.02) aligned with the 0.01 decision threshold
Covariate adjustmentThree adjusted models with smaller rate ratios and narrower intervals
Time-to-event endpointsTime to first exacerbation, first hospitalisation and death
Hazard ratio and log-rank testCox model for estimates; log-rank for p-values
MultiplicityTesting strategy protecting overall type I error across endpoints
Analysis populationTreated set with a data-integrity site exclusion

Statistical interpretation

The primary rate ratio favoured the combination numerically but did not meet the prespecified 2-sided 0.01 superiority criterion; secondary endpoint estimates were in a similar direction for exacerbations but with intervals including 1.00.

Clinical interpretation

Any exacerbation benefit of adding olodaterol to tiotropium in this population appears modest in size. Whether such an effect is clinically meaningful is a separate judgement from its statistical significance.

15. Related Tutorials

Learn more about the methods used in this trial:

16. Related Calculators

17. Sources

Explore more clinical trial statistics

Connect this trial's endpoints and methods to in-depth statistical tutorials and practical calculators.

18. Record Summary

DYNAGITO compared tiotropium + olodaterol with tiotropium alone in 7903 enrolled patients with severe to very severe COPD. Its primary negative binomial analysis estimated a rate ratio of 0.93 (99% CI 0.85–1.02; P = 0.0498) for moderate to severe exacerbations, which did not meet the trial's prespecified 2-sided 0.01 superiority threshold. Covariate-adjusted models gave estimates of 0.89 to 0.91 with 95% intervals excluding 1.00, and secondary exacerbation endpoints pointed in the same direction with intervals including 1.00. The most useful reading combines the size of the rate ratio, the confidence level matched to the testing strategy, the distinction between prespecified and supportive analyses, and the on-treatment, treated-set framing of the estimates.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry entry. The goal is to reconstruct the statistical story of the trial in a standardized format while clearly separating reported evidence from educational interpretation.