This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
TIOSPIR was a very large, double-blind, randomized, parallel-group phase 3 trial. Its purpose was not to show that tiotropium works (every arm received tiotropium) but to compare two delivery systems and doses: could the Respimat inhaler be shown to be no worse than the established HandiHaler for mortality, and was it better for exacerbations?
| Feature | TIOSPIR |
|---|---|
| Registry title | Comparison of Tiotropium in the HandiHaler Versus the Respimat in Chronic Obstructive Pulmonary Disease |
| Phase | Phase 3 |
| Condition | Pulmonary Disease, Chronic Obstructive |
| Design | Randomized, parallel-group, double-masked, three arms |
| Primary purpose | Treatment |
| Enrollment | 17,183 |
| Interventions | Tiotropium 18 mcg (HandiHaler); tiotropium 1.25 mcg, 2 actuations/day (Respimat 2.5 mcg); tiotropium 2.5 mcg, 2 actuations/day (Respimat 5 mcg) |
| Primary endpoints | Time to all-cause mortality; time to first COPD exacerbation (each up to 3 years) |
| Dates | Start May 2010; primary completion May 2013 |
| Sponsor | Boehringer Ingelheim (industry) |
| ClinicalTrials.gov | NCT01126437 |
2. Clinical Question
The central question was whether tiotropium delivered by the Respimat soft-mist inhaler, at either 2.5 mcg or 5 mcg daily, was non-inferior to tiotropium 18 mcg via HandiHaler for time to death from any cause, and, if so, whether Respimat 5 mcg was superior to HandiHaler 18 mcg for time to first COPD exacerbation.
Population
Patients with chronic obstructive pulmonary disease randomized across three parallel groups.
Intervention
Tiotropium via Respimat: 2.5 mcg daily (1.25 mcg, 2 actuations) or 5 mcg daily (2.5 mcg, 2 actuations), each with placebo.
Comparator
Tiotropium 18 mcg via HandiHaler, with placebo. This is an active comparator, not an untreated control.
Primary question
Is Respimat no worse than HandiHaler on mortality (hazard ratio margin 1.25), and is Respimat 5 mcg better on time to first exacerbation?
Framing matters here. Because the comparator is itself an active treatment, a "no difference" result is the hoped-for outcome for the mortality question, and the statistical burden is to show that any excess risk with Respimat is small enough to be ruled out. That is a fundamentally different logic from a placebo-controlled superiority trial, and it shapes every interpretation on this page.
3. Trial Design
Tiotropium 2.5 mcg and Placebo
- Tiotropium 1.25 mcg per actuation
- 2 actuations per day
- Serious adverse events: 1932 / 5724
Tiotropium 5 mcg and Placebo
- Tiotropium 2.5 mcg per actuation
- 2 actuations per day
- Serious adverse events: 1876 / 5705
Tiotropium 18 mcg and Placebo
- Tiotropium 18 mcg
- Active comparator
- Serious adverse events: 1928 / 5687
4. Analysis Populations
TIOSPIR used different analysis sets for different endpoints, and this is one of the most important details for interpreting the results. The mortality analysis retained patients after they stopped study drug; the exacerbation analyses did not.
| Analysis set | Registry definition | Used for |
|---|---|---|
| Death analysis set (DAS), including vital status follow-up | All randomized subjects excluding only subjects who were documented as not treated | Time to all-cause mortality; time to death from MACE |
| Treated set (TS), on-treatment only | Treated patients, with follow-up restricted to the on-treatment period | Exacerbation, hospitalization and MACE endpoints |
| Sub-study set, pulmonary function testing (SSS-PFT) | All subjects in the TS who consented to the spirometry sub-study and had at least baseline and one on-treatment trough FEV1 | Trough FEV1 over 120 weeks (sub-study of 1370 patients) |
For a mortality non-inferiority question, the DAS with vital status follow-up is a deliberately conservative choice: it keeps patients who discontinued treatment in the analysis, so deaths after discontinuation still count. The on-treatment exacerbation analyses answer a narrower question, namely what happens while patients are actually taking the assigned inhaler, and they rely on the assumption that stopping treatment is unrelated to the underlying exacerbation risk.
5. Primary Endpoints
| Endpoint | Registry description | Time frame | Hypothesis |
|---|---|---|---|
| Time to All-Cause Mortality | Number of patients with all-cause mortality | Up to 3 years | Non-inferiority, HR margin 1.25 |
| Time to First COPD Exacerbation | See definition below | Up to 3 years | Superiority (Respimat 5 mcg vs HandiHaler 18 mcg) |
How the registry defines a COPD exacerbation
An exacerbation was defined as "a complex of lower respiratory events/symptoms (increase of new onset) related to the underlying COPD, with duration of three days or more, requiring a change in treatment."
- Complex of symptoms: at least two of shortness of breath, sputum production (volume), occurrence of purulent sputum, cough, wheezing, and chest tightness.
- Required change in treatment: prescription of antibiotics and/or systemic steroids, and/or a newly prescribed maintenance respiratory medication (bronchodilators including theophyllines).
- Onset and end: onset was the first recorded symptom; the end was decided by the investigator based on clinical judgement.
- Severity: mild (a new prescription of maintenance bronchodilator only), moderate (antibiotics or systemic steroids without hospitalization), severe (hospitalization).
The definition combines symptoms with a treatment decision. That makes it clinically meaningful but also dependent on prescribing behaviour, which is one reason double masking is so valuable here: if investigators knew the device assignment, their threshold for prescribing antibiotics or steroids could differ between arms.
6. Statistical Methodology
Hierarchical testing of the primary hypotheses
The registry describes three tests conducted in a fixed order:
- Non-inferiority of time to death from any cause, Respimat 5 mcg vs HandiHaler 18 mcg.
- If non-inferiority was achieved, non-inferiority of Respimat 2.5 mcg vs HandiHaler 18 mcg for time to death.
- If successful, superiority of Respimat 5 mcg over HandiHaler 18 mcg for time to first COPD exacerbation.
Non-inferiority tests were performed at a one-sided α = 0.025 level; the superiority test at a two-sided α = 0.05. A fixed-sequence procedure controls the familywise type I error without splitting alpha: each hypothesis is tested at the full level, but only if every hypothesis before it has been rejected. Once a test in the chain fails, all later tests become descriptive.
Rejecting H0 at one-sided α = 0.025 is equivalent to the upper bound of the two-sided 95% confidence interval lying below 1.25.
Cox proportional-hazards regression
All time-to-event endpoints (mortality, first exacerbation, first moderate-to-severe exacerbation, first exacerbation-related hospitalization, first MACE and death from MACE) were analysed with Cox regression, with the hazard ratio as the effect measure. Patients without an event are censored at the end of their follow-up, which for the on-treatment analyses means the end of the treatment period.
The baseline hazard h0(t) is left unspecified; the model assumes only that the ratio of hazards between arms is constant over time.
Negative binomial regression for event counts
The number of COPD exacerbations and the number of exacerbation-related hospitalizations were analysed with negative binomial regression, reporting a rate ratio. Unlike time to first event, this uses every exacerbation a patient has, and the negative binomial distribution allows for overdispersion: some patients exacerbate repeatedly while many have few or none, which a Poisson model would understate.
Mixed model for repeated measures (MMRM) for FEV1
Trough FEV1 in the spirometry sub-study was analysed through Week 120 using REML-based repeated measures. The registry lists the model as including fixed categorical effects of treatment, investigative site, visit and treatment-by-visit interaction; continuous fixed covariates of baseline and baseline-by-visit interaction; and a random term of patient. The effect measure was an adjusted mean difference in litres, judged against a non-inferiority delta of 50 mL, again in hierarchical order (Respimat 5 mcg first, then Respimat 2.5 mcg).
7. Primary Result: Time to All-Cause Mortality
Both Respimat doses were compared with HandiHaler 18 mcg in the death analysis set including vital status follow-up, using Cox regression.
Respimat 5 mcg vs HandiHaler 18 mcg
95% CI: 0.837–1.094 · Non-inferiority margin: 1.25
First test in the hierarchy
Respimat 2.5 mcg vs HandiHaler 18 mcg
95% CI: 0.872–1.136 · Non-inferiority margin: 1.25
Second test in the hierarchy
In both comparisons the upper confidence bound (1.094 and 1.136) lies below the prespecified margin of 1.25, which is the condition for non-inferiority at one-sided α = 0.025. Because the first test succeeded, the second was permitted, and it also met the criterion. The hierarchy therefore passed to the exacerbation superiority test. The registry reports no p-values for these non-inferiority analyses; the decision rests on the confidence interval.
What the estimate means. An HR of 0.957 for Respimat 5 mcg means the estimated instantaneous rate of death was about 4.3% lower than with HandiHaler 18 mcg over follow-up of up to 3 years; the HR of 0.996 for Respimat 2.5 mcg is essentially 1. Both point estimates sit very close to no difference.
What it does not mean. It does not show that Respimat reduces mortality. Both intervals include 1, so the data are compatible with slightly lower or slightly higher mortality than HandiHaler. The trial was designed to exclude a meaningful excess, not to demonstrate a benefit, and a hazard ratio is not a difference in the proportion of patients who died.
Precision. The intervals are narrow for a mortality endpoint, reflecting the very large sample. The upper bounds rule out, with 95% confidence, a hazard increase of more than about 9.4% (Respimat 5 mcg) or 13.6% (Respimat 2.5 mcg). Whether a margin of 1.25, i.e. tolerating up to a 25% relative increase, is clinically acceptable is a judgement about the design, not something the data answer.
Why the p-value is not the criterion. In non-inferiority testing, the conventional p-value for "HR = 1" is irrelevant; a large p-value there would not demonstrate non-inferiority. What matters is where the confidence interval sits relative to the margin. Even a p-value against the margin would describe the strength of evidence, not the size of any effect.
Cautions. The Cox HR summarizes the whole follow-up under a proportional-hazards assumption; if the relative risk changed over time, a single HR averages over that. The DAS with vital status follow-up is conservative for non-inferiority, because it counts deaths after treatment stops, but in a trial where every arm received active tiotropium, dilution toward "no difference" from discontinuation or switching is still a general concern for any non-inferiority comparison.
8. Primary Result: Time to First COPD Exacerbation
The confirmatory test was Respimat 5 mcg vs HandiHaler 18 mcg in the treated set (on-treatment only), under H0: HR = 1 against a two-sided alternative at α = 0.05. The registry also reports the two other pairwise comparisons.
Respimat 5 mcg vs HandiHaler 18 mcg (confirmatory)
95% CI: 0.928–1.032 · P = 0.4194
| Comparison | Hazard ratio | 95% CI | P-value | Role |
|---|---|---|---|---|
| Respimat 5 mcg vs HandiHaler 18 mcg | 0.978 | 0.928–1.032 | 0.4194 | Third test in the hierarchy |
| Respimat 2.5 mcg vs HandiHaler 18 mcg | 1.016 | 0.964–1.070 | 0.5593 | Supportive |
| Respimat 2.5 mcg vs Respimat 5 mcg | 1.038 | 0.985–1.094 | 0.1639 | Supportive |
With P = 0.4194 and a confidence interval spanning 1, superiority of Respimat 5 mcg over HandiHaler 18 mcg for time to first exacerbation was not demonstrated. As this was the last test in the prespecified hierarchy, the failure does not affect the mortality conclusions, which were reached earlier in the sequence.
What the estimate means. An HR of 0.978 corresponds to an estimated 2.2% lower hazard of a first exacerbation with Respimat 5 mcg, a very small relative difference.
What it does not mean. A non-significant superiority test is not proof of equivalence. The trial did not prespecify an equivalence or non-inferiority margin for exacerbations, so "no significant difference" should not be restated as "the devices are equivalent for exacerbations", even though the interval is narrow.
Precision. The 95% CI of 0.928–1.032 is tight: it excludes a relative reduction larger than about 7.2% and a relative increase larger than about 3.2%. In practical terms, the data make any large difference between the devices on this endpoint implausible, which is informative even without a formal equivalence claim.
Why the p-value does not measure effect size. P = 0.4194 says only that data like these would be unremarkable if the true HR were 1. With a sample this large, even a trivially small true difference could produce a small p-value; conversely, the p-value here gives no indication of how large or small any true difference is. The interval does that.
Cautions. This analysis used the treated set, on-treatment only, so patients were censored when they stopped treatment. If discontinuation was related to worsening disease, the censoring is informative and could bias the comparison in either direction. The two supportive comparisons were outside the confirmatory hierarchy, so their p-values are nominal.
9. Secondary Endpoint Results
The registry reports formal analyses for several secondary endpoints. None of these carry confirmatory weight; they describe consistency and help rule out large differences in other outcomes.
Trough FEV1 over 120 weeks (sub-study of 1370 patients)
| Comparison | Adjusted mean difference (L) | 95% CI (L) | Hierarchy position |
|---|---|---|---|
| Respimat 5 mcg vs HandiHaler 18 mcg | -.010 | -.038 to 0.018 | First |
| Respimat 2.5 mcg vs HandiHaler 18 mcg | -.037 | -.065 to -.009 | Second |
The registry states the non-inferiority delta as 50 mL, with the 95% CI for each contrast compared against it. Since lower FEV1 with Respimat is the unfavourable direction, the relevant check is whether the lower confidence bound stays above −0.050 L. For Respimat 5 mcg the lower bound (−.038 L) does so, and the interval also includes 0. For Respimat 2.5 mcg the lower bound (−.065 L) extends beyond −0.050 L, so the interval does not exclude a deficit larger than the margin; the whole interval also lies below 0, indicating lower trough FEV1 than HandiHaler 18 mcg in this sub-study.
Exacerbation and hospitalization endpoints
| Endpoint | Comparison | Measure | Estimate (95% CI) | P-value |
|---|---|---|---|---|
| Number of COPD exacerbations | R 2.5 vs HH 18 | Rate ratio | 1.01 (0.95–1.06) | 0.8330 |
| R 5 vs HH 18 | Rate ratio | 0.99 (0.94–1.05) | 0.8047 | |
| R 2.5 vs R 5 | Rate ratio | 1.01 (0.96–1.07) | 0.6468 | |
| Time to first moderate to severe exacerbation | R 2.5 vs HH 18 | HR | 1.011 (0.959–1.066) | 0.6823 |
| R 5 vs HH 18 | HR | 0.983 (0.932–1.037) | 0.5377 | |
| R 2.5 vs R 5 | HR | 1.028 (0.975–1.084) | 0.3048 | |
| Time to first hospitalization associated with COPD exacerbation | R 2.5 vs HH 18 | HR | 1.068 (0.971–1.176) | 0.1762 |
| R 5 vs HH 18 | HR | 1.024 (0.929–1.128) | 0.6384 | |
| R 2.5 vs R 5 | HR | 1.044 (0.949–1.148) | 0.3784 | |
| Number of hospitalizations associated with COPD exacerbation | R 2.5 vs HH 18 | Rate ratio | 1.09 (0.98–1.22) | 0.1255 |
| R 5 vs HH 18 | Rate ratio | 1.06 (0.94–1.18) | 0.3441 | |
| R 2.5 vs R 5 | Rate ratio | 1.03 (0.92–1.16) | 0.5573 |
All of these analyses used the treated set, on-treatment only. Every estimate is close to 1 and every interval includes 1. Intervals for hospitalization endpoints are wider than for exacerbation endpoints, as expected for rarer events: the number of events, not the number of patients, drives precision in time-to-event and count models.
Major adverse cardiovascular events (MACE)
| Endpoint | Comparison | HR (95% CI) | P-value | Population |
|---|---|---|---|---|
| Time to onset of first MACE | R 2.5 vs HH 18 | 1.105 (0.913–1.336) | 0.3043 | Treated set, on-treatment only |
| R 5 vs HH 18 | 1.100 (0.909–1.331) | 0.3263 | ||
| R 2.5 vs R 5 | 1.004 (0.834–1.209) | 0.9644 | ||
| Time to death from MACE | R 2.5 vs HH 18 | 1.171 (0.898–1.526) | 0.2439 | TS including vital status follow-up, DAS |
| R 5 vs HH 18 | 1.111 (0.850–1.453) | 0.4413 | ||
| R 2.5 vs R 5 | 1.054 (0.814–1.363) | 0.6910 |
The registry notes that causes of death from MACE were determined by adjudication, and that time to death from MACE was analysed over the full study duration including vital status follow-up. The point estimates for both Respimat doses versus HandiHaler lie above 1, but the intervals are wide and include 1. These are the least precise results in the trial: an upper bound of 1.526 for death from MACE means the data cannot exclude a clinically important relative increase, just as a lower bound of 0.898 cannot exclude a modest decrease. Because cardiovascular safety was not framed as a non-inferiority hypothesis with a margin, the correct reading is "inconclusive in both directions", not "no difference".
10. Safety: Serious Adverse Events
| Arm | Affected | At risk |
|---|---|---|
| Tiotropium 2.5 mcg and Placebo (Respimat) | 1932 | 5724 |
| Tiotropium 5 mcg and Placebo (Respimat) | 1876 | 5705 |
| Tiotropium 18 mcg and Placebo (HandiHaler) | 1928 | 5687 |
Serious adverse event counts were similar across the three arms, with roughly one third of patients in each arm affected over up to 3 years. The registry notes that prospectively defined outcome events, serious adverse events, adverse events leading to discontinuation, and investigator-determined drug-related adverse events were required for collection in this trial. Crude proportions of this kind do not account for differing exposure time between patients, so for a multi-year trial they are best read as a descriptive overview rather than as a formal comparison.
11. Statistical Methods Explained
Why was mortality tested for non-inferiority rather than superiority?
The comparator was an established active treatment, HandiHaler 18 mcg. The key question was whether switching to the Respimat device could carry excess mortality risk. A non-inferiority design addresses this directly: it asks whether the data can rule out an increase in hazard of 25% or more. A superiority test would only ask whether Respimat was better, and a non-significant superiority result could not reassure anyone about safety.
Why is non-inferiority judged against the margin rather than the p-value?
The null hypothesis in a non-inferiority test is "Respimat is worse by at least the margin" (HR ≥ 1.25), not "there is no difference". A two-sided 95% CI whose upper bound is below 1.25 is equivalent to rejecting that null at one-sided α = 0.025. The usual p-value against HR = 1 answers a different question and cannot establish non-inferiority, which is why the registry reports none for these analyses.
What does the fixed-sequence hierarchy protect against?
With three confirmatory questions, testing each at full alpha independently would inflate the chance of at least one false-positive claim. The fixed sequence (R 5 mortality NI, then R 2.5 mortality NI, then R 5 exacerbation superiority) spends the full alpha at each step but stops the chain at the first failure. Here the first two tests succeeded and the third did not, so both mortality non-inferiority claims are protected and no exacerbation superiority claim can be made.
Why analyse both time to first exacerbation and the number of exacerbations?
Time to first exacerbation, analysed with a Cox model, uses only each patient's first event and is robust but discards repeat events. The number of exacerbations, analysed with negative binomial regression, uses the full event burden and accounts for patients who exacerbate much more often than others. Agreement between the HR of 0.978 and the rate ratio of 0.99 for Respimat 5 mcg vs HandiHaler suggests no difference is hidden in either repeated events or first events.
Why was an MMRM used for trough FEV1?
FEV1 was measured repeatedly up to Week 120, and some patients inevitably miss later visits. An MMRM uses all available measurements from each patient, models the correlation between repeated measures on the same person, and adjusts for baseline FEV1 and its interaction with visit. It gives valid estimates when missingness is "at random" given the observed data, a weaker assumption than the complete-case analysis that simply drops patients with gaps.
Why does the analysis population differ between mortality and exacerbation endpoints?
Mortality was analysed in the death analysis set with vital status follow-up, retaining patients who stopped treatment. This is appropriate for a safety-driven non-inferiority question, where excluding post-discontinuation deaths could hide harm. Exacerbations were analysed on-treatment, because an exacerbation after a patient has stopped the inhaler says little about the device's effect. The trade-off is that on-treatment analyses depend on discontinuation being unrelated to exacerbation risk.
12. Limitations
- Choice of margin: non-inferiority on mortality is only as meaningful as the margin. An HR of 1.25 allows a sizeable relative increase to count as "non-inferior", even though the observed upper bounds were well inside it.
- Absence of a placebo arm: with only active arms, the trial cannot say how either device performs relative to no bronchodilator. Non-inferiority assumes HandiHaler's effect is stable and present in this population (assay sensitivity).
- On-treatment censoring: exacerbation, hospitalization and first-MACE analyses censor patients when treatment stops. If stopping is linked to disease status, estimates can be biased.
- Proportional hazards: each Cox HR summarizes up to 3 years in one number. Time-varying effects, if present, are averaged.
- Multiplicity outside the hierarchy: the secondary endpoints and the non-hierarchical pairwise comparisons each carry nominal p-values. With many comparisons, occasional small p-values would be expected by chance.
- Imprecision for cardiovascular outcomes: MACE and death from MACE intervals are wide, so the trial is not definitive on cardiovascular safety differences between devices.
- Sub-study selection: FEV1 results come from 1370 consenting patients with sufficient spirometry and may not represent the whole enrolled population.
- Crude safety counts: serious adverse event counts are not exposure-adjusted and are descriptive.
13. Why This Trial Matters Statistically
TIOSPIR is a clear teaching example of how an active-controlled trial is built around a safety question. It shows a mortality non-inferiority margin applied to a hazard ratio, a fixed-sequence hierarchy that links non-inferiority and superiority tests, and the practical consequence of a failed final test in the chain.
| Concept | How it appears in TIOSPIR |
|---|---|
| Non-inferiority design | Mortality HR margin of 1.25, tested at one-sided α = 0.025 |
| Confidence intervals | Upper bounds 1.094 and 1.136 compared with the margin |
| Hierarchical testing | R 5 NI → R 2.5 NI → R 5 exacerbation superiority |
| Cox model / hazard ratio | All time-to-event endpoints, including mortality and first exacerbation |
| Negative binomial regression | Numbers of exacerbations and exacerbation-related hospitalizations |
| Rate ratio | Effect measure for recurrent-event counts |
| MMRM | Trough FEV1 through Week 120 in a spirometry sub-study |
| Analysis populations | DAS with vital status follow-up vs on-treatment treated set |
| Superiority vs equivalence | A non-significant exacerbation test that does not by itself imply equivalence |
Statistical interpretation
Both Respimat doses met the prespecified mortality non-inferiority criterion versus HandiHaler 18 mcg. Superiority of Respimat 5 mcg for time to first exacerbation was not shown (P = 0.4194). Secondary estimates were close to 1, with the widest uncertainty for cardiovascular outcomes.
Clinical interpretation
Within the trial's margin, the two devices showed similar mortality and exacerbation outcomes. The lower trough FEV1 with Respimat 2.5 mcg in the sub-study and the imprecise MACE estimates are the points that remain least settled.
14. Related Tutorials
Learn more about the methods used in this trial:
15. Related Calculators
16. Sources
- ClinicalTrials.gov: NCT01126437, Comparison of Tiotropium in the HandiHaler Versus the Respimat in Chronic Obstructive Pulmonary Disease.
- PubMed: PMID 23992515.
- PubMed: PMID 26369563.
- PubMed: PMID 26715479.
- PubMed: PMID 28957643.
- PubMed: PMID 29497289.
Keep exploring clinical trial statistics
Connect the methods used in TIOSPIR to step-by-step tutorials and hands-on calculators.