← Clinical Trials
COPD Phase 3 Non-Inferiority & Superiority NCT01126437

TIOSPIR: Complete Statistical Analysis of Tiotropium Respimat vs HandiHaler in COPD

An independent statistical review of TIOSPIR, a randomized, double-blind, three-arm phase 3 trial comparing tiotropium delivered by the Respimat inhaler at two doses (2.5 mcg and 5 mcg daily) with tiotropium 18 mcg delivered by the HandiHaler in chronic obstructive pulmonary disease, with all-cause mortality and first COPD exacerbation as primary endpoints.

Sponsor: Boehringer Ingelheim  ·  Start: May 2010  ·  Primary completion: May 2013  ·  Status: Completed
About this page

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

TIOSPIR was a very large, double-blind, randomized, parallel-group phase 3 trial. Its purpose was not to show that tiotropium works (every arm received tiotropium) but to compare two delivery systems and doses: could the Respimat inhaler be shown to be no worse than the established HandiHaler for mortality, and was it better for exacerbations?

17,183
Enrolled
Three parallel arms
1.25
NI margin (HR)
Time to death, any cause
0.957
Mortality HR
R 5 vs HH 18; 95% CI 0.837–1.094
0.978
Exacerbation HR
R 5 vs HH 18; 95% CI 0.928–1.032
FeatureTIOSPIR
Registry titleComparison of Tiotropium in the HandiHaler Versus the Respimat in Chronic Obstructive Pulmonary Disease
PhasePhase 3
ConditionPulmonary Disease, Chronic Obstructive
DesignRandomized, parallel-group, double-masked, three arms
Primary purposeTreatment
Enrollment17,183
InterventionsTiotropium 18 mcg (HandiHaler); tiotropium 1.25 mcg, 2 actuations/day (Respimat 2.5 mcg); tiotropium 2.5 mcg, 2 actuations/day (Respimat 5 mcg)
Primary endpointsTime to all-cause mortality; time to first COPD exacerbation (each up to 3 years)
DatesStart May 2010; primary completion May 2013
SponsorBoehringer Ingelheim (industry)
ClinicalTrials.govNCT01126437

2. Clinical Question

The central question was whether tiotropium delivered by the Respimat soft-mist inhaler, at either 2.5 mcg or 5 mcg daily, was non-inferior to tiotropium 18 mcg via HandiHaler for time to death from any cause, and, if so, whether Respimat 5 mcg was superior to HandiHaler 18 mcg for time to first COPD exacerbation.

Population

Patients with chronic obstructive pulmonary disease randomized across three parallel groups.

Intervention

Tiotropium via Respimat: 2.5 mcg daily (1.25 mcg, 2 actuations) or 5 mcg daily (2.5 mcg, 2 actuations), each with placebo.

Comparator

Tiotropium 18 mcg via HandiHaler, with placebo. This is an active comparator, not an untreated control.

Primary question

Is Respimat no worse than HandiHaler on mortality (hazard ratio margin 1.25), and is Respimat 5 mcg better on time to first exacerbation?

Framing matters here. Because the comparator is itself an active treatment, a "no difference" result is the hoped-for outcome for the mortality question, and the statistical burden is to show that any excess risk with Respimat is small enough to be ruled out. That is a fundamentally different logic from a placebo-controlled superiority trial, and it shapes every interpretation on this page.

3. Trial Design

01
Enroll17,183 patients with COPD
02
RandomizeThree parallel, double-masked arms
03
TreatRespimat 2.5 or 5 mcg, or HandiHaler 18 mcg
04
FollowUp to 3 years, including vital status
05
TestHierarchical NI then superiority
RESPIMAT 2.5 MCG

Tiotropium 2.5 mcg and Placebo

  • Tiotropium 1.25 mcg per actuation
  • 2 actuations per day
  • Serious adverse events: 1932 / 5724
RESPIMAT 5 MCG

Tiotropium 5 mcg and Placebo

  • Tiotropium 2.5 mcg per actuation
  • 2 actuations per day
  • Serious adverse events: 1876 / 5705
HANDIHALER 18 MCG

Tiotropium 18 mcg and Placebo

  • Tiotropium 18 mcg
  • Active comparator
  • Serious adverse events: 1928 / 5687
Why every arm includes placebo. The registry labels each arm "and Placebo". When two physically different inhalers are compared under double masking, this labelling is typical of a double-dummy arrangement: each patient uses both device types, one containing active drug and the other placebo, so neither patients nor investigators can infer the assignment from the device. This protects symptom-driven endpoints such as exacerbations from assessment bias.

4. Analysis Populations

TIOSPIR used different analysis sets for different endpoints, and this is one of the most important details for interpreting the results. The mortality analysis retained patients after they stopped study drug; the exacerbation analyses did not.

Analysis setRegistry definitionUsed for
Death analysis set (DAS), including vital status follow-upAll randomized subjects excluding only subjects who were documented as not treatedTime to all-cause mortality; time to death from MACE
Treated set (TS), on-treatment onlyTreated patients, with follow-up restricted to the on-treatment periodExacerbation, hospitalization and MACE endpoints
Sub-study set, pulmonary function testing (SSS-PFT)All subjects in the TS who consented to the spirometry sub-study and had at least baseline and one on-treatment trough FEV1Trough FEV1 over 120 weeks (sub-study of 1370 patients)

For a mortality non-inferiority question, the DAS with vital status follow-up is a deliberately conservative choice: it keeps patients who discontinued treatment in the analysis, so deaths after discontinuation still count. The on-treatment exacerbation analyses answer a narrower question, namely what happens while patients are actually taking the assigned inhaler, and they rely on the assumption that stopping treatment is unrelated to the underlying exacerbation risk.

5. Primary Endpoints

EndpointRegistry descriptionTime frameHypothesis
Time to All-Cause MortalityNumber of patients with all-cause mortalityUp to 3 yearsNon-inferiority, HR margin 1.25
Time to First COPD ExacerbationSee definition belowUp to 3 yearsSuperiority (Respimat 5 mcg vs HandiHaler 18 mcg)

How the registry defines a COPD exacerbation

An exacerbation was defined as "a complex of lower respiratory events/symptoms (increase of new onset) related to the underlying COPD, with duration of three days or more, requiring a change in treatment."

The definition combines symptoms with a treatment decision. That makes it clinically meaningful but also dependent on prescribing behaviour, which is one reason double masking is so valuable here: if investigators knew the device assignment, their threshold for prescribing antibiotics or steroids could differ between arms.

6. Statistical Methodology

Hierarchical testing of the primary hypotheses

The registry describes three tests conducted in a fixed order:

  1. Non-inferiority of time to death from any cause, Respimat 5 mcg vs HandiHaler 18 mcg.
  2. If non-inferiority was achieved, non-inferiority of Respimat 2.5 mcg vs HandiHaler 18 mcg for time to death.
  3. If successful, superiority of Respimat 5 mcg over HandiHaler 18 mcg for time to first COPD exacerbation.

Non-inferiority tests were performed at a one-sided α = 0.025 level; the superiority test at a two-sided α = 0.05. A fixed-sequence procedure controls the familywise type I error without splitting alpha: each hypothesis is tested at the full level, but only if every hypothesis before it has been rejected. Once a test in the chain fails, all later tests become descriptive.

Mortality non-inferiority hypotheses (Respimat / HandiHaler)
H0: HR ≥ 1.25   versus   Ha: HR < 1.25

Rejecting H0 at one-sided α = 0.025 is equivalent to the upper bound of the two-sided 95% confidence interval lying below 1.25.

Cox proportional-hazards regression

All time-to-event endpoints (mortality, first exacerbation, first moderate-to-severe exacerbation, first exacerbation-related hospitalization, first MACE and death from MACE) were analysed with Cox regression, with the hazard ratio as the effect measure. Patients without an event are censored at the end of their follow-up, which for the on-treatment analyses means the end of the treatment period.

Model structure
h(t | treatment) = h0(t) · exp(β · treatment),   HR = exp(β)

The baseline hazard h0(t) is left unspecified; the model assumes only that the ratio of hazards between arms is constant over time.

Negative binomial regression for event counts

The number of COPD exacerbations and the number of exacerbation-related hospitalizations were analysed with negative binomial regression, reporting a rate ratio. Unlike time to first event, this uses every exacerbation a patient has, and the negative binomial distribution allows for overdispersion: some patients exacerbate repeatedly while many have few or none, which a Poisson model would understate.

Mixed model for repeated measures (MMRM) for FEV1

Trough FEV1 in the spirometry sub-study was analysed through Week 120 using REML-based repeated measures. The registry lists the model as including fixed categorical effects of treatment, investigative site, visit and treatment-by-visit interaction; continuous fixed covariates of baseline and baseline-by-visit interaction; and a random term of patient. The effect measure was an adjusted mean difference in litres, judged against a non-inferiority delta of 50 mL, again in hierarchical order (Respimat 5 mcg first, then Respimat 2.5 mcg).

7. Primary Result: Time to All-Cause Mortality

Both Respimat doses were compared with HandiHaler 18 mcg in the death analysis set including vital status follow-up, using Cox regression.

Respimat 5 mcg vs HandiHaler 18 mcg

HR 0.957

95% CI: 0.837–1.094  ·  Non-inferiority margin: 1.25

First test in the hierarchy

Respimat 2.5 mcg vs HandiHaler 18 mcg

HR 0.996

95% CI: 0.872–1.136  ·  Non-inferiority margin: 1.25

Second test in the hierarchy

Mortality hazard ratios with 95% CI (dashed line = 1.0; red line = NI margin 1.25)
R 5 vs HH 18
0.957 (0.837–1.094)
R 2.5 vs HH 18
0.996 (0.872–1.136)
0.81.01.21.25

In both comparisons the upper confidence bound (1.094 and 1.136) lies below the prespecified margin of 1.25, which is the condition for non-inferiority at one-sided α = 0.025. Because the first test succeeded, the second was permitted, and it also met the criterion. The hierarchy therefore passed to the exacerbation superiority test. The registry reports no p-values for these non-inferiority analyses; the decision rests on the confidence interval.

Clinical Biostats interpretation

What the estimate means. An HR of 0.957 for Respimat 5 mcg means the estimated instantaneous rate of death was about 4.3% lower than with HandiHaler 18 mcg over follow-up of up to 3 years; the HR of 0.996 for Respimat 2.5 mcg is essentially 1. Both point estimates sit very close to no difference.

What it does not mean. It does not show that Respimat reduces mortality. Both intervals include 1, so the data are compatible with slightly lower or slightly higher mortality than HandiHaler. The trial was designed to exclude a meaningful excess, not to demonstrate a benefit, and a hazard ratio is not a difference in the proportion of patients who died.

Precision. The intervals are narrow for a mortality endpoint, reflecting the very large sample. The upper bounds rule out, with 95% confidence, a hazard increase of more than about 9.4% (Respimat 5 mcg) or 13.6% (Respimat 2.5 mcg). Whether a margin of 1.25, i.e. tolerating up to a 25% relative increase, is clinically acceptable is a judgement about the design, not something the data answer.

Why the p-value is not the criterion. In non-inferiority testing, the conventional p-value for "HR = 1" is irrelevant; a large p-value there would not demonstrate non-inferiority. What matters is where the confidence interval sits relative to the margin. Even a p-value against the margin would describe the strength of evidence, not the size of any effect.

Cautions. The Cox HR summarizes the whole follow-up under a proportional-hazards assumption; if the relative risk changed over time, a single HR averages over that. The DAS with vital status follow-up is conservative for non-inferiority, because it counts deaths after treatment stops, but in a trial where every arm received active tiotropium, dilution toward "no difference" from discontinuation or switching is still a general concern for any non-inferiority comparison.

8. Primary Result: Time to First COPD Exacerbation

The confirmatory test was Respimat 5 mcg vs HandiHaler 18 mcg in the treated set (on-treatment only), under H0: HR = 1 against a two-sided alternative at α = 0.05. The registry also reports the two other pairwise comparisons.

Respimat 5 mcg vs HandiHaler 18 mcg (confirmatory)

HR 0.978

95% CI: 0.928–1.032  ·  P = 0.4194

ComparisonHazard ratio95% CIP-valueRole
Respimat 5 mcg vs HandiHaler 18 mcg0.9780.928–1.0320.4194Third test in the hierarchy
Respimat 2.5 mcg vs HandiHaler 18 mcg1.0160.964–1.0700.5593Supportive
Respimat 2.5 mcg vs Respimat 5 mcg1.0380.985–1.0940.1639Supportive
Time to first exacerbation, hazard ratios with 95% CI (dashed line = 1.0)
R 5 vs HH 18
0.978 (0.928–1.032)
R 2.5 vs HH 18
1.016 (0.964–1.070)
R 2.5 vs R 5
1.038 (0.985–1.094)
0.81.01.2

With P = 0.4194 and a confidence interval spanning 1, superiority of Respimat 5 mcg over HandiHaler 18 mcg for time to first exacerbation was not demonstrated. As this was the last test in the prespecified hierarchy, the failure does not affect the mortality conclusions, which were reached earlier in the sequence.

Clinical Biostats interpretation

What the estimate means. An HR of 0.978 corresponds to an estimated 2.2% lower hazard of a first exacerbation with Respimat 5 mcg, a very small relative difference.

What it does not mean. A non-significant superiority test is not proof of equivalence. The trial did not prespecify an equivalence or non-inferiority margin for exacerbations, so "no significant difference" should not be restated as "the devices are equivalent for exacerbations", even though the interval is narrow.

Precision. The 95% CI of 0.928–1.032 is tight: it excludes a relative reduction larger than about 7.2% and a relative increase larger than about 3.2%. In practical terms, the data make any large difference between the devices on this endpoint implausible, which is informative even without a formal equivalence claim.

Why the p-value does not measure effect size. P = 0.4194 says only that data like these would be unremarkable if the true HR were 1. With a sample this large, even a trivially small true difference could produce a small p-value; conversely, the p-value here gives no indication of how large or small any true difference is. The interval does that.

Cautions. This analysis used the treated set, on-treatment only, so patients were censored when they stopped treatment. If discontinuation was related to worsening disease, the censoring is informative and could bias the comparison in either direction. The two supportive comparisons were outside the confirmatory hierarchy, so their p-values are nominal.

9. Secondary Endpoint Results

The registry reports formal analyses for several secondary endpoints. None of these carry confirmatory weight; they describe consistency and help rule out large differences in other outcomes.

Trough FEV1 over 120 weeks (sub-study of 1370 patients)

ComparisonAdjusted mean difference (L)95% CI (L)Hierarchy position
Respimat 5 mcg vs HandiHaler 18 mcg-.010-.038 to 0.018First
Respimat 2.5 mcg vs HandiHaler 18 mcg-.037-.065 to -.009Second
Adjusted mean difference in trough FEV1, litres (dashed line = 0; red line = −0.050 L, the 50 mL margin)
R 5 vs HH 18
-.010 (-.038 to 0.018)
R 2.5 vs HH 18
-.037 (-.065 to -.009)
-0.08-0.050+0.04

The registry states the non-inferiority delta as 50 mL, with the 95% CI for each contrast compared against it. Since lower FEV1 with Respimat is the unfavourable direction, the relevant check is whether the lower confidence bound stays above −0.050 L. For Respimat 5 mcg the lower bound (−.038 L) does so, and the interval also includes 0. For Respimat 2.5 mcg the lower bound (−.065 L) extends beyond −0.050 L, so the interval does not exclude a deficit larger than the margin; the whole interval also lies below 0, indicating lower trough FEV1 than HandiHaler 18 mcg in this sub-study.

Reading the FEV1 result carefully: this was a sub-study of 1370 patients, restricted to those who consented to spirometry and had baseline and at least one on-treatment measurement. It is therefore smaller and more selected than the main trial, and the MMRM estimate relies on a missing-at-random assumption for patients whose later visits are absent. A lung-function difference in this subset does not by itself translate into a difference in exacerbations or mortality, which were assessed in the full population.

Exacerbation and hospitalization endpoints

EndpointComparisonMeasureEstimate (95% CI)P-value
Number of COPD exacerbationsR 2.5 vs HH 18Rate ratio1.01 (0.95–1.06)0.8330
R 5 vs HH 18Rate ratio0.99 (0.94–1.05)0.8047
R 2.5 vs R 5Rate ratio1.01 (0.96–1.07)0.6468
Time to first moderate to severe exacerbationR 2.5 vs HH 18HR1.011 (0.959–1.066)0.6823
R 5 vs HH 18HR0.983 (0.932–1.037)0.5377
R 2.5 vs R 5HR1.028 (0.975–1.084)0.3048
Time to first hospitalization associated with COPD exacerbationR 2.5 vs HH 18HR1.068 (0.971–1.176)0.1762
R 5 vs HH 18HR1.024 (0.929–1.128)0.6384
R 2.5 vs R 5HR1.044 (0.949–1.148)0.3784
Number of hospitalizations associated with COPD exacerbationR 2.5 vs HH 18Rate ratio1.09 (0.98–1.22)0.1255
R 5 vs HH 18Rate ratio1.06 (0.94–1.18)0.3441
R 2.5 vs R 5Rate ratio1.03 (0.92–1.16)0.5573

All of these analyses used the treated set, on-treatment only. Every estimate is close to 1 and every interval includes 1. Intervals for hospitalization endpoints are wider than for exacerbation endpoints, as expected for rarer events: the number of events, not the number of patients, drives precision in time-to-event and count models.

Major adverse cardiovascular events (MACE)

EndpointComparisonHR (95% CI)P-valuePopulation
Time to onset of first MACER 2.5 vs HH 181.105 (0.913–1.336)0.3043Treated set, on-treatment only
R 5 vs HH 181.100 (0.909–1.331)0.3263
R 2.5 vs R 51.004 (0.834–1.209)0.9644
Time to death from MACER 2.5 vs HH 181.171 (0.898–1.526)0.2439TS including vital status follow-up, DAS
R 5 vs HH 181.111 (0.850–1.453)0.4413
R 2.5 vs R 51.054 (0.814–1.363)0.6910

The registry notes that causes of death from MACE were determined by adjudication, and that time to death from MACE was analysed over the full study duration including vital status follow-up. The point estimates for both Respimat doses versus HandiHaler lie above 1, but the intervals are wide and include 1. These are the least precise results in the trial: an upper bound of 1.526 for death from MACE means the data cannot exclude a clinically important relative increase, just as a lower bound of 0.898 cannot exclude a modest decrease. Because cardiovascular safety was not framed as a non-inferiority hypothesis with a margin, the correct reading is "inconclusive in both directions", not "no difference".

10. Safety: Serious Adverse Events

ArmAffectedAt risk
Tiotropium 2.5 mcg and Placebo (Respimat)19325724
Tiotropium 5 mcg and Placebo (Respimat)18765705
Tiotropium 18 mcg and Placebo (HandiHaler)19285687

Serious adverse event counts were similar across the three arms, with roughly one third of patients in each arm affected over up to 3 years. The registry notes that prospectively defined outcome events, serious adverse events, adverse events leading to discontinuation, and investigator-determined drug-related adverse events were required for collection in this trial. Crude proportions of this kind do not account for differing exposure time between patients, so for a multi-year trial they are best read as a descriptive overview rather than as a formal comparison.

11. Statistical Methods Explained

Why was mortality tested for non-inferiority rather than superiority?

The comparator was an established active treatment, HandiHaler 18 mcg. The key question was whether switching to the Respimat device could carry excess mortality risk. A non-inferiority design addresses this directly: it asks whether the data can rule out an increase in hazard of 25% or more. A superiority test would only ask whether Respimat was better, and a non-significant superiority result could not reassure anyone about safety.

Why is non-inferiority judged against the margin rather than the p-value?

The null hypothesis in a non-inferiority test is "Respimat is worse by at least the margin" (HR ≥ 1.25), not "there is no difference". A two-sided 95% CI whose upper bound is below 1.25 is equivalent to rejecting that null at one-sided α = 0.025. The usual p-value against HR = 1 answers a different question and cannot establish non-inferiority, which is why the registry reports none for these analyses.

What does the fixed-sequence hierarchy protect against?

With three confirmatory questions, testing each at full alpha independently would inflate the chance of at least one false-positive claim. The fixed sequence (R 5 mortality NI, then R 2.5 mortality NI, then R 5 exacerbation superiority) spends the full alpha at each step but stops the chain at the first failure. Here the first two tests succeeded and the third did not, so both mortality non-inferiority claims are protected and no exacerbation superiority claim can be made.

Why analyse both time to first exacerbation and the number of exacerbations?

Time to first exacerbation, analysed with a Cox model, uses only each patient's first event and is robust but discards repeat events. The number of exacerbations, analysed with negative binomial regression, uses the full event burden and accounts for patients who exacerbate much more often than others. Agreement between the HR of 0.978 and the rate ratio of 0.99 for Respimat 5 mcg vs HandiHaler suggests no difference is hidden in either repeated events or first events.

Why was an MMRM used for trough FEV1?

FEV1 was measured repeatedly up to Week 120, and some patients inevitably miss later visits. An MMRM uses all available measurements from each patient, models the correlation between repeated measures on the same person, and adjusts for baseline FEV1 and its interaction with visit. It gives valid estimates when missingness is "at random" given the observed data, a weaker assumption than the complete-case analysis that simply drops patients with gaps.

Why does the analysis population differ between mortality and exacerbation endpoints?

Mortality was analysed in the death analysis set with vital status follow-up, retaining patients who stopped treatment. This is appropriate for a safety-driven non-inferiority question, where excluding post-discontinuation deaths could hide harm. Exacerbations were analysed on-treatment, because an exacerbation after a patient has stopped the inhaler says little about the device's effect. The trade-off is that on-treatment analyses depend on discontinuation being unrelated to exacerbation risk.

12. Limitations

13. Why This Trial Matters Statistically

TIOSPIR is a clear teaching example of how an active-controlled trial is built around a safety question. It shows a mortality non-inferiority margin applied to a hazard ratio, a fixed-sequence hierarchy that links non-inferiority and superiority tests, and the practical consequence of a failed final test in the chain.

ConceptHow it appears in TIOSPIR
Non-inferiority designMortality HR margin of 1.25, tested at one-sided α = 0.025
Confidence intervalsUpper bounds 1.094 and 1.136 compared with the margin
Hierarchical testingR 5 NI → R 2.5 NI → R 5 exacerbation superiority
Cox model / hazard ratioAll time-to-event endpoints, including mortality and first exacerbation
Negative binomial regressionNumbers of exacerbations and exacerbation-related hospitalizations
Rate ratioEffect measure for recurrent-event counts
MMRMTrough FEV1 through Week 120 in a spirometry sub-study
Analysis populationsDAS with vital status follow-up vs on-treatment treated set
Superiority vs equivalenceA non-significant exacerbation test that does not by itself imply equivalence

Statistical interpretation

Both Respimat doses met the prespecified mortality non-inferiority criterion versus HandiHaler 18 mcg. Superiority of Respimat 5 mcg for time to first exacerbation was not shown (P = 0.4194). Secondary estimates were close to 1, with the widest uncertainty for cardiovascular outcomes.

Clinical interpretation

Within the trial's margin, the two devices showed similar mortality and exacerbation outcomes. The lower trough FEV1 with Respimat 2.5 mcg in the sub-study and the imprecise MACE estimates are the points that remain least settled.

14. Related Tutorials

Learn more about the methods used in this trial:

15. Related Calculators

16. Sources

Keep exploring clinical trial statistics

Connect the methods used in TIOSPIR to step-by-step tutorials and hands-on calculators.

Summary: TIOSPIR combined a mortality non-inferiority test with a hazard-ratio margin of 1.25, a hierarchical path to an exacerbation superiority test, negative binomial models for recurrent events, and an MMRM for lung function. The most useful reading keeps the confirmatory results (mortality non-inferiority for both Respimat doses; no demonstrated exacerbation superiority) separate from the supportive secondary estimates, and reads each confidence interval for what it can and cannot exclude.