This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
BREEZE was a completed phase 1, open-label, single-group clinical study evaluating Treprostinil Inhalation Powder (TreT) in subjects with pulmonary arterial hypertension who were currently using Tyvaso. The study enrolled 51 subjects and used no randomized comparator arm.
| Feature | BREEZE |
|---|---|
| Phase | Phase 1 |
| Condition | Pulmonary Arterial Hypertension |
| Study title | Open-label, Clinical Study to Evaluate the Safety and Tolerability of TreT in Subjects With PAH Currently Using Tyvaso |
| Design | Single-group, open-label |
| Allocation | NA |
| Primary purpose | Treatment |
| Enrollment | 51 |
| Number of arms | 1 |
| Intervention | Treprostinil Inhalation Powder |
| Status | Completed |
| Start date | 2019-09-17 |
| Primary completion date | 2023-08-22 |
| Lead sponsor | United Therapeutics |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03950739 |
2. Clinical Question
The clinical question in BREEZE is fundamentally different from that of a randomized superiority trial. The study evaluated Treprostinil Inhalation Powder in subjects with pulmonary arterial hypertension who were already using Tyvaso. Its primary purpose was treatment, and the design was single-group and open-label.
Population
Subjects with pulmonary arterial hypertension currently using Tyvaso.
Intervention
Treprostinil Inhalation Powder, identified in the registry as TreT.
Comparator
No separate randomized comparator arm was specified. The study has a single-group design.
Primary questions
How did 6MWD and patient-reported PAH symptoms and impact change after treatment, and how did subjects report satisfaction with and preference for the inhaled treprostinil device?
This design determines what can be learned statistically. A within-subject change can describe how measurements changed after treatment exposure, but it does not provide the same causal contrast as a randomized treatment-versus-control comparison. Without a concurrent randomized comparator, a change from baseline cannot by itself establish how much of the observed change was caused by TreT rather than by time, repeated measurement, background treatment, regression to the mean, or other factors.
3. Trial Design
The study's statistical structure is therefore primarily longitudinal and within-subject. Baseline measurements provide the reference point, and follow-up measurements after exposure to TreT describe change over time. This is a natural framework for endpoints such as 6-Minute Walk Distance and patient-reported symptom or impact scores, because each subject can serve as their own baseline reference.
4. Treatment and Dose Groups
The safety information is reported in treatment-phase and optional-extension groupings corresponding to the transition from Tyvaso to TreT. The registry identifies three treatment-phase dose groups and three optional-extension groupings.
Tyvaso to TreT 32 mcg
- Serious adverse events: 0/2 affected / at risk.
Tyvaso to TreT 48 mcg
- Serious adverse events: 1/27 affected / at risk.
Tyvaso to TreT 64 mcg
- Serious adverse events: 1/22 affected / at risk.
Tyvaso to TreT 32 mcg
- Serious adverse events: 1/2 affected / at risk.
| Phase / grouping | TreT dose | Serious adverse events |
|---|---|---|
| Treatment Phase | 32 mcg | 0/2 |
| Treatment Phase | 48 mcg | 1/27 |
| Treatment Phase | 64 mcg | 1/22 |
| Optional Extensio | 32 mcg | 1/2 |
| Optional Extensio | 48 mcg | 12/26 |
| Optional Extensio | 64 mcg | 12/21 |
The denominator is important when interpreting these safety summaries. The treatment-phase groups and optional-extension groups should not automatically be treated as independent randomized arms. The registry describes the overall study as single-group, while the safety display separates subjects according to dose and study phase. Consequently, the affected/at-risk counts are best read as phase- and dose-specific safety summaries rather than as evidence from a randomized dose-comparison experiment.
5. Endpoints
| Registered primary endpoint | Time frame | Registry definition |
|---|---|---|
| Change in 6-Minute Walk Distance (6MWD) From Baseline to Week 3 | From Baseline to 3 weeks of treatment with TreT | 6MWD was evaluated at study entry and after 3 weeks of treatment with TreT. |
| Subject Satisfaction With and Preference for Inhaled Treprostinil Devices | After 3 weeks of treatment with TreT, after switching from the Tyvaso Inhalation System | Subject satisfaction with and preference for the inhaled treprostinil device was evaluated with the Preference Questionnaire for Inhaled Treprostinil Devices (PQ-ITD). The PQ-ITD is a questionnaire given to evaluate subject satisfaction with and preference for inhaled treprostinil devices. The questionnaire provides 12 different statements around inhaled device satisfaction and allows for 5 response options: strongly disagree, disagree, neutral, agree, and strongly agree. |
| Change in Patient-reported PAH Symptoms and Impact From Baseline to Week 3 | From Baseline to 3 weeks of treatment with TreT | Patient-reported PAH symptoms and impact were evaluated with the PAH Symptoms and Impact (PAH-SYMPACT) Questionnaire. The PAH-SYMPACT is a 23-item patient-reported outcome questionnaire that consists of 11 symptom items, 11 impact items, and 1 item on oxygen use. The symptom items are divided into cardiopulmonary and cardiovascular domains, and the impact items are divided into physical and emotional/cognitive domains. Symptom and impact domain scores (range 0 to 4) are calculated as the sum of the scores for the items included in the domain divided by the number of items in the domain. For all domains, a higher score indicates more severe symptoms/impacts. |
| Change in Patient-reported PAH Symptoms and Impact From Baseline to Week 11 (for Subjects Participating in the OEP) | From Baseline to 11 weeks of treatment with TreT | Patient-reported PAH symptoms and impact were evaluated with the PAH Symptoms and Impact (PAH-SYMPACT) Questionnaire. The PAH-SYMPACT is a 23-item patient-reported outcome questionnaire that consists of 11 symptom items, 11 impact items, and 1 item on oxygen use. The symptom items are divided into cardiopulmonary and cardiovascular domains, and the impact items are divided into physical and emotional/cognitive domains. Symptom and impact domain scores (range 0 to 4) are calculated as the sum of the scores for the items included in the domain divided by the number of items in the domain. For all domains, a higher score indicates more severe symptoms/impacts. |
These endpoints represent three related but statistically distinct types of information. 6MWD is an objective functional measure expressed on a continuous scale. PAH-SYMPACT is a patient-reported outcome with domain scores ranging from 0 to 4, where higher values indicate more severe symptoms or impacts. The PQ-ITD is a structured preference and satisfaction questionnaire with categorical response options.
That distinction matters because the most appropriate statistical summary depends on the measurement scale. A continuous change score can be summarized using a mean and standard deviation or a median and interquartile range, while ordinal questionnaire responses are more naturally summarized using response distributions, proportions, or other methods that respect their categorical structure.
6. Statistical Methodology
The registry identifies the endpoints and their assessment time points, but no formal statistical analyses are posted in the ClinicalTrials.gov record. The appropriate statistical framework can therefore be described based on the structure of each endpoint without attributing an unreported analysis to the study.
Change from baseline for 6MWD
The 6MWD endpoint is defined as the change from baseline to Week 3. For each subject, the basic change score is:
A positive change indicates a higher Week 3 distance than the subject's baseline measurement; a negative change indicates a lower Week 3 distance.
For a single-group study, a natural descriptive analysis is the distribution of these subject-level changes. The mean change summarizes the average numerical change, while the median change describes the middle of the distribution and is less sensitive to unusually large or small observations.
If a formal inferential analysis were prespecified for a continuous within-subject change endpoint, a one-sample method could compare the mean change with zero, provided its assumptions were appropriate. A nonparametric alternative could be considered when the distribution of changes is strongly non-normal. The choice between these approaches should be determined from the prespecified statistical analysis plan rather than selected after viewing the results.
PAH-SYMPACT domain scores
PAH-SYMPACT domain scores are constructed from item responses and range from 0 to 4. The registry explicitly states that higher scores indicate more severe symptoms or impacts. Thus, unlike 6MWD, a decrease in a domain score represents movement toward fewer or less severe reported symptoms or impacts.
The baseline-to-Week-3 and baseline-to-Week-11 endpoints can therefore be represented as subject-level changes in each applicable domain. The primary interpretive task is to preserve the direction of the scale: an apparently negative numerical change is not inherently unfavorable when the measured construct is symptom severity.
Because the endpoint is patient-reported and contains multiple domains, analysis should also distinguish the statistical unit from the questionnaire item. A subject contributes observations across domains, but the domains should not automatically be treated as independent observations. Correlation among repeated patient-reported measurements is an important consideration when formal modeling is undertaken.
PQ-ITD satisfaction and preference
The PQ-ITD contains 12 statements concerning inhaled device satisfaction and provides five response options: strongly disagree, disagree, neutral, agree, and strongly agree. These are ordered categorical responses rather than measurements on a conventional continuous scale.
A straightforward statistical presentation would show the distribution of responses for each statement. For example, the proportion selecting each of the five response categories can provide more information than collapsing all responses into a single average score. If a composite score were defined in a prespecified analysis plan, its construction and interpretation would need to respect the questionnaire's measurement properties.
Safety summaries
Safety is naturally summarized using counts of subjects experiencing adverse events, with the number affected presented alongside the number at risk. For serious adverse events, the registry provides affected/at-risk counts separately for the treatment phase and optional extension and by TreT dose grouping.
Because the study is single-group and the reported safety groupings arise from treatment phase and dose rather than randomization, these counts describe observed safety experience. They do not establish a comparative risk ratio against a concurrent control group.
7. Planned Analysis
6MWD: change from baseline to Week 3
The registry defines this endpoint as the change in 6MWD from baseline to Week 3 of treatment with TreT. For an endpoint of this type in a single-group study, the principal analysis would ordinarily begin with subject-level changes and descriptive statistics for those changes. A formal one-sample analysis could test whether the mean or another prespecified measure of change differs from zero, depending on the analysis plan and distributional assumptions.
The important distinction is between describing a change and estimating a comparative treatment effect. In this study design, a baseline-to-Week-3 change is a within-subject quantity. It does not have the same interpretation as a difference in mean change between TreT and a randomized control group.
PQ-ITD: satisfaction and preference
The PQ-ITD endpoint is based on 12 statements and five ordered response categories. An appropriate analysis would ordinarily preserve the five-category response structure and summarize the distribution of responses for each statement. If a prespecified composite or summary measure were used, its definition would determine the appropriate inferential method.
The questionnaire is also evaluating satisfaction and preference rather than a direct physiologic measurement. Statistical interpretation should therefore focus on the pattern of reported responses rather than treating the response categories as though they were equally spaced numerical measurements.
PAH-SYMPACT: baseline to Week 3
The Week-3 PAH-SYMPACT endpoint is a change score derived from the questionnaire's symptom and impact domains. For this type of endpoint, the analysis would ordinarily summarize the baseline score, Week-3 score, and subject-level change for each domain. Because higher scores represent more severe symptoms or impacts, the direction of change must be interpreted accordingly.
PAH-SYMPACT: baseline to Week 11
The Week-11 endpoint applies to subjects participating in the OEP. This is an important analysis-population distinction. The Week-11 endpoint is not defined for every enrolled subject simply because 51 subjects were enrolled; it is specifically restricted to subjects participating in the OEP.
For a longitudinal endpoint assessed at both baseline and Week 11, a natural descriptive analysis would report the distribution of subject-level changes. If a formal repeated-measures model were prespecified, it could account for the correlation among measurements within the same subject, but the registry does not identify such a model.
8. How to Interpret the Single-Group Design
The most important statistical feature of BREEZE is not a particular test statistic. It is the absence of a randomized concurrent comparator. That feature changes the meaning of almost every efficacy-related result.
What the study can describe
It can describe how subjects' measurements changed after switching from Tyvaso to TreT, as well as reported device satisfaction, preference, patient-reported symptoms and impacts, and observed safety events.
What it cannot isolate
A within-subject change does not isolate the causal effect of TreT from temporal changes, repeated measurement effects, background treatment, or other influences occurring during follow-up.
Why baseline is still useful
Each subject's baseline measurement provides a natural reference for quantifying the direction and magnitude of subsequent change.
Why a comparator matters
A randomized comparator would provide a separate contemporaneous trajectory against which the TreT trajectory could be contrasted.
This does not make a single-group analysis statistically uninformative. It simply means the estimand is different. The central quantity is generally a within-subject change rather than a between-group treatment effect. That distinction should remain visible whenever the results are interpreted.
For example, suppose a subject's 6MWD increased between baseline and Week 3. The increase is real as a descriptive change if measured correctly. But the statistical question "Did the subject improve?" is not equivalent to the causal question "Did TreT cause the improvement?" The second question requires a counterfactual: what would the same subject's Week-3 6MWD have been under an alternative treatment or under no change in treatment? A single-group design does not observe that counterfactual directly.
9. Statistical Methods Explained
Why analyze change from baseline?
A change score puts each subject's follow-up measurement on a common individual reference point. This can be especially useful when baseline measurements differ across subjects. Rather than comparing raw Week-3 values alone, the analysis asks how much each subject moved relative to their own starting measurement.
Why does the single-group design matter?
With one treatment group, there is no randomized between-group contrast. A mean change, median change, or other summary describes the observed trajectory of the treated subjects. It does not directly estimate the difference that would have been observed between TreT and a concurrent randomized comparator.
How should a PAH-SYMPACT score be interpreted?
The registry defines domain scores from 0 to 4 and states that higher scores indicate more severe symptoms or impacts. Therefore, a lower follow-up score indicates less severe reported symptoms or impacts relative to a higher score. The sign of a change must always be interpreted in the context of what the scale measures.
Why should PQ-ITD responses not automatically be averaged?
The five PQ-ITD response options are ordered categories: strongly disagree, disagree, neutral, agree, and strongly agree. The ordering contains information, but the numerical distance between adjacent categories is not necessarily equivalent. Reporting the response distribution preserves information that can be lost when ordinal responses are converted into a simple average.
Why are Week-3 and Week-11 PAH-SYMPACT analyses different?
The Week-3 endpoint is defined for the treatment period, whereas the Week-11 endpoint is specifically for subjects participating in the OEP. The Week-11 population therefore has a different scope. Analyses at the two time points should not automatically be interpreted as though they were based on exactly the same set of subjects.
What does 0/2 for serious adverse events mean?
It means that 0 subjects were affected among 2 subjects at risk in that particular treatment-phase dose grouping. It is an observed count, not evidence that the true probability of a serious adverse event is zero. With only 2 subjects at risk, the estimate is intrinsically very uncertain.
Why should the optional-extension safety counts be separated from the treatment-phase counts?
The registry reports the treatment-phase and optional-extension groupings separately. The study is single-group, and the extension groupings reflect a later phase of exposure rather than an independent randomized cohort. Combining the counts without regard to study phase could obscure differences in exposure period and population composition.
10. Safety Results
The ClinicalTrials.gov record reports serious adverse events by dose grouping and study phase. These are the principal safety counts available for the registered record.
| Study phase | TreT dose | Affected | At risk |
|---|---|---|---|
| Treatment Phase | 32 mcg | 0 | 2 |
| Treatment Phase | 48 mcg | 1 | 27 |
| Treatment Phase | 64 mcg | 1 | 22 |
| Optional Extensio | 32 mcg | 1 | 2 |
| Optional Extensio | 48 mcg | 12 | 26 |
| Optional Extensio | 64 mcg | 12 | 21 |
The affected/at-risk format is useful because the numerator alone does not provide the context needed to interpret an event count. One serious adverse event among 27 subjects has a different statistical meaning from one event among 2 subjects. Small denominators produce substantial uncertainty even when the observed count is simple to report.
The optional-extension data also illustrate why exposure time and study phase matter in safety analysis. The registry separates the extension from the treatment phase rather than presenting one undifferentiated safety count. That distinction should be retained when describing the observed events.
The serious-adverse-event counts describe observed events in the listed dose and phase groupings. They do not establish comparative safety because BREEZE has no randomized comparator arm. They also should not be converted into conclusions about dose-related risk without an analysis that accounts for the underlying exposure structure and uncertainty.
11. Understanding Uncertainty in Small Safety Groups
Small denominators are particularly important in the BREEZE safety data. The treatment-phase 32 mcg grouping has 2 subjects at risk, while the treatment-phase 48 mcg and 64 mcg groupings have 27 and 22 subjects at risk, respectively. The optional-extension groupings have 2, 26, and 21 subjects at risk.
A zero count in a small group does not demonstrate absence of risk. If no serious adverse event is observed among a small number of subjects, the data are compatible with a wide range of underlying event probabilities. Conversely, a relatively large observed fraction in a small denominator can be highly unstable as an estimate of a broader population rate.
This is a general lesson in clinical-trial statistics: observed frequency is not the same thing as estimated underlying risk. Confidence intervals are particularly valuable for sparse event data because they communicate how much uncertainty surrounds an observed proportion.
The same principle applies to dose-specific comparisons. Even if one dose grouping has a higher observed event proportion than another, the difference may reflect random variation, differences in exposure, differences in the subjects represented in each grouping, or the small number of observations. A formal comparison would require an appropriate analysis plan and sufficiently informative data.
12. Patient-Reported Outcomes: Why Direction Matters
The PAH-SYMPACT questionnaire provides an instructive example of why statistical interpretation must begin with the measurement scale rather than with the sign of a numerical result.
| Endpoint | Scale / structure | Meaning of a higher score | Interpretive direction |
|---|---|---|---|
| 6MWD | Distance measure | Greater walking distance | An increase represents a higher measured distance. |
| PAH-SYMPACT symptom and impact domains | 0 to 4 | More severe symptoms / impacts | A decrease represents less severe reported symptoms or impacts. |
| PQ-ITD | 5 ordered response categories across 12 statements | Depends on the statement wording | Interpretation requires attention to the individual statement. |
This is more than a presentation issue. It prevents a common statistical error: treating every negative change as unfavorable and every positive change as favorable. The direction of a numerical effect has no inherent meaning until the underlying variable is understood.
For PAH-SYMPACT, the registry explicitly defines higher scores as more severe symptoms or impacts. Therefore, the clinically meaningful direction is toward lower domain scores. A statistical analysis should preserve that interpretation in tables, confidence intervals, and any graphical display.
The PQ-ITD requires another layer of care because it consists of 12 statements rather than a single continuously measured outcome. Agreement may indicate greater satisfaction for some statements, while the interpretation of a particular response depends on what the statement says. A response distribution is therefore often more transparent than reducing the questionnaire to one number without a prespecified scoring rule.
13. Longitudinal Interpretation
BREEZE contains assessments at multiple time points, including baseline, Week 3, and Week 11 for subjects participating in the OEP. Longitudinal data create opportunities for more informative analysis than a simple comparison of group averages, because repeated observations from the same subject are correlated.
Within-subject correlation
Measurements from the same subject tend to be related. Statistical models should account for this dependence when multiple time points are analyzed jointly.
Baseline adjustment
Change scores directly use baseline as the reference. Alternative longitudinal models can include baseline as a covariate when justified by the prespecified analysis plan.
Time-point interpretation
Week 3 and Week 11 answer different temporal questions and should not automatically be combined into a single effect estimate.
OEP population
The Week-11 PAH-SYMPACT endpoint is specifically restricted to subjects participating in the OEP.
A repeated-measures model can be useful when the scientific question concerns the trajectory over time rather than a single follow-up measurement. Depending on the endpoint and analysis plan, possibilities include mixed-effects models for continuous outcomes or models designed for ordinal questionnaire responses. But the registry does not specify such a model, so it should not be represented as the analysis actually used in BREEZE.
The distinction between observed follow-up and longitudinal modeling is important. Simply having multiple observations does not require a repeated-measures model. The analysis method should follow the endpoint definition, the intended estimand, the measurement schedule, and the prespecified statistical plan.
14. Missing Data and Analysis Population
The registered Week-11 PAH-SYMPACT endpoint is explicitly limited to subjects participating in the OEP. That wording creates an analysis-population boundary that is different from the overall enrollment of 51 subjects.
For a baseline-to-follow-up change endpoint, missing Week-3 or Week-11 measurements create an incomplete change score. The statistical consequences depend on why measurements are missing and on the analysis method used. A complete-case analysis uses only subjects with the required measurements, whereas longitudinal models or multiple-imputation approaches can make different assumptions about incomplete observations.
The choice is not merely technical. If subjects with missing follow-up measurements differ systematically from those with observed measurements, a complete-case summary may not represent the full enrolled population. Conversely, an imputation model introduces assumptions of its own. The registry does not specify a missing-data or imputation method in the information available here, so no particular approach should be treated as the registered method.
15. Statistical Interpretation of a Within-Subject Change
A baseline-to-Week-3 change estimate describes how the measured outcome differed from baseline among the subjects included in that analysis. For 6MWD, the unit is the same distance scale as the original measurement. For PAH-SYMPACT, the change is expressed on the domain's 0-to-4 scale.
A within-subject change is not automatically a treatment effect. Without a concurrent randomized comparator, the observed change cannot by itself separate the effect of TreT from other changes occurring between baseline and follow-up.
If a formal confidence interval were reported for a change estimate, it would quantify statistical uncertainty around the estimated average or other prespecified parameter. It would not describe the range of changes experienced by individual subjects.
If a hypothesis test were used, its p-value would describe the compatibility of the observed data with the null hypothesis under the specified model. It would not tell the reader whether a change is clinically large, nor would it establish causality in a single-group study.
16. Statistical Interpretation of the PQ-ITD
The PQ-ITD endpoint is fundamentally categorical. Each of its 12 statements has five possible responses: strongly disagree, disagree, neutral, agree, and strongly agree.
For an ordinal endpoint, the most transparent statistical display is often a response table showing the number or proportion of subjects selecting each category for each statement. This preserves the shape of the response distribution and avoids assuming that the distance from strongly disagree to disagree is numerically identical to the distance from agree to strongly agree.
If a formal ordinal regression model were prespecified, it could use the ordered structure of the responses directly. Such a model would require an explicit definition of the response variable, the analysis population, and any assumptions governing the relationship among response categories. Those details are not identified in the registry information available here.
Preference is also conceptually different from satisfaction. A subject can report favorable satisfaction characteristics without necessarily expressing the same degree of preference across all device-related statements. Keeping the individual questionnaire statements visible can therefore make the statistical interpretation more faithful to the endpoint.
17. What BREEZE Does Statistically — and Does Not — Establish
| Question | What the design can address | What requires caution |
|---|---|---|
| Did measurements change after switching to TreT? | Yes. Baseline-to-follow-up changes can be described. | The change is not automatically attributable solely to TreT. |
| What were subjects' reported symptoms and impacts? | PAH-SYMPACT provides structured patient-reported measurements. | Interpretation must respect the 0-to-4 domain scale and its direction. |
| How satisfied were subjects with the inhaled device? | PQ-ITD provides 12 statements with five response options. | Ordinal responses should not automatically be treated as equally spaced continuous values. |
| How many serious adverse events occurred? | The registry reports affected/at-risk counts by dose and study phase. | There is no randomized comparator for a comparative safety estimate. |
| Is one dose safer than another? | Dose-specific observed counts can be described. | A formal dose comparison is not established by the observed counts alone. |
| Does the study demonstrate a causal treatment effect? | It provides longitudinal observations after treatment exposure. | The single-group design lacks a randomized concurrent counterfactual. |
This distinction is central to responsible interpretation of phase 1 and single-group clinical research. A study can provide useful information about tolerability, patient experience, and changes over time without being designed to provide the same type of confirmatory causal evidence as a randomized controlled trial.
18. Important Limitations and Interpretation Issues
- No randomized comparator: BREEZE is registered as a single-group study with one arm. Baseline-to-follow-up changes therefore do not provide a randomized estimate of comparative treatment effect.
- Open-label design: masking is listed as none. Patient-reported outcomes such as satisfaction, preference, symptoms, and impacts can be influenced by knowledge of treatment exposure.
- Single-group efficacy interpretation: an observed change after switching from Tyvaso to TreT cannot by itself distinguish treatment effects from temporal or other influences.
- Ordinal questionnaire responses: the five PQ-ITD response categories are ordered categorical data and should not automatically be treated as equally spaced numerical measurements.
- Patient-reported outcome direction: PAH-SYMPACT uses a 0-to-4 scale in which higher scores indicate more severe symptoms or impacts. The direction of a change must therefore be interpreted accordingly.
- Week-11 population: the Week-11 PAH-SYMPACT endpoint is specifically for subjects participating in the OEP, so its population is not synonymous with overall enrollment.
- Small safety denominators: some dose-phase safety groupings contain only 2 subjects at risk, making their observed event frequencies highly uncertain.
- Phase-specific safety: treatment-phase and optional-extension counts should not be combined without considering their different study periods and populations.
- Unreported formal analyses: the registry provides no posted statistical analysis objects for the listed endpoints, so the exact inferential procedures and model specifications cannot be attributed to the registry.
- Missing-data considerations: the interpretation of longitudinal changes depends on which subjects have evaluable follow-up and on the assumptions used to address incomplete observations.
19. Why This Trial Matters Statistically
BREEZE is a useful teaching example because it demonstrates a type of clinical-trial question that is sometimes misunderstood when the statistical framework developed for randomized trials is applied mechanically to a single-group study.
| Statistical concept | How it appears in BREEZE |
|---|---|
| Single-group design | One registered study arm with no randomized comparator. |
| Within-subject change | 6MWD and PAH-SYMPACT are defined using change from baseline. |
| Continuous outcome | 6MWD is evaluated at study entry and after 3 weeks of treatment. |
| Patient-reported outcomes | PAH-SYMPACT measures symptoms and impacts across multiple domains. |
| Ordinal categorical data | PQ-ITD uses five ordered response options across 12 statements. |
| Analysis populations | The Week-11 PAH-SYMPACT endpoint is restricted to subjects participating in the OEP. |
| Safety denominators | Serious adverse events are reported as affected / at risk by dose and study phase. |
| Longitudinal analysis | Measurements are defined at baseline, Week 3, and Week 11 for the applicable endpoint. |
| Missing-data interpretation | Incomplete follow-up can affect baseline-to-follow-up change analyses. |
| Causal inference | The absence of a randomized comparator limits causal attribution of observed changes. |
The most important lesson is that design determines estimand. In a randomized two-arm trial, the principal question may be a treatment-versus-control contrast. In BREEZE, the registered endpoints are centered on changes observed after treatment exposure in a single group. The statistical analysis must therefore preserve the distinction between a descriptive longitudinal estimate and a comparative causal effect.
20. Trial Timeline
Study start
The BREEZE clinical study began.
Treprostinil Inhalation Powder exposure
Subjects with pulmonary arterial hypertension currently using Tyvaso were evaluated after switching to TreT.
Primary follow-up assessments
6MWD, PQ-ITD satisfaction and preference, and PAH-SYMPACT change were assessed at the registered Week-3 time point.
OEP patient-reported assessment
PAH-SYMPACT change from baseline to Week 11 was evaluated for subjects participating in the OEP.
Primary completion
The registry lists 2023-08-22 as the primary completion date.
21. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The study provides within-subject longitudinal endpoints and phase- and dose-specific safety counts. The appropriate statistical focus is on change, response distributions, uncertainty, and the structure of the observed data.
Clinical interpretation
The findings concern functional capacity, patient-reported PAH symptoms and impact, satisfaction and preference with an inhaled treprostinil device, and observed serious adverse events after transition to TreT.
The two perspectives should remain connected but distinct. A statistical change estimate describes the data numerically. Clinical interpretation asks what that change means in the context of pulmonary arterial hypertension and the patient experience.
22. A Practical Statistical Reading of the BREEZE Endpoints
Step 1: Identify the estimand
For 6MWD and PAH-SYMPACT, the registered estimand is a change from baseline at a specified follow-up time. This is fundamentally a within-subject estimand. For PQ-ITD, the estimand concerns subject satisfaction and preference responses after treatment with TreT.
Step 2: Identify the measurement scale
6MWD is a continuous functional measure. PAH-SYMPACT domain scores occupy a bounded 0-to-4 scale. PQ-ITD responses are ordinal categories. The statistical method should follow that measurement structure.
Step 3: Identify the analysis population
The overall enrollment is 51 subjects, but not every endpoint necessarily uses every enrolled subject. The Week-11 PAH-SYMPACT endpoint specifically concerns subjects participating in the OEP. Safety summaries likewise use phase- and dose-specific denominators.
Step 4: Describe the observed data
For continuous changes, useful summaries include the central tendency and variability of individual changes. For ordinal responses, the distribution across response categories is central. For safety, affected and at-risk counts should remain paired.
Step 5: Quantify uncertainty
Confidence intervals can quantify uncertainty around a prespecified population parameter. Their interpretation depends on the estimator and model used. In small groups, uncertainty can be substantial even when the observed count appears straightforward.
Step 6: Respect the study design
Because BREEZE is single-group and open-label, statistical results should be interpreted as observations associated with treatment exposure and follow-up rather than as randomized comparative treatment effects.
23. Sources
- ClinicalTrials.gov: BREEZE, NCT03950739.
- PubMed: PubMed record.
- PubMed: PubMed record.
Continue through the Clinical Biostats knowledge graph
Clinical trial statistics become easier to interpret when study design, endpoint measurement, analysis populations, uncertainty, and causal inference are considered together.
24. Record Summary
BREEZE is a completed phase 1, single-group, open-label study of Treprostinil Inhalation Powder in subjects with pulmonary arterial hypertension currently using Tyvaso. The registry reports an enrollment of 51 and defines primary endpoints around change in 6-Minute Walk Distance from baseline to Week 3, subject satisfaction and preference for inhaled treprostinil devices after 3 weeks, PAH-SYMPACT change from baseline to Week 3, and PAH-SYMPACT change from baseline to Week 11 for subjects participating in the OEP.
Statistically, the defining feature is the single-group design. The central efficacy quantities are within-subject changes rather than randomized between-group contrasts. That makes baseline-to-follow-up analysis appropriate for describing observed trajectories while limiting causal interpretation. The PAH-SYMPACT endpoints additionally require careful attention to scale direction, because higher domain scores indicate more severe symptoms or impacts. The PQ-ITD endpoint requires attention to the ordered categorical structure of its five response options across 12 statements.
The safety record provides affected/at-risk counts for serious adverse events across treatment-phase and optional-extension dose groupings. These counts are useful descriptions of observed safety experience, but their interpretation is constrained by the absence of a randomized comparator and, in some groupings, very small denominators.