This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
TRANSFORM was a randomized, parallel-design trial evaluating endoscopic lung volume reduction with Zephyr valves in participants with COPD and heterogeneous emphysema. The registered primary endpoint was the percentage of participants meeting a minimally clinically important difference in FEV1 improvement at 3 months.
| Feature | TRANSFORM |
|---|---|
| Trial acronym | TRANSFORM |
| ClinicalTrials.gov identifier | NCT02022683 |
| Title | To Improve Lung Function and Symptoms for Emphysema Patients Using Zephyr Valves |
| Conditions | COPD; Heterogeneous Emphysema |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 97 |
| Number of arms | 2 |
| Intervention | ELVR (Endoscopic Lung Volume Reduction) with Zephyr Valves (device) |
| Lead sponsor | Pulmonx Corporation |
| Sponsor type | Industry |
| Trial status | Completed |
2. Clinical Question
The central question was whether participants receiving endoscopic lung volume reduction with Zephyr valves had a higher percentage of FEV1 responders at 3 months than participants in the control arm.
Population
Trial participants with COPD and heterogeneous emphysema.
Intervention
ELVR (Endoscopic Lung Volume Reduction) with Zephyr Valves, a device intervention.
Comparator
A control arm within the randomized two-arm parallel trial. The registry record identifies the trial as having two arms but does not provide a more detailed intervention description for the control arm in the available record.
Primary question
What percentage of participants in the EBV treatment arm met the prespecified FEV1 responder definition compared with the control arm at 3 months?
3. Trial Design
Randomization is important because it establishes the treatment assignment before the outcome is assessed. In a randomized comparison, differences in the primary endpoint can be interpreted in relation to treatment assignment without relying solely on adjustment for measured baseline characteristics.
The absence of masking is also statistically relevant. The primary endpoint is based on FEV1, which is an objective physiological measurement, but the overall trial experience can still involve expectations about treatment assignment. An unmasked design therefore differs from a blinded trial in the ways treatment knowledge could influence behavior, follow-up, or other aspects of outcome ascertainment.
4. Endpoints
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Forced Expiratory Volume in 1-second (FEV1) - Responders | The percentage of trial participants in the EBV treatment arm meeting the minimally clinically important difference (MCID) of >12% improved forced expiratory volume in one second (FEV1), obtained immediately following bronchodilator therapy, as compared to the percentage in the control arm at 3 months post-procedure. | Between baseline and 3 months |
What the primary endpoint measures
The endpoint converts a continuous physiological measurement into a binary responder outcome. Each participant is classified according to whether the participant meets the prespecified threshold of >12% improvement in FEV1 between baseline and 3 months, with FEV1 obtained immediately following bronchodilator therapy.
The treatment comparison is therefore a comparison of proportions, rather than a direct comparison of mean FEV1 change. This distinction matters because a responder analysis focuses on the fraction of participants crossing a clinically defined threshold. Two treatment groups could have similar average changes while differing in the percentage who cross the responder threshold, or they could have similar responder percentages while differing in the magnitude of improvement among responders.
The registry specifies that FEV1 is obtained immediately following bronchodilator therapy and that the comparison is between the EBV treatment arm and the control arm at 3 months.
5. Statistical Methodology
The registry identifies a single primary endpoint based on the percentage of participants who meet a binary FEV1 responder definition. The record available here does not post a formal statistical analysis for that endpoint. The appropriate statistical framework therefore follows from the endpoint's structure rather than from a reported analysis result.
| Statistical feature | How it applies to the primary endpoint |
|---|---|
| Outcome type | Binary responder status at 3 months |
| Primary comparison | Responder percentage in the EBV treatment arm versus the control arm |
| Typical descriptive analysis | Responder count and responder percentage within each randomized arm |
| Typical inferential analysis | Comparison of two independent proportions, with a confidence interval for the between-arm difference or ratio |
| Potential model-based approach | Binary regression, such as logistic regression, if adjustment for prespecified covariates is appropriate |
| Analysis time point | 3 months post-procedure |
For a simple unadjusted analysis, the most direct effect measure is the difference in responder proportions:
Here, pEBV is the proportion meeting the >12% FEV1 improvement criterion in the EBV treatment arm, and pControl is the corresponding proportion in the control arm.
An absolute difference is particularly natural for a responder endpoint because it describes the additional percentage of participants meeting the clinically defined threshold in one arm compared with the other. A relative measure, such as a risk ratio, can also be informative, but it answers a different question and can behave differently when the control-group responder percentage is small.
For a randomized trial, the primary analysis would ordinarily preserve the randomized comparison. If every participant has an evaluable responder status at 3 months, the analysis can be expressed directly as a two-group comparison of proportions. If some participants do not have an evaluable 3-month FEV1 measurement, the handling of missing endpoint data becomes an important component of the statistical analysis.
Confidence intervals
A confidence interval should accompany the estimated treatment difference or another prespecified effect measure. The interval communicates the statistical precision of the estimate and helps distinguish a precise estimate from one that remains uncertain because of limited information.
For a difference in proportions, the confidence interval describes uncertainty around the estimated difference between the two randomized groups. It does not describe the range of responses among individual participants.
Binary regression as an optional model
If prespecified covariate adjustment were used, a logistic regression model could represent the probability of being an FEV1 responder as a function of randomized treatment and selected baseline covariates. The resulting odds ratio would describe a relative change in the odds of response, not a relative change in FEV1 itself.
For a primary randomized comparison, however, model complexity should serve a clear statistical purpose. Adjustment should not be introduced simply because baseline variables happen to differ numerically between randomized groups. The strength of randomization is that the treatment comparison can be made without requiring every prognostic variable to be modeled.
6. Planned Analysis
The registered primary analysis is a comparison of the percentage of participants meeting the FEV1 responder criterion in the EBV treatment arm with the percentage in the control arm at 3 months. Because the endpoint is binary, a standard analysis would summarize the number and percentage of responders in each randomized group and compare the two proportions.
The most interpretable primary effect measure would generally be the absolute difference in responder percentages. For example, if one group had a responder proportion of 60% and the other had 40%, the absolute difference would be 20 percentage points. That example illustrates the statistical meaning of the measure only; it is not an outcome of TRANSFORM.
A confidence interval around the treatment difference would quantify uncertainty. A hypothesis test could also be used to evaluate a prespecified null hypothesis of equal responder proportions. The p-value from such a test would describe the compatibility of the observed data with that null hypothesis under the specified statistical model; it would not measure the size or clinical importance of the FEV1 effect.
The choice between a Wald-type interval, an exact method, a score-based method, or a model-based interval should normally be specified before the analysis because the behavior of proportion-based intervals can differ, particularly with smaller samples or proportions close to 0 or 1.
Missing 3-month FEV1 measurements would require an explicit analysis convention. A complete-case responder analysis is straightforward but can become sensitive to whether missingness is related to treatment assignment or outcome. More elaborate approaches may be appropriate depending on the amount and mechanism of missing data. The registry record available here does not specify a missing-data or imputation strategy.
7. Statistical Methods Explained
What is an FEV1 responder?
An FEV1 responder is a participant who satisfies the registered criterion of >12% improvement in FEV1 between baseline and 3 months, with FEV1 measured immediately following bronchodilator therapy. The endpoint therefore transforms a continuous measurement into a yes/no clinical-response classification.
Why use a responder endpoint instead of only mean FEV1 change?
A responder endpoint focuses on the proportion of participants who reach a clinically defined threshold. This can make the result easier to interpret at the individual-participant level: the analysis asks how many participants achieved the specified improvement rather than only how much the average FEV1 changed.
Why does the >12% threshold matter?
The threshold defines who counts as a responder. Changing the threshold would change the classification of participants and therefore could change the estimated responder percentages and the treatment comparison. A prespecified threshold prevents the cutoff from being chosen after looking at the observed data.
What is the most direct treatment effect for this endpoint?
The absolute difference in responder proportions is a direct measure. If the EBV arm has responder proportion pEBV and the control arm has responder proportion pControl, the difference is pEBV − pControl. Expressed in percentage points, this tells readers how much the responder percentage differs between randomized groups.
What would a risk ratio mean?
A risk ratio would divide the responder probability in the EBV arm by the responder probability in the control arm. A ratio of 1 would represent equal responder probabilities, while a ratio above 1 would indicate a higher probability of meeting the responder definition in the EBV arm. It would not describe the magnitude of FEV1 improvement among individual responders.
Why is the p-value not the effect size?
A p-value is a measure associated with a statistical test under a specified null hypothesis. It depends on both the observed difference and the amount of information available. The effect size itself should be reported using a quantity such as an absolute responder difference or risk ratio, preferably with a confidence interval.
Why does missing data matter for a 3-month responder endpoint?
A participant without an evaluable 3-month FEV1 measurement may not be classifiable as a responder. If missingness is related to treatment assignment, clinical status, or the eventual outcome, simply excluding those participants can alter the interpretation of the comparison. A prespecified missing-data strategy is therefore part of a rigorous analysis of a binary endpoint.
8. Interpreting a Responder Analysis
The primary endpoint asks whether the percentage of participants crossing the >12% FEV1 improvement threshold differs between randomized groups at 3 months. This is not an analysis of the average FEV1 change and should not be described as one.
An absolute responder difference is expressed in percentage points. If the estimated difference were 15 percentage points, it would mean that the estimated proportion meeting the responder definition was 15 percentage points higher in one randomized group than the other. It would not mean that each participant experienced a 15% improvement in FEV1.
A risk ratio compares probabilities multiplicatively, while an odds ratio compares odds. Neither is interchangeable with an absolute percentage-point difference. The choice of effect measure should reflect the endpoint and the clinical question rather than being driven by which measure appears numerically larger.
A confidence interval around the responder difference indicates how precisely the treatment effect has been estimated. A wide interval signals substantial statistical uncertainty, while a narrower interval indicates greater precision. The interval does not represent the range of individual FEV1 responses.
A small p-value would not by itself establish that the difference is clinically large, and a larger p-value would not prove that the treatments are clinically identical. The effect estimate and its confidence interval provide the essential information about magnitude and uncertainty.
9. Continuous FEV1 Versus Responder FEV1
The primary endpoint uses FEV1 in a threshold-based way. That creates a useful but important distinction from analyzing FEV1 as a continuous outcome.
| Approach | Question answered | Statistical consequence |
|---|---|---|
| Continuous FEV1 change | How much did FEV1 change, on average, between baseline and 3 months? | Retains the numerical magnitude of each participant's change. |
| FEV1 responder status | What percentage achieved >12% improvement? | Converts the measurement into a binary outcome. |
Turning a continuous variable into a binary responder endpoint can make interpretation more clinically intuitive, but it also discards information. A participant with a 12.1% improvement and a participant with a much larger improvement are both classified as responders, while participants just below and just above the threshold are treated differently despite potentially similar measurements.
This does not make a responder endpoint invalid. It means that the analysis should be interpreted according to the question it was designed to answer. The registered endpoint is specifically about the percentage crossing the >12% threshold.
10. Timing of the Primary Endpoint
The registered time frame is between baseline and 3 months, with the endpoint evaluated at 3 months post-procedure. This fixed assessment window is important because responder status depends on when the FEV1 measurement is taken.
Reference FEV1
FEV1 establishes the participant-specific starting point against which the subsequent measurement is compared.
Primary endpoint assessment
FEV1 is obtained immediately following bronchodilator therapy, and the participant is classified according to whether improvement exceeds 12%.
A fixed endpoint time reduces ambiguity about the treatment comparison. It also means that the analysis is specifically a 3-month response comparison, rather than a time-to-event analysis in which participants are followed until an event occurs.
11. Randomization and Causal Interpretation
The randomized allocation is the central design feature supporting a causal comparison between the two arms. Randomization aims to distribute both measured and unmeasured prognostic factors across treatment assignments in expectation.
What randomization supports
It provides the basis for comparing outcomes between treatment assignments without assuming that all relevant prognostic variables have been measured and adjusted for.
What randomization does not guarantee
It does not guarantee numerically identical baseline characteristics in a particular sample, and it does not eliminate problems caused by missing outcome measurements, protocol deviations, or differential follow-up.
Because the trial was unmasked, interpretation should also consider the possibility that knowledge of assignment could influence aspects of participant or investigator behavior. The more objective the primary measurement, the less directly such influences map onto the measured value, but unmasking remains a design feature worth recognizing.
12. Missing Data and Analysis Population
The primary endpoint requires both a baseline reference and a 3-month FEV1 measurement to determine whether the participant meets the >12% responder threshold. Consequently, missing measurements can affect the denominator and the classification of responders.
| Issue | Why it matters |
|---|---|
| Missing baseline FEV1 | The percentage improvement cannot be calculated without the baseline reference. |
| Missing 3-month FEV1 | Responder status cannot be directly determined at the registered endpoint time. |
| Differential missingness | If missingness differs by randomized arm, the simple observed responder percentages may not represent comparable populations. |
| Imputation | An imputation strategy can influence the estimated treatment effect and should therefore be prespecified and justified. |
The registry record available here does not specify a missing-data or imputation method. That absence is important because a binary responder analysis can be sensitive to how participants without evaluable 3-month measurements are handled.
A rigorous statistical report would distinguish the randomized population from the population with observed primary endpoint data. It would also state how many participants were included in the denominator for each arm and whether any prespecified sensitivity analyses addressed missing observations.
13. Statistical Issues Specific to the Responder Threshold
The >12% cutoff creates several statistical features that are easy to overlook.
Threshold sensitivity
Participants close to the cutoff can change classification with relatively small measurement differences. The threshold therefore has a direct influence on the responder rate.
Loss of magnitude information
Once classified, the analysis does not distinguish a modest responder from a participant with a substantially larger improvement.
Denominator sensitivity
The responder percentage depends on which participants are included in the analysis denominator, making missing-data rules important.
Effect-measure choice
Risk difference, risk ratio, and odds ratio summarize different aspects of the same binary outcome and should not be treated as interchangeable.
These considerations do not undermine the endpoint. They define what the endpoint can and cannot tell us. The responder analysis is specifically informative about the probability of crossing a clinically defined FEV1 improvement threshold at the specified follow-up time.
14. Multiplicity and Interim Analysis
The registry information available here identifies one registered primary endpoint: FEV1 responders between baseline and 3 months. It does not provide a multiplicity strategy or an interim-analysis plan in the statistical information available for this record.
Multiplicity becomes important when investigators test several endpoints, several time points, or several subgroups and then interpret the most favorable result as though it were the only prespecified test. A single registered primary endpoint reduces one potential source of multiplicity, but the complete statistical analysis plan would be needed to determine whether additional confirmatory hypotheses or interim looks were part of the formal design.
Similarly, no formal interim-analysis method is reported in the available registry information. An interim analysis would normally require prespecified rules governing the timing of the look and the type I error spent at that stage.
15. Bayesian Methods and Model-Based Analysis
No Bayesian analysis is identified in the available registry information. The primary endpoint itself does not require a Bayesian model: the basic statistical problem is a comparison of two independent responder proportions.
A Bayesian analysis could theoretically model the responder probability in each arm and produce a posterior probability that one treatment exceeds the other by a specified amount. Such an analysis would require explicit prior distributions and prespecified decision criteria. Those elements are not reported in the registry information available here and therefore are not part of the documented statistical analysis of TRANSFORM.
16. Safety and Secondary Endpoints
The available registry information identifies the primary FEV1 responder endpoint but does not provide posted statistical analyses or numerical results for secondary endpoints or serious adverse events by arm.
Safety and efficacy are conceptually distinct dimensions of a clinical trial. A statistical report would ordinarily summarize adverse events by treatment assignment using counts and percentages, with serious adverse events and other prespecified safety categories analyzed according to the trial's safety definitions. The interpretation would depend on exposure, follow-up, event definitions, and the number of participants contributing safety information.
For TRANSFORM, the available ClinicalTrials.gov record does not report arm-specific serious adverse-event counts in the information used for this analysis. No numerical safety comparison is therefore presented here.
17. What a Complete Primary Results Table Would Contain
For this endpoint, a statistically informative results table would contain more than a single p-value. At minimum, readers would want the responder count and denominator in each randomized arm, the responder percentage, a treatment-effect measure, a confidence interval, and the statistical test or model used.
| Quantity | Purpose |
|---|---|
| Responders / analyzed participants | Shows the numerator and denominator underlying each percentage. |
| Responder percentage | Describes the proportion meeting the >12% FEV1 criterion. |
| Absolute difference | Shows the percentage-point difference between randomized arms. |
| Confidence interval | Quantifies statistical uncertainty around the treatment effect. |
| Hypothesis-test result | Assesses the prespecified null hypothesis under the selected statistical method. |
| Missing endpoint data | Shows how many randomized participants were not directly classifiable at 3 months. |
This structure is useful because it keeps the statistical evidence transparent. A responder percentage without its denominator can obscure how much information contributed to the estimate. A p-value without an effect estimate can obscure the magnitude of the observed difference. A point estimate without a confidence interval obscures precision.
18. Limitations
- No posted statistical analyses: the ClinicalTrials.gov record does not contain formal statistical analysis results for the registered primary endpoint.
- No numerical primary outcome: the available registry information identifies the endpoint definition and time frame but does not report the responder percentages or treatment-effect estimates.
- Binary endpoint: converting continuous FEV1 improvement into responder status loses information about the magnitude of improvement within each category.
- Threshold dependence: classification depends on the prespecified >12% improvement criterion, so participants close to the threshold may be particularly influential.
- Missing-data sensitivity: the endpoint requires evaluable baseline and 3-month measurements, while the available record does not specify a missing-data or imputation strategy.
- Unmasked design: the registry identifies no masking, so knowledge of assignment is a relevant design consideration when interpreting the trial.
- Incomplete statistical specification: the available record does not identify the exact inferential test, confidence-interval method, multiplicity strategy, or interim-analysis procedure.
- Limited endpoint scope: the primary endpoint addresses FEV1 response at 3 months and does not by itself characterize longer-term clinical outcomes.
19. Why This Trial Matters Statistically
TRANSFORM is a useful teaching example because it illustrates how clinical-trial statistics change when a continuous physiological measurement is converted into a clinically defined responder endpoint.
| Concept | How it appears in TRANSFORM |
|---|---|
| Randomization | Participants were randomized to 2 parallel arms. |
| Binary endpoint | The primary outcome classifies participants according to whether FEV1 improvement exceeds 12%. |
| Clinical threshold | The registered MCID is >12% improvement. |
| Fixed follow-up | The primary comparison occurs at 3 months post-procedure. |
| Continuous-to-binary transformation | FEV1, a continuous measurement, becomes responder versus non-responder. |
| Proportion comparison | The treatment groups are compared on the percentage meeting the responder criterion. |
| Confidence intervals | Useful for expressing precision around the between-arm treatment effect. |
| Missing data | Missing baseline or 3-month FEV1 measurements can affect responder classification. |
| Unmasked design | The registry identifies no masking. |
The trial also demonstrates why statistical interpretation should begin with the exact endpoint definition. A statement such as "FEV1 improved" is not equivalent to the registered primary endpoint. The endpoint is specifically the percentage of participants whose FEV1 improved by more than 12% from baseline to 3 months, with the measurement obtained immediately following bronchodilator therapy.
20. Statistical Interpretation of the Primary Question
The primary treatment effect would be a difference between the percentage of FEV1 responders in the EBV treatment arm and the percentage in the control arm. The most direct absolute measure would be expressed in percentage points.
A confidence interval around the estimated treatment difference would show how precisely the responder-rate difference had been estimated. The width of that interval would depend on the number of participants and the observed responder proportions.
A prespecified test could evaluate whether the observed responder proportions are compatible with a null hypothesis of no difference. The test would not establish how clinically important the difference is.
It would not describe the average magnitude of FEV1 improvement, the durability of improvement beyond the 3-month endpoint, or outcomes not included in the registered primary definition.
This distinction is central to responsible interpretation. The endpoint is clinically focused, but its statistical meaning remains precisely defined: a comparison of the probability of crossing a prespecified FEV1 threshold at a specified time.
21. Trial Timeline
Trial start
The registry lists January 28, 2014 as the study start date.
Primary completion
The registry lists January 20, 2017 as the primary completion date.
Study status
The ClinicalTrials.gov record identifies the study as completed.
22. Registry Record: What Is Established
The registry establishes a clear primary statistical question: compare the percentage of participants who meet a >12% improvement criterion for FEV1 between the EBV treatment arm and the control arm at 3 months. It also establishes the randomized, parallel, unmasked structure of the study and the enrollment of 97 participants.
The available registry information does not provide the numerical responder results, confidence intervals, p-values, secondary endpoint results, or arm-specific serious adverse-event results. It also does not identify a formal posted statistical analysis for the primary endpoint.
23. Sources
- ClinicalTrials.gov: TRANSFORM, NCT02022683.
- PubMed: PubMed record 39515624.
- PubMed: PubMed record 28885054.
Continue through the Clinical Biostats knowledge graph
Explore statistical methods, clinical-trial designs, and analysis workflows through the Clinical Biostats resources.
24. Record Summary
TRANSFORM is a randomized, parallel, unmasked trial enrolling 97 participants with COPD and heterogeneous emphysema. Its registered primary endpoint is the percentage of participants meeting a minimally clinically important difference of >12% improvement in FEV1 between baseline and 3 months, with FEV1 obtained immediately following bronchodilator therapy. The statistical structure is therefore a two-group comparison of binary responder proportions at a fixed follow-up time.
The most informative analysis would report responder counts and percentages in each randomized arm, an absolute difference in responder percentages with an appropriate confidence interval, and a prespecified hypothesis test. Interpretation would also require attention to missing 3-month measurements, because participants without evaluable FEV1 data cannot be directly classified under the registered responder definition.
Most importantly, the endpoint should not be confused with a comparison of mean FEV1 change. A responder analysis answers a narrower question: what percentage of participants cross the >12% improvement threshold by 3 months? That distinction provides the central statistical lesson of the TRANSFORM design.