← Clinical Trials
COPD Heterogeneous Emphysema Randomized NCT02022683

TRANSFORM: Complete Statistical Analysis of Zephyr Valves in Heterogeneous Emphysema

An independent statistical review of the randomized TRANSFORM trial evaluating endoscopic lung volume reduction with Zephyr valves in participants with COPD and heterogeneous emphysema, with emphasis on the primary FEV1 responder endpoint at 3 months.

Trial status: Completed  ·  Enrollment: 97  ·  Primary completion: January 20, 2017
ClinicalTrials.gov record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

TRANSFORM was a randomized, parallel-design trial evaluating endoscopic lung volume reduction with Zephyr valves in participants with COPD and heterogeneous emphysema. The registered primary endpoint was the percentage of participants meeting a minimally clinically important difference in FEV1 improvement at 3 months.

97
Enrollment
Total participants
2
Arms
Randomized parallel design
3 mo
Primary time frame
Baseline to 3 months
>12%
FEV1 MCID
Responder threshold
FeatureTRANSFORM
Trial acronymTRANSFORM
ClinicalTrials.gov identifierNCT02022683
TitleTo Improve Lung Function and Symptoms for Emphysema Patients Using Zephyr Valves
ConditionsCOPD; Heterogeneous Emphysema
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment97
Number of arms2
InterventionELVR (Endoscopic Lung Volume Reduction) with Zephyr Valves (device)
Lead sponsorPulmonx Corporation
Sponsor typeIndustry
Trial statusCompleted

2. Clinical Question

The central question was whether participants receiving endoscopic lung volume reduction with Zephyr valves had a higher percentage of FEV1 responders at 3 months than participants in the control arm.

Population

Trial participants with COPD and heterogeneous emphysema.

Intervention

ELVR (Endoscopic Lung Volume Reduction) with Zephyr Valves, a device intervention.

Comparator

A control arm within the randomized two-arm parallel trial. The registry record identifies the trial as having two arms but does not provide a more detailed intervention description for the control arm in the available record.

Primary question

What percentage of participants in the EBV treatment arm met the prespecified FEV1 responder definition compared with the control arm at 3 months?

3. Trial Design

01
Enroll 97 participants
02
Randomize 2 parallel arms
03
Intervene Zephyr-valve ELVR in treatment arm
04
Measure FEV1 after bronchodilator therapy
05
Compare Responder percentage at 3 months
Allocation
Randomized allocation to 2 arms in a parallel design.
Masking
None.
Primary purpose
Treatment.
Intervention
Endoscopic lung volume reduction with Zephyr valves.

Randomization is important because it establishes the treatment assignment before the outcome is assessed. In a randomized comparison, differences in the primary endpoint can be interpreted in relation to treatment assignment without relying solely on adjustment for measured baseline characteristics.

The absence of masking is also statistically relevant. The primary endpoint is based on FEV1, which is an objective physiological measurement, but the overall trial experience can still involve expectations about treatment assignment. An unmasked design therefore differs from a blinded trial in the ways treatment knowledge could influence behavior, follow-up, or other aspects of outcome ascertainment.

4. Endpoints

EndpointRegistry definitionTime frame
Forced Expiratory Volume in 1-second (FEV1) - Responders The percentage of trial participants in the EBV treatment arm meeting the minimally clinically important difference (MCID) of >12% improved forced expiratory volume in one second (FEV1), obtained immediately following bronchodilator therapy, as compared to the percentage in the control arm at 3 months post-procedure. Between baseline and 3 months

What the primary endpoint measures

The endpoint converts a continuous physiological measurement into a binary responder outcome. Each participant is classified according to whether the participant meets the prespecified threshold of >12% improvement in FEV1 between baseline and 3 months, with FEV1 obtained immediately following bronchodilator therapy.

The treatment comparison is therefore a comparison of proportions, rather than a direct comparison of mean FEV1 change. This distinction matters because a responder analysis focuses on the fraction of participants crossing a clinically defined threshold. Two treatment groups could have similar average changes while differing in the percentage who cross the responder threshold, or they could have similar responder percentages while differing in the magnitude of improvement among responders.

Responder definition
Responder = 1 if improvement in FEV1 from baseline to 3 months > 12%; otherwise Responder = 0

The registry specifies that FEV1 is obtained immediately following bronchodilator therapy and that the comparison is between the EBV treatment arm and the control arm at 3 months.

5. Statistical Methodology

The registry identifies a single primary endpoint based on the percentage of participants who meet a binary FEV1 responder definition. The record available here does not post a formal statistical analysis for that endpoint. The appropriate statistical framework therefore follows from the endpoint's structure rather than from a reported analysis result.

Statistical featureHow it applies to the primary endpoint
Outcome typeBinary responder status at 3 months
Primary comparisonResponder percentage in the EBV treatment arm versus the control arm
Typical descriptive analysisResponder count and responder percentage within each randomized arm
Typical inferential analysisComparison of two independent proportions, with a confidence interval for the between-arm difference or ratio
Potential model-based approachBinary regression, such as logistic regression, if adjustment for prespecified covariates is appropriate
Analysis time point3 months post-procedure

For a simple unadjusted analysis, the most direct effect measure is the difference in responder proportions:

Absolute responder difference
Δ = pEBV − pControl

Here, pEBV is the proportion meeting the >12% FEV1 improvement criterion in the EBV treatment arm, and pControl is the corresponding proportion in the control arm.

An absolute difference is particularly natural for a responder endpoint because it describes the additional percentage of participants meeting the clinically defined threshold in one arm compared with the other. A relative measure, such as a risk ratio, can also be informative, but it answers a different question and can behave differently when the control-group responder percentage is small.

For a randomized trial, the primary analysis would ordinarily preserve the randomized comparison. If every participant has an evaluable responder status at 3 months, the analysis can be expressed directly as a two-group comparison of proportions. If some participants do not have an evaluable 3-month FEV1 measurement, the handling of missing endpoint data becomes an important component of the statistical analysis.

Confidence intervals

A confidence interval should accompany the estimated treatment difference or another prespecified effect measure. The interval communicates the statistical precision of the estimate and helps distinguish a precise estimate from one that remains uncertain because of limited information.

Conceptual interpretation
Estimated effect ± uncertainty from sampling variability

For a difference in proportions, the confidence interval describes uncertainty around the estimated difference between the two randomized groups. It does not describe the range of responses among individual participants.

Binary regression as an optional model

If prespecified covariate adjustment were used, a logistic regression model could represent the probability of being an FEV1 responder as a function of randomized treatment and selected baseline covariates. The resulting odds ratio would describe a relative change in the odds of response, not a relative change in FEV1 itself.

For a primary randomized comparison, however, model complexity should serve a clear statistical purpose. Adjustment should not be introduced simply because baseline variables happen to differ numerically between randomized groups. The strength of randomization is that the treatment comparison can be made without requiring every prognostic variable to be modeled.

6. Planned Analysis

ClinicalTrials.gov reporting status: No formal statistical analyses were posted to ClinicalTrials.gov for the primary endpoint in the registry record.

The registered primary analysis is a comparison of the percentage of participants meeting the FEV1 responder criterion in the EBV treatment arm with the percentage in the control arm at 3 months. Because the endpoint is binary, a standard analysis would summarize the number and percentage of responders in each randomized group and compare the two proportions.

The most interpretable primary effect measure would generally be the absolute difference in responder percentages. For example, if one group had a responder proportion of 60% and the other had 40%, the absolute difference would be 20 percentage points. That example illustrates the statistical meaning of the measure only; it is not an outcome of TRANSFORM.

A confidence interval around the treatment difference would quantify uncertainty. A hypothesis test could also be used to evaluate a prespecified null hypothesis of equal responder proportions. The p-value from such a test would describe the compatibility of the observed data with that null hypothesis under the specified statistical model; it would not measure the size or clinical importance of the FEV1 effect.

The choice between a Wald-type interval, an exact method, a score-based method, or a model-based interval should normally be specified before the analysis because the behavior of proportion-based intervals can differ, particularly with smaller samples or proportions close to 0 or 1.

Missing 3-month FEV1 measurements would require an explicit analysis convention. A complete-case responder analysis is straightforward but can become sensitive to whether missingness is related to treatment assignment or outcome. More elaborate approaches may be appropriate depending on the amount and mechanism of missing data. The registry record available here does not specify a missing-data or imputation strategy.

7. Statistical Methods Explained

What is an FEV1 responder?

An FEV1 responder is a participant who satisfies the registered criterion of >12% improvement in FEV1 between baseline and 3 months, with FEV1 measured immediately following bronchodilator therapy. The endpoint therefore transforms a continuous measurement into a yes/no clinical-response classification.

Why use a responder endpoint instead of only mean FEV1 change?

A responder endpoint focuses on the proportion of participants who reach a clinically defined threshold. This can make the result easier to interpret at the individual-participant level: the analysis asks how many participants achieved the specified improvement rather than only how much the average FEV1 changed.

Why does the >12% threshold matter?

The threshold defines who counts as a responder. Changing the threshold would change the classification of participants and therefore could change the estimated responder percentages and the treatment comparison. A prespecified threshold prevents the cutoff from being chosen after looking at the observed data.

What is the most direct treatment effect for this endpoint?

The absolute difference in responder proportions is a direct measure. If the EBV arm has responder proportion pEBV and the control arm has responder proportion pControl, the difference is pEBV − pControl. Expressed in percentage points, this tells readers how much the responder percentage differs between randomized groups.

What would a risk ratio mean?

A risk ratio would divide the responder probability in the EBV arm by the responder probability in the control arm. A ratio of 1 would represent equal responder probabilities, while a ratio above 1 would indicate a higher probability of meeting the responder definition in the EBV arm. It would not describe the magnitude of FEV1 improvement among individual responders.

Why is the p-value not the effect size?

A p-value is a measure associated with a statistical test under a specified null hypothesis. It depends on both the observed difference and the amount of information available. The effect size itself should be reported using a quantity such as an absolute responder difference or risk ratio, preferably with a confidence interval.

Why does missing data matter for a 3-month responder endpoint?

A participant without an evaluable 3-month FEV1 measurement may not be classifiable as a responder. If missingness is related to treatment assignment, clinical status, or the eventual outcome, simply excluding those participants can alter the interpretation of the comparison. A prespecified missing-data strategy is therefore part of a rigorous analysis of a binary endpoint.

8. Interpreting a Responder Analysis

The estimand is a proportion-based treatment comparison

The primary endpoint asks whether the percentage of participants crossing the >12% FEV1 improvement threshold differs between randomized groups at 3 months. This is not an analysis of the average FEV1 change and should not be described as one.

Absolute difference is clinically transparent

An absolute responder difference is expressed in percentage points. If the estimated difference were 15 percentage points, it would mean that the estimated proportion meeting the responder definition was 15 percentage points higher in one randomized group than the other. It would not mean that each participant experienced a 15% improvement in FEV1.

Relative measures answer a different question

A risk ratio compares probabilities multiplicatively, while an odds ratio compares odds. Neither is interchangeable with an absolute percentage-point difference. The choice of effect measure should reflect the endpoint and the clinical question rather than being driven by which measure appears numerically larger.

The confidence interval describes precision

A confidence interval around the responder difference indicates how precisely the treatment effect has been estimated. A wide interval signals substantial statistical uncertainty, while a narrower interval indicates greater precision. The interval does not represent the range of individual FEV1 responses.

The p-value does not measure clinical importance

A small p-value would not by itself establish that the difference is clinically large, and a larger p-value would not prove that the treatments are clinically identical. The effect estimate and its confidence interval provide the essential information about magnitude and uncertainty.

9. Continuous FEV1 Versus Responder FEV1

The primary endpoint uses FEV1 in a threshold-based way. That creates a useful but important distinction from analyzing FEV1 as a continuous outcome.

ApproachQuestion answeredStatistical consequence
Continuous FEV1 changeHow much did FEV1 change, on average, between baseline and 3 months?Retains the numerical magnitude of each participant's change.
FEV1 responder statusWhat percentage achieved >12% improvement?Converts the measurement into a binary outcome.

Turning a continuous variable into a binary responder endpoint can make interpretation more clinically intuitive, but it also discards information. A participant with a 12.1% improvement and a participant with a much larger improvement are both classified as responders, while participants just below and just above the threshold are treated differently despite potentially similar measurements.

This does not make a responder endpoint invalid. It means that the analysis should be interpreted according to the question it was designed to answer. The registered endpoint is specifically about the percentage crossing the >12% threshold.

10. Timing of the Primary Endpoint

The registered time frame is between baseline and 3 months, with the endpoint evaluated at 3 months post-procedure. This fixed assessment window is important because responder status depends on when the FEV1 measurement is taken.

Baseline

Reference FEV1

FEV1 establishes the participant-specific starting point against which the subsequent measurement is compared.

3 months

Primary endpoint assessment

FEV1 is obtained immediately following bronchodilator therapy, and the participant is classified according to whether improvement exceeds 12%.

A fixed endpoint time reduces ambiguity about the treatment comparison. It also means that the analysis is specifically a 3-month response comparison, rather than a time-to-event analysis in which participants are followed until an event occurs.

11. Randomization and Causal Interpretation

The randomized allocation is the central design feature supporting a causal comparison between the two arms. Randomization aims to distribute both measured and unmeasured prognostic factors across treatment assignments in expectation.

What randomization supports

It provides the basis for comparing outcomes between treatment assignments without assuming that all relevant prognostic variables have been measured and adjusted for.

What randomization does not guarantee

It does not guarantee numerically identical baseline characteristics in a particular sample, and it does not eliminate problems caused by missing outcome measurements, protocol deviations, or differential follow-up.

Because the trial was unmasked, interpretation should also consider the possibility that knowledge of assignment could influence aspects of participant or investigator behavior. The more objective the primary measurement, the less directly such influences map onto the measured value, but unmasking remains a design feature worth recognizing.

12. Missing Data and Analysis Population

The primary endpoint requires both a baseline reference and a 3-month FEV1 measurement to determine whether the participant meets the >12% responder threshold. Consequently, missing measurements can affect the denominator and the classification of responders.

IssueWhy it matters
Missing baseline FEV1The percentage improvement cannot be calculated without the baseline reference.
Missing 3-month FEV1Responder status cannot be directly determined at the registered endpoint time.
Differential missingnessIf missingness differs by randomized arm, the simple observed responder percentages may not represent comparable populations.
ImputationAn imputation strategy can influence the estimated treatment effect and should therefore be prespecified and justified.

The registry record available here does not specify a missing-data or imputation method. That absence is important because a binary responder analysis can be sensitive to how participants without evaluable 3-month measurements are handled.

A rigorous statistical report would distinguish the randomized population from the population with observed primary endpoint data. It would also state how many participants were included in the denominator for each arm and whether any prespecified sensitivity analyses addressed missing observations.

13. Statistical Issues Specific to the Responder Threshold

The >12% cutoff creates several statistical features that are easy to overlook.

Threshold sensitivity

Participants close to the cutoff can change classification with relatively small measurement differences. The threshold therefore has a direct influence on the responder rate.

Loss of magnitude information

Once classified, the analysis does not distinguish a modest responder from a participant with a substantially larger improvement.

Denominator sensitivity

The responder percentage depends on which participants are included in the analysis denominator, making missing-data rules important.

Effect-measure choice

Risk difference, risk ratio, and odds ratio summarize different aspects of the same binary outcome and should not be treated as interchangeable.

These considerations do not undermine the endpoint. They define what the endpoint can and cannot tell us. The responder analysis is specifically informative about the probability of crossing a clinically defined FEV1 improvement threshold at the specified follow-up time.

14. Multiplicity and Interim Analysis

The registry information available here identifies one registered primary endpoint: FEV1 responders between baseline and 3 months. It does not provide a multiplicity strategy or an interim-analysis plan in the statistical information available for this record.

Multiplicity becomes important when investigators test several endpoints, several time points, or several subgroups and then interpret the most favorable result as though it were the only prespecified test. A single registered primary endpoint reduces one potential source of multiplicity, but the complete statistical analysis plan would be needed to determine whether additional confirmatory hypotheses or interim looks were part of the formal design.

Similarly, no formal interim-analysis method is reported in the available registry information. An interim analysis would normally require prespecified rules governing the timing of the look and the type I error spent at that stage.

15. Bayesian Methods and Model-Based Analysis

No Bayesian analysis is identified in the available registry information. The primary endpoint itself does not require a Bayesian model: the basic statistical problem is a comparison of two independent responder proportions.

A Bayesian analysis could theoretically model the responder probability in each arm and produce a posterior probability that one treatment exceeds the other by a specified amount. Such an analysis would require explicit prior distributions and prespecified decision criteria. Those elements are not reported in the registry information available here and therefore are not part of the documented statistical analysis of TRANSFORM.

16. Safety and Secondary Endpoints

The available registry information identifies the primary FEV1 responder endpoint but does not provide posted statistical analyses or numerical results for secondary endpoints or serious adverse events by arm.

Safety and efficacy are conceptually distinct dimensions of a clinical trial. A statistical report would ordinarily summarize adverse events by treatment assignment using counts and percentages, with serious adverse events and other prespecified safety categories analyzed according to the trial's safety definitions. The interpretation would depend on exposure, follow-up, event definitions, and the number of participants contributing safety information.

For TRANSFORM, the available ClinicalTrials.gov record does not report arm-specific serious adverse-event counts in the information used for this analysis. No numerical safety comparison is therefore presented here.

17. What a Complete Primary Results Table Would Contain

For this endpoint, a statistically informative results table would contain more than a single p-value. At minimum, readers would want the responder count and denominator in each randomized arm, the responder percentage, a treatment-effect measure, a confidence interval, and the statistical test or model used.

QuantityPurpose
Responders / analyzed participantsShows the numerator and denominator underlying each percentage.
Responder percentageDescribes the proportion meeting the >12% FEV1 criterion.
Absolute differenceShows the percentage-point difference between randomized arms.
Confidence intervalQuantifies statistical uncertainty around the treatment effect.
Hypothesis-test resultAssesses the prespecified null hypothesis under the selected statistical method.
Missing endpoint dataShows how many randomized participants were not directly classifiable at 3 months.

This structure is useful because it keeps the statistical evidence transparent. A responder percentage without its denominator can obscure how much information contributed to the estimate. A p-value without an effect estimate can obscure the magnitude of the observed difference. A point estimate without a confidence interval obscures precision.

18. Limitations

19. Why This Trial Matters Statistically

TRANSFORM is a useful teaching example because it illustrates how clinical-trial statistics change when a continuous physiological measurement is converted into a clinically defined responder endpoint.

ConceptHow it appears in TRANSFORM
RandomizationParticipants were randomized to 2 parallel arms.
Binary endpointThe primary outcome classifies participants according to whether FEV1 improvement exceeds 12%.
Clinical thresholdThe registered MCID is >12% improvement.
Fixed follow-upThe primary comparison occurs at 3 months post-procedure.
Continuous-to-binary transformationFEV1, a continuous measurement, becomes responder versus non-responder.
Proportion comparisonThe treatment groups are compared on the percentage meeting the responder criterion.
Confidence intervalsUseful for expressing precision around the between-arm treatment effect.
Missing dataMissing baseline or 3-month FEV1 measurements can affect responder classification.
Unmasked designThe registry identifies no masking.

The trial also demonstrates why statistical interpretation should begin with the exact endpoint definition. A statement such as "FEV1 improved" is not equivalent to the registered primary endpoint. The endpoint is specifically the percentage of participants whose FEV1 improved by more than 12% from baseline to 3 months, with the measurement obtained immediately following bronchodilator therapy.

20. Statistical Interpretation of the Primary Question

Question 1 · What would constitute the treatment effect?

The primary treatment effect would be a difference between the percentage of FEV1 responders in the EBV treatment arm and the percentage in the control arm. The most direct absolute measure would be expressed in percentage points.

Question 2 · What would establish precision?

A confidence interval around the estimated treatment difference would show how precisely the responder-rate difference had been estimated. The width of that interval would depend on the number of participants and the observed responder proportions.

Question 3 · What would a statistical test establish?

A prespecified test could evaluate whether the observed responder proportions are compatible with a null hypothesis of no difference. The test would not establish how clinically important the difference is.

Question 4 · What would the responder endpoint not establish?

It would not describe the average magnitude of FEV1 improvement, the durability of improvement beyond the 3-month endpoint, or outcomes not included in the registered primary definition.

This distinction is central to responsible interpretation. The endpoint is clinically focused, but its statistical meaning remains precisely defined: a comparison of the probability of crossing a prespecified FEV1 threshold at a specified time.

21. Trial Timeline

January 28, 2014

Trial start

The registry lists January 28, 2014 as the study start date.

January 20, 2017

Primary completion

The registry lists January 20, 2017 as the primary completion date.

Completed

Study status

The ClinicalTrials.gov record identifies the study as completed.

22. Registry Record: What Is Established

The registry establishes a clear primary statistical question: compare the percentage of participants who meet a >12% improvement criterion for FEV1 between the EBV treatment arm and the control arm at 3 months. It also establishes the randomized, parallel, unmasked structure of the study and the enrollment of 97 participants.

The available registry information does not provide the numerical responder results, confidence intervals, p-values, secondary endpoint results, or arm-specific serious adverse-event results. It also does not identify a formal posted statistical analysis for the primary endpoint.

Statistical reading of the record: the design and endpoint definition are sufficiently specific to explain the appropriate statistical framework, but the numerical treatment effect cannot be characterized from the available ClinicalTrials.gov analysis record because no formal statistical analyses are posted there.

23. Sources

Continue through the Clinical Biostats knowledge graph

Explore statistical methods, clinical-trial designs, and analysis workflows through the Clinical Biostats resources.

24. Record Summary

TRANSFORM is a randomized, parallel, unmasked trial enrolling 97 participants with COPD and heterogeneous emphysema. Its registered primary endpoint is the percentage of participants meeting a minimally clinically important difference of >12% improvement in FEV1 between baseline and 3 months, with FEV1 obtained immediately following bronchodilator therapy. The statistical structure is therefore a two-group comparison of binary responder proportions at a fixed follow-up time.

The most informative analysis would report responder counts and percentages in each randomized arm, an absolute difference in responder percentages with an appropriate confidence interval, and a prespecified hypothesis test. Interpretation would also require attention to missing 3-month measurements, because participants without evaluable FEV1 data cannot be directly classified under the registered responder definition.

Most importantly, the endpoint should not be confused with a comparison of mean FEV1 change. A responder analysis answers a narrower question: what percentage of participants cross the >12% improvement threshold by 3 months? That distinction provides the central statistical lesson of the TRANSFORM design.

Clinical Biostats methodology: The statistical interpretation begins with the estimand implied by the registered endpoint, then distinguishes the effect measure, uncertainty, missing-data implications, and limitations of the design. This keeps the clinical question and the statistical question aligned.