This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
EMPROVE was a randomized, parallel-group clinical trial evaluating the Spiration Valve System for emphysema. The registry lists 172 enrolled participants and 3 arms, with treatment and control interventions consisting of the Spiration Valve System and Medical Management.
| Feature | EMPROVE |
|---|---|
| Trial acronym | EMPROVE |
| ClinicalTrials.gov identifier | NCT01812447 |
| Official title | Evaluation of the Spiration® Valve System for Emphysema to Improve Lung Function |
| Condition | Emphysema |
| Status | COMPLETED |
| Phase | NA |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | NONE |
| Primary purpose | TREATMENT |
| Enrollment | 172 |
| Number of arms | 3 |
| Lead sponsor | Olympus Corporation of the Americas |
| Sponsor type | INDUSTRY |
2. Clinical Question
The central statistical question is whether the treatment and control groups differ in the mean change in forced expiratory volume in 1 second (FEV1) from baseline to 6 months.
Population
Participants enrolled in the EMPROVE trial with the registry-listed condition of emphysema.
Intervention
Spiration Valve System, classified in the registry as a device intervention.
Comparator
Medical Management, classified in the registry as an other intervention.
Primary question
What is the difference between the treatment and control groups in mean change in FEV1 from baseline to 6 months?
3. Trial Design
The registry describes EMPROVE as a randomized, parallel-group trial with no masking. It has 3 arms and enrolled 172 participants. The primary purpose is listed as treatment.
Three-arm structure
The registry identifies 3 study arms, while the registered primary effectiveness endpoint is expressed as a difference between the treatment and control groups. The registry information presented here does not provide arm-level enrollment counts or additional arm-specific descriptions beyond the listed interventions.
Spiration Valve System
- Intervention type: device
- Evaluated for emphysema
- Primary endpoint measured using change in FEV1
Medical Management
- Intervention type: other
- Comparator in the registered primary effectiveness endpoint
- Primary endpoint measured using change in FEV1
4. Endpoints
The registry lists one registered primary endpoint. It is a continuous pulmonary-function outcome defined using the change in FEV1 from baseline to 6 months.
| Endpoint | Registry definition | Time frame |
|---|---|---|
| Primary effectiveness endpoint | The primary effectiveness endpoint will be the difference between the treatment and control groups in the mean change in forced expiratory volume in 1 second (FEV1) | Baseline and 6 Months |
What the endpoint measures
FEV1 is a continuous measurement of forced expiratory volume in 1 second. For this trial, the registered endpoint is not simply the FEV1 value at 6 months. It is the change from baseline, followed by a comparison of the mean change between the treatment and control groups.
Here, ΔFEV1 represents the participant's change between the baseline and 6-month measurements. The treatment effect therefore compares average within-participant changes across randomized groups.
5. Statistical Methodology
Primary estimand
The registered endpoint naturally defines a continuous between-group contrast: the difference in mean change in FEV1 between the treatment and control groups. This is an absolute difference on the FEV1 measurement scale rather than a ratio, hazard ratio, or odds ratio.
For a participant with baseline FEV1 denoted by \(Y_0\) and 6-month FEV1 denoted by \(Y_6\), the individual change can be written conceptually as:
The group-level comparison then asks whether the mean value of ΔY differs between randomized treatment groups.
Difference in means
A straightforward analysis of the registered endpoint would compare the average change in FEV1 between the treatment and control groups. The resulting estimate has the same units as FEV1 and directly describes the absolute separation in mean change.
A positive δ would indicate a larger mean increase in FEV1 in the treatment group relative to control; a negative δ would indicate a smaller mean change. The direction should always be interpreted using the actual FEV1 coding and analysis specification.
ANCOVA as a common analysis of change-oriented continuous endpoints
For randomized trials with a continuous outcome measured at baseline and follow-up, an analysis of covariance (ANCOVA) can model the follow-up FEV1 while adjusting for baseline FEV1 and treatment group. This approach can improve precision when baseline measurements explain meaningful variation in the follow-up measurement.
Conceptually, such a model can be represented as:
The coefficient associated with treatment represents the adjusted difference in mean 6-month FEV1 between randomized groups, conditional on baseline FEV1. The registry record presented here does not specify that ANCOVA was the analysis method used.
Analysis of change scores
Another direct approach is to calculate each participant's change from baseline and compare the resulting change scores between groups. The statistical target is then exactly aligned with the registry's wording: the difference in mean change.
Confidence intervals
A confidence interval around the difference in mean change would quantify statistical uncertainty around the estimated treatment effect. Its width depends on factors including variability in FEV1 changes, sample size, and the statistical model used.
Hypothesis testing
For a conventional two-group comparison, the null hypothesis would be that the population mean change is the same in the treatment and control groups. A formal test would assess whether the observed difference is compatible with that null hypothesis under the selected statistical model.
6. Statistical Methods Explained
What does "mean change from baseline" mean?
Each participant has a baseline FEV1 measurement and a measurement at 6 months. The participant-level change is calculated first. The mean change is then the average of those individual changes within a study group.
Why compare change rather than only the 6-month FEV1?
Change from baseline incorporates each participant's starting value. Two groups could have similar 6-month averages while differing in how much their FEV1 changed from their respective baselines. A change endpoint directly targets improvement or deterioration over the specified period.
Why might baseline adjustment improve precision?
FEV1 at baseline can explain some of the variation in FEV1 at follow-up. An ANCOVA-style analysis uses that baseline information to account for between-participant variation while estimating the treatment-group difference at 6 months. Greater explanatory power can produce a more precise estimate than an unadjusted comparison when the modeling assumptions are appropriate.
What does a difference in mean change tell us?
It describes the average separation in FEV1 change between randomized groups. Unlike a hazard ratio, it is not a relative rate of an event over time. Unlike a percentage change, it remains on the original FEV1 measurement scale.
Why does randomization matter for this endpoint?
Randomization creates the basis for comparing groups under the trial design. In expectation, treatment assignment is independent of measured and unmeasured baseline factors, so differences in outcomes can be interpreted within the causal framework established by the randomized comparison, subject to adherence, missing data, post-randomization events, and other design considerations.
What does an unmasked trial change statistically?
Because the registry lists masking as none, participants and investigators were not described as masked. For an objective physiological measurement such as FEV1, the measurement itself can be less subjective than an outcome based on symptoms or investigator judgment, but lack of masking can still affect behavior, treatment decisions, follow-up, adherence, and other aspects of trial conduct.
7. Missing Data and Analysis Considerations
The primary endpoint requires FEV1 measurements at baseline and 6 months. Consequently, missing measurements can affect the analysis because an individual change score cannot be calculated directly when one of the required measurements is unavailable.
Complete paired measurements
Participants with both required measurements can contribute directly to a change-from-baseline analysis.
Missing follow-up
A missing 6-month FEV1 prevents direct calculation of the participant's observed 6-month change without an additional modeling or imputation strategy.
Imputation
Methods such as multiple imputation can address missing continuous outcomes under explicit assumptions, but the registry information does not specify an imputation method.
Model assumptions
Any regression-based analysis depends on assumptions concerning the outcome distribution, covariance structure, and relationship between baseline and follow-up measurements.
Missing-data handling is particularly important because the estimand concerns a comparison of mean change at a specific follow-up time. The assumptions used to handle missing FEV1 measurements can influence both the estimate and its uncertainty.
8. Statistical Interpretation of the Primary Endpoint
The primary treatment effect is the difference between the treatment and control groups in mean change in FEV1. This is an absolute between-group difference in a continuous pulmonary-function measure.
A positive estimated difference would mean that the treatment group's average change was greater than the control group's average change, given the direction of the contrast. It would not mean that every participant improved by that amount.
A mean difference is a group-level summary. It does not describe the distribution of individual responses, the proportion of participants who improved, or whether every participant experienced the same direction or magnitude of change.
If a confidence interval were reported around the mean difference, it would describe uncertainty around the estimated group-level treatment effect under the specified statistical framework. A narrow interval would indicate greater precision; a wider interval would indicate greater uncertainty.
A p-value addresses compatibility with a specified null hypothesis under the statistical model. It does not tell the reader how large the treatment effect is. Effect size, confidence interval, and clinical context should be considered separately.
9. Statistical Issues Specific to a Three-Arm Trial
EMPROVE has 3 arms, while the registry's primary effectiveness endpoint is stated as a difference between treatment and control groups. This distinction matters because a multi-arm trial can generate more than one possible pairwise comparison.
| Issue | Statistical implication |
|---|---|
| Three randomized arms | More than one treatment-group comparison can potentially be defined within the overall randomized design. |
| Registered primary contrast | The primary endpoint specifically describes a difference between the treatment and control groups. |
| Multiplicity | If multiple pairwise comparisons are formally tested, the prespecified error-control strategy becomes important. |
| Interpretation | A result for the registered treatment-versus-control contrast should not automatically be generalized to every possible pairwise comparison. |
The registry information presented here does not specify a multiplicity adjustment, alpha allocation, or hierarchical testing procedure for the 3-arm structure. Those details would determine how multiple formal comparisons should be interpreted.
10. Statistical Issues Specific to an Unmasked Device Trial
The registry lists masking as none. This creates a different set of considerations from a double-blind pharmacologic trial.
- Objective measurement: FEV1 is a quantitative physiological endpoint, which can reduce some forms of subjective outcome assessment.
- Behavior and adherence: awareness of treatment assignment can influence participant behavior or engagement with follow-up procedures.
- Clinical management: investigators who know treatment assignment may make subsequent clinical decisions differently across groups.
- Assessment procedures: standardized pulmonary-function testing remains important because measurement variability can affect the estimated mean change.
- Interpretation: the randomized design remains central to the comparison, but the absence of masking should be considered when evaluating potential sources of bias.
11. Trial Timeline
Trial start
The registry lists June 2013 as the study start date.
Randomized clinical trial period
The study was conducted as a randomized, parallel-group trial evaluating the Spiration Valve System and Medical Management in emphysema.
Primary endpoint assessment window
The registered primary effectiveness endpoint compares treatment and control groups according to mean change in FEV1 from baseline to 6 months.
Primary completion
The registry lists November 2017 as the primary completion date.
12. Trial Status and Sponsorship
| Registry characteristic | EMPROVE |
|---|---|
| Status | COMPLETED |
| Lead sponsor | Olympus Corporation of the Americas |
| Sponsor type | INDUSTRY |
| Start date | 2013-06 |
| Primary completion date | 2017-11 |
Sponsorship is a descriptive feature of the trial record. It does not, by itself, determine the validity of an endpoint estimate. Statistical interpretation should instead consider the randomized design, endpoint definition, analysis population, missing-data handling, prespecified analysis, and precision of the resulting estimates.
13. Planned Analysis
The registry's primary effectiveness endpoint is the difference between the treatment and control groups in the mean change in FEV1 from baseline to 6 months. This is a continuous outcome and would typically be analyzed by comparing mean changes between randomized groups, with an appropriate estimate of uncertainty and a prespecified hypothesis test.
A confidence interval would quantify uncertainty around this difference, while a hypothesis test could assess the null hypothesis of no between-group difference.
Possible model-based approaches
For a baseline-and-follow-up continuous endpoint, common approaches include direct comparison of change scores and ANCOVA of the follow-up measurement with baseline FEV1 as a covariate. The latter can be particularly useful when baseline FEV1 is strongly associated with the follow-up measurement.
If repeated measurements beyond the primary 6-month endpoint were incorporated into a longitudinal analysis, a mixed-effects model could account for within-participant correlation. However, the registry endpoint specified here is defined at baseline and 6 months, and the registry information does not specify such a longitudinal model.
14. Limitations
- No posted statistical analysis: the ClinicalTrials.gov record does not provide a formal statistical analysis for the registered primary effectiveness endpoint.
- No posted effect estimate: the registry information does not provide the numerical difference in mean change in FEV1 between treatment and control.
- No confidence interval reported: the registry information does not provide a confidence interval for the primary treatment effect.
- No p-value reported: the registry information does not provide a formal hypothesis-test result for the primary endpoint.
- Three-arm structure: the trial includes 3 arms, but the registered primary endpoint describes the treatment-versus-control comparison without providing arm-level enrollment counts in the record summarized here.
- Unmasked design: masking is listed as none, which can introduce sources of bias unrelated to the statistical comparison itself.
- Missing-data uncertainty: the primary endpoint requires baseline and 6-month FEV1 measurements, while the registry information does not specify a missing-data or imputation strategy.
- Analysis-model uncertainty: the registry defines the endpoint but does not specify the statistical model used to estimate the treatment effect.
- Limited endpoint information: the registry information presented here identifies one primary effectiveness endpoint and does not provide additional outcome estimates for interpretation.
15. Why This Trial Matters Statistically
EMPROVE is a useful statistical teaching example because its primary endpoint is a continuous change-from-baseline measure rather than a time-to-event endpoint. The design therefore illustrates a different branch of clinical-trial methodology from trials whose primary outcomes are survival, progression, or binary response.
| Concept | How it appears in EMPROVE |
|---|---|
| Randomization | The trial is listed as RANDOMIZED, providing the framework for a treatment-group comparison. |
| Parallel-group design | The registry lists PARALLEL as the design model. |
| Three-arm design | The registry lists 3 study arms. |
| Continuous endpoint | The primary effectiveness endpoint is mean change in FEV1. |
| Change from baseline | The primary endpoint compares change between baseline and 6 months. |
| Difference in means | The treatment effect is naturally expressed as a between-group difference in mean change. |
| Baseline adjustment | ANCOVA is a common framework for continuous follow-up outcomes when baseline measurements are available. |
| Missing data | Loss of a baseline or 6-month FEV1 measurement can affect calculation of individual change scores. |
| Multiplicity | A 3-arm design can create multiple potential pairwise comparisons, making prespecified error control important. |
| Unmasked design | The registry lists masking as NONE, creating potential sources of bias that should be considered separately from randomization. |
16. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The registered primary endpoint is a between-group comparison of mean change in FEV1 from baseline to 6 months. Its natural effect measure is an absolute difference in mean change.
Clinical interpretation
The clinical meaning of an observed FEV1 difference depends on the magnitude of the change, variability among participants, precision of the estimate, and the clinical context in which the pulmonary-function measurement is interpreted.
These two perspectives should remain distinct. Statistical significance, if formally established, would address evidence against a null hypothesis; clinical interpretation asks whether the magnitude and distribution of the observed change are meaningful for patients and clinicians.
17. Understanding the Primary Endpoint Without a Reported Result
Because the registry record does not post a numerical estimate for the primary endpoint, the most useful statistical interpretation is to understand exactly what an eventual result would represent.
If the difference is positive
Under the stated treatment-minus-control direction, a positive difference would indicate a larger average increase in FEV1 in the treatment group.
If the difference is near zero
A difference near zero would indicate similar average changes between groups, although the confidence interval would be needed to distinguish a precise estimate near zero from a highly uncertain estimate.
If the difference is negative
Under the same direction, a negative difference would indicate a smaller average change in FEV1 in the treatment group.
If the interval is wide
A wide confidence interval would indicate limited precision and would make the magnitude of the treatment effect more uncertain even if the point estimate appears substantial.
18. What Would Make the Statistical Result Complete?
A complete report of the primary effectiveness analysis would ordinarily allow the reader to identify at least four components of the treatment comparison:
- Effect estimate: the difference in mean change in FEV1 between treatment and control.
- Uncertainty: a confidence interval around that difference.
- Hypothesis test: a p-value or equivalent inferential result under the prespecified testing framework.
- Analysis population and method: the participants included in the analysis and the statistical model used to obtain the estimate.
These components answer different questions. The effect estimate describes magnitude, the confidence interval describes precision, the p-value addresses compatibility with a null hypothesis, and the analysis population and model establish how the numerical result was obtained.
19. Sources
- ClinicalTrials.gov: EMPROVE (NCT01812447), Evaluation of the Spiration® Valve System for Emphysema to Improve Lung Function.
Continue through the Clinical Biostats trial library
Explore additional clinical trial analyses and statistical methods across the Clinical Biostats site.
20. Record Summary
EMPROVE is a completed randomized, parallel-group trial of the Spiration Valve System in emphysema, with 172 participants and 3 study arms. Its registered primary effectiveness endpoint is the difference between treatment and control groups in the mean change in FEV1 from baseline to 6 months.
Statistically, this endpoint is a continuous between-group treatment contrast. The most direct interpretation is an absolute difference in average change in pulmonary function. Appropriate analysis would ordinarily require a prespecified method for estimating that difference, quantifying its uncertainty, and conducting the corresponding hypothesis test. Baseline adjustment, missing-data handling, and the treatment of multiple comparisons are important design and analysis considerations.
The registry identifies the trial as randomized and parallel, with no masking, and lists Olympus Corporation of the Americas as the lead sponsor. The study began in 2013-06 and reached primary completion in 2017-11. The ClinicalTrials.gov record does not post a formal statistical analysis or numerical result for the primary effectiveness endpoint.