← Clinical Trials
Emphysema Randomized FEV1 Endpoint NCT01812447

EMPROVE: Complete Statistical Analysis of the Spiration Valve System in Emphysema

An educational statistical review of the randomized EMPROVE trial evaluating the Spiration Valve System for emphysema, with emphasis on its primary effectiveness endpoint: the difference between treatment and control groups in mean change in forced expiratory volume in 1 second (FEV1).

EMPROVE  ·  NCT01812447  ·  Completed  ·  Enrollment 172
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

EMPROVE was a randomized, parallel-group clinical trial evaluating the Spiration Valve System for emphysema. The registry lists 172 enrolled participants and 3 arms, with treatment and control interventions consisting of the Spiration Valve System and Medical Management.

172
Enrollment
Total participants
3
Arms
Parallel design
6 mo
Primary endpoint
Baseline to 6 months
2013–2017
Trial period
Start to primary completion
FeatureEMPROVE
Trial acronymEMPROVE
ClinicalTrials.gov identifierNCT01812447
Official titleEvaluation of the Spiration® Valve System for Emphysema to Improve Lung Function
ConditionEmphysema
StatusCOMPLETED
PhaseNA
AllocationRANDOMIZED
Design modelPARALLEL
MaskingNONE
Primary purposeTREATMENT
Enrollment172
Number of arms3
Lead sponsorOlympus Corporation of the Americas
Sponsor typeINDUSTRY

2. Clinical Question

The central statistical question is whether the treatment and control groups differ in the mean change in forced expiratory volume in 1 second (FEV1) from baseline to 6 months.

Population

Participants enrolled in the EMPROVE trial with the registry-listed condition of emphysema.

Intervention

Spiration Valve System, classified in the registry as a device intervention.

Comparator

Medical Management, classified in the registry as an other intervention.

Primary question

What is the difference between the treatment and control groups in mean change in FEV1 from baseline to 6 months?

3. Trial Design

The registry describes EMPROVE as a randomized, parallel-group trial with no masking. It has 3 arms and enrolled 172 participants. The primary purpose is listed as treatment.

01
Enroll 172 participants
02
Randomize 3-arm randomized design
03
Treat Spiration Valve System or Medical Management
04
Measure FEV1 at baseline and 6 months
05
Compare Mean change between groups
Allocation
RANDOMIZED
Design model
PARALLEL
Masking
NONE
Primary purpose
TREATMENT

Three-arm structure

The registry identifies 3 study arms, while the registered primary effectiveness endpoint is expressed as a difference between the treatment and control groups. The registry information presented here does not provide arm-level enrollment counts or additional arm-specific descriptions beyond the listed interventions.

INTERVENTION

Spiration Valve System

  • Intervention type: device
  • Evaluated for emphysema
  • Primary endpoint measured using change in FEV1
CONTROL

Medical Management

  • Intervention type: other
  • Comparator in the registered primary effectiveness endpoint
  • Primary endpoint measured using change in FEV1
Interpretation of the design: randomization establishes the basis for comparing treatment groups, while the parallel structure means participants are assigned to study groups rather than sequentially receiving each intervention. Because masking is listed as none, the trial was not described in the registry as a masked trial.

4. Endpoints

The registry lists one registered primary endpoint. It is a continuous pulmonary-function outcome defined using the change in FEV1 from baseline to 6 months.

EndpointRegistry definitionTime frame
Primary effectiveness endpoint The primary effectiveness endpoint will be the difference between the treatment and control groups in the mean change in forced expiratory volume in 1 second (FEV1) Baseline and 6 Months

What the endpoint measures

FEV1 is a continuous measurement of forced expiratory volume in 1 second. For this trial, the registered endpoint is not simply the FEV1 value at 6 months. It is the change from baseline, followed by a comparison of the mean change between the treatment and control groups.

Conceptual treatment effect
Difference in mean change = Mean(ΔFEV1treatment) − Mean(ΔFEV1control)

Here, ΔFEV1 represents the participant's change between the baseline and 6-month measurements. The treatment effect therefore compares average within-participant changes across randomized groups.

5. Statistical Methodology

Primary estimand

The registered endpoint naturally defines a continuous between-group contrast: the difference in mean change in FEV1 between the treatment and control groups. This is an absolute difference on the FEV1 measurement scale rather than a ratio, hazard ratio, or odds ratio.

For a participant with baseline FEV1 denoted by \(Y_0\) and 6-month FEV1 denoted by \(Y_6\), the individual change can be written conceptually as:

Change from baseline
ΔY = Y6 − Y0

The group-level comparison then asks whether the mean value of ΔY differs between randomized treatment groups.

Difference in means

A straightforward analysis of the registered endpoint would compare the average change in FEV1 between the treatment and control groups. The resulting estimate has the same units as FEV1 and directly describes the absolute separation in mean change.

Generic two-group contrast
δ = \bar{ΔY}T − \bar{ΔY}C

A positive δ would indicate a larger mean increase in FEV1 in the treatment group relative to control; a negative δ would indicate a smaller mean change. The direction should always be interpreted using the actual FEV1 coding and analysis specification.

ANCOVA as a common analysis of change-oriented continuous endpoints

For randomized trials with a continuous outcome measured at baseline and follow-up, an analysis of covariance (ANCOVA) can model the follow-up FEV1 while adjusting for baseline FEV1 and treatment group. This approach can improve precision when baseline measurements explain meaningful variation in the follow-up measurement.

Conceptually, such a model can be represented as:

Illustrative ANCOVA structure
Y6 = β0 + β1Treatment + β2Y0 + ε

The coefficient associated with treatment represents the adjusted difference in mean 6-month FEV1 between randomized groups, conditional on baseline FEV1. The registry record presented here does not specify that ANCOVA was the analysis method used.

Analysis of change scores

Another direct approach is to calculate each participant's change from baseline and compare the resulting change scores between groups. The statistical target is then exactly aligned with the registry's wording: the difference in mean change.

Confidence intervals

A confidence interval around the difference in mean change would quantify statistical uncertainty around the estimated treatment effect. Its width depends on factors including variability in FEV1 changes, sample size, and the statistical model used.

Hypothesis testing

For a conventional two-group comparison, the null hypothesis would be that the population mean change is the same in the treatment and control groups. A formal test would assess whether the observed difference is compatible with that null hypothesis under the selected statistical model.

Important distinction: the registry defines the primary endpoint but does not provide a posted statistical analysis for it. The choice between an analysis of change scores, ANCOVA, or another prespecified method should therefore be taken from the applicable statistical analysis plan when that document is available.

6. Statistical Methods Explained

What does "mean change from baseline" mean?

Each participant has a baseline FEV1 measurement and a measurement at 6 months. The participant-level change is calculated first. The mean change is then the average of those individual changes within a study group.

Why compare change rather than only the 6-month FEV1?

Change from baseline incorporates each participant's starting value. Two groups could have similar 6-month averages while differing in how much their FEV1 changed from their respective baselines. A change endpoint directly targets improvement or deterioration over the specified period.

Why might baseline adjustment improve precision?

FEV1 at baseline can explain some of the variation in FEV1 at follow-up. An ANCOVA-style analysis uses that baseline information to account for between-participant variation while estimating the treatment-group difference at 6 months. Greater explanatory power can produce a more precise estimate than an unadjusted comparison when the modeling assumptions are appropriate.

What does a difference in mean change tell us?

It describes the average separation in FEV1 change between randomized groups. Unlike a hazard ratio, it is not a relative rate of an event over time. Unlike a percentage change, it remains on the original FEV1 measurement scale.

Why does randomization matter for this endpoint?

Randomization creates the basis for comparing groups under the trial design. In expectation, treatment assignment is independent of measured and unmeasured baseline factors, so differences in outcomes can be interpreted within the causal framework established by the randomized comparison, subject to adherence, missing data, post-randomization events, and other design considerations.

What does an unmasked trial change statistically?

Because the registry lists masking as none, participants and investigators were not described as masked. For an objective physiological measurement such as FEV1, the measurement itself can be less subjective than an outcome based on symptoms or investigator judgment, but lack of masking can still affect behavior, treatment decisions, follow-up, adherence, and other aspects of trial conduct.

7. Missing Data and Analysis Considerations

The primary endpoint requires FEV1 measurements at baseline and 6 months. Consequently, missing measurements can affect the analysis because an individual change score cannot be calculated directly when one of the required measurements is unavailable.

Complete paired measurements

Participants with both required measurements can contribute directly to a change-from-baseline analysis.

Missing follow-up

A missing 6-month FEV1 prevents direct calculation of the participant's observed 6-month change without an additional modeling or imputation strategy.

Imputation

Methods such as multiple imputation can address missing continuous outcomes under explicit assumptions, but the registry information does not specify an imputation method.

Model assumptions

Any regression-based analysis depends on assumptions concerning the outcome distribution, covariance structure, and relationship between baseline and follow-up measurements.

Missing-data handling is particularly important because the estimand concerns a comparison of mean change at a specific follow-up time. The assumptions used to handle missing FEV1 measurements can influence both the estimate and its uncertainty.

8. Statistical Interpretation of the Primary Endpoint

Effect estimate

The primary treatment effect is the difference between the treatment and control groups in mean change in FEV1. This is an absolute between-group difference in a continuous pulmonary-function measure.

A positive estimated difference would mean that the treatment group's average change was greater than the control group's average change, given the direction of the contrast. It would not mean that every participant improved by that amount.

What the estimate does not mean

A mean difference is a group-level summary. It does not describe the distribution of individual responses, the proportion of participants who improved, or whether every participant experienced the same direction or magnitude of change.

Why the confidence interval matters

If a confidence interval were reported around the mean difference, it would describe uncertainty around the estimated group-level treatment effect under the specified statistical framework. A narrow interval would indicate greater precision; a wider interval would indicate greater uncertainty.

Why a p-value does not measure effect size

A p-value addresses compatibility with a specified null hypothesis under the statistical model. It does not tell the reader how large the treatment effect is. Effect size, confidence interval, and clinical context should be considered separately.

9. Statistical Issues Specific to a Three-Arm Trial

EMPROVE has 3 arms, while the registry's primary effectiveness endpoint is stated as a difference between treatment and control groups. This distinction matters because a multi-arm trial can generate more than one possible pairwise comparison.

IssueStatistical implication
Three randomized arms More than one treatment-group comparison can potentially be defined within the overall randomized design.
Registered primary contrast The primary endpoint specifically describes a difference between the treatment and control groups.
Multiplicity If multiple pairwise comparisons are formally tested, the prespecified error-control strategy becomes important.
Interpretation A result for the registered treatment-versus-control contrast should not automatically be generalized to every possible pairwise comparison.

The registry information presented here does not specify a multiplicity adjustment, alpha allocation, or hierarchical testing procedure for the 3-arm structure. Those details would determine how multiple formal comparisons should be interpreted.

10. Statistical Issues Specific to an Unmasked Device Trial

The registry lists masking as none. This creates a different set of considerations from a double-blind pharmacologic trial.

11. Trial Timeline

2013-06

Trial start

The registry lists June 2013 as the study start date.

2013–2017

Randomized clinical trial period

The study was conducted as a randomized, parallel-group trial evaluating the Spiration Valve System and Medical Management in emphysema.

Baseline and 6 Months

Primary endpoint assessment window

The registered primary effectiveness endpoint compares treatment and control groups according to mean change in FEV1 from baseline to 6 months.

2017-11

Primary completion

The registry lists November 2017 as the primary completion date.

12. Trial Status and Sponsorship

Registry characteristicEMPROVE
StatusCOMPLETED
Lead sponsorOlympus Corporation of the Americas
Sponsor typeINDUSTRY
Start date2013-06
Primary completion date2017-11

Sponsorship is a descriptive feature of the trial record. It does not, by itself, determine the validity of an endpoint estimate. Statistical interpretation should instead consider the randomized design, endpoint definition, analysis population, missing-data handling, prespecified analysis, and precision of the resulting estimates.

13. Planned Analysis

The registry's primary effectiveness endpoint is the difference between the treatment and control groups in the mean change in FEV1 from baseline to 6 months. This is a continuous outcome and would typically be analyzed by comparing mean changes between randomized groups, with an appropriate estimate of uncertainty and a prespecified hypothesis test.

Typical analysis framework
Treatment effect = Mean change in FEV1T − Mean change in FEV1C

A confidence interval would quantify uncertainty around this difference, while a hypothesis test could assess the null hypothesis of no between-group difference.

Possible model-based approaches

For a baseline-and-follow-up continuous endpoint, common approaches include direct comparison of change scores and ANCOVA of the follow-up measurement with baseline FEV1 as a covariate. The latter can be particularly useful when baseline FEV1 is strongly associated with the follow-up measurement.

If repeated measurements beyond the primary 6-month endpoint were incorporated into a longitudinal analysis, a mixed-effects model could account for within-participant correlation. However, the registry endpoint specified here is defined at baseline and 6 months, and the registry information does not specify such a longitudinal model.

Results status: no formal statistical analyses are posted in the ClinicalTrials.gov record used for this analysis. The registry therefore identifies the primary effectiveness endpoint and its measurement window, but it does not provide an estimate, confidence interval, or p-value for that endpoint in the record described here.

14. Limitations

15. Why This Trial Matters Statistically

EMPROVE is a useful statistical teaching example because its primary endpoint is a continuous change-from-baseline measure rather than a time-to-event endpoint. The design therefore illustrates a different branch of clinical-trial methodology from trials whose primary outcomes are survival, progression, or binary response.

ConceptHow it appears in EMPROVE
Randomization The trial is listed as RANDOMIZED, providing the framework for a treatment-group comparison.
Parallel-group design The registry lists PARALLEL as the design model.
Three-arm design The registry lists 3 study arms.
Continuous endpoint The primary effectiveness endpoint is mean change in FEV1.
Change from baseline The primary endpoint compares change between baseline and 6 months.
Difference in means The treatment effect is naturally expressed as a between-group difference in mean change.
Baseline adjustment ANCOVA is a common framework for continuous follow-up outcomes when baseline measurements are available.
Missing data Loss of a baseline or 6-month FEV1 measurement can affect calculation of individual change scores.
Multiplicity A 3-arm design can create multiple potential pairwise comparisons, making prespecified error control important.
Unmasked design The registry lists masking as NONE, creating potential sources of bias that should be considered separately from randomization.

16. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The registered primary endpoint is a between-group comparison of mean change in FEV1 from baseline to 6 months. Its natural effect measure is an absolute difference in mean change.

Clinical interpretation

The clinical meaning of an observed FEV1 difference depends on the magnitude of the change, variability among participants, precision of the estimate, and the clinical context in which the pulmonary-function measurement is interpreted.

These two perspectives should remain distinct. Statistical significance, if formally established, would address evidence against a null hypothesis; clinical interpretation asks whether the magnitude and distribution of the observed change are meaningful for patients and clinicians.

17. Understanding the Primary Endpoint Without a Reported Result

Because the registry record does not post a numerical estimate for the primary endpoint, the most useful statistical interpretation is to understand exactly what an eventual result would represent.

If the difference is positive

Under the stated treatment-minus-control direction, a positive difference would indicate a larger average increase in FEV1 in the treatment group.

If the difference is near zero

A difference near zero would indicate similar average changes between groups, although the confidence interval would be needed to distinguish a precise estimate near zero from a highly uncertain estimate.

If the difference is negative

Under the same direction, a negative difference would indicate a smaller average change in FEV1 in the treatment group.

If the interval is wide

A wide confidence interval would indicate limited precision and would make the magnitude of the treatment effect more uncertain even if the point estimate appears substantial.

18. What Would Make the Statistical Result Complete?

A complete report of the primary effectiveness analysis would ordinarily allow the reader to identify at least four components of the treatment comparison:

  1. Effect estimate: the difference in mean change in FEV1 between treatment and control.
  2. Uncertainty: a confidence interval around that difference.
  3. Hypothesis test: a p-value or equivalent inferential result under the prespecified testing framework.
  4. Analysis population and method: the participants included in the analysis and the statistical model used to obtain the estimate.

These components answer different questions. The effect estimate describes magnitude, the confidence interval describes precision, the p-value addresses compatibility with a null hypothesis, and the analysis population and model establish how the numerical result was obtained.

Interpretation principle: a statistical result should be read as a package rather than as a single number. For a continuous endpoint such as mean change in FEV1, the treatment difference and its confidence interval are especially important because they preserve information about both magnitude and uncertainty.

19. Sources

Continue through the Clinical Biostats trial library

Explore additional clinical trial analyses and statistical methods across the Clinical Biostats site.

20. Record Summary

EMPROVE is a completed randomized, parallel-group trial of the Spiration Valve System in emphysema, with 172 participants and 3 study arms. Its registered primary effectiveness endpoint is the difference between treatment and control groups in the mean change in FEV1 from baseline to 6 months.

Statistically, this endpoint is a continuous between-group treatment contrast. The most direct interpretation is an absolute difference in average change in pulmonary function. Appropriate analysis would ordinarily require a prespecified method for estimating that difference, quantifying its uncertainty, and conducting the corresponding hypothesis test. Baseline adjustment, missing-data handling, and the treatment of multiple comparisons are important design and analysis considerations.

The registry identifies the trial as randomized and parallel, with no masking, and lists Olympus Corporation of the Americas as the lead sponsor. The study began in 2013-06 and reached primary completion in 2017-11. The ClinicalTrials.gov record does not post a formal statistical analysis or numerical result for the primary effectiveness endpoint.

Statistical interpretation: For a randomized trial with a continuous baseline-and-follow-up endpoint, the central statistical question is not simply whether the follow-up measurements differ. It is whether the randomized treatment groups differ in their mean change from baseline, how precisely that difference is estimated, and whether the analysis appropriately accounts for the trial's design and missing-data structure.