← Clinical Trials
Pulmonary Arterial Hypertension Phase 3 Time-to-Event Analysis NCT01560624

FREEDOM-EV: Complete Statistical Analysis of Treprostinil Diolamine in Pulmonary Arterial Hypertension

An independent statistical review of the randomized phase 3 FREEDOM-EV trial evaluating treprostinil diolamine (UT-15C) versus placebo in subjects with pulmonary arterial hypertension receiving background oral monotherapy.

Trial status: COMPLETED  ·  Enrollment: 690  ·  Primary completion: June 24, 2018
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results presented here are restricted to the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for FREEDOM-EV. ClinicalTrials.gov provides the official trial registry record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

FREEDOM-EV was a randomized, parallel, quadruple-masked phase 3 trial comparing treprostinil diolamine (UT-15C) with placebo in subjects with pulmonary arterial hypertension receiving background oral monotherapy. The primary endpoint was time to first clinical worsening event, analyzed from randomization to approximately 4 years.

690
Enrollment
Randomized phase 3 trial
2
Treatment arms
UT-15C vs placebo
0.74
Primary HR
95% CI 0.56–0.97
0.0275
Primary P-value
Cox proportional-hazards model
FeatureFREEDOM-EV
Trial nameFREEDOM-EV
PhasePhase 3
ConditionPulmonary Arterial Hypertension
DesignRandomized, parallel, quadruple-masked
Primary purposeTreatment
Enrollment690
InterventionsTreprostinil Diolamine (drug); Placebo (drug)
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Lead sponsorUnited Therapeutics
Trial statusCOMPLETED
ClinicalTrials.govNCT01560624

2. Clinical Question

The primary statistical question was whether treprostinil diolamine, compared with placebo, changed the time from randomization to the first clinical worsening event in subjects with pulmonary arterial hypertension receiving background oral monotherapy.

Population

Subjects with pulmonary arterial hypertension receiving background oral monotherapy.

Intervention

Treprostinil Diolamine (UT-15C).

Comparator

Placebo.

Primary question

Does UT-15C improve the time to first clinical worsening event compared with placebo?

3. Trial Design

01
Randomize690 subjects
02
Parallel armsUT-15C vs placebo
03
Quadruple maskingMasked trial design
04
FollowFrom randomization
05
AssessClinical worsening and secondary outcomes
Allocation
RANDOMIZED
Design model
PARALLEL
Masking
QUADRUPLE
Primary purpose
TREATMENT
ARM A

Treprostinil diolamine

  • Treprostinil Diolamine (UT-15C)
  • Drug intervention
  • Compared with placebo
ARM B

Placebo

  • Placebo drug intervention
  • Comparator arm
  • Compared with UT-15C

The trial was completed after starting on June 26, 2012, with primary completion on June 24, 2018. The registry reports an enrollment of 690 subjects across two treatment arms.

4. Endpoints

The registered primary endpoint was a time-to-event outcome assessed from randomization to approximately 4 years.

EndpointRegistry definition / time frameEndpoint type
Time to First Clinical Worsening Event From randomization to approximately 4 years Time-to-event

Primary endpoint definition

Clinical worsening was assessed continuously from randomization until the subject's last study visit. Clinical worsening events were defined as death (all causes), hospitalizations due to worsening pulmonary arterial hypertension (PAH), initiation of an inhaled or infused prostacyclin (PGI2) for the treatment of worsening PAH, disease progression, or unsatisfactory long-term clinical response.

Secondary endpoints with posted analyses

Secondary endpointTime frameOutcome unitMethod
Change in 6-Minute Walk Distance From Baseline to Week 24 meters ANCOVA
Change in Plasma N-Terminal Pro-brain Natriuretic Peptide (NT-proBNP) From Baseline to Week 24 From Baseline to Week 24 Ratio to baseline ANCOVA
Change in World Health Organization Functional Class (WHO FC) From Baseline to Week 48 Baseline to Week 48 Participants Fisher exact test

5. Statistical Results

Primary endpoint: Time to First Clinical Worsening Event

The primary endpoint was analyzed as a time-to-event outcome using a Cox proportional-hazards model and a log-rank test. Both analyses compared UT-15C with placebo and used a superiority hypothesis.

Cox proportional-hazards analysis

HR 0.74

95% CI: 0.56–0.97   ·   P = 0.0275

Comparison: UT-15C vs Placebo

Clinical Biostats interpretation

An HR of 0.74 means that, under the fitted Cox proportional-hazards model, the estimated instantaneous rate of experiencing the first clinical worsening event in the UT-15C group was approximately 26% lower than in the placebo group.

The hazard ratio does not mean that 26% of patients avoided clinical worsening, nor does it mean that every individual patient had exactly a 26% reduction in risk. It is a relative time-to-event measure based on the observed follow-up and model.

The 95% CI of 0.56–0.97 describes uncertainty around the estimated hazard ratio. Because the interval remains below 1, the reported interval is consistent with a lower estimated hazard of clinical worsening for UT-15C relative to placebo under this analysis. It does not describe the range of individual patient effects.

The P = 0.0275 addresses the statistical evidence against the null hypothesis represented by the model comparison. A p-value is not a measure of effect size and does not indicate the probability that the treatment works. Interpretation of the Cox result also depends on the proportional-hazards framework and on how censoring is handled.

Primary endpoint: log-rank analysis

Log-rank test

P = 0.0391

Comparison: UT-15C vs Placebo

Endpoint: Time to First Clinical Worsening Event

Clinical Biostats interpretation

The log-rank test provides a separate hypothesis test comparing the time-to-event distributions between UT-15C and placebo. The reported P = 0.0391 indicates that the trial registry's formal log-rank comparison produced evidence against the corresponding null hypothesis under the reported superiority framework.

The log-rank p-value does not quantify how large the treatment effect is. The magnitude of the observed relative effect is described by the Cox hazard ratio, 0.74, together with its 95% CI of 0.56–0.97.

The two primary analyses answer related but different statistical questions: the log-rank test evaluates separation between time-to-event distributions, while the Cox model provides a relative hazard estimate with a confidence interval. Neither analysis by itself establishes how an individual patient will experience clinical worsening.

Secondary endpoint: Change in 6-Minute Walk Distance

EndpointMethodEffect measureEstimate95% CIP-value
Change in 6-Minute Walk Distance from Baseline to Week 24 ANCOVA Hodges Lehmann estimate location shift 7.0 meters 0–16.0 0.0913
Statistical interpretation

The registry reports a Hodges-Lehmann estimate of location shift of 7.0 meters for the UT-15C versus placebo comparison, with a two-sided 95% CI of 0–16.0 meters and P = 0.0913.

The estimate describes the reported between-group location shift on the 6-minute walk distance outcome; it is not a hazard ratio and should not be interpreted as a relative treatment effect. The confidence interval expresses uncertainty around the estimated shift, while the p-value addresses the associated hypothesis test rather than the clinical magnitude of the shift.

Secondary endpoint: Change in Plasma NT-proBNP

EndpointMethodOutcome unitP-value
Change in Plasma N-Terminal Pro-brain Natriuretic Peptide (NT-proBNP) From Baseline to Week 24 ANCOVA Ratio to baseline <0.0001
Analysis population: Subjects who did not have a valid NT-proBNP measurement at Baseline or Week 24 were excluded from the analysis. The registry reports the ANCOVA p-value but does not provide an effect estimate or confidence interval for this posted analysis.

The P < 0.0001 provides evidence against the null hypothesis represented by the reported ANCOVA comparison. It does not, by itself, describe the magnitude of the difference between treatment groups. Because the ClinicalTrials.gov record does not provide the corresponding estimate or confidence interval, the statistical evidence and the size of the effect should be kept conceptually separate.

Secondary endpoint: Change in WHO Functional Class

EndpointMethodTime frameP-value
Change in World Health Organization Functional Class (WHO FC) From Baseline to Week 48 Fisher exact test Baseline to Week 48 0.0028

The registry reports a Fisher exact test with P = 0.0028 for the UT-15C versus placebo comparison. The ClinicalTrials.gov record does not provide a corresponding effect estimate, confidence interval, or cell counts, so the result is appropriately reported as a hypothesis-test result rather than converted into an unsupported measure of effect size.

6. Secondary Results in Context

OutcomeAnalysisReported resultWhat the statistic tells us
Time to First Clinical Worsening Event Cox proportional-hazards model HR 0.74; 95% CI 0.56–0.97; P = 0.0275 Relative time-to-event effect estimate with uncertainty interval
Time to First Clinical Worsening Event Log-rank test P = 0.0391 Hypothesis test comparing time-to-event distributions
Change in 6-Minute Walk Distance ANCOVA Hodges-Lehmann estimate location shift 7.0; 95% CI 0–16.0; P = 0.0913 Estimated between-group location shift with uncertainty
Change in NT-proBNP ANCOVA P < 0.0001 Hypothesis-test evidence; no effect estimate reported
Change in WHO Functional Class Fisher exact test P = 0.0028 Hypothesis-test evidence for the categorical comparison

This table illustrates an important statistical distinction: a trial can report several p-values across different endpoint types, but those p-values are not interchangeable. A hazard ratio summarizes a relative time-to-event effect, a location-shift estimate summarizes a continuous outcome comparison, and Fisher's exact test addresses a categorical comparison.

7. Statistical Methodology

Kaplan-Meier estimation and time-to-event endpoints

The primary endpoint is a time-to-event outcome. For this type of endpoint, Kaplan-Meier estimation is commonly used to describe the probability of remaining event-free over time while accounting for right-censored observations.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

where di is the number of events at time ti and ni is the number at risk immediately before that time.

The ClinicalTrials.gov record specifically reports a Cox proportional-hazards model and a log-rank test for the primary endpoint. Kaplan-Meier estimation is the natural descriptive companion to those methods for a time-to-event endpoint, although no Kaplan-Meier estimates are reported in the ClinicalTrials.gov record.

Cox proportional-hazards model

The Cox proportional-hazards model expresses the relative hazard between treatment groups through a hazard ratio. The reported primary estimate was 0.74 for UT-15C versus placebo.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the treatment group

A hazard ratio is a relative time-to-event measure. It is not a relative risk, an absolute risk difference, a probability of benefit, or a statement about the outcome of every individual patient.

Log-rank test

The log-rank test compares the observed pattern of events over time between treatment groups. It is particularly suited to randomized time-to-event comparisons because it incorporates the ordering of event times and accommodates censoring under its usual assumptions.

ANCOVA

ANCOVA, or analysis of covariance, is a linear-model framework used to compare a continuous outcome between groups while incorporating covariate information. In FREEDOM-EV, ANCOVA was the reported method for the change in 6-minute walk distance and the change in plasma NT-proBNP.

For a change-from-baseline outcome, the model can be viewed conceptually as comparing treatment groups after accounting for relevant baseline information. The exact covariates and model specification are not provided in the ClinicalTrials.gov record, so they are not inferred here.

Hodges-Lehmann estimate of location shift

The Hodges-Lehmann approach provides an estimate of the location shift between two distributions. Here, the registry reports a 7.0 estimate for the change in 6-minute walk distance, with a two-sided 95% CI of 0–16.0.

Fisher exact test

Fisher's exact test is designed for categorical data and evaluates the association between treatment group and categorical outcome counts using the exact distribution conditional on the relevant margins. It was the reported method for change in WHO Functional Class from baseline to Week 48.

8. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

The primary endpoint measures the time until the first clinical worsening event rather than simply whether an event occurred. A Cox model uses the timing of events and censoring information to estimate a relative hazard between the randomized groups. Its principal effect measure is the hazard ratio.

What does an HR of 0.74 mean?

Under the fitted model, an HR of 0.74 means that the estimated instantaneous rate of the first clinical worsening event in the UT-15C group was about 26% lower than in the placebo group. It does not mean that 26% of patients were protected from worsening or that each patient's risk was reduced by exactly 26%.

Why use a log-rank test as well as a Cox model?

The log-rank test provides a hypothesis test comparing the time-to-event experience of the two randomized groups, whereas the Cox model supplies an interpretable relative effect estimate and confidence interval. Reporting both gives complementary information about statistical evidence and effect magnitude.

What does the 95% confidence interval of 0.56–0.97 tell us?

The interval quantifies statistical uncertainty around the estimated hazard ratio under the model and sampling framework. It is not a prediction interval for individual patients and does not mean that 95% of patients have hazard ratios somewhere between 0.56 and 0.97.

Why does P = 0.0275 not measure treatment effect size?

A p-value describes the evidence against a specified null hypothesis under the statistical model. It depends on both the observed difference and the amount of information in the data. The hazard ratio and its confidence interval are therefore necessary to understand the magnitude and precision of the time-to-event effect.

Why is ANCOVA appropriate for the continuous secondary endpoints?

ANCOVA is a standard method for comparing continuous outcomes between randomized groups while incorporating baseline information through a linear-model framework. In this trial, the registry identifies ANCOVA as the analysis method for change in 6-minute walk distance and change in NT-proBNP.

Why is Fisher exact testing appropriate for WHO Functional Class?

The registry identifies the WHO Functional Class endpoint as binary and reports Fisher exact testing. Fisher's exact test is appropriate for categorical comparisons when the analysis is based on contingency-table counts and an exact rather than large-sample approximation is desired.

9. Primary Endpoint: What the Hazard Ratio Does — and Does Not — Mean

Effect size

The primary Cox estimate of HR 0.74 corresponds to an estimated 26% lower hazard of the first clinical worsening event for UT-15C relative to placebo, because 1 − 0.74 = 0.26.

Not an absolute risk reduction

The HR does not tell us the absolute probability that a subject will experience a clinical worsening event, nor does it provide a number needed to treat. Those quantities require absolute event probabilities or other information not contained in the registry-reported primary analysis.

Precision

The 95% CI of 0.56–0.97 indicates that the point estimate is not known without uncertainty. The upper endpoint is close to 1, so the interval includes relative effects that are smaller than the point estimate while remaining below the null value of 1.

Model assumptions

The Cox interpretation relies on the proportional-hazards framework. A single hazard ratio is most naturally interpreted as a relative hazard under that model rather than as a constant individual-level risk reduction throughout follow-up.

10. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These counts provide an arm-specific safety summary but do not by themselves establish a causal comparison of safety event rates without additional statistical context.

Treatment armSerious adverse events affectedAt risk
UT-15C116346
Placebo110344
Serious adverse events: affected / at risk
UT-15C
116 / 346
Placebo
110 / 344

The denominators are shown because a raw event count alone cannot be interpreted as a treatment-group rate. The ClinicalTrials.gov record identifies the number affected and the number at risk in each arm but does not provide a formal statistical analysis of the serious-adverse-event comparison.

11. Analysis Population and Missing Data

The ClinicalTrials.gov record provides one explicit analysis-population rule: for the NT-proBNP endpoint, subjects who did not have a valid NT-proBNP measurement at Baseline or Week 24 were excluded from the analysis.

NT-proBNP analysis population: Subjects without a valid NT-proBNP measurement at Baseline or Week 24 were excluded from the analysis. This means the reported ANCOVA result applies to the defined analysis population rather than automatically to every enrolled subject.

This distinction matters because exclusions based on missing measurements can affect the population contributing information to a secondary analysis. The ClinicalTrials.gov record does not provide the number excluded, the reasons for missing measurements, or an imputation procedure, so no additional missing-data mechanism is inferred.

12. Design Features That Shape Statistical Interpretation

Design featureWhat the ClinicalTrials.gov record supportsStatistical relevance
Randomization Allocation was RANDOMIZED. Randomization supports comparison of treatment groups under the trial's allocation framework.
Parallel design Design model was PARALLEL. Subjects are compared according to their randomized treatment groups rather than through a crossover sequence.
Quadruple masking Masking was QUADRUPLE. Masking can reduce opportunities for knowledge of treatment assignment to influence trial conduct or outcome assessment.
Superiority hypothesis Primary and secondary posted analyses were designated SUPERIORITY. The inferential question is whether the treatment produces evidence of a difference favoring the intervention rather than whether it meets a non-inferiority margin.
Time-to-event primary endpoint Primary endpoint type was time-to-event. Timing of events and censoring are central to the primary analysis, motivating Cox and log-rank methods.

13. Understanding the Primary Endpoint Definition

The clinical worsening endpoint is a composite time-to-first-event outcome. Its definition includes several clinically distinct event types: death from all causes, hospitalization due to worsening PAH, initiation of an inhaled or infused prostacyclin for worsening PAH, disease progression, and unsatisfactory long-term clinical response.

Why a composite endpoint?

A composite endpoint can capture several ways in which a patient's disease status may worsen while maintaining a single time-to-first-event analysis framework.

Why time to first event?

The endpoint is defined around the first clinical worsening event, so the analysis focuses on when the first qualifying event occurs rather than counting all subsequent events.

Why the definition matters

A hazard ratio for a composite endpoint describes the relative hazard of experiencing the first qualifying event, not the hazard of each component considered separately.

Why censoring matters

Subjects who do not experience a qualifying event during observed follow-up contribute information up to their last available observation under the assumptions of the time-to-event analysis.

14. Reading the Secondary Continuous Outcomes

6-Minute Walk Distance

The 6-minute walk distance endpoint was evaluated as a change from baseline to Week 24 using ANCOVA. The reported Hodges-Lehmann estimate of location shift was 7.0, with a two-sided 95% CI of 0–16.0 and P = 0.0913.

Three pieces of information should be read together. The estimate describes the observed direction and magnitude of the reported location shift; the confidence interval describes uncertainty around that estimate; and the p-value addresses the associated hypothesis test. None of these quantities should be interpreted as a direct probability that an individual subject will improve.

NT-proBNP

The NT-proBNP endpoint was defined as change from baseline to Week 24, with outcome expressed as a ratio to baseline. The reported method was ANCOVA and the reported p-value was <0.0001.

Because no effect estimate or confidence interval is in the ClinicalTrials.gov record, the p-value cannot be converted into a quantitative statement about the magnitude of the treatment difference. The ratio-to-baseline scale also means that the underlying statistical interpretation differs from the meter-based 6-minute walk distance endpoint.

WHO Functional Class

The WHO Functional Class endpoint was evaluated from baseline to Week 48 and analyzed using Fisher exact testing. The reported p-value was 0.0028.

Because the ClinicalTrials.gov record does not provide the underlying cell counts or an effect estimate, the result is appropriately described as evidence from the reported categorical hypothesis test rather than as a calculated risk ratio, odds ratio, or absolute difference.

15. Statistical Significance vs Effect Size

StatisticFREEDOM-EV examplePrimary interpretation
Hazard ratio 0.74 Relative time-to-event effect estimate
95% CI for HR 0.56–0.97 Uncertainty around the hazard-ratio estimate
Cox p-value 0.0275 Evidence against the relevant null hypothesis
Log-rank p-value 0.0391 Evidence from a time-to-event distribution comparison
6-minute walk location shift 7.0 Estimated location shift in meters
6-minute walk 95% CI 0–16.0 Uncertainty around the location-shift estimate
6-minute walk p-value 0.0913 Evidence from the reported ANCOVA analysis

This distinction is central to clinical-trial interpretation. A small p-value can coexist with a modest effect estimate, and a numerically large effect estimate can have substantial uncertainty. The most informative interpretation therefore combines the estimated effect, confidence interval, and hypothesis-test result rather than relying on the p-value alone.

16. Limitations

17. Why This Trial Matters Statistically

FREEDOM-EV is a useful teaching case because it combines randomized treatment allocation, quadruple masking, a composite time-to-event primary endpoint, Cox proportional-hazards modeling, log-rank testing, continuous secondary outcomes analyzed by ANCOVA, a categorical endpoint analyzed by Fisher exact testing, and arm-specific safety reporting.

ConceptHow it appears in FREEDOM-EV
Randomization Randomized allocation to UT-15C or placebo
Blinding Quadruple masking
Time-to-event analysis Time to first clinical worsening event
Hazard ratio Primary Cox estimate of 0.74
Confidence interval 95% CI of 0.56–0.97 for the primary hazard ratio
Log-rank test Primary endpoint p-value of 0.0391
ANCOVA 6-minute walk distance and NT-proBNP analyses
Hodges-Lehmann estimate 7.0 location-shift estimate for 6-minute walk distance
Fisher exact test WHO Functional Class analysis
Missing-data handling Subjects without valid baseline or Week 24 NT-proBNP measurements were excluded from that analysis
Safety analysis Serious adverse events reported as affected / at risk by arm

18. Related Tutorials

Learn more about the methods used in this trial:

19. Related Statistical Calculators

20. Sources

Continue with the underlying statistical methods

Explore the survival-analysis, continuous-outcome, categorical-data, and clinical-trial methods represented in FREEDOM-EV.

21. Record Summary

FREEDOM-EV provides a compact example of how several statistical frameworks can coexist within one randomized phase 3 clinical trial. The primary endpoint was time to first clinical worsening event, analyzed with both a Cox proportional-hazards model and a log-rank test. The Cox analysis reported an HR of 0.74 with a 95% CI of 0.56–0.97 and P = 0.0275, while the log-rank analysis reported P = 0.0391. Secondary analyses used ANCOVA for change in 6-minute walk distance and change in NT-proBNP and Fisher exact testing for change in WHO Functional Class.

The most informative statistical reading is not to treat all five reported analyses as interchangeable. The primary hazard ratio provides a relative time-to-event effect estimate; its confidence interval describes precision; the log-rank test supplies complementary evidence about the time-to-event distributions; ANCOVA provides a framework for continuous outcomes; and Fisher exact testing addresses the categorical endpoint. Together, these results illustrate why clinical-trial interpretation requires attention to endpoint definition, analysis population, effect measure, uncertainty, and model assumptions.

Clinical Biostats methodology: A trial-results page should distinguish reported numerical evidence from statistical interpretation. The purpose is to explain what each analysis measures, what its estimate means, what it does not mean, and which limitations affect interpretation without extending the registry record beyond the registry-reported evidence.