This page separates reported trial results from statistical interpretation. Numerical results presented here are restricted to the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for FREEDOM-EV. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
FREEDOM-EV was a randomized, parallel, quadruple-masked phase 3 trial comparing treprostinil diolamine (UT-15C) with placebo in subjects with pulmonary arterial hypertension receiving background oral monotherapy. The primary endpoint was time to first clinical worsening event, analyzed from randomization to approximately 4 years.
| Feature | FREEDOM-EV |
|---|---|
| Trial name | FREEDOM-EV |
| Phase | Phase 3 |
| Condition | Pulmonary Arterial Hypertension |
| Design | Randomized, parallel, quadruple-masked |
| Primary purpose | Treatment |
| Enrollment | 690 |
| Interventions | Treprostinil Diolamine (drug); Placebo (drug) |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| Lead sponsor | United Therapeutics |
| Trial status | COMPLETED |
| ClinicalTrials.gov | NCT01560624 |
2. Clinical Question
The primary statistical question was whether treprostinil diolamine, compared with placebo, changed the time from randomization to the first clinical worsening event in subjects with pulmonary arterial hypertension receiving background oral monotherapy.
Population
Subjects with pulmonary arterial hypertension receiving background oral monotherapy.
Intervention
Treprostinil Diolamine (UT-15C).
Comparator
Placebo.
Primary question
Does UT-15C improve the time to first clinical worsening event compared with placebo?
3. Trial Design
Treprostinil diolamine
- Treprostinil Diolamine (UT-15C)
- Drug intervention
- Compared with placebo
Placebo
- Placebo drug intervention
- Comparator arm
- Compared with UT-15C
The trial was completed after starting on June 26, 2012, with primary completion on June 24, 2018. The registry reports an enrollment of 690 subjects across two treatment arms.
4. Endpoints
The registered primary endpoint was a time-to-event outcome assessed from randomization to approximately 4 years.
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Time to First Clinical Worsening Event | From randomization to approximately 4 years | Time-to-event |
Primary endpoint definition
Clinical worsening was assessed continuously from randomization until the subject's last study visit. Clinical worsening events were defined as death (all causes), hospitalizations due to worsening pulmonary arterial hypertension (PAH), initiation of an inhaled or infused prostacyclin (PGI2) for the treatment of worsening PAH, disease progression, or unsatisfactory long-term clinical response.
Secondary endpoints with posted analyses
| Secondary endpoint | Time frame | Outcome unit | Method |
|---|---|---|---|
| Change in 6-Minute Walk Distance | From Baseline to Week 24 | meters | ANCOVA |
| Change in Plasma N-Terminal Pro-brain Natriuretic Peptide (NT-proBNP) From Baseline to Week 24 | From Baseline to Week 24 | Ratio to baseline | ANCOVA |
| Change in World Health Organization Functional Class (WHO FC) From Baseline to Week 48 | Baseline to Week 48 | Participants | Fisher exact test |
5. Statistical Results
Primary endpoint: Time to First Clinical Worsening Event
The primary endpoint was analyzed as a time-to-event outcome using a Cox proportional-hazards model and a log-rank test. Both analyses compared UT-15C with placebo and used a superiority hypothesis.
Cox proportional-hazards analysis
95% CI: 0.56–0.97 · P = 0.0275
Comparison: UT-15C vs Placebo
An HR of 0.74 means that, under the fitted Cox proportional-hazards model, the estimated instantaneous rate of experiencing the first clinical worsening event in the UT-15C group was approximately 26% lower than in the placebo group.
The hazard ratio does not mean that 26% of patients avoided clinical worsening, nor does it mean that every individual patient had exactly a 26% reduction in risk. It is a relative time-to-event measure based on the observed follow-up and model.
The 95% CI of 0.56–0.97 describes uncertainty around the estimated hazard ratio. Because the interval remains below 1, the reported interval is consistent with a lower estimated hazard of clinical worsening for UT-15C relative to placebo under this analysis. It does not describe the range of individual patient effects.
The P = 0.0275 addresses the statistical evidence against the null hypothesis represented by the model comparison. A p-value is not a measure of effect size and does not indicate the probability that the treatment works. Interpretation of the Cox result also depends on the proportional-hazards framework and on how censoring is handled.
Primary endpoint: log-rank analysis
Log-rank test
Comparison: UT-15C vs Placebo
Endpoint: Time to First Clinical Worsening Event
The log-rank test provides a separate hypothesis test comparing the time-to-event distributions between UT-15C and placebo. The reported P = 0.0391 indicates that the trial registry's formal log-rank comparison produced evidence against the corresponding null hypothesis under the reported superiority framework.
The log-rank p-value does not quantify how large the treatment effect is. The magnitude of the observed relative effect is described by the Cox hazard ratio, 0.74, together with its 95% CI of 0.56–0.97.
The two primary analyses answer related but different statistical questions: the log-rank test evaluates separation between time-to-event distributions, while the Cox model provides a relative hazard estimate with a confidence interval. Neither analysis by itself establishes how an individual patient will experience clinical worsening.
Secondary endpoint: Change in 6-Minute Walk Distance
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Change in 6-Minute Walk Distance from Baseline to Week 24 | ANCOVA | Hodges Lehmann estimate location shift | 7.0 meters | 0–16.0 | 0.0913 |
The registry reports a Hodges-Lehmann estimate of location shift of 7.0 meters for the UT-15C versus placebo comparison, with a two-sided 95% CI of 0–16.0 meters and P = 0.0913.
The estimate describes the reported between-group location shift on the 6-minute walk distance outcome; it is not a hazard ratio and should not be interpreted as a relative treatment effect. The confidence interval expresses uncertainty around the estimated shift, while the p-value addresses the associated hypothesis test rather than the clinical magnitude of the shift.
Secondary endpoint: Change in Plasma NT-proBNP
| Endpoint | Method | Outcome unit | P-value |
|---|---|---|---|
| Change in Plasma N-Terminal Pro-brain Natriuretic Peptide (NT-proBNP) From Baseline to Week 24 | ANCOVA | Ratio to baseline | <0.0001 |
The P < 0.0001 provides evidence against the null hypothesis represented by the reported ANCOVA comparison. It does not, by itself, describe the magnitude of the difference between treatment groups. Because the ClinicalTrials.gov record does not provide the corresponding estimate or confidence interval, the statistical evidence and the size of the effect should be kept conceptually separate.
Secondary endpoint: Change in WHO Functional Class
| Endpoint | Method | Time frame | P-value |
|---|---|---|---|
| Change in World Health Organization Functional Class (WHO FC) From Baseline to Week 48 | Fisher exact test | Baseline to Week 48 | 0.0028 |
The registry reports a Fisher exact test with P = 0.0028 for the UT-15C versus placebo comparison. The ClinicalTrials.gov record does not provide a corresponding effect estimate, confidence interval, or cell counts, so the result is appropriately reported as a hypothesis-test result rather than converted into an unsupported measure of effect size.
6. Secondary Results in Context
| Outcome | Analysis | Reported result | What the statistic tells us |
|---|---|---|---|
| Time to First Clinical Worsening Event | Cox proportional-hazards model | HR 0.74; 95% CI 0.56–0.97; P = 0.0275 | Relative time-to-event effect estimate with uncertainty interval |
| Time to First Clinical Worsening Event | Log-rank test | P = 0.0391 | Hypothesis test comparing time-to-event distributions |
| Change in 6-Minute Walk Distance | ANCOVA | Hodges-Lehmann estimate location shift 7.0; 95% CI 0–16.0; P = 0.0913 | Estimated between-group location shift with uncertainty |
| Change in NT-proBNP | ANCOVA | P < 0.0001 | Hypothesis-test evidence; no effect estimate reported |
| Change in WHO Functional Class | Fisher exact test | P = 0.0028 | Hypothesis-test evidence for the categorical comparison |
This table illustrates an important statistical distinction: a trial can report several p-values across different endpoint types, but those p-values are not interchangeable. A hazard ratio summarizes a relative time-to-event effect, a location-shift estimate summarizes a continuous outcome comparison, and Fisher's exact test addresses a categorical comparison.
7. Statistical Methodology
Kaplan-Meier estimation and time-to-event endpoints
The primary endpoint is a time-to-event outcome. For this type of endpoint, Kaplan-Meier estimation is commonly used to describe the probability of remaining event-free over time while accounting for right-censored observations.
where di is the number of events at time ti and ni is the number at risk immediately before that time.
The ClinicalTrials.gov record specifically reports a Cox proportional-hazards model and a log-rank test for the primary endpoint. Kaplan-Meier estimation is the natural descriptive companion to those methods for a time-to-event endpoint, although no Kaplan-Meier estimates are reported in the ClinicalTrials.gov record.
Cox proportional-hazards model
The Cox proportional-hazards model expresses the relative hazard between treatment groups through a hazard ratio. The reported primary estimate was 0.74 for UT-15C versus placebo.
A hazard ratio is a relative time-to-event measure. It is not a relative risk, an absolute risk difference, a probability of benefit, or a statement about the outcome of every individual patient.
Log-rank test
The log-rank test compares the observed pattern of events over time between treatment groups. It is particularly suited to randomized time-to-event comparisons because it incorporates the ordering of event times and accommodates censoring under its usual assumptions.
ANCOVA
ANCOVA, or analysis of covariance, is a linear-model framework used to compare a continuous outcome between groups while incorporating covariate information. In FREEDOM-EV, ANCOVA was the reported method for the change in 6-minute walk distance and the change in plasma NT-proBNP.
For a change-from-baseline outcome, the model can be viewed conceptually as comparing treatment groups after accounting for relevant baseline information. The exact covariates and model specification are not provided in the ClinicalTrials.gov record, so they are not inferred here.
Hodges-Lehmann estimate of location shift
The Hodges-Lehmann approach provides an estimate of the location shift between two distributions. Here, the registry reports a 7.0 estimate for the change in 6-minute walk distance, with a two-sided 95% CI of 0–16.0.
Fisher exact test
Fisher's exact test is designed for categorical data and evaluates the association between treatment group and categorical outcome counts using the exact distribution conditional on the relevant margins. It was the reported method for change in WHO Functional Class from baseline to Week 48.
8. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
The primary endpoint measures the time until the first clinical worsening event rather than simply whether an event occurred. A Cox model uses the timing of events and censoring information to estimate a relative hazard between the randomized groups. Its principal effect measure is the hazard ratio.
What does an HR of 0.74 mean?
Under the fitted model, an HR of 0.74 means that the estimated instantaneous rate of the first clinical worsening event in the UT-15C group was about 26% lower than in the placebo group. It does not mean that 26% of patients were protected from worsening or that each patient's risk was reduced by exactly 26%.
Why use a log-rank test as well as a Cox model?
The log-rank test provides a hypothesis test comparing the time-to-event experience of the two randomized groups, whereas the Cox model supplies an interpretable relative effect estimate and confidence interval. Reporting both gives complementary information about statistical evidence and effect magnitude.
What does the 95% confidence interval of 0.56–0.97 tell us?
The interval quantifies statistical uncertainty around the estimated hazard ratio under the model and sampling framework. It is not a prediction interval for individual patients and does not mean that 95% of patients have hazard ratios somewhere between 0.56 and 0.97.
Why does P = 0.0275 not measure treatment effect size?
A p-value describes the evidence against a specified null hypothesis under the statistical model. It depends on both the observed difference and the amount of information in the data. The hazard ratio and its confidence interval are therefore necessary to understand the magnitude and precision of the time-to-event effect.
Why is ANCOVA appropriate for the continuous secondary endpoints?
ANCOVA is a standard method for comparing continuous outcomes between randomized groups while incorporating baseline information through a linear-model framework. In this trial, the registry identifies ANCOVA as the analysis method for change in 6-minute walk distance and change in NT-proBNP.
Why is Fisher exact testing appropriate for WHO Functional Class?
The registry identifies the WHO Functional Class endpoint as binary and reports Fisher exact testing. Fisher's exact test is appropriate for categorical comparisons when the analysis is based on contingency-table counts and an exact rather than large-sample approximation is desired.
9. Primary Endpoint: What the Hazard Ratio Does — and Does Not — Mean
The primary Cox estimate of HR 0.74 corresponds to an estimated 26% lower hazard of the first clinical worsening event for UT-15C relative to placebo, because 1 − 0.74 = 0.26.
The HR does not tell us the absolute probability that a subject will experience a clinical worsening event, nor does it provide a number needed to treat. Those quantities require absolute event probabilities or other information not contained in the registry-reported primary analysis.
The 95% CI of 0.56–0.97 indicates that the point estimate is not known without uncertainty. The upper endpoint is close to 1, so the interval includes relative effects that are smaller than the point estimate while remaining below the null value of 1.
The Cox interpretation relies on the proportional-hazards framework. A single hazard ratio is most naturally interpreted as a relative hazard under that model rather than as a constant individual-level risk reduction throughout follow-up.
10. Safety
The ClinicalTrials.gov record reports serious adverse events by treatment arm. These counts provide an arm-specific safety summary but do not by themselves establish a causal comparison of safety event rates without additional statistical context.
| Treatment arm | Serious adverse events affected | At risk |
|---|---|---|
| UT-15C | 116 | 346 |
| Placebo | 110 | 344 |
The denominators are shown because a raw event count alone cannot be interpreted as a treatment-group rate. The ClinicalTrials.gov record identifies the number affected and the number at risk in each arm but does not provide a formal statistical analysis of the serious-adverse-event comparison.
11. Analysis Population and Missing Data
The ClinicalTrials.gov record provides one explicit analysis-population rule: for the NT-proBNP endpoint, subjects who did not have a valid NT-proBNP measurement at Baseline or Week 24 were excluded from the analysis.
This distinction matters because exclusions based on missing measurements can affect the population contributing information to a secondary analysis. The ClinicalTrials.gov record does not provide the number excluded, the reasons for missing measurements, or an imputation procedure, so no additional missing-data mechanism is inferred.
12. Design Features That Shape Statistical Interpretation
| Design feature | What the ClinicalTrials.gov record supports | Statistical relevance |
|---|---|---|
| Randomization | Allocation was RANDOMIZED. | Randomization supports comparison of treatment groups under the trial's allocation framework. |
| Parallel design | Design model was PARALLEL. | Subjects are compared according to their randomized treatment groups rather than through a crossover sequence. |
| Quadruple masking | Masking was QUADRUPLE. | Masking can reduce opportunities for knowledge of treatment assignment to influence trial conduct or outcome assessment. |
| Superiority hypothesis | Primary and secondary posted analyses were designated SUPERIORITY. | The inferential question is whether the treatment produces evidence of a difference favoring the intervention rather than whether it meets a non-inferiority margin. |
| Time-to-event primary endpoint | Primary endpoint type was time-to-event. | Timing of events and censoring are central to the primary analysis, motivating Cox and log-rank methods. |
13. Understanding the Primary Endpoint Definition
The clinical worsening endpoint is a composite time-to-first-event outcome. Its definition includes several clinically distinct event types: death from all causes, hospitalization due to worsening PAH, initiation of an inhaled or infused prostacyclin for worsening PAH, disease progression, and unsatisfactory long-term clinical response.
Why a composite endpoint?
A composite endpoint can capture several ways in which a patient's disease status may worsen while maintaining a single time-to-first-event analysis framework.
Why time to first event?
The endpoint is defined around the first clinical worsening event, so the analysis focuses on when the first qualifying event occurs rather than counting all subsequent events.
Why the definition matters
A hazard ratio for a composite endpoint describes the relative hazard of experiencing the first qualifying event, not the hazard of each component considered separately.
Why censoring matters
Subjects who do not experience a qualifying event during observed follow-up contribute information up to their last available observation under the assumptions of the time-to-event analysis.
14. Reading the Secondary Continuous Outcomes
6-Minute Walk Distance
The 6-minute walk distance endpoint was evaluated as a change from baseline to Week 24 using ANCOVA. The reported Hodges-Lehmann estimate of location shift was 7.0, with a two-sided 95% CI of 0–16.0 and P = 0.0913.
Three pieces of information should be read together. The estimate describes the observed direction and magnitude of the reported location shift; the confidence interval describes uncertainty around that estimate; and the p-value addresses the associated hypothesis test. None of these quantities should be interpreted as a direct probability that an individual subject will improve.
NT-proBNP
The NT-proBNP endpoint was defined as change from baseline to Week 24, with outcome expressed as a ratio to baseline. The reported method was ANCOVA and the reported p-value was <0.0001.
Because no effect estimate or confidence interval is in the ClinicalTrials.gov record, the p-value cannot be converted into a quantitative statement about the magnitude of the treatment difference. The ratio-to-baseline scale also means that the underlying statistical interpretation differs from the meter-based 6-minute walk distance endpoint.
WHO Functional Class
The WHO Functional Class endpoint was evaluated from baseline to Week 48 and analyzed using Fisher exact testing. The reported p-value was 0.0028.
Because the ClinicalTrials.gov record does not provide the underlying cell counts or an effect estimate, the result is appropriately described as evidence from the reported categorical hypothesis test rather than as a calculated risk ratio, odds ratio, or absolute difference.
15. Statistical Significance vs Effect Size
| Statistic | FREEDOM-EV example | Primary interpretation |
|---|---|---|
| Hazard ratio | 0.74 | Relative time-to-event effect estimate |
| 95% CI for HR | 0.56–0.97 | Uncertainty around the hazard-ratio estimate |
| Cox p-value | 0.0275 | Evidence against the relevant null hypothesis |
| Log-rank p-value | 0.0391 | Evidence from a time-to-event distribution comparison |
| 6-minute walk location shift | 7.0 | Estimated location shift in meters |
| 6-minute walk 95% CI | 0–16.0 | Uncertainty around the location-shift estimate |
| 6-minute walk p-value | 0.0913 | Evidence from the reported ANCOVA analysis |
This distinction is central to clinical-trial interpretation. A small p-value can coexist with a modest effect estimate, and a numerically large effect estimate can have substantial uncertainty. The most informative interpretation therefore combines the estimated effect, confidence interval, and hypothesis-test result rather than relying on the p-value alone.
16. Limitations
- Composite endpoint: the primary endpoint combines several types of clinical worsening. A single hazard ratio therefore summarizes the time to the first qualifying component rather than separately estimating every component.
- Time-to-event assumptions: interpretation of the Cox hazard ratio depends on the proportional-hazards framework and appropriate handling of censoring.
- Limited primary-effect information: the registry supplies a hazard ratio and confidence interval but does not provide median time-to-event estimates in the ClinicalTrials.gov record.
- Secondary endpoint reporting: the NT-proBNP and WHO Functional Class analyses have reported p-values without corresponding effect estimates and confidence intervals in the ClinicalTrials.gov record.
- NT-proBNP exclusions: subjects without valid baseline or Week 24 measurements were excluded from that analysis. The ClinicalTrials.gov record does not describe the extent or mechanism of missingness.
- Safety comparison: serious adverse-event counts by arm are reported, but the ClinicalTrials.gov record does not provide a formal statistical comparison of these events.
- Multiple endpoints: several outcomes were analyzed, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment strategy. The reported p-values should therefore be interpreted according to the information actually registry-reported rather than assuming a particular multiplicity procedure.
- Generalizability: the trial population is defined as subjects with pulmonary arterial hypertension receiving background oral monotherapy. Application of results outside that population requires consideration of differences in patient characteristics and treatment context.
17. Why This Trial Matters Statistically
FREEDOM-EV is a useful teaching case because it combines randomized treatment allocation, quadruple masking, a composite time-to-event primary endpoint, Cox proportional-hazards modeling, log-rank testing, continuous secondary outcomes analyzed by ANCOVA, a categorical endpoint analyzed by Fisher exact testing, and arm-specific safety reporting.
| Concept | How it appears in FREEDOM-EV |
|---|---|
| Randomization | Randomized allocation to UT-15C or placebo |
| Blinding | Quadruple masking |
| Time-to-event analysis | Time to first clinical worsening event |
| Hazard ratio | Primary Cox estimate of 0.74 |
| Confidence interval | 95% CI of 0.56–0.97 for the primary hazard ratio |
| Log-rank test | Primary endpoint p-value of 0.0391 |
| ANCOVA | 6-minute walk distance and NT-proBNP analyses |
| Hodges-Lehmann estimate | 7.0 location-shift estimate for 6-minute walk distance |
| Fisher exact test | WHO Functional Class analysis |
| Missing-data handling | Subjects without valid baseline or Week 24 NT-proBNP measurements were excluded from that analysis |
| Safety analysis | Serious adverse events reported as affected / at risk by arm |
18. Related Tutorials
Learn more about the methods used in this trial:
19. Related Statistical Calculators
20. Sources
- ClinicalTrials.gov: FREEDOM-EV, NCT01560624.
- Linked publication: PubMed record for PMID 31765604: https://pubmed.ncbi.nlm.nih.gov/31765604/.
Continue with the underlying statistical methods
Explore the survival-analysis, continuous-outcome, categorical-data, and clinical-trial methods represented in FREEDOM-EV.
21. Record Summary
FREEDOM-EV provides a compact example of how several statistical frameworks can coexist within one randomized phase 3 clinical trial. The primary endpoint was time to first clinical worsening event, analyzed with both a Cox proportional-hazards model and a log-rank test. The Cox analysis reported an HR of 0.74 with a 95% CI of 0.56–0.97 and P = 0.0275, while the log-rank analysis reported P = 0.0391. Secondary analyses used ANCOVA for change in 6-minute walk distance and change in NT-proBNP and Fisher exact testing for change in WHO Functional Class.
The most informative statistical reading is not to treat all five reported analyses as interchangeable. The primary hazard ratio provides a relative time-to-event effect estimate; its confidence interval describes precision; the log-rank test supplies complementary evidence about the time-to-event distributions; ANCOVA provides a framework for continuous outcomes; and Fisher exact testing addresses the categorical endpoint. Together, these results illustrate why clinical-trial interpretation requires attention to endpoint definition, analysis population, effect measure, uncertainty, and model assumptions.