← Clinical Trials
Pulmonary Arterial Hypertension Phase 3 Monotherapy NCT00325403

FREEDOM-M: Complete Statistical Analysis of Oral Treprostinil in Pulmonary Arterial Hypertension

An independent statistical review of the randomized phase 3 FREEDOM-M trial evaluating oral treprostinil (UT-15C) sustained release tablets versus placebo as monotherapy for pulmonary arterial hypertension, with particular focus on six-minute walk distance and the statistical methods used to compare treatment groups.

Trial start: 2006-10  ·  Primary completion: 2011-04  ·  Enrollment: 349  ·  Sponsor: United Therapeutics
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data reported for FREEDOM-M.

1. Trial at a Glance

FREEDOM-M was a randomized, double-blind, parallel phase 3 trial evaluating oral treprostinil (UT-15C) sustained release tablets versus placebo as monotherapy for pulmonary arterial hypertension. The primary endpoint was the placebo-corrected change in six-minute walk distance from Baseline to Week 12.

349
Enrollment
Phase 3
2
Arms
Randomized, parallel
228
mITT group
Access to 0.25 mg tablets
23
Primary estimate
95% CI 4–41 meters
FeatureFREEDOM-M
Trial nameFREEDOM-M
NCT identifierNCT00325403
PhasePhase 3
StatusCOMPLETED
ConditionPulmonary Hypertension
Brief titleFREEDOM - M: Oral Treprostinil as Monotherapy for the Treatment of Pulmonary Arterial Hypertension (PAH)
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment349
InterventionsOral treprostinil (UT-15C) Sustained Release Tablets; Placebo
Hypothesis typeSuperiority
Results postedYes

2. Clinical Question

The central statistical question was whether oral treprostinil produced a difference in the change in six-minute walk distance from Baseline to Week 12 compared with placebo in patients with pulmonary arterial hypertension.

Population

Participants enrolled in the FREEDOM-M phase 3 trial for pulmonary arterial hypertension.

Intervention

Oral treprostinil (UT-15C) sustained release tablets.

Comparator

Placebo.

Primary question

Does oral treprostinil produce a superior placebo-corrected change in six-minute walk distance from Baseline to Week 12?

3. Trial Design

01
Randomize349 enrolled
02
Double-blindTwo parallel arms
03
TreatmentOral treprostinil or placebo
04
Assess6MWD and functional measures
05
CompareWeek 12 treatment difference
Allocation
Randomized, with an allocation ratio of 2:1 between oral treprostinil and placebo according to the primary analysis notes.
Masking
Double-blind.
Design
Parallel-group phase 3 treatment trial.
Primary analysis
ANCOVA comparing the placebo versus UT-15C groups using the modified intention-to-treat group.
ARM A

UT-15C / Oral Treprostinil

  • Oral treprostinil (UT-15C)
  • Sustained release tablets
  • Included in the primary comparison against placebo
ARM B

Placebo

  • Placebo
  • Included in the primary comparison against oral treprostinil

4. Randomization and Analysis Population

The primary analysis was conducted in a modified intention-to-treat (mITT) group of n=228. The registry defines this group as subjects with access to 0.25 mg tablets at randomization. The analysis notes state that all alpha was spent on this subgroup.

Population / featureRegistry-supported description
Overall enrollment349 participants
Primary efficacy populationModified intention-to-treat group
mITT sizen=228
mITT definitionSubjects with access to 0.25 mg tablets at randomization
Alpha allocationAll alpha was spent on this subgroup
Analysis comparisonPlacebo vs UT-15C (Oral Treprostinil)
Why the mITT population matters: the headline primary result is not described as an analysis of all 349 enrolled participants. It is based on the mITT group of 228 subjects defined by tablet access at randomization. That distinction matters when interpreting the size, precision, and generalizability of the primary comparison.

5. Primary Endpoint

EndpointDefinition / assessmentAnalysis
Six Minute Walk Distance (6MWD) Placebo corrected change in six minute walk distance (6MWD) from Baseline to Week 12, which correlates with the historical clinical standard for assessing patient functional status in the treatment of PAH and is considered an objective measure of patient functional status by the American Thoracic Society (ATS). ANCOVA

The registered time frame was Baseline and Week 12. The endpoint was continuous and measured in meters. The statistical analysis used the Hodges-Lehmann (H-L) estimate as the reported effect measure, with a two-sided 95% confidence interval.

6. Primary Result: Six Minute Walk Distance

Placebo-corrected change in 6MWD

23 meters

95% CI: 4–41 meters   ·   P = 0.0125

ANCOVA  ·  modified intention-to-treat population, n=228  ·  superiority hypothesis

Primary endpointTime frameEffect measureEstimate95% CIP-value
Six Minute Walk Distance (6MWD) Baseline and Week 12 Hodges-Lehmann (H-L) 23 meters 4 to 41 meters 0.0125
Primary analysis framework

Method: ANCOVA

Treatment comparison → change in 6MWD from Baseline to Week 12

The registry reports a Hodges-Lehmann estimate of 23 with a two-sided 95% confidence interval of 4 to 41 and a P-value of 0.0125.

Clinical Biostats interpretation

The reported estimate of 23 meters represents the placebo-corrected treatment difference in the change in six-minute walk distance from Baseline to Week 12, as estimated using the registry-reported analysis framework. Because the estimate is positive, the comparison favors the oral-treprostinil group in terms of the direction of change in 6MWD.

The 95% confidence interval of 4 to 41 meters expresses statistical uncertainty around the estimated treatment difference. It does not mean that individual patients experienced changes only within this interval, nor does it describe the distribution of individual treatment responses.

The P-value of 0.0125 quantifies the statistical evidence against the null hypothesis under the specified testing framework. It is not a measure of how large or clinically important the treatment effect is. Effect magnitude and uncertainty are better communicated by the estimate and its confidence interval.

The analysis is also population-specific: the primary result was obtained in the mITT group of n=228, rather than simply in all 349 enrolled participants. The registry also states that all alpha was spent on this subgroup.

Sample-size and power context

The registry analysis notes state that, using an allocation ratio of 2:1 between oral treprostinil and placebo, a fixed sample size of approximately 195 subjects with access to 0.25 mg tablets at randomization would provide at least 90% power at a significance level of 0.01 using a two-sided hypothesis to detect a between-treatment difference in the change from Baseline to Week 12 in distance traversed during the six-minute walk, under the assumptions specified in the registry analysis text.

This provides important context for interpreting the primary result: the trial's statistical planning was tied specifically to the subgroup with access to 0.25 mg tablets, and the reported primary analysis likewise used the mITT population of 228.

7. Secondary Endpoint Results

The registry also reports several secondary analyses. These use different endpoints and, in some cases, different statistical methods. They should therefore not be treated as interchangeable with the prespecified primary analysis.

6MWD at Week 11

Placebo-corrected change in 6MWD

13 meters

95% CI: −2 to 33 meters   ·   P = 0.0653

Baseline and Week 11  ·  ANCOVA  ·  Hodges-Lehmann estimate

Clinical Biostats interpretation

The estimated treatment difference at Week 11 was 13 meters, with a two-sided 95% confidence interval from −2 to 33 meters. The interval includes zero, so the registry-reported estimate is less precise than the Week 12 primary result with respect to excluding no treatment difference.

The P-value of 0.0653 is not an effect-size measure. It should be read together with the estimate and confidence interval rather than used as a substitute for them.

6MWD at Week 8

Placebo-corrected change in 6MWD

17 meters

95% CI: 1–33 meters   ·   P = 0.0307

Baseline and Week 8  ·  ANCOVA  ·  Hodges-Lehmann estimate

Clinical Biostats interpretation

The Week 8 estimate was 17 meters, with a 95% confidence interval from 1 to 33 meters. The positive estimate and interval provide a treatment-group difference in the same direction as the Week 12 primary analysis, while the interval indicates the statistical precision of this earlier assessment.

Because this is a secondary time point rather than the registered primary endpoint, it should not replace the Week 12 analysis when describing the primary trial question.

6MWD at Week 4

Placebo-corrected change in 6MWD

12 meters

95% CI: 0–24 meters   ·   P = 0.0518

Baseline and Week 4  ·  ANCOVA  ·  Hodges-Lehmann estimate

Clinical Biostats interpretation

The Week 4 estimate was 12 meters, with a two-sided 95% confidence interval from 0 to 24 meters. The estimate is positive, but the interval reaches zero. The P-value of 0.0518 describes the statistical evidence at this time point; it does not determine the magnitude of the estimated difference.

World Health Organization Functional Classification for PAH

EndpointTime frameMethodEstimate95% CIP-value
World Health Organization Functional Classification for PAH Baseline and Week 12 Wilcoxon rank sum test 0 0 to 0 0.7380

The analysis population was subjects with a WHO functional classification assessment at Week 12. The registry reports a Hodges-Lehmann estimate of 0 and a two-sided 95% confidence interval of 0 to 0. The registry notes that the estimated parameter and confidence interval were calculated.

Borg Dyspnea Score

EndpointTime frameMethodEstimate95% CIP-value
Borg Dyspnea Score Baseline and Week 12 Wilcoxon rank-sum test 0 −1 to 0 0.4887

The endpoint was measured in units on a scale. The nonparametric Wilcoxon rank-sum approach produced a Hodges-Lehmann estimate of 0, with a two-sided 95% confidence interval from −1 to 0 and a P-value of 0.4887.

Dyspnea-Fatigue Index

EndpointTime frameMethodEstimate95% CIP-value
Dyspnea-Fatigue Index Baseline and Week 12 Wilcoxon sum-rank test 0 0 to 1 0.6116

Two subjects, one in the placebo arm and one in the oral-treprostinil arm, from the primary analysis population of n=228 did not have a Baseline dyspnea-fatigue index score and were not included in this analysis. The reported Hodges-Lehmann estimate was 0, with a two-sided 95% confidence interval of 0 to 1.

Clinical Worsening Assessment

EndpointTime frameMethod95% CIP-value
Clinical Worsening Assessment Baseline and Week 12 Fisher Exact 95% 1.000

The Clinical Worsening Assessment was analyzed as a binary endpoint using Fisher Exact. The registry reports a 1.000 P-value and a 95% confidence interval level, but does not provide an effect estimate in the posted statistical analysis.

8. Post-Hoc Analyses

The registry also reports post-hoc analyses of six-minute walk distance. These analyses can help describe patterns within particular baseline categories or across the entire study population, but they are distinct from the registered primary endpoint.

Post-hoc analysisTime framePopulation / subgroupEstimate95% CIP-value
6MWD by baseline WHO Functional Classification III or IV Baseline and Week 12 Baseline WHO Functional Classification III or IV 26 meters 1–49 meters 0.0326
6MWD by baseline WHO Functional Classification: I or II Baseline and Week 12 Baseline WHO Functional Classification I or II 16.0 meters −15 to 47 meters 0.2275
6MWD by PAH Etiology: Idiopathic or Heritable PAH Baseline and Week 12 Idiopathic or Heritable PAH 32 meters 10–55 meters 0.0024
6MWD for the Entire Study Population Baseline and Week 12 All subjects enrolled, regardless of tablet strength availability at randomization 25.5 meters 10–41 meters 0.0001
6MWD for the Entire Study Population Baseline and Week 11 All subjects enrolled, regardless of tablet strength availability at randomization 17 meters 3–33 meters 0.0025
6MWD for the Entire Study Population Baseline and Week 8 All subjects enrolled, regardless of tablet strength availability at randomization 20 meters 7–34 meters 0.0008
6MWD for the Entire Study Population Baseline and Week 4 All subjects enrolled, regardless of tablet strength availability at randomization 14 meters 3.9–25 meters 0.0025

All of these post-hoc 6MWD analyses used ANCOVA and reported a Hodges-Lehmann estimate with a two-sided 95% confidence interval.

How to read the post-hoc analyses

The subgroup estimates are descriptive extensions of the primary analysis rather than replacements for it. For example, the estimated difference was 26 meters for participants with baseline WHO Functional Classification III or IV and 16.0 meters for those with baseline WHO Functional Classification I or II. Their confidence intervals differ substantially in width, illustrating how subgroup analysis can reduce statistical precision.

The two subgroup P-values also differ, but that alone does not establish that the treatment effect differs between the two functional-classification groups. Demonstrating treatment-effect heterogeneity would ordinarily require an appropriate interaction analysis rather than comparing P-values from separate subgroups.

The entire-study-population analyses are also analytically distinct from the primary mITT analysis because the registry explicitly states that they included all enrolled subjects regardless of tablet strength availability at randomization.

9. Statistical Methodology

MethodRole in FREEDOM-MEndpoint(s)
ANCOVA Primary linear-model comparison of treatment groups Primary and post-hoc 6MWD analyses
Wilcoxon / Mann-Whitney approach Nonparametric comparison between treatment groups WHO Functional Classification, Borg Dyspnea Score, Dyspnea-Fatigue Index
Fisher exact test Exact categorical comparison Clinical Worsening Assessment
Hodges-Lehmann estimate Reported effect measure for the continuous/nonparametric analyses where provided 6MWD and several functional measures

ANCOVA for the primary endpoint

The registry reports ANCOVA for the primary 6MWD analysis. In a clinical trial, ANCOVA is commonly used to compare treatment groups on a continuous outcome while accounting for baseline measurement. For a change-from-baseline endpoint, the underlying goal is to estimate the treatment-group difference while improving statistical efficiency relative to an analysis that ignores baseline information.

Conceptual model
Outcome at follow-up = treatment effect + baseline information + residual variation

The exact model specification beyond the registry's reported ANCOVA method is not provided in the ClinicalTrials.gov record, so this page does not add unreported covariates or model terms.

Nonparametric analyses

The WHO Functional Classification, Borg Dyspnea Score, and Dyspnea-Fatigue Index analyses used Wilcoxon-based methods. These methods compare the distributions or ranks of observations rather than relying on the same normal-error assumptions as a conventional parametric mean comparison.

Fisher exact test

Clinical Worsening Assessment was analyzed using Fisher Exact. This is appropriate for categorical comparisons when an exact test of the treatment-group association is desired, particularly when cell counts may be small.

Hodges-Lehmann estimates

The registry reports the Hodges-Lehmann (H-L) estimate for the primary 6MWD analysis and several other analyses. In a two-group setting, the Hodges-Lehmann approach provides a robust location estimate based on pairwise treatment-group comparisons. Its interpretation should follow the endpoint and analysis framework reported by the registry rather than being treated automatically as a conventional arithmetic mean difference.

10. Statistical Methods Explained

Why was ANCOVA used for 6MWD?

6MWD is a continuous outcome, making a linear-model framework such as ANCOVA a natural choice. ANCOVA can incorporate baseline information when estimating the treatment difference at the follow-up assessment. The ClinicalTrials.gov record identifies ANCOVA as the method used, but do not provide the complete model specification, so additional covariates should not be assumed.

What does a 23-meter estimate mean?

The primary Hodges-Lehmann estimate was 23 meters. In the context of this trial, that is the reported placebo-corrected treatment difference in change in 6MWD from Baseline to Week 12. It is not the same thing as saying that every participant walked 23 meters farther, nor does it describe the response of any particular patient.

What does the 95% confidence interval of 4 to 41 mean?

The interval quantifies statistical uncertainty around the estimated treatment difference. The lower and upper limits are 4 and 41 meters. A confidence interval is about the uncertainty of the estimated treatment effect under the statistical framework; it is not a prediction interval for individual patients.

Why doesn't the P-value measure treatment effect size?

The primary P-value was 0.0125. A P-value measures the strength of evidence against a null hypothesis under a specified testing framework. It depends on both the magnitude of an observed difference and the amount of information available. The estimate and confidence interval are therefore necessary to understand the size and precision of the treatment effect.

Why are the Week 8, Week 11, and Week 12 results not interchangeable?

They answer related but distinct time-point questions. The primary endpoint was defined at Week 12, while Week 8 and Week 11 were secondary endpoints. The estimates were 17, 13, and 23 meters, respectively. Even though they concern the same underlying measure, their inferential roles in the trial are different.

Why does the analysis population matter?

The primary result was based on the mITT group of n=228, defined by access to 0.25 mg tablets at randomization. Post-hoc analyses of the entire study population used all enrolled subjects regardless of tablet strength availability. Comparing these results therefore involves more than comparing two estimates: the populations being analyzed are different.

Why should subgroup P-values be interpreted cautiously?

The post-hoc analyses by baseline WHO Functional Classification illustrate the issue. A subgroup estimate can be statistically significant in one category and not in another without demonstrating that the underlying treatment effects truly differ. A formal interaction test is generally needed to evaluate treatment-effect heterogeneity.

11. Secondary and Post-Hoc Results in Context

AnalysisEstimate95% CIP-valueInterpretive role
Primary 6MWD, Week 12, mITT23 meters4–41 meters0.0125Registered primary comparison
6MWD, Week 11, mITT13 meters−2–33 meters0.0653Secondary time point
6MWD, Week 8, mITT17 meters1–33 meters0.0307Secondary time point
6MWD, Week 4, mITT12 meters0–24 meters0.0518Secondary time point
6MWD, WHO III or IV26 meters1–49 meters0.0326Post-hoc subgroup
6MWD, WHO I or II16.0 meters−15–47 meters0.2275Post-hoc subgroup
6MWD, idiopathic or heritable PAH32 meters10–55 meters0.0024Post-hoc subgroup
6MWD, entire study population, Week 1225.5 meters10–41 meters0.0001Post-hoc full population
6MWD, entire study population, Week 1117 meters3–33 meters0.0025Post-hoc full population
6MWD, entire study population, Week 820 meters7–34 meters0.0008Post-hoc full population
6MWD, entire study population, Week 414 meters3.9–25 meters0.0025Post-hoc full population

The pattern across the reported 6MWD analyses is statistically informative because the estimates are not identical at each time point or population definition. That is expected: treatment effects are estimated with sampling variability, and the analyzed populations differ. The correct comparison is therefore not to look for a single universal number, but to identify which estimate answers the prespecified primary question.

12. Safety

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk for the reported safety category.

ArmSerious adverse events affected / at risk
Placebo15/77
UT-15C (Oral Treprostinil)27/151
Clinical Biostats interpretation

The reported serious-adverse-event counts are 15/77 for placebo and 27/151 for UT-15C. These denominators are the affected/at-risk figures reported in the ClinicalTrials.gov record and should not be replaced with the overall enrollment of 349 or the primary mITT size of 228.

The ClinicalTrials.gov record does not provide a formal between-arm hypothesis test for these serious-adverse-event figures, so this page does not calculate one or infer a safety comparison beyond the reported counts.

13. Multiplicity and Alpha

The primary analysis notes state that all alpha was spent on the mITT subgroup. They also describe a sample-size framework using a significance level of 0.01 with a two-sided hypothesis.

This matters because a clinical trial can contain many endpoints and repeated time points, while its confirmatory statistical framework may allocate the type I error budget to a particular primary comparison. The fact that later secondary and post-hoc analyses have P-values does not automatically give them the same confirmatory status as the registered primary endpoint.

Do not treat every reported P-value as equally confirmatory. The Week 12 primary 6MWD analysis was the registered primary analysis. Week 4, Week 8, and Week 11 6MWD results were secondary analyses, while the subgroup and entire-study-population 6MWD analyses were post-hoc. Their statistical roles are therefore different even when the same ANCOVA method was used.

14. Missing Data and Measurement Considerations

The ClinicalTrials.gov record identifies one explicit missing-baseline issue: two subjects, one in the placebo arm and one in the oral-treprostinil arm, from the primary analysis population of n=228 did not have a Baseline dyspnea-fatigue index score and were not included in that analysis.

The ClinicalTrials.gov record do not specify an imputation method for that missing Dyspnea-Fatigue Index information, nor do they provide a broader missing-data strategy for the primary 6MWD analysis. This page therefore does not infer an imputation procedure.

Statistical principle: excluding observations with missing baseline information and imputing missing follow-up outcomes are different decisions. The ClinicalTrials.gov record documents the former for the Dyspnea-Fatigue Index but does not establish a general imputation strategy for the trial.

15. Blinding and Randomization

FREEDOM-M is identified in the registry as randomized, parallel, and double-masked. Randomization provides the basis for comparing treatment groups without deliberately assigning treatment according to baseline prognosis, while double masking is intended to reduce knowledge of treatment assignment during the trial.

The analysis notes specify a 2:1 allocation ratio between oral treprostinil and placebo. The ClinicalTrials.gov record does not identify the complete randomization stratification scheme, so no additional stratification factors are presented here.

16. What the Different Endpoints Measure

6MWD

A continuous measure reported in meters. The primary endpoint was the placebo-corrected change from Baseline to Week 12.

WHO Functional Classification

A functional classification endpoint analyzed with a Wilcoxon rank sum test in the posted analysis.

Borg Dyspnea Score

A scale-based endpoint analyzed using a Wilcoxon rank-sum approach.

Clinical Worsening

A binary endpoint analyzed using Fisher Exact.

These endpoints should not be interpreted as interchangeable. The six-minute walk distance provides a continuous measure of walking performance, while the other endpoints capture different dimensions of functional status, symptoms, or clinical worsening. A treatment effect can therefore differ in magnitude and statistical evidence across endpoints without the results being mathematically inconsistent.

17. Reading the Primary Result Correctly

Estimate

The primary estimated placebo-corrected difference was 23 meters. This is the central estimate of the treatment-group difference reported by the registry.

Precision

The two-sided 95% CI was 4 to 41 meters. The width of the interval communicates the uncertainty surrounding the estimated difference.

Statistical evidence

The primary P-value was 0.0125. It addresses the statistical testing question; it does not quantify clinical importance or the probability that the treatment is effective.

Analysis population

The primary analysis used the modified intention-to-treat population of n=228. This is essential context for the reported estimate and confidence interval.

18. Primary vs Secondary vs Post-Hoc Analysis

Analysis classExample in FREEDOM-MInterpretive status
Primary 6MWD, Baseline and Week 12 Registered primary efficacy question
Secondary 6MWD at Weeks 4, 8, and 11 Additional prespecified outcome analyses
Secondary WHO Functional Classification, Borg Dyspnea Score, Dyspnea-Fatigue Index, Clinical Worsening Assessment Additional outcome analyses
Post-hoc 6MWD by baseline WHO Functional Classification and PAH etiology Exploratory subgroup analyses
Post-hoc 6MWD for the entire study population Exploratory analysis using all enrolled subjects

This hierarchy is important because statistical significance is not the only dimension of interpretation. The prespecified primary endpoint has a different evidentiary role from a post-hoc subgroup analysis, even if the latter has a smaller P-value.

19. Why This Trial Matters Statistically

FREEDOM-M is a useful teaching case because the posted analyses bring several core clinical-trial concepts together in a single randomized phase 3 study: a continuous primary endpoint, ANCOVA, a modified intention-to-treat population, a two-sided confidence interval, a prespecified superiority hypothesis, nonparametric secondary analyses, an exact categorical test, and post-hoc population and subgroup analyses.

ConceptHow it appears in FREEDOM-M
RandomizationRandomized parallel-group design
BlindingDouble masking
Allocation ratio2:1 oral treprostinil to placebo in the primary analysis notes
Modified ITTPrimary analysis in the mITT group, n=228
ANCOVAPrimary and post-hoc 6MWD analyses
Hodges-Lehmann estimateReported effect measure for the primary 6MWD analysis and several other analyses
Confidence intervalTwo-sided 95% CI for the primary treatment estimate
P-value0.0125 for the primary Week 12 analysis
Wilcoxon testWHO Functional Classification, Borg Dyspnea Score, Dyspnea-Fatigue Index
Fisher exact testClinical Worsening Assessment
Multiplicity / alphaAnalysis notes state that all alpha was spent on the mITT subgroup
Post-hoc analysisWHO functional-classification subgroups, PAH etiology, and entire study population
Missing dataTwo subjects lacked Baseline Dyspnea-Fatigue Index scores

20. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary analysis estimated a 23-meter placebo-corrected difference in change in 6MWD from Baseline to Week 12, with a two-sided 95% CI of 4 to 41 meters and a P-value of 0.0125 in the mITT population of n=228.

Clinical interpretation

The primary endpoint was six-minute walk distance, an objective measure of patient functional status described in the registry as correlating with the historical clinical standard for assessing functional status in PAH. Clinical meaning should consider the magnitude of the estimated difference alongside uncertainty and the other reported endpoints.

The distinction matters. Statistical evidence tells us how compatible the observed data are with a specified null hypothesis under the analysis framework. Clinical interpretation asks what the size and nature of the observed difference mean in the context of patient function. Neither question is fully answered by the P-value alone.

21. Important Limitations and Interpretation Issues

22. A Practical Statistical Reading of FREEDOM-M

Design

Randomized, double-blind, parallel phase 3 trial

The trial enrolled 349 participants and compared oral treprostinil with placebo.

Primary population

Modified intention-to-treat group

The primary analysis used n=228 subjects with access to 0.25 mg tablets at randomization, with all alpha spent on this subgroup.

Primary endpoint

6MWD at Week 12

The registered endpoint was the placebo-corrected change in six-minute walk distance from Baseline to Week 12.

Primary analysis

ANCOVA

The treatment comparison produced a Hodges-Lehmann estimate of 23 meters, with a two-sided 95% CI of 4 to 41 meters and P=0.0125.

Additional evidence

Secondary and post-hoc analyses

Additional 6MWD time points, functional measures, clinical worsening, subgroup analyses, and entire-study-population analyses were also reported.

23. What the Primary Estimate Does — and Does Not — Mean

What it means

The 23-meter estimate summarizes the reported placebo-corrected difference in change in 6MWD from Baseline to Week 12 under the primary analysis framework.

What it does not mean

It does not mean that every participant receiving oral treprostinil increased their walking distance by exactly 23 meters. It is a population-level treatment comparison, not an individual prediction.

What the confidence interval adds

The 4 to 41 meter interval shows the uncertainty around the estimated treatment difference. The interval is more informative about precision than the point estimate alone.

What the P-value adds

The 0.0125 P-value provides the statistical-test result under the specified superiority framework. It should not be interpreted as the probability that the treatment effect is exactly 23 meters or as a measure of clinical importance.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Calculators

26. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods behind randomized trials, continuous outcomes, confidence intervals, nonparametric tests, and categorical comparisons.

27. Record Summary

FREEDOM-M provides a compact example of how a randomized phase 3 trial can combine several statistical approaches around a clinically meaningful functional endpoint. The registered primary endpoint was the placebo-corrected change in six-minute walk distance from Baseline to Week 12, analyzed by ANCOVA in a modified intention-to-treat population of n=228. The reported Hodges-Lehmann estimate was 23 meters, with a two-sided 95% CI of 4 to 41 meters and a P-value of 0.0125.

The broader statistical record includes secondary 6MWD assessments at Weeks 4, 8, and 11, nonparametric analyses of functional and symptom measures, a Fisher exact analysis of clinical worsening, and post-hoc analyses by baseline functional classification, PAH etiology, and the entire study population. Reading these results correctly requires keeping the endpoint hierarchy, analysis population, confidence intervals, P-values, and post-hoc status distinct rather than reducing the trial to a single numerical result.

Clinical Biostats methodology: A trial-results page should separate the reported statistical evidence from its educational interpretation. For FREEDOM-M, the most important distinctions are the registered Week 12 primary endpoint, the mITT population of n=228, the ANCOVA/Hodges-Lehmann analysis, and the different inferential roles of secondary and post-hoc analyses.