← Clinical Trials
Pulmonary Arterial Hypertension Phase 4 Randomized NCT00303459

COMPASS-2: Complete Statistical Analysis of Bosentan in Pulmonary Arterial Hypertension

An independent statistical analysis of the randomized COMPASS-2 trial evaluating the effects of the combination of bosentan and sildenafil versus sildenafil monotherapy on pulmonary arterial hypertension, with emphasis on the primary morbidity/mortality time-to-event endpoint and the reported functional, biomarker, quality-of-life, and safety outcomes.

Phase 4  ·  Completed  ·  Enrollment 334  ·  Start May 2006  ·  Primary completion December 2013
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

COMPASS-2 was a randomized, parallel, quadruple-masked phase 4 trial in pulmonary arterial hypertension. The registered primary endpoint was time to first confirmed morbidity/mortality event up to the end of study, analyzed with a log-rank test and expressed as a hazard ratio.

334
Enrollment
All randomized set
2
Arms
Bosentan vs placebo
0.831
Primary HR
97.31% CI 0.582–1.187
0.2508
Primary P-value
Two-sided superiority analysis
FeatureCOMPASS-2
Trial nameCOMPASS-2
ClinicalTrials.gov identifierNCT00303459
PhasePhase 4
ConditionPulmonary Arterial Hypertension
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment334
Arms2
Interventions recordedBosentan (drug); placebo (drug)
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
StatusCompleted
Start2006-05
Primary completion2013-12
Lead sponsorActelion
Sponsor typeIndustry

2. Clinical Question

The central question was whether adding bosentan to sildenafil produced a different time to first confirmed morbidity/mortality event than sildenafil monotherapy in participants with pulmonary arterial hypertension.

Population

Participants with pulmonary arterial hypertension enrolled in the COMPASS-2 randomized phase 4 trial.

Intervention

The trial evaluated the combination of bosentan and sildenafil; the intervention recorded in the ClinicalTrials.gov record was bosentan.

Comparator

Sildenafil monotherapy, represented in the randomized comparison by placebo as the additional study drug.

Primary question

Does the bosentan combination change the time to first confirmed morbidity/mortality event compared with sildenafil monotherapy?

3. Trial Design

01
Randomize334 participants
02
Two armsBosentan vs placebo
03
Quadruple maskParallel design
04
Follow-upPrimary endpoint to study end
05
AnalyzeLog-rank and reported secondary methods
STUDY ARM · BOSENTAN

Bosentan combination

  • Bosentan was the recorded active intervention.
  • The trial title describes the comparison as bosentan plus sildenafil versus sildenafil monotherapy.
  • Serious adverse events: 73/159 affected participants.
CONTROL ARM · PLACEBO

Sildenafil monotherapy

  • Placebo was the recorded comparator intervention.
  • The trial title describes this group as sildenafil monotherapy.
  • Serious adverse events: 102/174 affected participants.
Masking is statistically relevant. The registry describes COMPASS-2 as quadruple-masked. Masking can reduce the potential for knowledge of treatment assignment to influence participant-reported outcomes, investigator assessment, or other aspects of trial conduct. The ClinicalTrials.gov record identifies the masking structure but do not specify which four parties were masked.

4. Endpoints

The registry contains one registered primary endpoint and multiple posted secondary outcomes. The primary endpoint is a time-to-event measure, while the secondary outcomes span time-to-event, continuous, binary, biomarker, and quality-of-life measures.

Primary Endpoint

EndpointDefinition / time frameAnalysis
Time to First Confirmed Morbidity/Mortality Event up to the End of Study From baseline to end of study, approximately 86 months. Kaplan-Meier estimate of percentage of participants without a morbidity/mortality event. A morbidity/mortality event is defined as the occurrence of death; hospitalization for worsening or complication of PAH or intravenous prostanoid initiation; atrial septostomy; lung transplantation; or worsening PAH, defined as "moderately" or "markedly" worsened PAH symptoms using a patient global self-assessment (PGSA) scale AND initiation of inhaled or subcutaneous prostanoids or the disease progression package (open-label bosentan). If a patient replied "no change" or "mildly worse" on the PGSA, a decrease in 6MWT of 20% versus last visit or 30% versus baseline is also required to confirm the event. Log-rank test; hazard ratio

Secondary Endpoints

EndpointTime frameMethodEffect measure
Time to First Confirmed Death, Hospitalization for Worsening or Complication of PAH or Initiation of Intravenous Prostanoids, Atrial Septostomy, or Lung TransplantationBaseline to end of study, approximately 86 monthsLog-rank testHazard ratio
Change From Baseline to Week 16 in 6 Minute Walk Test (6MWT)From baseline to week 16Wilcoxon / Mann-Whitney testMean Difference (Net)
Number of Participants With Improved, No Change, or Worsened World Health Organisation Functional Class From Baseline to Week 16From baseline to Week 16Fisher exact testRelative risk of improvement
Time to Death of All Causes From Baseline to End of StudyBaseline to end of study, approximately 86 monthsLog-rank testHazard ratio
Adjusted Percentage Ratio From Baseline in N-terminal Pro-B-type Natriuretic Peptide (NT-pro-BNP)Baseline to Month 20MMRM (mixed model for repeated measures)Percentage change over placebo
Change From Baseline to Week 16 in Borg Dyspnea IndexBaseline to Week 16Wilcoxon / Mann-Whitney testMean Difference (Net)
Change From Baseline to Week 16 in the EuroQol 5 Dimensions (EQ-5D) Questionnaire Calculated ScoreFrom baseline to Week 16Wilcoxon / Mann-Whitney testMean Difference (Net)
Change From Baseline to Week 16 in the EuroQol 5 Dimensions (EQ-5D) Visual Analogue Scale ScoreBaseline to Week 16Wilcoxon / Mann-Whitney testMean Difference (Net)

5. Statistical Methodology

Primary time-to-event analysis

The primary endpoint was analyzed in the all randomized set using a log-rank test. The effect measure was a hazard ratio comparing bosentan with placebo, with a two-sided 97.31% confidence interval and a superiority hypothesis.

Primary analysis structure
Randomized participants → time-to-event follow-up → Kaplan-Meier framework → log-rank comparison → hazard ratio

The registry reports the outcome unit as percentage of participants-Kaplan Meier and identifies the formal comparison as Log Rank. The ClinicalTrials.gov record does not report a Cox model for the primary analysis.

Kaplan-Meier estimation

The primary endpoint is naturally represented as a time-to-event distribution because participants can experience the first morbidity/mortality event at different times and participants who have not experienced the event by the relevant observation point contribute follow-up information without necessarily having an event.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni represents the number at risk immediately before that time. The registry describes the primary outcome as a Kaplan-Meier estimate of the percentage of participants without a morbidity/mortality event.

Hazard ratio

The primary effect measure was a hazard ratio of 0.831. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the bosentan group relative to placebo within the framework of the time-to-event analysis.

Interpretation of the primary hazard ratio
HR = 0.831  →  estimated hazard in bosentan group relative to placebo

The hazard ratio is a relative time-to-event measure. It is not an absolute risk difference, a probability that a participant will benefit, or the percentage of participants who avoid an event.

Nonparametric analysis of continuous outcomes

The 6MWT, Borg Dyspnea Index, EQ-5D calculated score, and EQ-5D Visual Analogue Scale were analyzed with the Wilcoxon (Mann-Whitney) method. The registry reports a mean difference (net) as the effect measure. This pairing is worth noting because the inferential test is nonparametric while the effect measure reported by the registry is a mean difference.

Fisher exact test for functional class

The WHO functional-class outcome classified participants as improved, unchanged, or worsened from baseline to Week 16. The reported method was Fisher exact testing, with relative risk of improvement as the effect measure.

Repeated-measures analysis of NT-pro-BNP

The NT-pro-BNP endpoint used a repeated-measures analysis, normalized here as an MMRM (mixed model for repeated measures). The analysis population was all randomized patients with a baseline and at least one post-baseline value, with assessments considered where at least 60% of patients had a post-baseline value. The reported effect measure was percentage change over placebo.

Analysis populations

OutcomeAnalysis population
Primary endpointAll randomized set
Secondary time-to-event outcomesAll randomized set
6MWTAll randomized set
WHO functional classAll randomized set
Borg Dyspnea IndexAll randomized set
EQ-5D calculated scoreAll randomized set
EQ-5D Visual Analogue ScaleAll randomized set
NT-pro-BNPAll randomized patients with a baseline and at least one post-baseline value; assessments considered where at least 60% of patients have a post-baseline value

6. Results: Primary Endpoint

Time to First Confirmed Morbidity/Mortality Event

Hazard ratio for first confirmed morbidity/mortality event

0.831

97.31% CI: 0.582–1.187   ·   P = 0.2508

Analysis population: all randomized set  ·  Method: log-rank test  ·  Hypothesis: superiority

Primary endpointEstimateConfidence intervalP-valueMethod
Time to First Confirmed Morbidity/Mortality Event up to the End of Study HR 0.831 97.31% two-sided CI 0.582–1.187 0.2508 Log-rank test
Clinical Biostats interpretation

The estimated hazard ratio of 0.831 means that the estimated instantaneous rate of the first confirmed morbidity/mortality event was about 83.1% as high in the bosentan group as in the placebo group, within the time-to-event analysis. Expressed as a simple relative interpretation, this corresponds to an estimated hazard about 16.9% lower with bosentan.

The hazard ratio does not mean that 16.9% of participants avoided an event, that 16.9% fewer participants experienced an event, or that every participant experienced the same proportional reduction. It is a relative measure of event rate over time.

The two-sided 97.31% confidence interval of 0.582–1.187 indicates substantial statistical uncertainty around the point estimate. Because the interval includes 1, values compatible with a lower event hazard as well as values compatible with a higher event hazard remain within the reported interval.

The P-value of 0.2508 describes the evidence against the null hypothesis under the specified superiority testing framework; it does not measure the magnitude or clinical importance of the treatment effect. The point estimate and confidence interval are therefore essential parts of the result.

The analysis is based on time-to-event data, so censoring and the construction of the risk sets matter. The ClinicalTrials.gov record identifies a log-rank analysis and a hazard-ratio effect measure but do not report a proportional-hazards assessment or a separate Cox-model analysis. The primary analysis population was the all randomized set.

Educational note: a valid Kaplan-Meier curve cannot be reconstructed from the reported hazard ratio, confidence interval, and P-value alone. The ClinicalTrials.gov record does not contain the underlying event and censoring times needed to construct a participant-level curve.

7. Secondary Results

Time to First Confirmed Death, Hospitalization, Intravenous Prostanoids, Atrial Septostomy, or Lung Transplantation

Hazard ratio

0.963

95% CI: 0.673–1.380   ·   P = 0.8385

Log-rank test  ·  All randomized set  ·  Superiority

This secondary time-to-event endpoint was narrower than the primary morbidity/mortality composite because the registry definition excludes the worsening-PAH component included in the primary endpoint. The estimated hazard ratio of 0.963 is close to 1, while the confidence interval extends on both sides of 1.

Change From Baseline to Week 16 in 6 Minute Walk Test

EndpointNet mean difference95% CIP-valueMethod
6MWT change from baseline to Week 1621.8 m5.9 to 37.80.0106Wilcoxon (Mann-Whitney)

The reported net mean difference was 21.8 m, with a 95% confidence interval from 5.9 to 37.8 and P = 0.0106. The registry reports the analysis as a Wilcoxon (Mann-Whitney) comparison and labels the effect measure as mean difference (net).

WHO Functional Class

EndpointRelative risk of improvement95% CIP-valueMethod
Improved, no change, or worsened WHO Functional Class from baseline to Week 160.980.60 to 1.611.0000Fisher exact test

The reported relative risk of improvement was 0.98, with a 95% confidence interval of 0.60 to 1.61 and P = 1.0000. Fisher's exact test was used for the categorical outcome.

Time to Death of All Causes

Hazard ratio for all-cause death

0.855

95% CI: 0.544–1.344   ·   P = 0.4974

Log-rank test  ·  All randomized set  ·  Superiority

The all-cause mortality analysis estimated a hazard ratio of 0.855. Its 95% confidence interval was 0.544–1.344, with P = 0.4974. As with the primary time-to-event endpoint, the hazard ratio is a relative event-rate measure rather than an absolute mortality probability.

Adjusted Percentage Ratio From Baseline in NT-pro-BNP

Percentage change over placebo

−23.52

95% CI: −33.69 to −11.79   ·   P = 0.0003

MMRM (mixed model for repeated measures)  ·  Baseline to Month 20

The reported percentage change over placebo was −23.52, with a 95% confidence interval from −33.69 to −11.79 and P = 0.0003. The repeated-measures analysis used all randomized patients with a baseline and at least one post-baseline value, with assessments considered where at least 60% of patients had a post-baseline value.

Borg Dyspnea Index

EndpointNet mean difference95% CIP-valueMethod
Change from baseline to Week 16 in Borg Dyspnea Index−0.01−0.42 to 0.390.9566Wilcoxon (Mann-Whitney)

The reported net mean difference was −0.01, with a 95% confidence interval of −0.42 to 0.39 and P = 0.9566. The interval spans zero, which is the null value for a difference measure.

EQ-5D Calculated Score

EndpointNet mean difference95% CIP-valueMethod
Change from baseline to Week 16 in EQ-5D Questionnaire Calculated Score0.020−0.036 to 0.0760.5571Wilcoxon (Mann-Whitney)

The reported net mean difference was 0.020, with a 95% confidence interval of −0.036 to 0.076 and P = 0.5571.

EQ-5D Visual Analogue Scale

EndpointNet mean difference95% CIP-valueMethod
Change from baseline to Week 16 in EQ-5D Visual Analogue Scale Score0.2−3.7 to 4.00.4086Wilcoxon (Mann-Whitney)

The reported net mean difference was 0.2, with a 95% confidence interval of −3.7 to 4.0 and P = 0.4086.

8. Secondary Results in One View

OutcomeEffect estimateConfidence intervalP-valueStatistical method
Primary morbidity/mortality eventHR 0.83197.31% CI 0.582–1.1870.2508Log-rank
Death / hospitalization / IV prostanoids / septostomy / transplantHR 0.96395% CI 0.673–1.3800.8385Log-rank
6MWTMean difference 21.895% CI 5.9–37.80.0106Wilcoxon (Mann-Whitney)
WHO functional-class improvementRR 0.9895% CI 0.60–1.611.0000Fisher exact
All-cause deathHR 0.85595% CI 0.544–1.3440.4974Log-rank
NT-pro-BNP−23.5295% CI −33.69 to −11.790.0003MMRM
Borg Dyspnea IndexMean difference −0.0195% CI −0.42 to 0.390.9566Wilcoxon (Mann-Whitney)
EQ-5D calculated scoreMean difference 0.02095% CI −0.036 to 0.0760.5571Wilcoxon (Mann-Whitney)
EQ-5D Visual Analogue ScaleMean difference 0.295% CI −3.7 to 4.00.4086Wilcoxon (Mann-Whitney)
Multiple endpoints require context. COMPASS-2 reports one primary endpoint and several secondary endpoints. The ClinicalTrials.gov record identifies all of these hypothesis tests as superiority analyses, but do not provide an alpha-allocation or multiplicity-adjustment strategy for the secondary endpoints. Therefore, the individual P-values should not be interpreted as though the ClinicalTrials.gov record establishes a separate familywise error-control procedure for all secondary outcomes.

9. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm. These figures are counts of affected participants over the corresponding at-risk denominators.

Safety measureBosentanPlacebo
Participants with serious adverse events / at risk73/159102/174

Bosentan

73/159 participants were reported as affected by serious adverse events among those at risk.

Placebo

102/174 participants were reported as affected by serious adverse events among those at risk.

The serious-adverse-event data are presented separately from the efficacy analyses because safety and efficacy answer different questions. The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates, confidence interval, or P-value.

10. Statistical Methods Explained

Why was a log-rank test used for the primary endpoint?

The primary outcome is a time-to-event endpoint: participants can experience the first morbidity/mortality event at different points during follow-up, while some observations may remain event-free at the end of their available follow-up. The log-rank test is designed to compare the survival experience of two groups across follow-up rather than comparing a single binary event rate at one fixed time.

What does a hazard ratio of 0.831 mean?

A hazard ratio of 0.831 represents the estimated event rate in the bosentan group relative to the placebo group within the time-to-event analysis. A value below 1 corresponds to a lower estimated instantaneous event rate. It does not mean that 16.9% fewer participants necessarily experienced the event, nor does it directly provide an absolute risk reduction.

Why is the confidence interval important?

The primary 97.31% confidence interval ranges from 0.582 to 1.187. The point estimate alone could suggest a lower event hazard, but the interval shows that the statistical uncertainty is broad enough to include 1. A confidence interval therefore gives information about precision that cannot be obtained from the point estimate alone.

Why doesn't the P-value measure effect size?

The primary P-value of 0.2508 summarizes evidence against the null hypothesis under the specified testing framework. It does not quantify how large the treatment effect is. Effect size is described by the hazard ratio, while the confidence interval describes uncertainty around that estimate.

Why was Fisher's exact test used for functional class?

The WHO functional-class endpoint is categorical: participants were classified as improved, unchanged, or worsened. Fisher's exact test is a method for comparing categorical outcomes between two groups and is particularly useful when cell counts may be limited. The registry pairs this test with a relative risk of improvement.

What does the MMRM analysis contribute for NT-pro-BNP?

The NT-pro-BNP endpoint involves measurements over time rather than a single follow-up observation. The reported repeated-measures analysis, normalized as MMRM, uses longitudinal information while restricting the analysis to randomized patients with a baseline and at least one post-baseline value. The registry additionally states that assessments were considered where at least 60% of patients had a post-baseline value.

11. Understanding the Primary Endpoint More Deeply

Composite endpoint

The primary endpoint combines several clinically meaningful events, including death, hospitalization for worsening or complication of PAH or intravenous prostanoid initiation, atrial septostomy, lung transplantation, and worsening PAH.

Time matters

Because the endpoint is time-to-event, the analysis considers not only whether an event occurred but also when the first confirmed event occurred.

Randomized analysis

The primary analysis used the all randomized set, preserving the randomized comparison as the basis for the primary efficacy result.

Relative effect

The hazard ratio describes relative event rates. It should be interpreted alongside the confidence interval and the underlying Kaplan-Meier framework rather than as an absolute probability.

The primary endpoint illustrates an important distinction in clinical-trial statistics: a composite outcome can be statistically analyzed as a single time-to-event endpoint even though it contains several different clinical events. The resulting hazard ratio therefore describes the first qualifying component of the composite, not a separate effect estimate for each component.

The ClinicalTrials.gov record also make clear why the analysis population matters. All randomized participants form the analysis population for the primary endpoint, so the comparison remains anchored to randomized treatment assignment rather than being restricted to participants who remained on treatment or completed every assessment.

12. Interpreting the Secondary Endpoint Pattern

The secondary results span several statistical scales. The time-to-event outcomes use hazard ratios, the functional-class analysis uses a relative risk, the 6MWT and quality-of-life outcomes use mean differences, and the NT-pro-BNP analysis reports a percentage change over placebo. These effect measures cannot be placed on one common numerical scale.

Effect measureUsed in COMPASS-2 forNull valueCore interpretation
Hazard ratioPrimary morbidity/mortality event; composite secondary time-to-event endpoint; all-cause death1Relative instantaneous event rate
Mean difference6MWT; Borg Dyspnea Index; EQ-5D calculated score; EQ-5D Visual Analogue Scale0Difference between groups on the reported scale
Risk ratioWHO functional-class improvement1Relative probability of improvement
Percentage change over placeboNT-pro-BNP0Reported relative change compared with placebo

This distinction is important when reading the results. For example, a hazard ratio of 0.831 and a mean difference of 21.8 are not alternative ways of reporting the same quantity. They answer different statistical questions about different endpoints.

13. Confidence Intervals and What They Show

Several COMPASS-2 results demonstrate why confidence intervals should be read alongside P-values. The primary hazard ratio is 0.831 with a 97.31% confidence interval of 0.582–1.187. The secondary all-cause mortality hazard ratio is 0.855 with a 95% confidence interval of 0.544–1.344. Both intervals cross the null value of 1.

Difference measures
Null difference = 0

For the 6MWT, the 95% confidence interval is 5.9 to 37.8, while the Borg Dyspnea Index interval is −0.42 to 0.39. For difference measures, zero is the corresponding null value.

The intervals also illustrate how the amount of uncertainty depends on the endpoint. The primary time-to-event result has a relatively broad range around the hazard ratio, while the continuous and longitudinal endpoints have intervals expressed in their own measurement units or relative-change scale.

14. Planned Statistical Reasoning for Time-to-Event Outcomes

The primary endpoint can be understood as a sequence of statistical operations rather than as a single calculation.

Step 1

Define the first qualifying event

The registry specifies a composite morbidity/mortality definition and measures time from baseline to the first confirmed event.

Step 2

Retain timing information

Participants contribute information according to their available follow-up, rather than reducing the entire endpoint to a single yes/no measurement without regard to timing.

Step 3

Estimate event-free experience

The registry specifies a Kaplan-Meier estimate of the percentage of participants without a morbidity/mortality event.

Step 4

Compare randomized groups

The reported formal comparison is a log-rank test between bosentan and placebo.

Step 5

Quantify relative effect

The treatment effect is summarized by a hazard ratio with a confidence interval.

This structure helps explain why a time-to-event result cannot be interpreted solely from the final percentage of participants who experienced an event. The timing of events and censoring information are part of the analysis.

15. Missing Data and Analysis Eligibility

The ClinicalTrials.gov record provides a specific eligibility rule for the NT-pro-BNP repeated-measures analysis: patients needed a baseline and at least one post-baseline value, and assessments were considered where at least 60% of patients had a post-baseline value.

Important distinction: this is not the same analysis population definition as the primary endpoint, which used the all randomized set. A reader should therefore avoid assuming that every reported endpoint includes exactly the same participants or observations.

The ClinicalTrials.gov record does not specify an imputation method for missing NT-pro-BNP values, nor do they provide a separate missing-data sensitivity analysis. No additional imputation strategy is therefore attributed to COMPASS-2 on this page.

16. What the Primary Result Does — and Does Not — Establish

Statistical interpretation

The primary analysis estimated a hazard ratio of 0.831 for time to first confirmed morbidity/mortality event, with a two-sided 97.31% confidence interval of 0.582–1.187 and P = 0.2508.

What the estimate means

The point estimate is below 1, corresponding to a lower estimated instantaneous event rate in the bosentan group relative to placebo within the analyzed time-to-event framework.

What it does not mean

The result does not provide a percentage of patients who benefited, an absolute reduction in morbidity/mortality events, or an individual-level prediction. Those quantities require different information and effect measures.

Why uncertainty matters

The confidence interval extends from below 1 to above 1. Consequently, the point estimate should not be treated as though it were a precise estimate of the treatment effect.

Why the P-value is secondary to the effect estimate

The P-value of 0.2508 is evidence under the specified hypothesis-testing framework; it is not a measure of effect magnitude. A small P-value would not by itself establish a large clinical effect, just as a larger P-value does not turn the hazard ratio into a null effect.

17. Limitations

18. Why This Trial Matters Statistically

COMPASS-2 is a useful teaching example because it places several common clinical-trial methods in the same randomized study. The primary endpoint requires time-to-event reasoning, while the secondary endpoints demonstrate how statistical methods change when the outcome changes from survival time to continuous measurements, categorical improvement, repeated biomarker measurements, and quality-of-life scales.

ConceptHow it appears in COMPASS-2
RandomizationRandomized parallel comparison of two study arms
BlindingQuadruple masking
Kaplan-Meier estimationPrimary endpoint reported as percentage of participants without a morbidity/mortality event
Log-rank testPrimary and secondary time-to-event comparisons
Hazard ratioPrimary morbidity/mortality endpoint, composite secondary time-to-event endpoint, and all-cause death
Confidence intervalQuantifies uncertainty around the reported effect estimates
Wilcoxon / Mann-Whitney6MWT, Borg Dyspnea Index, EQ-5D calculated score, and EQ-5D Visual Analogue Scale
Fisher exact testWHO functional-class improvement outcome
MMRMRepeated-measures analysis of NT-pro-BNP through Month 20
Risk ratioRelative risk of improvement in WHO functional class
Multiple endpointsOne primary endpoint accompanied by multiple secondary outcomes

19. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary hazard-ratio estimate was below 1, but the reported 97.31% confidence interval included 1 and the P-value was 0.2508. Secondary endpoints produced a range of estimates across time-to-event, continuous, categorical, and longitudinal analyses.

Clinical interpretation

The ClinicalTrials.gov record describes differences in several measured outcomes, including the reported 6MWT and NT-pro-BNP results, as well as serious adverse events. Clinical meaning requires considering the specific endpoint, its scale, its uncertainty, and the design of the comparison rather than reducing the trial to one statistic.

This distinction is particularly important for COMPASS-2 because its endpoints are heterogeneous. A change in 6MWT, a percentage change in NT-pro-BNP, a hazard ratio for mortality, and a relative risk of functional-class improvement represent different clinical quantities. They should be interpreted on their own scales before considering the broader evidence from the trial.

20. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the survival-analysis, categorical-data, nonparametric, longitudinal, and effect-measure methods represented in COMPASS-2.

23. Record Summary

COMPASS-2 provides a compact example of how a randomized clinical trial can combine several statistical frameworks around a single clinical question. Its primary endpoint was a time-to-first confirmed morbidity/mortality event analyzed in the all randomized set with Kaplan-Meier estimation and a log-rank comparison, producing a hazard ratio of 0.831, a two-sided 97.31% confidence interval of 0.582–1.187, and P = 0.2508.

The secondary results demonstrate the importance of matching the statistical method to the endpoint: log-rank testing for time-to-event outcomes, Wilcoxon (Mann-Whitney) testing for several Week 16 continuous outcomes, Fisher's exact testing for WHO functional-class categories, and MMRM for repeated NT-pro-BNP measurements. The reported serious-adverse-event figures were 73/159 for bosentan and 102/174 for placebo.

The most important statistical lesson is that the trial cannot be reduced to a single P-value. The treatment comparison has to be understood through the randomized design, the definition and timing of the primary composite endpoint, the hazard-ratio effect measure, the width and null-crossing of its confidence interval, the distinct methods used for secondary outcomes, the analysis populations, and the absence of a registry-reported multiplicity strategy for the secondary endpoints.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry. The goal is to reconstruct the statistical story of the trial while clearly separating reported numerical evidence from educational interpretation and avoiding unsupported assumptions about analyses that are not documented in the ClinicalTrials.gov record.