← Clinical Trials
Pulmonary Arterial Hypertension Phase 3 Completed NCT01178073

AMBITION: Complete Statistical Analysis of Ambrisentan and Tadalafil in Pulmonary Arterial Hypertension

An independent statistical review of the randomized phase 3 AMBITION trial evaluating first-line ambrisentan plus tadalafil combination therapy versus ambrisentan or tadalafil monotherapy in subjects with pulmonary arterial hypertension.

Trial start: 2010-10-01  ·  Primary completion: 2014-07-31  ·  Enrollment: 610
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the data posted on ClinicalTrials.gov for NCT01178073.

1. Trial at a Glance

AMBITION was a randomized, parallel-group, quadruple-masked phase 3 treatment trial in pulmonary arterial hypertension. The registry reports 610 enrolled participants, 3 arms, 1 registered primary endpoint, 6 posted outcome measures, and 18 posted statistical analyses.

610
Enrolled
Phase 3 trial
3
Arms
Combination + 2 monotherapies
0.502
Primary HR
Combination vs pooled monotherapy
0.0002
Primary P-value
Stratified log-rank analysis
FeatureAMBITION
Trial nameAMBITION
ClinicalTrials.gov identifierNCT01178073
PhasePhase 3
Therapeutic areaPulmonology
ConditionHypertension, Pulmonary
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment610
Arms3
Registered primary endpoints1
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted6
Statistical analyses posted18
Lead sponsorGlaxoSmithKline
Sponsor typeIndustry
StatusCompleted

2. Clinical Question

The central statistical question was whether first-line ambrisentan plus tadalafil reduced the time to the first adjudicated clinical failure event compared with first-line monotherapy with either ambrisentan or tadalafil in subjects with pulmonary arterial hypertension.

Population

Subjects with pulmonary arterial hypertension participating in the AMBITION phase 3 randomized trial.

Intervention

First-line combination therapy with ambrisentan and tadalafil.

Comparator

First-line monotherapy with ambrisentan or tadalafil, analyzed both as a pooled monotherapy group and separately by monotherapy arm.

Primary question

Does combination therapy change the time to first adjudicated clinical failure through the Final Assessment Visit relative to monotherapy?

The registry defines clinical failure as the first adjudicated event among death, hospitalization for worsening pulmonary arterial hypertension, disease progression, or unsatisfactory long-term clinical response. The registered time frame is from baseline through the Final Assessment Visit, with an average of 609 days.

3. Trial Design

01
Enroll610 participants
02
Randomize3 parallel arms
03
TreatCombination or monotherapy
04
AssessClinical failure and Week 24 measures
05
AnalyzemITT population
ARM 1 · 302 AT RISK FOR SERIOUS AEs

Ambrisentan + tadalafil

  • First-line combination therapy.
  • Compared with pooled monotherapy for the primary time-to-event analysis.
  • Also compared separately with ambrisentan monotherapy and tadalafil monotherapy.
MONOTHERAPY ARMS

Ambrisentan or tadalafil

  • Two first-line monotherapy groups.
  • Analyzed as a pooled monotherapy comparator for the primary comparison.
  • Also retained as separate comparators for additional primary analyses.
What the registry supports: the ClinicalTrials.gov record identifies the allocation as randomized, the design model as parallel, and masking as quadruple. They do not provide a treatment-sequence or crossover description, so no crossover effect is inferred here.

4. Analysis Population and Comparison Structure

The primary and posted secondary analyses in the ClinicalTrials.gov record use the modified intention-to-treat (mITT) population. The registry's secondary analyses also specify when only participants with available data at the relevant time points, or participants with a required response classification, were analyzed.

Analysis featureRegistry-supported description
Primary analysis populationmITT Population
Secondary biomarker analysismITT population; only participants with data available at the specified time points
Clinical response analysismITT population; only participants with a “Yes”/“No” response
6 Minute Walk Distance analysismITT population; only participants with baseline data
WHO Functional Class analysismITT population; only participants with baseline data
Primary pooled comparisonCombination therapy versus pooled ambrisentan or tadalafil monotherapy
Individual monotherapy comparisonsCombination therapy versus ambrisentan monotherapy; combination therapy versus tadalafil monotherapy

This distinction matters statistically. A randomized trial begins with treatment assignment, but an analysis population can impose additional eligibility conditions for a particular endpoint. For example, the Week 24 endpoints requiring observed measurements are explicitly restricted in the registry to participants with the necessary data. Such restrictions can change the population contributing information to that endpoint even when the overall trial remains randomized.

5. Primary Endpoint

EndpointRegistry definitionAnalysis
First adjudicated clinical failure Number of participants with first adjudicated clinical failure event, death, hospitalisation for worsening PAH, disease progression, unsatisfactory long-term clinical response, all through FAV Stratified log-rank test; hazard ratio from Cox proportional-hazards model

Time frame: From Baseline up to the Final Assessment Visit (FAV), with an average of 609 days.

The endpoint is a classic time-to-event outcome. The statistical target is not simply whether an event eventually occurred, but how the timing of the first qualifying event differed between treatment groups over follow-up. Participants without an observed clinical failure by the end of their available follow-up contribute censored information rather than being treated as if they experienced the event at the end of observation.

6. Primary Results

The registry reports three formal primary-endpoint analyses. All three use the mITT population, a stratified log-rank test, and a hazard ratio calculated using a Cox proportional-hazards model. Each comparison is a superiority analysis.

Combination Therapy vs Pooled Monotherapy

Hazard ratio for first adjudicated clinical failure

0.502

95% CI: 0.348–0.724   ·   P = 0.0002

Combination Therapy: Ambrisentan + Tadalafil / Monotherapy Pooled: Ambrisentan or Tadalafil

Clinical Biostats interpretation

The estimated hazard ratio of 0.502 means that, under the Cox model used for this analysis, the estimated instantaneous rate of first adjudicated clinical failure in the combination-therapy group was about 50.2% of the corresponding estimated rate in the pooled monotherapy group. Equivalently, 0.502 corresponds to an approximately 49.8% lower estimated hazard for the combination relative to pooled monotherapy.

The hazard ratio does not mean that exactly 49.8% of participants avoided clinical failure, nor does it mean that every participant had the same proportional reduction in event risk. It is a relative time-to-event measure derived from a Cox model.

The 95% confidence interval of 0.348 to 0.724 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not a range containing the individual treatment effects experienced by patients.

The p-value of 0.0002 addresses the evidence against the relevant null hypothesis under the prespecified superiority comparison. It does not measure the size or clinical importance of the treatment effect. Effect size is conveyed by the hazard ratio and its confidence interval.

Because the hazard ratio was calculated with a Cox proportional-hazards model, interpretation also depends on the model's proportional-hazards framework. The ClinicalTrials.gov record does not provide diagnostics establishing whether that assumption held throughout follow-up.

Combination Therapy vs Ambrisentan Monotherapy

Hazard ratio for first adjudicated clinical failure

0.477

95% CI: 0.314–0.723   ·   P = 0.0004

Combination Therapy: Ambrisentan + Tadalafil / Ambrisentan Monotherapy

Clinical Biostats interpretation

The estimated hazard ratio of 0.477 indicates that the estimated instantaneous rate of first adjudicated clinical failure under combination therapy was approximately 47.7% of that under ambrisentan monotherapy in the fitted Cox model. This corresponds to an approximately 52.3% lower estimated hazard for combination therapy.

The 95% confidence interval, 0.314 to 0.723, expresses uncertainty around that estimated relative hazard. Because the interval lies below 1, the direction of the estimated treatment effect is consistently toward a lower hazard for combination therapy within the interval reported by the registry.

The p-value of 0.0004 is evidence against the null hypothesis for this comparison under the reported superiority analysis. It should not be interpreted as a probability that the null hypothesis is true, nor as a measure of how large the treatment effect is.

This is a comparison with one specific monotherapy arm, not the pooled monotherapy analysis. The distinction matters because the comparator population changes between analyses. The hazard ratio should therefore not be transferred from one comparison to another simply because the intervention group is the same.

As with the pooled analysis, the estimate is model-based and concerns the relative hazard over the analyzed follow-up rather than an absolute probability of clinical failure.

Combination Therapy vs Tadalafil Monotherapy

Hazard ratio for first adjudicated clinical failure

0.528

95% CI: 0.338–0.827   ·   P = 0.0045

Combination Therapy: Ambrisentan + Tadalafil / Tadalafil Monotherapy

Clinical Biostats interpretation

The estimated hazard ratio of 0.528 indicates that the estimated instantaneous rate of first adjudicated clinical failure with combination therapy was about 52.8% of that with tadalafil monotherapy under the reported Cox model. This corresponds to an approximately 47.2% lower estimated hazard.

The 95% confidence interval is 0.338 to 0.827. The interval communicates the precision of the estimate and remains below 1 across its reported range. It does not indicate that the true effect for every patient must fall somewhere inside this interval.

The p-value of 0.0045 provides a measure of evidence against the null hypothesis for this statistical comparison. It does not quantify the magnitude of the benefit. The magnitude is better understood from the hazard ratio together with its confidence interval.

Again, this is a time-to-event comparison, so the analysis uses the ordering and timing of events as well as censoring information. It should not be reduced to a simple comparison of percentages unless an appropriate absolute-risk analysis is separately provided.

7. Comparing the Three Primary Hazard Ratios

ComparisonHR95% CIP-valueMethod
Combination vs pooled monotherapy0.5020.348–0.7240.0002Stratified log-rank; Cox HR
Combination vs ambrisentan0.4770.314–0.7230.0004Stratified log-rank; Cox HR
Combination vs tadalafil0.5280.338–0.8270.0045Stratified log-rank; Cox HR

The three estimates are numerically close in direction and magnitude, but the comparisons answer different questions because the comparator groups differ. The pooled comparison addresses combination therapy versus both monotherapies considered together. The other two analyses isolate each monotherapy comparator.

Do not treat the three hazard ratios as three estimates of exactly the same estimand. The intervention is constant, but the comparator changes. The appropriate interpretation is therefore comparison-specific rather than an attempt to identify one universal treatment-effect number.

8. Secondary Endpoint Results

The registry contains formal statistical analyses for secondary outcomes at Week 24. These analyses cover N-terminal pro-B-type natriuretic peptide, satisfactory clinical response, 6 Minute Walk Distance, World Health Organization Functional Class, and Borg Dyspnea Index.

N-Terminal Pro-B-Type Natriuretic Peptide

The registered endpoint is Percent Change From Baseline in the N-Terminal Pro-B-Type Natriuretic Peptide at Week 24, assessed from baseline to Week 24. ANCOVA was used in the mITT population, with only participants having data available at the specified time points analyzed.

ComparisonMean percent difference95% CIP-value
Combination vs pooled monotherapy-33.81-44.78 to -20.66<0.0001
Combination vs ambrisentan-25.09-40.04 to -6.400.0111
Combination vs tadalafil-41.51-53.16 to -26.97<0.0001
Clinical Biostats interpretation

The registry defines each estimate as the percent change under combination therapy minus the corresponding percent change under the comparator. Therefore, a negative estimate indicates a lower mean percent change in the combination group relative to that comparator under the reported ANCOVA analysis.

For the pooled comparison, the estimate is -33.81 with a 95% CI of -44.78 to -20.66. For the ambrisentan comparison it is -25.09 with a 95% CI of -40.04 to -6.40, and for the tadalafil comparison it is -41.51 with a 95% CI of -53.16 to -26.97.

These are not hazard ratios, odds ratios, or absolute percentage-point differences in the probability of an event. The registry labels the effect as a mean percent difference, so the numerical scale should be interpreted on that endpoint's percent-change scale.

The analysis is based only on participants with the required data at the specified time points. Consequently, the estimate describes the analyzed population rather than automatically representing every enrolled participant.

Satisfactory Clinical Response at Week 24

The registered endpoint is Percentage of Participants With a Satisfactory Clinical Response at Week 24, with a baseline and Week 24 time frame. Logistic regression was used in the mITT population, restricted to participants with a “Yes”/“No” response.

ComparisonOdds ratio95% CIP-value
Combination vs pooled monotherapy1.5631.054–2.3190.0264
Combination vs ambrisentan1.4240.878–2.3080.1518
Combination vs tadalafil1.7231.047–2.8330.0321
Clinical Biostats interpretation

An odds ratio compares the odds of the specified binary response between groups; it is not the same as a risk ratio or a difference in response probabilities. An odds ratio of 1.563 means that the estimated odds of a satisfactory clinical response were 1.563 times as high with combination therapy as with pooled monotherapy under the reported logistic regression analysis.

The corresponding confidence interval, 1.054 to 2.319, quantifies uncertainty around that odds-ratio estimate. For the ambrisentan comparison, the estimate is 1.424 with a 95% CI of 0.878 to 2.308. Because that interval includes 1, the registry's p-value of 0.1518 does not provide the same statistical evidence against the null hypothesis as the pooled comparison.

For tadalafil monotherapy, the odds ratio is 1.723, with a 95% CI of 1.047 to 2.833 and a p-value of 0.0321.

Importantly, an odds ratio should not be translated into a percentage increase in the probability of response without knowing the underlying response probability. The odds scale and probability scale are different.

6 Minute Walk Distance at Week 24

The registered endpoint is Change From Baseline in the 6 Minute Walk Distance Test at Week 24. The analysis was conducted in the mITT population among participants with baseline data, using a stratified Wilcoxon Rank Sum Test.

ComparisonMedian difference95% CIP-value
Combination vs pooled monotherapy22.75 meters12.00 to 33.50<0.0001
Combination vs ambrisentan24.75 meters11.00 to 38.500.0005
Combination vs tadalafil20.85 meters8.00 to 33.700.0030
Clinical Biostats interpretation

The reported median difference is defined as the change from baseline with combination therapy minus the change from baseline with the comparator. Thus, the pooled-monotherapy estimate of 22.75 meters indicates a higher reported change from baseline in the combination group by that median-difference measure.

The 95% confidence interval of 12.00 to 33.50 meters describes uncertainty around the reported effect estimate. It does not mean that every individual patient's improvement lies within those limits.

The Wilcoxon Rank Sum framework is useful when the analysis is based on rank information rather than requiring the outcome to follow a normal distribution. The registry specifically identifies the analysis as a stratified Wilcoxon Rank Sum Test and reports the effect measure as a median difference.

The separate comparisons give estimates of 24.75 meters versus ambrisentan and 20.85 meters versus tadalafil. These should be read as separate treatment comparisons rather than combined into one pooled effect.

World Health Organization Functional Class at Week 24

The registered endpoint is Change From Baseline in the World Health Organization Functional Class at Week 24. The reported analysis used a stratified Wilcoxon Rank Sum Test in the mITT population among participants with baseline data.

ComparisonMedian difference95% CIP-value / testing note
Combination vs pooled monotherapy0.00.0 to 0.00.2287
Combination vs ambrisentan0.00.0 to 0.0Not formally tested under predefined hierarchical procedure
Combination vs tadalafil0.00.0 to 0.0Not formally tested under predefined hierarchical procedure
Multiplicity matters here. The registry explicitly states that the ambrisentan and tadalafil comparisons were not formally tested according to the pre-defined hierarchical testing procedure. Their reported estimates and confidence intervals therefore should not be treated as if those comparisons had received the same formal confirmatory testing status as an endpoint within the hierarchy.
Clinical Biostats interpretation

The pooled-monotherapy analysis reports a median difference of 0.0 with a 95% confidence interval of 0.0 to 0.0 and a p-value of 0.2287. On the reported scale, the estimated median difference is therefore zero.

The two individual-monotherapy comparisons also report 0.0 for the estimate and both confidence limits. However, the registry explicitly identifies those comparisons as not formally tested under the predefined hierarchical testing procedure. That distinction is important: a numerical estimate can be displayed in a registry even when the comparison does not carry the same confirmatory interpretation as a formally tested hypothesis.

Borg Dyspnea Index at Week 24

The registered endpoint is Change From Baseline in Borg Dyspnea Index at Week 24, measured from baseline to Week 24. A stratified Wilcoxon Rank Sum Test was used in the mITT population among participants with baseline data.

ComparisonMedian difference95% CITesting status
Combination vs pooled monotherapy-0.38-0.75 to 0.00Not formally tested under predefined hierarchical procedure
Combination vs ambrisentan-0.50-1.00 to 0.00Not formally tested under predefined hierarchical procedure
Combination vs tadalafil-0.50-1.00 to 0.00Not formally tested under predefined hierarchical procedure
Clinical Biostats interpretation

The reported median differences are negative for all three comparisons. Because the effect is defined as combination therapy minus the comparator, the negative direction indicates a lower change score under combination therapy on the reported Borg Dyspnea Index scale.

The pooled estimate is -0.38 with a 95% CI of -0.75 to 0.00. The ambrisentan and tadalafil comparisons are both -0.50 with 95% CIs of -1.00 to 0.00.

The registry specifically states that these comparisons were not formally tested according to the pre-defined hierarchical testing procedure. Consequently, the numerical estimates should be treated as reported comparative estimates rather than as independent confirmatory findings.

9. Secondary Results: Integrated Statistical View

EndpointMethodPrimary pooled comparison95% CIP-value
N-terminal pro-B-type natriuretic peptide, Week 24ANCOVA-33.81 mean percent difference-44.78 to -20.66<0.0001
Satisfactory clinical response, Week 24Logistic regressionOR 1.5631.054–2.3190.0264
6 Minute Walk Distance, Week 24Stratified Wilcoxon Rank Sum22.75 meter median difference12.00–33.50<0.0001
WHO Functional Class, Week 24Stratified Wilcoxon Rank Sum0.0 median difference0.0–0.00.2287
Borg Dyspnea Index, Week 24Stratified Wilcoxon Rank Sum-0.38 median difference-0.75–0.00Not formally tested

The secondary analyses illustrate why a clinical-trial results page should not reduce the evidence to a single p-value. The endpoints are measured on different scales and analyzed with different statistical models. A hazard ratio, odds ratio, mean percent difference, and median difference cannot be compared numerically as though they were interchangeable effect measures.

10. Statistical Methodology

Stratified log-rank test

The primary endpoint is a time-to-event outcome, and the registry reports a stratified log-rank test for each primary comparison. The log-rank framework compares the event experience of randomized groups across follow-up while accounting for the fact that patients can enter the analysis set at risk and then either experience the event or become censored.

Time-to-event concept
At each event time: observed events are compared with the events expected under the null hypothesis

The stratified form performs this comparison across strata rather than treating all participants as belonging to one undifferentiated risk set.

The ClinicalTrials.gov record identifies the method as stratified but do not provide the individual stratification variables. They therefore should not be reconstructed or inferred from general clinical-trial practice.

Cox proportional-hazards model

The registry analysis notes state that each primary hazard ratio was calculated using a Cox proportional-hazards model. The hazard ratio is defined as the estimated hazard for the combination-therapy group divided by the hazard for the specified monotherapy comparator.

Hazard-ratio interpretation
HR = hazard in combination therapy / hazard in comparator

An HR below 1 indicates a lower estimated instantaneous event rate in the numerator group under the fitted model. It is not an absolute risk difference and is not a direct measure of the proportion of participants who benefit.

Why censoring matters

Time-to-event analysis is designed for situations in which not every participant experiences the event during observed follow-up. A participant who has not experienced clinical failure when observation ends contributes information about remaining event-free up to the censoring time. This is one reason a time-to-event analysis contains more information than simply classifying every participant as “event” or “no event.”

ANCOVA

The secondary N-terminal pro-B-type natriuretic peptide endpoint was analyzed using ANCOVA. In general, ANCOVA models a continuous outcome while allowing relevant covariate adjustment. For this registry analysis, the ClinicalTrials.gov record identifies the method and the resulting mean percent difference, but do not provide the complete model specification or covariate list. Those details should therefore not be invented.

Logistic regression

The satisfactory clinical response endpoint is binary because the registry specifies a “Yes”/“No” response. Logistic regression is therefore used to model the odds of response and report an odds ratio. The ClinicalTrials.gov record does not provide the full regression specification, so the interpretation is limited to the reported odds ratios and confidence intervals.

Wilcoxon Rank Sum testing

The 6 Minute Walk Distance, WHO Functional Class, and Borg Dyspnea Index analyses use a stratified Wilcoxon Rank Sum Test. This is a rank-based comparison that does not require the outcome to be modeled as normally distributed. The registry reports a median difference as the effect measure.

11. Statistical Methods Explained

Why was a stratified log-rank test used for the primary endpoint?

The primary endpoint is time to first adjudicated clinical failure, so both the timing of the event and the censoring process matter. A log-rank test is designed for comparing survival-type distributions between groups. The stratified version extends that comparison across predefined strata. The registry specifically reports the stratified log-rank test as the method for all three primary comparisons.

What does a hazard ratio of 0.502 mean?

It means that the fitted Cox model estimates the instantaneous clinical-failure hazard in the combination group to be 0.502 times the hazard in pooled monotherapy. That can be expressed as an approximately 49.8% lower estimated hazard. It does not mean a 49.8% absolute reduction in the number of participants experiencing an event, and it does not mean that every participant has the same risk reduction.

Why is the confidence interval important?

The point estimate is only one summary of the treatment comparison. The 95% confidence interval of 0.348 to 0.724 for the pooled primary comparison communicates how precisely the hazard ratio was estimated under the statistical framework. A narrower interval generally provides greater precision than a wider one, while the location of the interval indicates the range of relative effects compatible with the analysis.

Why does the p-value not measure effect size?

A p-value summarizes evidence against a null hypothesis under a specified statistical model and testing procedure. It does not tell us how large the treatment effect is. The AMBITION primary analyses demonstrate this distinction: the reported p-values are 0.0002, 0.0004, and 0.0045, while the corresponding hazard ratios are 0.502, 0.477, and 0.528. Effect magnitude and statistical evidence are different quantities.

Why is an odds ratio not the same as a risk ratio?

An odds ratio compares odds, not probabilities directly. For the satisfactory clinical response endpoint, an odds ratio of 1.563 means the estimated odds of a satisfactory response were 1.563 times the comparator odds. Without the underlying response probability, it is not valid to say that the probability of response increased by 56.3%.

Why does the hierarchical testing note matter?

The registry explicitly identifies the WHO Functional Class comparisons against the individual monotherapies and the Borg Dyspnea Index comparisons as not formally tested according to the predefined hierarchical testing procedure. This means the numerical estimates and confidence intervals should be distinguished from confirmatory hypothesis-testing claims. Multiplicity is therefore part of the interpretation, not an afterthought.

12. Confidence Intervals Across Effect Measures

The AMBITION registry results use several different effect measures. Reading each confidence interval requires first identifying what quantity is being estimated.

Effect measureTrial exampleInterpretation
Hazard ratio0.502Relative instantaneous event rate under the Cox model
Odds ratio1.563Relative odds of a satisfactory clinical response
Mean percent difference-33.81Difference between mean percent changes as defined by the registry
Median difference22.75 metersReported difference on the median-difference scale for change in 6 Minute Walk Distance

This distinction prevents a common statistical error: treating every number greater than or less than 1 as though it were a risk ratio. The numerical value has meaning only within its effect-measure definition.

13. Multiplicity and the Testing Hierarchy

The ClinicalTrials.gov record explicitly identify multiplicity adjustment in the analysis text for the WHO Functional Class and Borg Dyspnea Index endpoints. In particular, the registry states that certain individual-monotherapy comparisons were not formally tested according to the pre-defined hierarchical testing procedure.

Formally analyzed

The primary clinical-failure endpoint has three posted formal analyses, each with a hazard ratio, confidence interval, and p-value.

Hierarchy-sensitive

Some secondary comparisons are accompanied by an explicit statement that they were not formally tested under the predefined hierarchical procedure.

A hierarchical testing procedure can determine which hypotheses receive formal confirmatory testing and how later hypotheses inherit or do not inherit the available type I error. The key point for interpretation is that a reported numerical comparison does not automatically carry the same evidentiary status as a formally tested hypothesis.

Statistical caution: the ClinicalTrials.gov record does not provide the complete hierarchical sequence or alpha allocation. The page therefore describes the registry's explicit multiplicity statements without reconstructing an unreported testing hierarchy.

14. Missing Data and Analysis Populations

Several secondary outcomes explicitly restrict analysis to participants with the necessary observations:

These restrictions are important because the number enrolled in the trial is not necessarily the number contributing to every secondary analysis. The ClinicalTrials.gov record does not provide missing-data counts, imputation procedures, or sensitivity analyses. Accordingly, no imputation method is attributed to AMBITION here.

Why this matters: an analysis based only on participants with observed measurements can be appropriate, but the statistical interpretation depends on why observations are missing. Without the missing-data mechanism and the prespecified handling strategy, the magnitude of any potential bias cannot be determined from the posted analysis summary alone.

15. Censoring and the Time-to-Event Endpoint

The primary endpoint is defined as time to the first adjudicated clinical failure event through the Final Assessment Visit. The statistical analysis therefore differs fundamentally from a simple binary endpoint assessed at one fixed time.

Conceptual survival-analysis structure
Observed follow-up = time to event, or time to censoring if no event is observed

A censored participant still contributes information for the period during which the participant is known to remain event-free.

The registry reports an average follow-up horizon of 609 days for the primary endpoint. That is the registered time frame reported here; it should not be converted into a median follow-up or interpreted as if every participant was observed for exactly 609 days.

The clinical-failure endpoint is also a composite of several possible first events: death, hospitalization for worsening PAH, disease progression, and unsatisfactory long-term clinical response. A composite endpoint can increase the number of observed events, but its interpretation depends on the clinical meaning and relative contribution of its components. The ClinicalTrials.gov record does not provide component-specific event counts, so no component-level conclusion is drawn.

16. Safety Results

The ClinicalTrials.gov record provides serious adverse-event counts by arm. These are reported as affected participants divided by the participants at risk for the serious-adverse-event analysis.

Treatment groupSerious adverse eventsAffected / at risk
Combination Therapy: Ambrisentan + TadalafilSerious adverse events124 / 302
Ambrisentan MonotherapySerious adverse events63 / 152
Tadalafil MonotherapySerious adverse events68 / 151

These are arm-level serious-adverse-event counts, not estimates of the primary clinical-failure endpoint. They should therefore be kept separate from the efficacy hazard ratios.

Clinical Biostats interpretation

The ClinicalTrials.gov record shows that serious adverse events affected 124 of 302 participants in the combination-therapy group, 63 of 152 in the ambrisentan-monotherapy group, and 68 of 151 in the tadalafil-monotherapy group.

The denominator is important: the reported quantity is affected participants divided by participants at risk, rather than the number of serious adverse-event episodes. A participant can potentially experience more than one adverse event, but the registry-reported summary is expressed as affected participants.

No formal between-arm statistical test for serious adverse events is reported in the ClinicalTrials.gov record. Therefore, the figures are presented descriptively rather than converted into an unsupported p-value, risk ratio, or comparative safety conclusion.

17. Primary vs Secondary Statistical Questions

QuestionEndpointEffect measureMethod
Does combination therapy change time to first adjudicated clinical failure?Primary clinical failure endpointHazard ratioStratified log-rank; Cox model
Does combination therapy change NT-proBNP percent change at Week 24?NT-proBNPMean percent differenceANCOVA
Does combination therapy change the odds of satisfactory clinical response?Satisfactory clinical responseOdds ratioLogistic regression
Does combination therapy change 6 Minute Walk Distance?6 Minute Walk DistanceMedian differenceStratified Wilcoxon Rank Sum
Does combination therapy change WHO Functional Class?WHO Functional ClassMedian differenceStratified Wilcoxon Rank Sum
Does combination therapy change Borg Dyspnea Index?Borg Dyspnea IndexMedian differenceStratified Wilcoxon Rank Sum

This structure illustrates an important feature of clinical-trial statistics: the endpoint determines the statistical question, and the statistical question determines the appropriate effect measure. A time-to-event endpoint calls for a different framework from a binary response or a continuous change score.

18. How to Read the AMBITION Primary Result

Relative effect

The primary pooled comparison produced an HR of 0.502. This is a relative time-to-event measure indicating a lower estimated instantaneous rate of first adjudicated clinical failure under combination therapy in the Cox model.

Absolute effect

The ClinicalTrials.gov record does not provide absolute event rates, median time to clinical failure, or Kaplan-Meier estimates for the primary endpoint. Those quantities should therefore not be reconstructed from the hazard ratio.

Precision

The 95% CI of 0.348 to 0.724 shows the uncertainty around the estimated hazard ratio. It is more informative than the point estimate alone because it indicates the range of relative effects compatible with the statistical analysis at the stated confidence level.

Statistical evidence

The p-value of 0.0002 is evidence against the null hypothesis for the reported superiority comparison. It does not establish the magnitude of clinical benefit, and it should not be interpreted as a probability that the treatment effect is real.

What remains unknown from the ClinicalTrials.gov record

The posted summary does not provide median time to clinical failure, the number of primary-endpoint events by arm, Kaplan-Meier survival probabilities, subgroup estimates, or the detailed Cox-model specification. Those quantities are therefore not presented as if they were available.

19. Statistical Interpretation of the Secondary Endpoints

The secondary endpoints show why statistical interpretation must remain endpoint-specific.

Biomarker

NT-proBNP uses ANCOVA and a mean percent difference. Its negative estimates describe the direction of the difference in percent change as defined by the registry.

Binary response

Satisfactory clinical response uses logistic regression and odds ratios. The odds ratio cannot be interpreted as a direct probability difference.

Functional capacity

6 Minute Walk Distance uses a stratified Wilcoxon Rank Sum test and a median difference, with the pooled estimate of 22.75 meters.

Patient-reported / clinical scales

WHO Functional Class and Borg Dyspnea Index use the same rank-based framework, with explicit hierarchy-related cautions for several comparisons.

A statistically significant result on one endpoint does not automatically imply that every other endpoint has the same effect. Conversely, a nonsignificant endpoint does not prove that the treatment has no effect. Each estimate must be interpreted in the context of its scale, analysis population, confidence interval, and testing status.

20. Limitations

21. Why This Trial Matters Statistically

AMBITION is a useful statistical teaching case because it combines several core clinical-trial methods within a single randomized phase 3 design. The same trial moves from a time-to-event primary endpoint to continuous percent-change, binary response, and rank-based functional outcomes.

ConceptHow it appears in AMBITION
RandomizationRandomized, parallel-group phase 3 design with 3 arms
MaskingQuadruple masking
Time-to-event analysisFirst adjudicated clinical failure through FAV
Stratified log-rank testReported for all three primary comparisons
Cox modelUsed to calculate the primary hazard ratios
Hazard ratio0.502, 0.477, and 0.528 for the three primary comparisons
ANCOVAUsed for NT-proBNP percent change at Week 24
Logistic regressionUsed for satisfactory clinical response at Week 24
Wilcoxon Rank SumUsed for 6 Minute Walk Distance, WHO Functional Class, and Borg Dyspnea Index
MultiplicityExplicit hierarchy-related limitations for selected secondary comparisons
Confidence intervalsReported for all registry-reported statistical effect estimates
mITT populationUsed for the primary and registry-reported secondary analyses

The educational value is especially strong because the trial demonstrates that “the statistical analysis” is not one procedure. A well-designed trial can require several different methods, each matched to the endpoint's data type and scientific question.

22. What the Primary Hazard Ratio Does — and Does Not — Mean

Primary pooled comparison
HR = 0.502   |   95% CI = 0.348–0.724

The estimated hazard of first adjudicated clinical failure under combination therapy was 0.502 times the estimated hazard under pooled monotherapy in the reported Cox analysis.

There are several ways this result can be misunderstood.

The confidence interval is equally important. A point estimate can appear precise simply because it is displayed to three decimal places, but numerical formatting is not statistical precision. The interval from 0.348 to 0.724 is the appropriate companion to the 0.502 estimate for communicating uncertainty.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through Clinical Biostats

Explore the statistical methods behind randomized trials, then move from methodology to practical calculation and analysis workflows.

26. Record Summary

AMBITION provides a compact example of how several statistical frameworks can coexist within one randomized phase 3 trial. Its primary endpoint is a time-to-event outcome analyzed with stratified log-rank testing and Cox-model hazard ratios, while its secondary outcomes use ANCOVA, logistic regression, and stratified Wilcoxon Rank Sum testing.

The primary pooled comparison produced an HR of 0.502 with a 95% CI of 0.348 to 0.724 and a p-value of 0.0002. The corresponding comparisons with ambrisentan and tadalafil monotherapy produced HRs of 0.477 and 0.528, respectively. The secondary results provide additional effect measures on biomarker, response, functional-capacity, and symptom scales, while the registry explicitly identifies multiplicity-related limits for selected comparisons.

The most important statistical lesson is that these estimates should be interpreted according to their underlying estimands. A hazard ratio describes relative event hazards; an odds ratio describes relative odds; a mean percent difference describes a difference in mean percent change; and a median difference describes a difference on the reported median-difference scale. None of these measures should be substituted for another.

Clinical Biostats methodology: This page separates the numerical results reported by the trial registry from the statistical interpretation needed to understand them. Where the ClinicalTrials.gov record does not report a quantity—such as absolute clinical-failure rates, median time to failure, detailed stratification variables, or imputation procedures—it is not reconstructed or inferred.