This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the data posted on ClinicalTrials.gov for NCT01178073.
1. Trial at a Glance
AMBITION was a randomized, parallel-group, quadruple-masked phase 3 treatment trial in pulmonary arterial hypertension. The registry reports 610 enrolled participants, 3 arms, 1 registered primary endpoint, 6 posted outcome measures, and 18 posted statistical analyses.
| Feature | AMBITION |
|---|---|
| Trial name | AMBITION |
| ClinicalTrials.gov identifier | NCT01178073 |
| Phase | Phase 3 |
| Therapeutic area | Pulmonology |
| Condition | Hypertension, Pulmonary |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 610 |
| Arms | 3 |
| Registered primary endpoints | 1 |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 6 |
| Statistical analyses posted | 18 |
| Lead sponsor | GlaxoSmithKline |
| Sponsor type | Industry |
| Status | Completed |
2. Clinical Question
The central statistical question was whether first-line ambrisentan plus tadalafil reduced the time to the first adjudicated clinical failure event compared with first-line monotherapy with either ambrisentan or tadalafil in subjects with pulmonary arterial hypertension.
Population
Subjects with pulmonary arterial hypertension participating in the AMBITION phase 3 randomized trial.
Intervention
First-line combination therapy with ambrisentan and tadalafil.
Comparator
First-line monotherapy with ambrisentan or tadalafil, analyzed both as a pooled monotherapy group and separately by monotherapy arm.
Primary question
Does combination therapy change the time to first adjudicated clinical failure through the Final Assessment Visit relative to monotherapy?
The registry defines clinical failure as the first adjudicated event among death, hospitalization for worsening pulmonary arterial hypertension, disease progression, or unsatisfactory long-term clinical response. The registered time frame is from baseline through the Final Assessment Visit, with an average of 609 days.
3. Trial Design
Ambrisentan + tadalafil
- First-line combination therapy.
- Compared with pooled monotherapy for the primary time-to-event analysis.
- Also compared separately with ambrisentan monotherapy and tadalafil monotherapy.
Ambrisentan or tadalafil
- Two first-line monotherapy groups.
- Analyzed as a pooled monotherapy comparator for the primary comparison.
- Also retained as separate comparators for additional primary analyses.
4. Analysis Population and Comparison Structure
The primary and posted secondary analyses in the ClinicalTrials.gov record use the modified intention-to-treat (mITT) population. The registry's secondary analyses also specify when only participants with available data at the relevant time points, or participants with a required response classification, were analyzed.
| Analysis feature | Registry-supported description |
|---|---|
| Primary analysis population | mITT Population |
| Secondary biomarker analysis | mITT population; only participants with data available at the specified time points |
| Clinical response analysis | mITT population; only participants with a “Yes”/“No” response |
| 6 Minute Walk Distance analysis | mITT population; only participants with baseline data |
| WHO Functional Class analysis | mITT population; only participants with baseline data |
| Primary pooled comparison | Combination therapy versus pooled ambrisentan or tadalafil monotherapy |
| Individual monotherapy comparisons | Combination therapy versus ambrisentan monotherapy; combination therapy versus tadalafil monotherapy |
This distinction matters statistically. A randomized trial begins with treatment assignment, but an analysis population can impose additional eligibility conditions for a particular endpoint. For example, the Week 24 endpoints requiring observed measurements are explicitly restricted in the registry to participants with the necessary data. Such restrictions can change the population contributing information to that endpoint even when the overall trial remains randomized.
5. Primary Endpoint
| Endpoint | Registry definition | Analysis |
|---|---|---|
| First adjudicated clinical failure | Number of participants with first adjudicated clinical failure event, death, hospitalisation for worsening PAH, disease progression, unsatisfactory long-term clinical response, all through FAV | Stratified log-rank test; hazard ratio from Cox proportional-hazards model |
Time frame: From Baseline up to the Final Assessment Visit (FAV), with an average of 609 days.
The endpoint is a classic time-to-event outcome. The statistical target is not simply whether an event eventually occurred, but how the timing of the first qualifying event differed between treatment groups over follow-up. Participants without an observed clinical failure by the end of their available follow-up contribute censored information rather than being treated as if they experienced the event at the end of observation.
6. Primary Results
The registry reports three formal primary-endpoint analyses. All three use the mITT population, a stratified log-rank test, and a hazard ratio calculated using a Cox proportional-hazards model. Each comparison is a superiority analysis.
Combination Therapy vs Pooled Monotherapy
Hazard ratio for first adjudicated clinical failure
95% CI: 0.348–0.724 · P = 0.0002
Combination Therapy: Ambrisentan + Tadalafil / Monotherapy Pooled: Ambrisentan or Tadalafil
The estimated hazard ratio of 0.502 means that, under the Cox model used for this analysis, the estimated instantaneous rate of first adjudicated clinical failure in the combination-therapy group was about 50.2% of the corresponding estimated rate in the pooled monotherapy group. Equivalently, 0.502 corresponds to an approximately 49.8% lower estimated hazard for the combination relative to pooled monotherapy.
The hazard ratio does not mean that exactly 49.8% of participants avoided clinical failure, nor does it mean that every participant had the same proportional reduction in event risk. It is a relative time-to-event measure derived from a Cox model.
The 95% confidence interval of 0.348 to 0.724 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not a range containing the individual treatment effects experienced by patients.
The p-value of 0.0002 addresses the evidence against the relevant null hypothesis under the prespecified superiority comparison. It does not measure the size or clinical importance of the treatment effect. Effect size is conveyed by the hazard ratio and its confidence interval.
Because the hazard ratio was calculated with a Cox proportional-hazards model, interpretation also depends on the model's proportional-hazards framework. The ClinicalTrials.gov record does not provide diagnostics establishing whether that assumption held throughout follow-up.
Combination Therapy vs Ambrisentan Monotherapy
Hazard ratio for first adjudicated clinical failure
95% CI: 0.314–0.723 · P = 0.0004
Combination Therapy: Ambrisentan + Tadalafil / Ambrisentan Monotherapy
The estimated hazard ratio of 0.477 indicates that the estimated instantaneous rate of first adjudicated clinical failure under combination therapy was approximately 47.7% of that under ambrisentan monotherapy in the fitted Cox model. This corresponds to an approximately 52.3% lower estimated hazard for combination therapy.
The 95% confidence interval, 0.314 to 0.723, expresses uncertainty around that estimated relative hazard. Because the interval lies below 1, the direction of the estimated treatment effect is consistently toward a lower hazard for combination therapy within the interval reported by the registry.
The p-value of 0.0004 is evidence against the null hypothesis for this comparison under the reported superiority analysis. It should not be interpreted as a probability that the null hypothesis is true, nor as a measure of how large the treatment effect is.
This is a comparison with one specific monotherapy arm, not the pooled monotherapy analysis. The distinction matters because the comparator population changes between analyses. The hazard ratio should therefore not be transferred from one comparison to another simply because the intervention group is the same.
As with the pooled analysis, the estimate is model-based and concerns the relative hazard over the analyzed follow-up rather than an absolute probability of clinical failure.
Combination Therapy vs Tadalafil Monotherapy
Hazard ratio for first adjudicated clinical failure
95% CI: 0.338–0.827 · P = 0.0045
Combination Therapy: Ambrisentan + Tadalafil / Tadalafil Monotherapy
The estimated hazard ratio of 0.528 indicates that the estimated instantaneous rate of first adjudicated clinical failure with combination therapy was about 52.8% of that with tadalafil monotherapy under the reported Cox model. This corresponds to an approximately 47.2% lower estimated hazard.
The 95% confidence interval is 0.338 to 0.827. The interval communicates the precision of the estimate and remains below 1 across its reported range. It does not indicate that the true effect for every patient must fall somewhere inside this interval.
The p-value of 0.0045 provides a measure of evidence against the null hypothesis for this statistical comparison. It does not quantify the magnitude of the benefit. The magnitude is better understood from the hazard ratio together with its confidence interval.
Again, this is a time-to-event comparison, so the analysis uses the ordering and timing of events as well as censoring information. It should not be reduced to a simple comparison of percentages unless an appropriate absolute-risk analysis is separately provided.
7. Comparing the Three Primary Hazard Ratios
| Comparison | HR | 95% CI | P-value | Method |
|---|---|---|---|---|
| Combination vs pooled monotherapy | 0.502 | 0.348–0.724 | 0.0002 | Stratified log-rank; Cox HR |
| Combination vs ambrisentan | 0.477 | 0.314–0.723 | 0.0004 | Stratified log-rank; Cox HR |
| Combination vs tadalafil | 0.528 | 0.338–0.827 | 0.0045 | Stratified log-rank; Cox HR |
The three estimates are numerically close in direction and magnitude, but the comparisons answer different questions because the comparator groups differ. The pooled comparison addresses combination therapy versus both monotherapies considered together. The other two analyses isolate each monotherapy comparator.
8. Secondary Endpoint Results
The registry contains formal statistical analyses for secondary outcomes at Week 24. These analyses cover N-terminal pro-B-type natriuretic peptide, satisfactory clinical response, 6 Minute Walk Distance, World Health Organization Functional Class, and Borg Dyspnea Index.
N-Terminal Pro-B-Type Natriuretic Peptide
The registered endpoint is Percent Change From Baseline in the N-Terminal Pro-B-Type Natriuretic Peptide at Week 24, assessed from baseline to Week 24. ANCOVA was used in the mITT population, with only participants having data available at the specified time points analyzed.
| Comparison | Mean percent difference | 95% CI | P-value |
|---|---|---|---|
| Combination vs pooled monotherapy | -33.81 | -44.78 to -20.66 | <0.0001 |
| Combination vs ambrisentan | -25.09 | -40.04 to -6.40 | 0.0111 |
| Combination vs tadalafil | -41.51 | -53.16 to -26.97 | <0.0001 |
The registry defines each estimate as the percent change under combination therapy minus the corresponding percent change under the comparator. Therefore, a negative estimate indicates a lower mean percent change in the combination group relative to that comparator under the reported ANCOVA analysis.
For the pooled comparison, the estimate is -33.81 with a 95% CI of -44.78 to -20.66. For the ambrisentan comparison it is -25.09 with a 95% CI of -40.04 to -6.40, and for the tadalafil comparison it is -41.51 with a 95% CI of -53.16 to -26.97.
These are not hazard ratios, odds ratios, or absolute percentage-point differences in the probability of an event. The registry labels the effect as a mean percent difference, so the numerical scale should be interpreted on that endpoint's percent-change scale.
The analysis is based only on participants with the required data at the specified time points. Consequently, the estimate describes the analyzed population rather than automatically representing every enrolled participant.
Satisfactory Clinical Response at Week 24
The registered endpoint is Percentage of Participants With a Satisfactory Clinical Response at Week 24, with a baseline and Week 24 time frame. Logistic regression was used in the mITT population, restricted to participants with a “Yes”/“No” response.
| Comparison | Odds ratio | 95% CI | P-value |
|---|---|---|---|
| Combination vs pooled monotherapy | 1.563 | 1.054–2.319 | 0.0264 |
| Combination vs ambrisentan | 1.424 | 0.878–2.308 | 0.1518 |
| Combination vs tadalafil | 1.723 | 1.047–2.833 | 0.0321 |
An odds ratio compares the odds of the specified binary response between groups; it is not the same as a risk ratio or a difference in response probabilities. An odds ratio of 1.563 means that the estimated odds of a satisfactory clinical response were 1.563 times as high with combination therapy as with pooled monotherapy under the reported logistic regression analysis.
The corresponding confidence interval, 1.054 to 2.319, quantifies uncertainty around that odds-ratio estimate. For the ambrisentan comparison, the estimate is 1.424 with a 95% CI of 0.878 to 2.308. Because that interval includes 1, the registry's p-value of 0.1518 does not provide the same statistical evidence against the null hypothesis as the pooled comparison.
For tadalafil monotherapy, the odds ratio is 1.723, with a 95% CI of 1.047 to 2.833 and a p-value of 0.0321.
Importantly, an odds ratio should not be translated into a percentage increase in the probability of response without knowing the underlying response probability. The odds scale and probability scale are different.
6 Minute Walk Distance at Week 24
The registered endpoint is Change From Baseline in the 6 Minute Walk Distance Test at Week 24. The analysis was conducted in the mITT population among participants with baseline data, using a stratified Wilcoxon Rank Sum Test.
| Comparison | Median difference | 95% CI | P-value |
|---|---|---|---|
| Combination vs pooled monotherapy | 22.75 meters | 12.00 to 33.50 | <0.0001 |
| Combination vs ambrisentan | 24.75 meters | 11.00 to 38.50 | 0.0005 |
| Combination vs tadalafil | 20.85 meters | 8.00 to 33.70 | 0.0030 |
The reported median difference is defined as the change from baseline with combination therapy minus the change from baseline with the comparator. Thus, the pooled-monotherapy estimate of 22.75 meters indicates a higher reported change from baseline in the combination group by that median-difference measure.
The 95% confidence interval of 12.00 to 33.50 meters describes uncertainty around the reported effect estimate. It does not mean that every individual patient's improvement lies within those limits.
The Wilcoxon Rank Sum framework is useful when the analysis is based on rank information rather than requiring the outcome to follow a normal distribution. The registry specifically identifies the analysis as a stratified Wilcoxon Rank Sum Test and reports the effect measure as a median difference.
The separate comparisons give estimates of 24.75 meters versus ambrisentan and 20.85 meters versus tadalafil. These should be read as separate treatment comparisons rather than combined into one pooled effect.
World Health Organization Functional Class at Week 24
The registered endpoint is Change From Baseline in the World Health Organization Functional Class at Week 24. The reported analysis used a stratified Wilcoxon Rank Sum Test in the mITT population among participants with baseline data.
| Comparison | Median difference | 95% CI | P-value / testing note |
|---|---|---|---|
| Combination vs pooled monotherapy | 0.0 | 0.0 to 0.0 | 0.2287 |
| Combination vs ambrisentan | 0.0 | 0.0 to 0.0 | Not formally tested under predefined hierarchical procedure |
| Combination vs tadalafil | 0.0 | 0.0 to 0.0 | Not formally tested under predefined hierarchical procedure |
The pooled-monotherapy analysis reports a median difference of 0.0 with a 95% confidence interval of 0.0 to 0.0 and a p-value of 0.2287. On the reported scale, the estimated median difference is therefore zero.
The two individual-monotherapy comparisons also report 0.0 for the estimate and both confidence limits. However, the registry explicitly identifies those comparisons as not formally tested under the predefined hierarchical testing procedure. That distinction is important: a numerical estimate can be displayed in a registry even when the comparison does not carry the same confirmatory interpretation as a formally tested hypothesis.
Borg Dyspnea Index at Week 24
The registered endpoint is Change From Baseline in Borg Dyspnea Index at Week 24, measured from baseline to Week 24. A stratified Wilcoxon Rank Sum Test was used in the mITT population among participants with baseline data.
| Comparison | Median difference | 95% CI | Testing status |
|---|---|---|---|
| Combination vs pooled monotherapy | -0.38 | -0.75 to 0.00 | Not formally tested under predefined hierarchical procedure |
| Combination vs ambrisentan | -0.50 | -1.00 to 0.00 | Not formally tested under predefined hierarchical procedure |
| Combination vs tadalafil | -0.50 | -1.00 to 0.00 | Not formally tested under predefined hierarchical procedure |
The reported median differences are negative for all three comparisons. Because the effect is defined as combination therapy minus the comparator, the negative direction indicates a lower change score under combination therapy on the reported Borg Dyspnea Index scale.
The pooled estimate is -0.38 with a 95% CI of -0.75 to 0.00. The ambrisentan and tadalafil comparisons are both -0.50 with 95% CIs of -1.00 to 0.00.
The registry specifically states that these comparisons were not formally tested according to the pre-defined hierarchical testing procedure. Consequently, the numerical estimates should be treated as reported comparative estimates rather than as independent confirmatory findings.
9. Secondary Results: Integrated Statistical View
| Endpoint | Method | Primary pooled comparison | 95% CI | P-value |
|---|---|---|---|---|
| N-terminal pro-B-type natriuretic peptide, Week 24 | ANCOVA | -33.81 mean percent difference | -44.78 to -20.66 | <0.0001 |
| Satisfactory clinical response, Week 24 | Logistic regression | OR 1.563 | 1.054–2.319 | 0.0264 |
| 6 Minute Walk Distance, Week 24 | Stratified Wilcoxon Rank Sum | 22.75 meter median difference | 12.00–33.50 | <0.0001 |
| WHO Functional Class, Week 24 | Stratified Wilcoxon Rank Sum | 0.0 median difference | 0.0–0.0 | 0.2287 |
| Borg Dyspnea Index, Week 24 | Stratified Wilcoxon Rank Sum | -0.38 median difference | -0.75–0.00 | Not formally tested |
The secondary analyses illustrate why a clinical-trial results page should not reduce the evidence to a single p-value. The endpoints are measured on different scales and analyzed with different statistical models. A hazard ratio, odds ratio, mean percent difference, and median difference cannot be compared numerically as though they were interchangeable effect measures.
10. Statistical Methodology
Stratified log-rank test
The primary endpoint is a time-to-event outcome, and the registry reports a stratified log-rank test for each primary comparison. The log-rank framework compares the event experience of randomized groups across follow-up while accounting for the fact that patients can enter the analysis set at risk and then either experience the event or become censored.
The stratified form performs this comparison across strata rather than treating all participants as belonging to one undifferentiated risk set.
The ClinicalTrials.gov record identifies the method as stratified but do not provide the individual stratification variables. They therefore should not be reconstructed or inferred from general clinical-trial practice.
Cox proportional-hazards model
The registry analysis notes state that each primary hazard ratio was calculated using a Cox proportional-hazards model. The hazard ratio is defined as the estimated hazard for the combination-therapy group divided by the hazard for the specified monotherapy comparator.
An HR below 1 indicates a lower estimated instantaneous event rate in the numerator group under the fitted model. It is not an absolute risk difference and is not a direct measure of the proportion of participants who benefit.
Why censoring matters
Time-to-event analysis is designed for situations in which not every participant experiences the event during observed follow-up. A participant who has not experienced clinical failure when observation ends contributes information about remaining event-free up to the censoring time. This is one reason a time-to-event analysis contains more information than simply classifying every participant as “event” or “no event.”
ANCOVA
The secondary N-terminal pro-B-type natriuretic peptide endpoint was analyzed using ANCOVA. In general, ANCOVA models a continuous outcome while allowing relevant covariate adjustment. For this registry analysis, the ClinicalTrials.gov record identifies the method and the resulting mean percent difference, but do not provide the complete model specification or covariate list. Those details should therefore not be invented.
Logistic regression
The satisfactory clinical response endpoint is binary because the registry specifies a “Yes”/“No” response. Logistic regression is therefore used to model the odds of response and report an odds ratio. The ClinicalTrials.gov record does not provide the full regression specification, so the interpretation is limited to the reported odds ratios and confidence intervals.
Wilcoxon Rank Sum testing
The 6 Minute Walk Distance, WHO Functional Class, and Borg Dyspnea Index analyses use a stratified Wilcoxon Rank Sum Test. This is a rank-based comparison that does not require the outcome to be modeled as normally distributed. The registry reports a median difference as the effect measure.
11. Statistical Methods Explained
Why was a stratified log-rank test used for the primary endpoint?
The primary endpoint is time to first adjudicated clinical failure, so both the timing of the event and the censoring process matter. A log-rank test is designed for comparing survival-type distributions between groups. The stratified version extends that comparison across predefined strata. The registry specifically reports the stratified log-rank test as the method for all three primary comparisons.
What does a hazard ratio of 0.502 mean?
It means that the fitted Cox model estimates the instantaneous clinical-failure hazard in the combination group to be 0.502 times the hazard in pooled monotherapy. That can be expressed as an approximately 49.8% lower estimated hazard. It does not mean a 49.8% absolute reduction in the number of participants experiencing an event, and it does not mean that every participant has the same risk reduction.
Why is the confidence interval important?
The point estimate is only one summary of the treatment comparison. The 95% confidence interval of 0.348 to 0.724 for the pooled primary comparison communicates how precisely the hazard ratio was estimated under the statistical framework. A narrower interval generally provides greater precision than a wider one, while the location of the interval indicates the range of relative effects compatible with the analysis.
Why does the p-value not measure effect size?
A p-value summarizes evidence against a null hypothesis under a specified statistical model and testing procedure. It does not tell us how large the treatment effect is. The AMBITION primary analyses demonstrate this distinction: the reported p-values are 0.0002, 0.0004, and 0.0045, while the corresponding hazard ratios are 0.502, 0.477, and 0.528. Effect magnitude and statistical evidence are different quantities.
Why is an odds ratio not the same as a risk ratio?
An odds ratio compares odds, not probabilities directly. For the satisfactory clinical response endpoint, an odds ratio of 1.563 means the estimated odds of a satisfactory response were 1.563 times the comparator odds. Without the underlying response probability, it is not valid to say that the probability of response increased by 56.3%.
Why does the hierarchical testing note matter?
The registry explicitly identifies the WHO Functional Class comparisons against the individual monotherapies and the Borg Dyspnea Index comparisons as not formally tested according to the predefined hierarchical testing procedure. This means the numerical estimates and confidence intervals should be distinguished from confirmatory hypothesis-testing claims. Multiplicity is therefore part of the interpretation, not an afterthought.
12. Confidence Intervals Across Effect Measures
The AMBITION registry results use several different effect measures. Reading each confidence interval requires first identifying what quantity is being estimated.
| Effect measure | Trial example | Interpretation |
|---|---|---|
| Hazard ratio | 0.502 | Relative instantaneous event rate under the Cox model |
| Odds ratio | 1.563 | Relative odds of a satisfactory clinical response |
| Mean percent difference | -33.81 | Difference between mean percent changes as defined by the registry |
| Median difference | 22.75 meters | Reported difference on the median-difference scale for change in 6 Minute Walk Distance |
This distinction prevents a common statistical error: treating every number greater than or less than 1 as though it were a risk ratio. The numerical value has meaning only within its effect-measure definition.
13. Multiplicity and the Testing Hierarchy
The ClinicalTrials.gov record explicitly identify multiplicity adjustment in the analysis text for the WHO Functional Class and Borg Dyspnea Index endpoints. In particular, the registry states that certain individual-monotherapy comparisons were not formally tested according to the pre-defined hierarchical testing procedure.
Formally analyzed
The primary clinical-failure endpoint has three posted formal analyses, each with a hazard ratio, confidence interval, and p-value.
Hierarchy-sensitive
Some secondary comparisons are accompanied by an explicit statement that they were not formally tested under the predefined hierarchical procedure.
A hierarchical testing procedure can determine which hypotheses receive formal confirmatory testing and how later hypotheses inherit or do not inherit the available type I error. The key point for interpretation is that a reported numerical comparison does not automatically carry the same evidentiary status as a formally tested hypothesis.
14. Missing Data and Analysis Populations
Several secondary outcomes explicitly restrict analysis to participants with the necessary observations:
- The N-terminal pro-B-type natriuretic peptide analysis includes participants with data available at the specified time points.
- The satisfactory clinical response analysis includes participants with a “Yes”/“No” response.
- The 6 Minute Walk Distance analysis includes participants with baseline data.
- The WHO Functional Class analysis includes participants with baseline data.
- The Borg Dyspnea Index analysis includes participants with baseline data.
These restrictions are important because the number enrolled in the trial is not necessarily the number contributing to every secondary analysis. The ClinicalTrials.gov record does not provide missing-data counts, imputation procedures, or sensitivity analyses. Accordingly, no imputation method is attributed to AMBITION here.
15. Censoring and the Time-to-Event Endpoint
The primary endpoint is defined as time to the first adjudicated clinical failure event through the Final Assessment Visit. The statistical analysis therefore differs fundamentally from a simple binary endpoint assessed at one fixed time.
A censored participant still contributes information for the period during which the participant is known to remain event-free.
The registry reports an average follow-up horizon of 609 days for the primary endpoint. That is the registered time frame reported here; it should not be converted into a median follow-up or interpreted as if every participant was observed for exactly 609 days.
The clinical-failure endpoint is also a composite of several possible first events: death, hospitalization for worsening PAH, disease progression, and unsatisfactory long-term clinical response. A composite endpoint can increase the number of observed events, but its interpretation depends on the clinical meaning and relative contribution of its components. The ClinicalTrials.gov record does not provide component-specific event counts, so no component-level conclusion is drawn.
16. Safety Results
The ClinicalTrials.gov record provides serious adverse-event counts by arm. These are reported as affected participants divided by the participants at risk for the serious-adverse-event analysis.
| Treatment group | Serious adverse events | Affected / at risk |
|---|---|---|
| Combination Therapy: Ambrisentan + Tadalafil | Serious adverse events | 124 / 302 |
| Ambrisentan Monotherapy | Serious adverse events | 63 / 152 |
| Tadalafil Monotherapy | Serious adverse events | 68 / 151 |
These are arm-level serious-adverse-event counts, not estimates of the primary clinical-failure endpoint. They should therefore be kept separate from the efficacy hazard ratios.
The ClinicalTrials.gov record shows that serious adverse events affected 124 of 302 participants in the combination-therapy group, 63 of 152 in the ambrisentan-monotherapy group, and 68 of 151 in the tadalafil-monotherapy group.
The denominator is important: the reported quantity is affected participants divided by participants at risk, rather than the number of serious adverse-event episodes. A participant can potentially experience more than one adverse event, but the registry-reported summary is expressed as affected participants.
No formal between-arm statistical test for serious adverse events is reported in the ClinicalTrials.gov record. Therefore, the figures are presented descriptively rather than converted into an unsupported p-value, risk ratio, or comparative safety conclusion.
17. Primary vs Secondary Statistical Questions
| Question | Endpoint | Effect measure | Method |
|---|---|---|---|
| Does combination therapy change time to first adjudicated clinical failure? | Primary clinical failure endpoint | Hazard ratio | Stratified log-rank; Cox model |
| Does combination therapy change NT-proBNP percent change at Week 24? | NT-proBNP | Mean percent difference | ANCOVA |
| Does combination therapy change the odds of satisfactory clinical response? | Satisfactory clinical response | Odds ratio | Logistic regression |
| Does combination therapy change 6 Minute Walk Distance? | 6 Minute Walk Distance | Median difference | Stratified Wilcoxon Rank Sum |
| Does combination therapy change WHO Functional Class? | WHO Functional Class | Median difference | Stratified Wilcoxon Rank Sum |
| Does combination therapy change Borg Dyspnea Index? | Borg Dyspnea Index | Median difference | Stratified Wilcoxon Rank Sum |
This structure illustrates an important feature of clinical-trial statistics: the endpoint determines the statistical question, and the statistical question determines the appropriate effect measure. A time-to-event endpoint calls for a different framework from a binary response or a continuous change score.
18. How to Read the AMBITION Primary Result
The primary pooled comparison produced an HR of 0.502. This is a relative time-to-event measure indicating a lower estimated instantaneous rate of first adjudicated clinical failure under combination therapy in the Cox model.
The ClinicalTrials.gov record does not provide absolute event rates, median time to clinical failure, or Kaplan-Meier estimates for the primary endpoint. Those quantities should therefore not be reconstructed from the hazard ratio.
The 95% CI of 0.348 to 0.724 shows the uncertainty around the estimated hazard ratio. It is more informative than the point estimate alone because it indicates the range of relative effects compatible with the statistical analysis at the stated confidence level.
The p-value of 0.0002 is evidence against the null hypothesis for the reported superiority comparison. It does not establish the magnitude of clinical benefit, and it should not be interpreted as a probability that the treatment effect is real.
The posted summary does not provide median time to clinical failure, the number of primary-endpoint events by arm, Kaplan-Meier survival probabilities, subgroup estimates, or the detailed Cox-model specification. Those quantities are therefore not presented as if they were available.
19. Statistical Interpretation of the Secondary Endpoints
The secondary endpoints show why statistical interpretation must remain endpoint-specific.
Biomarker
NT-proBNP uses ANCOVA and a mean percent difference. Its negative estimates describe the direction of the difference in percent change as defined by the registry.
Binary response
Satisfactory clinical response uses logistic regression and odds ratios. The odds ratio cannot be interpreted as a direct probability difference.
Functional capacity
6 Minute Walk Distance uses a stratified Wilcoxon Rank Sum test and a median difference, with the pooled estimate of 22.75 meters.
Patient-reported / clinical scales
WHO Functional Class and Borg Dyspnea Index use the same rank-based framework, with explicit hierarchy-related cautions for several comparisons.
A statistically significant result on one endpoint does not automatically imply that every other endpoint has the same effect. Conversely, a nonsignificant endpoint does not prove that the treatment has no effect. Each estimate must be interpreted in the context of its scale, analysis population, confidence interval, and testing status.
20. Limitations
- Registry-level detail: the ClinicalTrials.gov record provides the reported statistical analyses but not the complete statistical analysis plan, so unreported model details should not be reconstructed.
- Hazard-ratio interpretation: the primary effect measures are Cox-model hazard ratios. A hazard ratio is not an absolute risk difference, risk ratio, or probability of benefit.
- Proportional-hazards assumption: the Cox model is based on a proportional-hazards framework. The ClinicalTrials.gov record does not provide diagnostics demonstrating whether that assumption was satisfied.
- Composite endpoint: clinical failure combines death, hospitalization for worsening PAH, disease progression, and unsatisfactory long-term clinical response. Component-specific counts are not reported.
- Analysis populations: secondary analyses use endpoint-specific availability requirements, such as having baseline data or a Yes/No response.
- Missing-data handling: the ClinicalTrials.gov record does not report imputation methods or sensitivity analyses.
- Multiplicity: the registry explicitly states that some secondary comparisons were not formally tested under the predefined hierarchical testing procedure.
- Safety comparison: serious adverse-event counts are reported by arm, but no formal comparative safety analysis is reported in the ClinicalTrials.gov record.
- Limited absolute-risk information: the registry-reported primary results include hazard ratios and confidence intervals but not absolute clinical-failure probabilities or median time to failure.
- No unsupported subgroup interpretation: subgroup estimates and interaction analyses are not reported, so treatment-effect heterogeneity cannot be assessed from this record.
21. Why This Trial Matters Statistically
AMBITION is a useful statistical teaching case because it combines several core clinical-trial methods within a single randomized phase 3 design. The same trial moves from a time-to-event primary endpoint to continuous percent-change, binary response, and rank-based functional outcomes.
| Concept | How it appears in AMBITION |
|---|---|
| Randomization | Randomized, parallel-group phase 3 design with 3 arms |
| Masking | Quadruple masking |
| Time-to-event analysis | First adjudicated clinical failure through FAV |
| Stratified log-rank test | Reported for all three primary comparisons |
| Cox model | Used to calculate the primary hazard ratios |
| Hazard ratio | 0.502, 0.477, and 0.528 for the three primary comparisons |
| ANCOVA | Used for NT-proBNP percent change at Week 24 |
| Logistic regression | Used for satisfactory clinical response at Week 24 |
| Wilcoxon Rank Sum | Used for 6 Minute Walk Distance, WHO Functional Class, and Borg Dyspnea Index |
| Multiplicity | Explicit hierarchy-related limitations for selected secondary comparisons |
| Confidence intervals | Reported for all registry-reported statistical effect estimates |
| mITT population | Used for the primary and registry-reported secondary analyses |
The educational value is especially strong because the trial demonstrates that “the statistical analysis” is not one procedure. A well-designed trial can require several different methods, each matched to the endpoint's data type and scientific question.
22. What the Primary Hazard Ratio Does — and Does Not — Mean
The estimated hazard of first adjudicated clinical failure under combination therapy was 0.502 times the estimated hazard under pooled monotherapy in the reported Cox analysis.
There are several ways this result can be misunderstood.
- It is not a 49.8 percentage-point reduction. The 49.8% figure is a relative reduction in the estimated hazard derived from 1 − 0.502.
- It is not a probability. A hazard ratio does not state the probability that an individual patient will experience clinical failure.
- It is not necessarily constant risk at every time point. The Cox interpretation is tied to a proportional-hazards framework.
- It does not give the median time to failure. No median clinical-failure time is reported in the ClinicalTrials.gov record.
- It does not replace absolute outcomes. Absolute event probabilities or time-specific survival estimates would answer different questions.
The confidence interval is equally important. A point estimate can appear precise simply because it is displayed to three decimal places, but numerical formatting is not statistical precision. The interval from 0.348 to 0.724 is the appropriate companion to the 0.502 estimate for communicating uncertainty.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT01178073 — AMBITION.
- PubMed: PMID 32161055.
- PubMed: PMID 31511080.
- PubMed: PMID 28039187.
- PubMed: PMID 26308684.
- PubMed: PMID 26196225.
Continue through Clinical Biostats
Explore the statistical methods behind randomized trials, then move from methodology to practical calculation and analysis workflows.
26. Record Summary
AMBITION provides a compact example of how several statistical frameworks can coexist within one randomized phase 3 trial. Its primary endpoint is a time-to-event outcome analyzed with stratified log-rank testing and Cox-model hazard ratios, while its secondary outcomes use ANCOVA, logistic regression, and stratified Wilcoxon Rank Sum testing.
The primary pooled comparison produced an HR of 0.502 with a 95% CI of 0.348 to 0.724 and a p-value of 0.0002. The corresponding comparisons with ambrisentan and tadalafil monotherapy produced HRs of 0.477 and 0.528, respectively. The secondary results provide additional effect measures on biomarker, response, functional-capacity, and symptom scales, while the registry explicitly identifies multiplicity-related limits for selected comparisons.
The most important statistical lesson is that these estimates should be interpreted according to their underlying estimands. A hazard ratio describes relative event hazards; an odds ratio describes relative odds; a mean percent difference describes a difference in mean percent change; and a median difference describes a difference on the reported median-difference scale. None of these measures should be substituted for another.