← Clinical Trials
Pulmonary Arterial Hypertension Phase 3 Randomized NCT04576988

STELLAR: Complete Statistical Analysis of Sotatercept in Pulmonary Arterial Hypertension

An independent statistical review of the randomized phase 3 STELLAR trial evaluating sotatercept plus background pulmonary arterial hypertension therapy versus placebo plus background therapy, with emphasis on the reported nonparametric, categorical, and time-to-event analyses.

STELLAR  ·  Phase 3  ·  324 enrolled  ·  Completed
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record. Statistical explanations distinguish reported analyses from general interpretation of the methods.

1. Trial at a Glance

STELLAR was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating sotatercept plus background PAH therapy versus placebo plus background PAH therapy in participants with pulmonary arterial hypertension. The registry reports 324 participants, two treatment arms, and a completed study.

324
Enrolled
Randomized trial
2
Arms
Parallel design
40.8
6MWD difference
95% CI 27.53–54.14
0.163
Clinical worsening HR
95% CI 0.076–0.347
FeatureSTELLAR
Trial nameSTELLAR
ClinicalTrials.gov identifierNCT04576988
PhasePhase 3
ConditionPulmonary Arterial Hypertension
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment324
Arms2
StatusCompleted
StartJanuary 25, 2021
Primary completionAugust 26, 2022
Lead sponsorAcceleron Pharma, Inc., a wholly-owned subsidiary of Merck & Co., Inc., Rahway, NJ USA
Sponsor typeIndustry

2. Clinical Question

The central question was whether sotatercept added to background PAH therapy differed from placebo added to background PAH therapy with respect to functional capacity, clinical worsening, and other prespecified measures of pulmonary vascular status, biomarkers, symptoms, functional class, and risk status.

Population

Participants with pulmonary arterial hypertension enrolled in the randomized phase 3 STELLAR trial.

Intervention

Sotatercept plus background PAH therapy.

Comparator

Placebo plus background PAH therapy.

Primary question

How did the randomized treatment groups differ on the registered primary endpoints, particularly change from baseline in 6MWD at Week 24?

For statistical interpretation, the distinction between the randomized treatment contrast and the individual endpoint is important. The trial did not rely on a single universal effect measure. The reported analyses included a nonparametric treatment difference for continuous outcomes, a hazard ratio for a time-to-event outcome, and two-sided CMH tests for binary outcomes.

3. Trial Design

01
Randomize324 participants
02
Two armsSotatercept or placebo
03
Background therapyPAH therapy continued as part of both comparisons
04
Week 246MWD and multiple secondary outcomes
05
Longer follow-upClinical worsening endpoint up to approximately 18 months
Allocation
Randomized
Design model
Parallel
Masking
Quadruple
Primary purpose
Treatment
ARM · SOTATERCEPT

Sotatercept plus background PAH therapy

  • Sotatercept
  • Background PAH therapy
  • Primary comparison period: DBPC period
  • Week 24 assessments included functional, physiologic, biomarker, symptom, and risk-related outcomes
ARM · PLACEBO

Placebo plus background PAH therapy

  • Placebo
  • Background PAH therapy
  • Primary comparison period: DBPC period
  • Week 24 assessments included the same registered outcome framework

The registry describes the trial as randomized, parallel, and quadruple-masked. These design features matter statistically because randomization creates the basis for comparing treatment groups, while masking can reduce the opportunity for knowledge of treatment assignment to influence assessment or behavior. The ClinicalTrials.gov record does not provide a more detailed description of which four parties were masked, so this page does not infer that detail.

4. Endpoints

ClinicalTrials.gov lists three primary endpoints. One is a continuous functional endpoint, while the other two are binary safety-related endpoints.

Registered primary endpointTime frameDefinition / assessmentFormal analysis in the ClinicalTrials.gov record
Change From Baseline in 6-Minute Walk Distance (6MWD) at Week 24Baseline and Week 24The 6MWD was the distance walked in 6 minutes as a measure of functional capacity. This was assessed using the 6-minute walk test (6MWT). Per protocol, change from baseline in 6MWD at Week 24 was reported for DBPC period.Aligned Rank Stratified Wilcoxon (ARSW); Hodges-Lehmann treatment difference
Number of Participants Who Experienced an Adverse Event (AE)Up to approximately 24 weeksAn AE was any untoward medical occurrence in a study participant administered a study drug, which did not necessarily have a causal relationship with this treatment.No formal statistical comparison is reported in the primary analysis record
Number of Participants Who Discontinued Study Treatment Due to an AEUp to approximately 24 weeksAn AE was any untoward medical occurrence in a study participant administered a study drug, which did not necessarily have a causal relationship with this treatment.No formal statistical comparison is reported in the primary analysis record

The registry also reports nine secondary statistical analyses in the ClinicalTrials.gov record. These cover pulmonary vascular resistance, NT-proBNP, clinical worsening, PAH-SYMPACT domains, multicomponent improvement, WHO functional-class improvement, and maintenance or achievement of a low risk score.

5. Analysis Populations and Statistical Structure

The statistical analyses posted on ClinicalTrials.gov consistently define the main Week 24 analysis population as all randomized participants who received at least one dose of study treatment and had the relevant baseline measurement. The time-to-event analysis has a more specific outcome-based definition: all randomized participants who received at least one dose of study treatment and who died or experienced a first clinical worsening event.

Analysis populationRegistry definition / role
Week 24 continuous outcomesAll randomized participants who received at least one dose of study treatment and had the relevant baseline value.
Binary Week 24 outcomesAll randomized participants who received at least one dose of study treatment and had the relevant baseline measurement.
Death or first clinical worsening eventAll randomized participants who received at least one dose of study treatment and who died or experienced a first clinical worsening event.
Why this matters: The analysis population is part of the estimand. A treatment difference calculated among participants with a baseline measurement answers a somewhat different operational question from a calculation that includes every randomized participant regardless of whether the measurement was observed. Because the ClinicalTrials.gov record specifies these populations, they should be retained when interpreting the reported estimates.

6. Primary Result: Change From Baseline in 6MWD at Week 24

The primary efficacy analysis compared change from baseline in 6MWD at Week 24 between sotatercept plus background PAH therapy and placebo plus background PAH therapy during the DBPC period. The registry reports an Aligned Rank Stratified Wilcoxon (ARSW) analysis and a Hodges-Lehmann location-shift estimate.

Hodges-Lehmann treatment difference in 6MWD

40.8 meters

95% CI: 27.53–54.14   ·   P < 0.001

Two-sided 95% confidence interval; analysis used the ARSW framework.

FeatureReported result
OutcomeChange From Baseline in 6-Minute Walk Distance (6MWD) at Week 24
Time frameBaseline and Week 24
UnitMeters
AnalysisAligned Rank Stratified Wilcoxon (ARSW)
Effect measureTreatment difference
Estimate40.8
95% CI27.53 to 54.14
P-value<0.001
Missing-data / rank handlingDeaths assigned the worst rank; missing data due to non-fatal clinical worsening assigned the next worst rank
Clinical Biostats interpretation

The reported 40.8-meter treatment difference is a Hodges-Lehmann location-shift estimate from the nonparametric analysis. In practical statistical terms, it summarizes the location of the treatment-group difference in change from baseline under the rank-based analysis used by the registry.

It does not mean that every participant receiving sotatercept walked exactly 40.8 meters farther, nor does it mean that each individual participant experienced a 40.8-meter improvement. A treatment-group location estimate summarizes a distribution; individual responses can be above or below that value.

The 95% CI of 27.53 to 54.14 meters describes uncertainty around the estimated treatment difference under the analysis framework. It is not a range containing 95% of individual patient responses and should not be interpreted as an individual-response interval.

The P-value of <0.001 addresses evidence against the null hypothesis represented by the statistical test. It does not measure the magnitude or clinical importance of the effect. Effect size and precision are conveyed more directly by the 40.8-meter estimate and its confidence interval.

The analysis also has an important methodological feature: deaths were assigned the worst rank and missing observations attributable to non-fatal clinical worsening were assigned the next worst rank. Consequently, the analysis is not simply a comparison of observed 6MWD changes among participants with complete Week 24 measurements. The missing-data and intercurrent-event rules are part of what the reported estimate means.

7. How the 6MWD Analysis Works

The choice of an Aligned Rank Stratified Wilcoxon method indicates that the primary analysis was rank-based rather than a conventional mean-difference model. This is useful when the analysis is designed to compare the distributions or relative locations of treatment responses without relying on the same assumptions required for a standard parametric comparison of means.

Conceptual rank-based comparison
Observed responses  →  alignment  →  ranks  →  treatment comparison

The key idea is that the analysis operates on ordered information rather than treating the raw outcome values as if they must satisfy a particular normal-theory model.

The word aligned is important. Alignment is intended to remove the contribution of relevant nuisance or stratification structure before ranking, while the subsequent rank comparison evaluates the treatment contrast. The registry specifically identifies the method as ARSW, so it is more accurate to describe the analysis as a stratified rank procedure than simply to call it an ordinary Wilcoxon test.

The reported Hodges-Lehmann estimate complements the rank test by giving the analysis an interpretable effect-size scale: meters. This is preferable to reporting only a P-value because a hypothesis test can establish evidence for a difference without communicating the size or precision of that difference.

8. Secondary Result: Pulmonary Vascular Resistance

Pulmonary vascular resistance (PVR) was evaluated as a secondary continuous endpoint at Week 24. The registry reports an ARSW test with a Hodges-Lehmann location-shift estimate.

Treatment difference in PVR

−234.6 dynes*sec/cm5

95% CI: −288.37 to −180.75   ·   P < 0.001

Two-sided 95% confidence interval.

Statistical interpretation

The negative estimate indicates that the estimated treatment-group location shift in change from baseline was lower for the sotatercept group relative to the placebo group on the reported PVR scale. The entire reported 95% CI is below zero, from −288.37 to −180.75.

The result should not be read as a universal decrease of 234.6 dynes*sec/cm5 for every participant. It is a population-level treatment difference estimated through a rank-based method.

The P-value of <0.001 is evidence from the reported statistical test; it is not a measure of how large the PVR difference is. The estimate and confidence interval provide that magnitude-and-precision information.

9. Secondary Result: NT-proBNP

Change from baseline in NT-proBNP levels at Week 24 was another secondary continuous endpoint analyzed with an ARSW test and Hodges-Lehmann estimate.

Treatment difference in NT-proBNP

−441.6 pg/mL

95% CI: −573.54 to −309.61   ·   P <.001

Two-sided 95% confidence interval.

Clinical Biostats interpretation

The reported estimate of −441.6 pg/mL represents the estimated location shift in change from baseline between the randomized treatment groups under the specified rank-based analysis. The negative direction indicates a lower change on the NT-proBNP scale in the sotatercept group relative to the placebo group.

The 95% CI from −573.54 to −309.61 quantifies uncertainty around that estimated treatment difference. Because the interval remains below zero, the reported estimate is consistently in the same direction across the interval.

Again, the P-value of <.001 should not be interpreted as an effect-size measure. It answers a hypothesis-testing question; the estimated difference and confidence interval communicate magnitude and precision.

10. Secondary Result: Time to Death or First Clinical Worsening Event

STELLAR also included a time-to-event endpoint defined as time to death or the first occurrence of a clinical worsening event, with a time frame of up to approximately 18 months. The registry reports a log-rank analysis and a Cox proportional-hazards model for the hazard ratio.

Hazard ratio for death or first clinical worsening event

0.163

95% CI: 0.076–0.347   ·   P < 0.001

Two-sided 95% confidence interval; log-rank comparison with Cox proportional-hazards estimation.

FeatureReported result
EndpointTime to Death or the First Occurrence of Clinical Worsening Event
Time frameUp to approximately 18 months
UnitWeeks
AnalysisLog Rank
Effect measureHazard Ratio (HR)
HR0.163
95% CI0.076 to 0.347
P-value<0.001
StratificationWHO FC II/III and background PAH therapy
ModelCox proportional hazard model with treatment group as the covariate, stratified by WHO FC II/III and background PAH therapy
Clinical Biostats interpretation

A hazard ratio of 0.163 means that, under the fitted Cox model and over the analyzed follow-up, the estimated instantaneous hazard of the composite event was about 16.3% as large in the sotatercept group as in the placebo group. Equivalently, the point estimate corresponds to an approximately 83.7% lower estimated hazard because 1 − 0.163 = 0.837.

This does not mean that 83.7% of participants avoided the event, that the probability of an event was reduced by exactly 83.7%, or that every participant had the same proportional reduction. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.

The 95% CI of 0.076 to 0.347 communicates uncertainty around the estimated hazard ratio. It is not the range of individual treatment effects. The interval is entirely below 1, which is consistent with the direction of the reported treatment effect.

The P-value of <0.001 provides evidence from the reported log-rank test against the relevant null comparison. It does not quantify the magnitude of the treatment effect. The HR and its confidence interval provide that information.

Finally, the Cox interpretation depends on the proportional-hazards model used to estimate the HR. If the relative hazards change materially over time, a single HR can compress a more complicated time-dependent treatment effect into one summary number. The ClinicalTrials.gov record does not provide a formal proportional-hazards diagnostic, so the HR should be interpreted as the model-based summary that was reported rather than as a claim that proportional hazards were empirically demonstrated.

11. Why the Survival Analysis Is Different From the 6MWD Analysis

The 6MWD endpoint is a continuous change-from-baseline outcome measured at a specified time point. The death-or-clinical-worsening endpoint is a time-to-event outcome. These structures require different statistical summaries.

6MWD

The analysis compares change from baseline at Week 24 using a rank-based ARSW procedure and reports a treatment difference through a Hodges-Lehmann estimate.

Clinical worsening

The analysis uses event times, with a log-rank test for group comparison and a Cox model to estimate the hazard ratio.

Different effect scales

40.8 meters and HR 0.163 are not quantities that can be directly compared numerically. Each belongs to a different estimand and outcome scale.

Censoring

Time-to-event analysis can incorporate participants whose event status is not observed through the full follow-up by treating their available follow-up according to the censoring framework.

This distinction is fundamental. A strong clinical-trial analysis should not force every endpoint into the same statistical template. The outcome structure determines the appropriate analysis, and the interpretation must stay attached to the effect measure that was actually reported.

12. Secondary Patient-Reported and Functional Outcomes

Three PAH-SYMPACT® domain scores were analyzed as secondary continuous outcomes. All used the ARSW framework with Hodges-Lehmann treatment differences and the same general rank-handling rule for deaths and non-fatal clinical worsening-related missing data.

EndpointEstimate95% CIP-value
Physical Impacts Domain Score at Week 24−0.26−0.490 to −0.0400.010
Cardiopulmonary Symptoms Domain Score at Week 24−0.13−0.256 to −0.0140.028
Cognitive/Emotional Impacts Domain Score at Week 24−0.16−0.399 to 0.0840.156

Physical Impacts Domain

The reported Hodges-Lehmann treatment difference was −0.26, with a 95% CI of −0.490 to −0.040 and P = 0.010. The estimate is below zero, and the reported confidence interval also remains below zero.

Cardiopulmonary Symptoms Domain

The reported treatment difference was −0.13, with a 95% CI of −0.256 to −0.014 and P = 0.028. The estimate and interval are both negative.

Cognitive/Emotional Impacts Domain

The reported treatment difference was −0.16, with a 95% CI of −0.399 to 0.084 and P = 0.156. Unlike the other two reported PAH-SYMPACT domains, the confidence interval crosses zero.

Multiplicity matters: These are multiple secondary endpoints. A collection of nominal P-values should not automatically be interpreted as though each were an independently confirmatory test with its own unrestricted type I error budget. The ClinicalTrials.gov record identifies the analyses and P-values but do not provide a multiplicity-adjustment scheme for these secondary endpoints, so this page does not infer one.

13. Secondary Binary Outcomes

Four reported secondary outcomes used the Cochran-Mantel-Haenszel (CMH) method. The ClinicalTrials.gov record reports two-sided P-values and identifies WHO FC II/III and background PAH therapy as strata.

Secondary endpointAnalysisStrataP-value
Change From Baseline in the Percentage of Participants Achieving Multicomponent Improvement at Week 24Cochran-Mantel-HaenszelWHO FC II/III and background PAH therapy<0.001
Change From Baseline in the Percentage of Participants Who Improve in WHO FC at Week 24Cochran-Mantel-HaenszelWHO FC II/III and background PAH therapy<0.001
Change From Baseline in Percentage of Participants Who Maintain or Achieve a Low Risk Score Using the Simplified French Risk Score Calculator at Week 24Cochran-Mantel-HaenszelWHO FC II/III and background PAH therapy<0.001

The registry-reported statistical-analysis list contains three CMH outcomes. the ClinicalTrials.gov record identifies the overall registry as having 10 statistical analyses, with the remaining reported analyses consisting of the primary 6MWD analysis, PVR, NT-proBNP, time to death or clinical worsening, and the three PAH-SYMPACT domains.

For a CMH analysis, the important idea is that the treatment comparison is evaluated while accounting for prespecified categorical strata. Rather than collapsing all observations into a single unadjusted table, the method combines information across strata while preserving the stratified structure.

Conceptual CMH structure
Treatment effect = combined evidence across prespecified strata

The method is particularly useful when the analysis is intended to control for categorical stratification variables while assessing an association between treatment and a binary outcome.

14. Stratification and Covariate Adjustment

The ClinicalTrials.gov record explicitly identify WHO FC II/III and background PAH therapy as strata for the time-to-event and CMH analyses. For the clinical-worsening endpoint, the log-rank P-value was calculated with those variables as strata, and the Cox model was likewise stratified by them.

MethodStratification / adjustment described in the registry
Log-rank test for death or clinical worseningWHO FC II/III and background PAH therapy
Cox model for death or clinical worseningStratified by WHO FC II/III and background PAH therapy
CMH secondary binary analysesWHO FC II/III and background PAH therapy

Stratification does not mean that treatment effects are estimated separately for each stratum and then simply averaged. In a stratified analysis, the comparison is structured so that information from the strata contributes to an overall treatment comparison while respecting the stratification variables.

This can be important when the stratification variables are associated with the outcome or represent clinically meaningful baseline structure. It can also make the analysis more closely aligned with the randomized trial design when the same factors were used to structure the treatment comparison.

15. Missing Data and Rank Assignment

One of the most statistically distinctive features of the registry-reported STELLAR analyses is the explicit handling of missing Week 24 data associated with clinical outcomes. For the primary 6MWD analysis and the reported ARSW secondary continuous analyses, the registry states that participants who died were assigned the worst rank and participants with missing data due to a non-fatal clinical worsening event were assigned the next worst rank.

Death

Assigned the worst rank in the reported rank-based analyses.

Non-fatal clinical worsening

Missing data due to a non-fatal clinical worsening event were assigned the next worst rank.

Observed Week 24 value

Participants with the relevant baseline and Week 24 information contributed their observed outcome to the rank-based comparison.

Interpretive consequence

The analysis integrates important clinical events into the ranking rather than treating all missing observations as statistically interchangeable.

This is more than a technical detail. Missing-data handling can change the estimand. A simple complete-case analysis would implicitly focus on participants with available measurements. The rank assignment used here instead explicitly gives poor outcomes the worst ranks when death or clinical worsening prevents an informative Week 24 measurement.

Multiple imputation: the ClinicalTrials.gov record and individual analysis records identify “Multiple imputation / missing data” as an analysis concept, but the ClinicalTrials.gov record does not provide enough detail to reconstruct a specific imputation algorithm, number of imputations, imputation model, or sensitivity-analysis framework. This page therefore discusses the reported missing-data/rank rules without inventing those additional details.

16. Statistical Methods Explained

Why was an ARSW analysis used for 6MWD?

The registry reports an Aligned Rank Stratified Wilcoxon analysis. A rank-based approach compares ordered responses rather than requiring the primary comparison to be expressed solely as a difference in means. The alignment and stratification components allow the analysis to account for relevant structure before the rank comparison. The exact implementation details beyond the registry description are not reproduced here.

What does the Hodges-Lehmann estimate of 40.8 mean?

It is a location-shift estimate of the overall treatment difference from the rank-based analysis, expressed in the original 6MWD unit of meters. It is best understood as an effect-size summary accompanying the nonparametric hypothesis test. It does not mean that every individual patient experienced exactly a 40.8-meter difference.

Why is a hazard ratio used for clinical worsening?

The endpoint is defined by the time until death or the first clinical worsening event. A hazard ratio summarizes the relative instantaneous event rate between treatment groups under the Cox model. Unlike a simple proportion of participants with an event, it uses the timing of events and the available follow-up.

What does HR 0.163 mean?

Under the fitted model, the estimated instantaneous hazard in the sotatercept group is 0.163 times the hazard in the placebo group. The corresponding arithmetic interpretation is an approximately 83.7% lower estimated hazard. It is not an 83.7-percentage-point reduction in event probability and does not mean that every patient receives exactly the same relative reduction.

Why use a stratified log-rank test?

The registry analysis specifies WHO FC II/III and background PAH therapy as strata. A stratified log-rank test allows the time-to-event comparison to account for that stratified structure rather than ignoring it. The resulting P-value is therefore tied to the stratified comparison described in the registry.

Why use the CMH test for binary outcomes?

The CMH test is designed for categorical outcomes when the analysis includes strata. In STELLAR, the ClinicalTrials.gov record identifies WHO FC II/III and background PAH therapy as the strata. The method combines evidence across those strata while preserving the stratified design.

Why does missing-data handling matter so much here?

For the ARSW analyses, deaths and certain clinical-worsening-related missing observations were assigned explicitly poor ranks. That means the analysis encodes the clinical significance of these events rather than treating all missing measurements as equivalent. The resulting estimate therefore depends not only on observed 6MWD or other continuous values but also on the prespecified rank rules.

17. Confidence Intervals and P-values

The reported STELLAR analyses illustrate why confidence intervals and P-values should be read together but should not be treated as interchangeable statistics.

EndpointEstimate95% CIP-value
6MWD treatment difference40.827.53–54.14<0.001
PVR treatment difference−234.6−288.37 to −180.75<0.001
NT-proBNP treatment difference−441.6−573.54 to −309.61<.001
Death / first clinical worsening HR0.1630.076–0.347<0.001
PAH-SYMPACT physical impacts−0.26−0.490 to −0.0400.010
PAH-SYMPACT cardiopulmonary symptoms−0.13−0.256 to −0.0140.028
PAH-SYMPACT cognitive/emotional impacts−0.16−0.399 to 0.0840.156

A confidence interval gives an uncertainty range around the reported effect estimate under the relevant statistical framework. Its interpretation depends on the estimator and assumptions used to construct it. It should not be interpreted as a prediction interval for future individual patients.

A P-value instead measures how compatible the observed data are with a null hypothesis under the specified test. A very small P-value can accompany a modest effect when the data are sufficiently informative, while a larger P-value can occur with a potentially important point estimate when uncertainty is substantial.

A useful reading sequence

First: identify the outcome and effect measure. Second: determine the direction of the estimate. Third: examine the confidence interval for precision and whether it includes the null value. Fourth: read the P-value as evidence from the specified hypothesis test rather than as a measure of effect magnitude.

18. Comparing the Reported Effect Measures

The numerical estimates in STELLAR live on different scales. Treating them as if they were directly comparable would be a statistical error.

MeasureMeaningReported STELLAR example
Treatment differenceDifference in the location of the outcome distribution under the reported rank-based analysis6MWD: 40.8 meters
Treatment differenceLocation-shift estimate for a physiologic or biomarker outcomePVR: −234.6 dynes*sec/cm5
Treatment differenceLocation-shift estimate for a biomarker outcomeNT-proBNP: −441.6 pg/mL
Hazard ratioRelative instantaneous event rate under the Cox modelDeath/clinical worsening: 0.163
CMH P-valueEvidence from a stratified categorical comparisonMulticomponent improvement: <0.001

A difference of 40.8 meters cannot be compared with a hazard ratio of 0.163 by asking which number is “larger” or “smaller.” They represent fundamentally different quantities. Statistical interpretation begins with understanding the estimand before considering the numerical result.

19. Safety Results

The registered primary safety endpoints were the number of participants who experienced an adverse event and the number who discontinued study treatment due to an adverse event, both assessed up to approximately 24 weeks. The ClinicalTrials.gov record does not provide formal comparative estimates, confidence intervals, or P-values for these two primary safety endpoints.

Primary safety endpointTime frameFormal statistical comparison reported
Number of Participants Who Experienced an Adverse Event (AE)Up to approximately 24 weeksNot reported in the listed primary statistical analysis
Number of Participants Who Discontinued Study Treatment Due to an AEUp to approximately 24 weeksNot reported in the listed primary statistical analysis

Serious adverse events by arm are reported in the ClinicalTrials.gov record. The registry data provide affected participants over the corresponding risk sets in two labeled periods.

Period label in the ClinicalTrials.gov recordTreatment armSerious adverse events, affected / at risk
DBPSotatercept Plus Background PAH Therapy23/163
DBPPlacebo Plus Background PAH Therapy36/160
LTDSotatercept Plus Background PAH Therapy26/158
LTDPlacebo Plus Background PAH Therapy14/142

These serious-adverse-event figures are descriptive affected/at-risk counts as reported in the registry by the registry. They should not be transformed here into a new comparative risk estimate because the ClinicalTrials.gov record does not specify a formal statistical analysis for these safety counts, nor do they establish that the two labeled periods are directly interchangeable exposure windows.

Safety interpretation: Affected/at-risk counts are not the same as incidence rates, hazard ratios, or risk differences. The denominators shown above are the registry's reported at-risk counts for the labeled periods. No additional exposure-time calculation is warranted from the ClinicalTrials.gov record.

20. Design Topics That Are Not Supported by the Supplied Record

Several statistical topics commonly appear in clinical-trial analyses but are not documented sufficiently in the registry-reported STELLAR data to support a detailed trial-specific reconstruction.

TopicWhat can be stated from the ClinicalTrials.gov record
Non-inferiority marginNo non-inferiority margin is reported in the ClinicalTrials.gov record.
CrossoverNo crossover rule or crossover analysis is reported in the ClinicalTrials.gov record.
Factorial designThe design is described as parallel; no factorial structure is reported.
Bayesian methodsNo Bayesian analysis is identified in the registry-reported methods.
Interim analysisNo interim-analysis boundary, alpha-spending method, or interim decision rule is reported.
MultiplicityThe ClinicalTrials.gov record identifies multiple endpoints and P-values but do not provide a detailed multiplicity-control procedure.
Imputation modelMultiple imputation / missing data is identified as an analysis concept, but the ClinicalTrials.gov record does not specify an imputation model or number of imputations.

This distinction is important for a statistical teaching page. The absence of a documented method in the ClinicalTrials.gov record should not be replaced by a generic assumption about what a phase 3 trial “must have” done. Trial-specific methodology should be reported only when the available record supports it.

21. Primary Endpoint Interpretation in Context

The primary 6MWD result is statistically notable because it combines four layers of information: a prespecified Week 24 functional endpoint, a rank-based analysis, an explicit treatment-difference estimate, and a missing-data rule that assigns poor ranks to clinically serious intercurrent events.

Endpoint scale

The effect is expressed in meters, making the treatment contrast directly interpretable on the 6MWD scale.

Nonparametric method

The ARSW analysis uses ranks rather than relying solely on a conventional parametric comparison of means.

Precision

The 95% CI of 27.53–54.14 provides information about uncertainty around the 40.8 estimate.

Clinical events

Death and non-fatal clinical worsening are explicitly incorporated into the ranking rules for the analysis.

The result should therefore be read as a statistical treatment contrast under a particular estimand and analysis rule, rather than simply as “the average improvement in walking distance.” That distinction is especially important when outcomes can be affected by death, clinical worsening, or missing follow-up measurements.

22. Interpreting the Clinical Worsening Hazard Ratio

The time-to-event result adds a different dimension to the evidence. The 6MWD analysis asks about functional capacity at Week 24, while the clinical-worsening analysis asks about when a participant first experienced death or a clinical worsening event.

Hazard-ratio interpretation
HR = 0.163  →  estimated hazard ratio under the stratified Cox model

A hazard ratio below 1 indicates a lower estimated instantaneous event hazard in the sotatercept group relative to the placebo group under the fitted model. The numerical complement, 1 − HR, should not be confused with an absolute event-rate reduction.

The reported confidence interval, 0.076 to 0.347, is important because it shows that the point estimate is not the only plausible value supported by the statistical analysis. The interval is relatively broad on the multiplicative scale, even though it remains below 1.

Another important distinction is between hazard and . Hazard describes the instantaneous event rate conditional on having remained event-free up to a particular point in time. A hazard ratio compares those rates. It does not directly provide the probability that a randomly selected participant will experience the event by a specified date.

23. What the Secondary Results Add

The secondary analyses span several domains: pulmonary vascular physiology, biomarkers, time-to-event outcomes, patient-reported symptoms and impacts, functional class, multicomponent improvement, and risk status. This breadth is statistically useful because different endpoints capture different dimensions of the disease process.

DomainEndpoint examplesStatistical method
Functional capacity6MWDARSW / Hodges-Lehmann
Pulmonary vascular physiologyPVRARSW / Hodges-Lehmann
BiomarkerNT-proBNPARSW / Hodges-Lehmann
Clinical eventsDeath or first clinical worseningLog-rank / Cox HR
Patient-reported outcomesPAH-SYMPACT domainsARSW / Hodges-Lehmann
Categorical improvementMulticomponent improvement, WHO FC improvement, low risk scoreCMH

The statistical lesson is that consistency across different outcome domains can be informative, but it does not turn every secondary endpoint into an independent confirmatory conclusion. Each endpoint has its own measurement properties, analysis population, estimand, uncertainty, and potential multiplicity considerations.

24. Limitations and Interpretation Issues

25. Why This Trial Matters Statistically

STELLAR is a useful teaching case because it demonstrates how a single randomized clinical trial can require several distinct statistical frameworks depending on the endpoint. Its the ClinicalTrials.gov record includes rank-based continuous-outcome analyses, a stratified survival analysis, a Cox model, and stratified categorical testing.

ConceptHow it appears in STELLAR
RandomizationRandomized parallel-group phase 3 design with 324 enrolled participants.
MaskingQuadruple-masked design.
Nonparametric analysisARSW analyses for 6MWD and several continuous secondary outcomes.
Hodges-Lehmann estimateUsed to report treatment differences for rank-based continuous outcomes.
Missing-data handlingDeaths and non-fatal clinical-worsening-related missing values assigned poor ranks in the ARSW analyses.
Confidence intervalsReported for the primary 6MWD estimate and several secondary continuous and time-to-event estimates.
Log-rank testUsed for time to death or first clinical worsening.
Hazard ratioCox model estimate of the relative hazard for death or clinical worsening.
Stratified survival analysisWHO FC II/III and background PAH therapy used as strata.
Cochran-Mantel-Haenszel testUsed for reported binary secondary outcomes with the same strata.
MultiplicityMultiple primary and secondary endpoints require careful interpretation of nominal P-values.

26. Trial Timeline

January 25, 2021

Study start

The registry lists January 25, 2021 as the STELLAR study start date.

August 26, 2022

Primary completion

The registry lists August 26, 2022 as the primary completion date.

Completed

Registry status

The ClinicalTrials.gov record identifies the study status as completed.

the ClinicalTrials.gov record therefore provides a defined enrollment and completion history, while the statistical analyses cover both the Week 24 primary/secondary assessments and a clinical-worsening endpoint with follow-up extending to approximately 18 months.

27. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary 6MWD analysis produced a Hodges-Lehmann treatment difference of 40.8 meters with a two-sided 95% CI of 27.53–54.14 and P <0.001. The reported clinical-worsening analysis produced HR 0.163 with a 95% CI of 0.076–0.347 and P <0.001.

Clinical interpretation

The statistical results describe differences between randomized treatment groups across functional capacity and time to death or clinical worsening. Determining the clinical importance of any particular magnitude requires consideration of the endpoint's clinical context rather than relying on statistical significance alone.

The distinction is particularly important for the PAH-SYMPACT outcomes. The physical impacts and cardiopulmonary symptoms domains have reported P-values of 0.010 and 0.028, while the cognitive/emotional impacts domain has a P-value of 0.156. These results should be interpreted as separate endpoint-specific estimates, not collapsed into a single overall conclusion about patient-reported outcomes.

28. A Statistical Reading of the Complete Evidence Set

Viewed as a statistical system, the registry-reported STELLAR results have three complementary layers.

Functional outcome

6MWD provides a continuous measure at Week 24. The ARSW analysis reports the treatment contrast in meters through a Hodges-Lehmann estimate.

Physiologic and biomarker outcomes

PVR and NT-proBNP provide additional continuous endpoints, analyzed with the same general rank-based framework.

Clinical events

Time to death or clinical worsening captures event timing and is summarized through the log-rank test and Cox hazard ratio.

Categorical outcomes

CMH analyses evaluate reported binary secondary outcomes while accounting for WHO FC II/III and background PAH therapy strata.

The resulting statistical narrative is therefore broader than any one P-value. The primary endpoint supplies a treatment difference on a functional scale; secondary continuous endpoints provide additional treatment contrasts; the time-to-event endpoint provides a relative hazard measure; and the categorical analyses provide stratified evidence for binary outcomes.

29. Related Tutorials

Learn more about the methods used in this trial:

30. Related Statistical Calculators

31. Sources

The trial-specific numerical results and methodological descriptions on this page are restricted to the ClinicalTrials.gov record. The PubMed links are provided as navigation to the associated publications and are not used here to introduce additional numerical results.

Continue through the Clinical Biostats statistical learning pathway

Use the trial's endpoints as a starting point for deeper study of rank-based methods, stratified categorical analysis, survival analysis, hazard ratios, confidence intervals, and missing-data methods.

32. Record Summary

STELLAR provides a useful example of how a randomized phase 3 trial can use different statistical methods for different endpoint structures. The primary 6MWD endpoint was analyzed using an Aligned Rank Stratified Wilcoxon procedure with a Hodges-Lehmann treatment difference of 40.8 meters and a two-sided 95% CI of 27.53–54.14. Secondary continuous outcomes used the same general rank-based framework, while the time-to-death-or-clinical-worsening endpoint used a stratified log-rank test and Cox proportional-hazards model, producing an HR of 0.163 with a 95% CI of 0.076–0.347. Binary secondary outcomes used the Cochran-Mantel-Haenszel method with WHO FC II/III and background PAH therapy as strata.

The most important statistical lesson is that the numbers cannot be interpreted independently of their methods. A treatment difference in meters, a location shift in a biomarker, a hazard ratio, and a CMH P-value answer different questions. Confidence intervals describe precision around their respective estimates, while P-values provide hypothesis-testing evidence rather than measures of effect magnitude.

Clinical Biostats methodology: A trial-results page should not merely reproduce reported numbers. The goal is to reconstruct the statistical story of the trial: what was randomized, which endpoints were measured, which estimands and statistical methods were used, how missing observations were handled, what the reported estimates mean, and where the available registry information limits interpretation.