This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record. Statistical explanations distinguish reported analyses from general interpretation of the methods.
1. Trial at a Glance
STELLAR was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating sotatercept plus background PAH therapy versus placebo plus background PAH therapy in participants with pulmonary arterial hypertension. The registry reports 324 participants, two treatment arms, and a completed study.
| Feature | STELLAR |
|---|---|
| Trial name | STELLAR |
| ClinicalTrials.gov identifier | NCT04576988 |
| Phase | Phase 3 |
| Condition | Pulmonary Arterial Hypertension |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 324 |
| Arms | 2 |
| Status | Completed |
| Start | January 25, 2021 |
| Primary completion | August 26, 2022 |
| Lead sponsor | Acceleron Pharma, Inc., a wholly-owned subsidiary of Merck & Co., Inc., Rahway, NJ USA |
| Sponsor type | Industry |
2. Clinical Question
The central question was whether sotatercept added to background PAH therapy differed from placebo added to background PAH therapy with respect to functional capacity, clinical worsening, and other prespecified measures of pulmonary vascular status, biomarkers, symptoms, functional class, and risk status.
Population
Participants with pulmonary arterial hypertension enrolled in the randomized phase 3 STELLAR trial.
Intervention
Sotatercept plus background PAH therapy.
Comparator
Placebo plus background PAH therapy.
Primary question
How did the randomized treatment groups differ on the registered primary endpoints, particularly change from baseline in 6MWD at Week 24?
For statistical interpretation, the distinction between the randomized treatment contrast and the individual endpoint is important. The trial did not rely on a single universal effect measure. The reported analyses included a nonparametric treatment difference for continuous outcomes, a hazard ratio for a time-to-event outcome, and two-sided CMH tests for binary outcomes.
3. Trial Design
Sotatercept plus background PAH therapy
- Sotatercept
- Background PAH therapy
- Primary comparison period: DBPC period
- Week 24 assessments included functional, physiologic, biomarker, symptom, and risk-related outcomes
Placebo plus background PAH therapy
- Placebo
- Background PAH therapy
- Primary comparison period: DBPC period
- Week 24 assessments included the same registered outcome framework
The registry describes the trial as randomized, parallel, and quadruple-masked. These design features matter statistically because randomization creates the basis for comparing treatment groups, while masking can reduce the opportunity for knowledge of treatment assignment to influence assessment or behavior. The ClinicalTrials.gov record does not provide a more detailed description of which four parties were masked, so this page does not infer that detail.
4. Endpoints
ClinicalTrials.gov lists three primary endpoints. One is a continuous functional endpoint, while the other two are binary safety-related endpoints.
| Registered primary endpoint | Time frame | Definition / assessment | Formal analysis in the ClinicalTrials.gov record |
|---|---|---|---|
| Change From Baseline in 6-Minute Walk Distance (6MWD) at Week 24 | Baseline and Week 24 | The 6MWD was the distance walked in 6 minutes as a measure of functional capacity. This was assessed using the 6-minute walk test (6MWT). Per protocol, change from baseline in 6MWD at Week 24 was reported for DBPC period. | Aligned Rank Stratified Wilcoxon (ARSW); Hodges-Lehmann treatment difference |
| Number of Participants Who Experienced an Adverse Event (AE) | Up to approximately 24 weeks | An AE was any untoward medical occurrence in a study participant administered a study drug, which did not necessarily have a causal relationship with this treatment. | No formal statistical comparison is reported in the primary analysis record |
| Number of Participants Who Discontinued Study Treatment Due to an AE | Up to approximately 24 weeks | An AE was any untoward medical occurrence in a study participant administered a study drug, which did not necessarily have a causal relationship with this treatment. | No formal statistical comparison is reported in the primary analysis record |
The registry also reports nine secondary statistical analyses in the ClinicalTrials.gov record. These cover pulmonary vascular resistance, NT-proBNP, clinical worsening, PAH-SYMPACT domains, multicomponent improvement, WHO functional-class improvement, and maintenance or achievement of a low risk score.
5. Analysis Populations and Statistical Structure
The statistical analyses posted on ClinicalTrials.gov consistently define the main Week 24 analysis population as all randomized participants who received at least one dose of study treatment and had the relevant baseline measurement. The time-to-event analysis has a more specific outcome-based definition: all randomized participants who received at least one dose of study treatment and who died or experienced a first clinical worsening event.
| Analysis population | Registry definition / role |
|---|---|
| Week 24 continuous outcomes | All randomized participants who received at least one dose of study treatment and had the relevant baseline value. |
| Binary Week 24 outcomes | All randomized participants who received at least one dose of study treatment and had the relevant baseline measurement. |
| Death or first clinical worsening event | All randomized participants who received at least one dose of study treatment and who died or experienced a first clinical worsening event. |
6. Primary Result: Change From Baseline in 6MWD at Week 24
The primary efficacy analysis compared change from baseline in 6MWD at Week 24 between sotatercept plus background PAH therapy and placebo plus background PAH therapy during the DBPC period. The registry reports an Aligned Rank Stratified Wilcoxon (ARSW) analysis and a Hodges-Lehmann location-shift estimate.
Hodges-Lehmann treatment difference in 6MWD
95% CI: 27.53–54.14 · P < 0.001
Two-sided 95% confidence interval; analysis used the ARSW framework.
| Feature | Reported result |
|---|---|
| Outcome | Change From Baseline in 6-Minute Walk Distance (6MWD) at Week 24 |
| Time frame | Baseline and Week 24 |
| Unit | Meters |
| Analysis | Aligned Rank Stratified Wilcoxon (ARSW) |
| Effect measure | Treatment difference |
| Estimate | 40.8 |
| 95% CI | 27.53 to 54.14 |
| P-value | <0.001 |
| Missing-data / rank handling | Deaths assigned the worst rank; missing data due to non-fatal clinical worsening assigned the next worst rank |
The reported 40.8-meter treatment difference is a Hodges-Lehmann location-shift estimate from the nonparametric analysis. In practical statistical terms, it summarizes the location of the treatment-group difference in change from baseline under the rank-based analysis used by the registry.
It does not mean that every participant receiving sotatercept walked exactly 40.8 meters farther, nor does it mean that each individual participant experienced a 40.8-meter improvement. A treatment-group location estimate summarizes a distribution; individual responses can be above or below that value.
The 95% CI of 27.53 to 54.14 meters describes uncertainty around the estimated treatment difference under the analysis framework. It is not a range containing 95% of individual patient responses and should not be interpreted as an individual-response interval.
The P-value of <0.001 addresses evidence against the null hypothesis represented by the statistical test. It does not measure the magnitude or clinical importance of the effect. Effect size and precision are conveyed more directly by the 40.8-meter estimate and its confidence interval.
The analysis also has an important methodological feature: deaths were assigned the worst rank and missing observations attributable to non-fatal clinical worsening were assigned the next worst rank. Consequently, the analysis is not simply a comparison of observed 6MWD changes among participants with complete Week 24 measurements. The missing-data and intercurrent-event rules are part of what the reported estimate means.
7. How the 6MWD Analysis Works
The choice of an Aligned Rank Stratified Wilcoxon method indicates that the primary analysis was rank-based rather than a conventional mean-difference model. This is useful when the analysis is designed to compare the distributions or relative locations of treatment responses without relying on the same assumptions required for a standard parametric comparison of means.
The key idea is that the analysis operates on ordered information rather than treating the raw outcome values as if they must satisfy a particular normal-theory model.
The word aligned is important. Alignment is intended to remove the contribution of relevant nuisance or stratification structure before ranking, while the subsequent rank comparison evaluates the treatment contrast. The registry specifically identifies the method as ARSW, so it is more accurate to describe the analysis as a stratified rank procedure than simply to call it an ordinary Wilcoxon test.
The reported Hodges-Lehmann estimate complements the rank test by giving the analysis an interpretable effect-size scale: meters. This is preferable to reporting only a P-value because a hypothesis test can establish evidence for a difference without communicating the size or precision of that difference.
8. Secondary Result: Pulmonary Vascular Resistance
Pulmonary vascular resistance (PVR) was evaluated as a secondary continuous endpoint at Week 24. The registry reports an ARSW test with a Hodges-Lehmann location-shift estimate.
Treatment difference in PVR
95% CI: −288.37 to −180.75 · P < 0.001
Two-sided 95% confidence interval.
The negative estimate indicates that the estimated treatment-group location shift in change from baseline was lower for the sotatercept group relative to the placebo group on the reported PVR scale. The entire reported 95% CI is below zero, from −288.37 to −180.75.
The result should not be read as a universal decrease of 234.6 dynes*sec/cm5 for every participant. It is a population-level treatment difference estimated through a rank-based method.
The P-value of <0.001 is evidence from the reported statistical test; it is not a measure of how large the PVR difference is. The estimate and confidence interval provide that magnitude-and-precision information.
9. Secondary Result: NT-proBNP
Change from baseline in NT-proBNP levels at Week 24 was another secondary continuous endpoint analyzed with an ARSW test and Hodges-Lehmann estimate.
Treatment difference in NT-proBNP
95% CI: −573.54 to −309.61 · P <.001
Two-sided 95% confidence interval.
The reported estimate of −441.6 pg/mL represents the estimated location shift in change from baseline between the randomized treatment groups under the specified rank-based analysis. The negative direction indicates a lower change on the NT-proBNP scale in the sotatercept group relative to the placebo group.
The 95% CI from −573.54 to −309.61 quantifies uncertainty around that estimated treatment difference. Because the interval remains below zero, the reported estimate is consistently in the same direction across the interval.
Again, the P-value of <.001 should not be interpreted as an effect-size measure. It answers a hypothesis-testing question; the estimated difference and confidence interval communicate magnitude and precision.
10. Secondary Result: Time to Death or First Clinical Worsening Event
STELLAR also included a time-to-event endpoint defined as time to death or the first occurrence of a clinical worsening event, with a time frame of up to approximately 18 months. The registry reports a log-rank analysis and a Cox proportional-hazards model for the hazard ratio.
Hazard ratio for death or first clinical worsening event
95% CI: 0.076–0.347 · P < 0.001
Two-sided 95% confidence interval; log-rank comparison with Cox proportional-hazards estimation.
| Feature | Reported result |
|---|---|
| Endpoint | Time to Death or the First Occurrence of Clinical Worsening Event |
| Time frame | Up to approximately 18 months |
| Unit | Weeks |
| Analysis | Log Rank |
| Effect measure | Hazard Ratio (HR) |
| HR | 0.163 |
| 95% CI | 0.076 to 0.347 |
| P-value | <0.001 |
| Stratification | WHO FC II/III and background PAH therapy |
| Model | Cox proportional hazard model with treatment group as the covariate, stratified by WHO FC II/III and background PAH therapy |
A hazard ratio of 0.163 means that, under the fitted Cox model and over the analyzed follow-up, the estimated instantaneous hazard of the composite event was about 16.3% as large in the sotatercept group as in the placebo group. Equivalently, the point estimate corresponds to an approximately 83.7% lower estimated hazard because 1 − 0.163 = 0.837.
This does not mean that 83.7% of participants avoided the event, that the probability of an event was reduced by exactly 83.7%, or that every participant had the same proportional reduction. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.
The 95% CI of 0.076 to 0.347 communicates uncertainty around the estimated hazard ratio. It is not the range of individual treatment effects. The interval is entirely below 1, which is consistent with the direction of the reported treatment effect.
The P-value of <0.001 provides evidence from the reported log-rank test against the relevant null comparison. It does not quantify the magnitude of the treatment effect. The HR and its confidence interval provide that information.
Finally, the Cox interpretation depends on the proportional-hazards model used to estimate the HR. If the relative hazards change materially over time, a single HR can compress a more complicated time-dependent treatment effect into one summary number. The ClinicalTrials.gov record does not provide a formal proportional-hazards diagnostic, so the HR should be interpreted as the model-based summary that was reported rather than as a claim that proportional hazards were empirically demonstrated.
11. Why the Survival Analysis Is Different From the 6MWD Analysis
The 6MWD endpoint is a continuous change-from-baseline outcome measured at a specified time point. The death-or-clinical-worsening endpoint is a time-to-event outcome. These structures require different statistical summaries.
6MWD
The analysis compares change from baseline at Week 24 using a rank-based ARSW procedure and reports a treatment difference through a Hodges-Lehmann estimate.
Clinical worsening
The analysis uses event times, with a log-rank test for group comparison and a Cox model to estimate the hazard ratio.
Different effect scales
40.8 meters and HR 0.163 are not quantities that can be directly compared numerically. Each belongs to a different estimand and outcome scale.
Censoring
Time-to-event analysis can incorporate participants whose event status is not observed through the full follow-up by treating their available follow-up according to the censoring framework.
This distinction is fundamental. A strong clinical-trial analysis should not force every endpoint into the same statistical template. The outcome structure determines the appropriate analysis, and the interpretation must stay attached to the effect measure that was actually reported.
12. Secondary Patient-Reported and Functional Outcomes
Three PAH-SYMPACT® domain scores were analyzed as secondary continuous outcomes. All used the ARSW framework with Hodges-Lehmann treatment differences and the same general rank-handling rule for deaths and non-fatal clinical worsening-related missing data.
| Endpoint | Estimate | 95% CI | P-value |
|---|---|---|---|
| Physical Impacts Domain Score at Week 24 | −0.26 | −0.490 to −0.040 | 0.010 |
| Cardiopulmonary Symptoms Domain Score at Week 24 | −0.13 | −0.256 to −0.014 | 0.028 |
| Cognitive/Emotional Impacts Domain Score at Week 24 | −0.16 | −0.399 to 0.084 | 0.156 |
Physical Impacts Domain
The reported Hodges-Lehmann treatment difference was −0.26, with a 95% CI of −0.490 to −0.040 and P = 0.010. The estimate is below zero, and the reported confidence interval also remains below zero.
Cardiopulmonary Symptoms Domain
The reported treatment difference was −0.13, with a 95% CI of −0.256 to −0.014 and P = 0.028. The estimate and interval are both negative.
Cognitive/Emotional Impacts Domain
The reported treatment difference was −0.16, with a 95% CI of −0.399 to 0.084 and P = 0.156. Unlike the other two reported PAH-SYMPACT domains, the confidence interval crosses zero.
13. Secondary Binary Outcomes
Four reported secondary outcomes used the Cochran-Mantel-Haenszel (CMH) method. The ClinicalTrials.gov record reports two-sided P-values and identifies WHO FC II/III and background PAH therapy as strata.
| Secondary endpoint | Analysis | Strata | P-value |
|---|---|---|---|
| Change From Baseline in the Percentage of Participants Achieving Multicomponent Improvement at Week 24 | Cochran-Mantel-Haenszel | WHO FC II/III and background PAH therapy | <0.001 |
| Change From Baseline in the Percentage of Participants Who Improve in WHO FC at Week 24 | Cochran-Mantel-Haenszel | WHO FC II/III and background PAH therapy | <0.001 |
| Change From Baseline in Percentage of Participants Who Maintain or Achieve a Low Risk Score Using the Simplified French Risk Score Calculator at Week 24 | Cochran-Mantel-Haenszel | WHO FC II/III and background PAH therapy | <0.001 |
The registry-reported statistical-analysis list contains three CMH outcomes. the ClinicalTrials.gov record identifies the overall registry as having 10 statistical analyses, with the remaining reported analyses consisting of the primary 6MWD analysis, PVR, NT-proBNP, time to death or clinical worsening, and the three PAH-SYMPACT domains.
For a CMH analysis, the important idea is that the treatment comparison is evaluated while accounting for prespecified categorical strata. Rather than collapsing all observations into a single unadjusted table, the method combines information across strata while preserving the stratified structure.
The method is particularly useful when the analysis is intended to control for categorical stratification variables while assessing an association between treatment and a binary outcome.
14. Stratification and Covariate Adjustment
The ClinicalTrials.gov record explicitly identify WHO FC II/III and background PAH therapy as strata for the time-to-event and CMH analyses. For the clinical-worsening endpoint, the log-rank P-value was calculated with those variables as strata, and the Cox model was likewise stratified by them.
| Method | Stratification / adjustment described in the registry |
|---|---|
| Log-rank test for death or clinical worsening | WHO FC II/III and background PAH therapy |
| Cox model for death or clinical worsening | Stratified by WHO FC II/III and background PAH therapy |
| CMH secondary binary analyses | WHO FC II/III and background PAH therapy |
Stratification does not mean that treatment effects are estimated separately for each stratum and then simply averaged. In a stratified analysis, the comparison is structured so that information from the strata contributes to an overall treatment comparison while respecting the stratification variables.
This can be important when the stratification variables are associated with the outcome or represent clinically meaningful baseline structure. It can also make the analysis more closely aligned with the randomized trial design when the same factors were used to structure the treatment comparison.
15. Missing Data and Rank Assignment
One of the most statistically distinctive features of the registry-reported STELLAR analyses is the explicit handling of missing Week 24 data associated with clinical outcomes. For the primary 6MWD analysis and the reported ARSW secondary continuous analyses, the registry states that participants who died were assigned the worst rank and participants with missing data due to a non-fatal clinical worsening event were assigned the next worst rank.
Death
Assigned the worst rank in the reported rank-based analyses.
Non-fatal clinical worsening
Missing data due to a non-fatal clinical worsening event were assigned the next worst rank.
Observed Week 24 value
Participants with the relevant baseline and Week 24 information contributed their observed outcome to the rank-based comparison.
Interpretive consequence
The analysis integrates important clinical events into the ranking rather than treating all missing observations as statistically interchangeable.
This is more than a technical detail. Missing-data handling can change the estimand. A simple complete-case analysis would implicitly focus on participants with available measurements. The rank assignment used here instead explicitly gives poor outcomes the worst ranks when death or clinical worsening prevents an informative Week 24 measurement.
16. Statistical Methods Explained
Why was an ARSW analysis used for 6MWD?
The registry reports an Aligned Rank Stratified Wilcoxon analysis. A rank-based approach compares ordered responses rather than requiring the primary comparison to be expressed solely as a difference in means. The alignment and stratification components allow the analysis to account for relevant structure before the rank comparison. The exact implementation details beyond the registry description are not reproduced here.
What does the Hodges-Lehmann estimate of 40.8 mean?
It is a location-shift estimate of the overall treatment difference from the rank-based analysis, expressed in the original 6MWD unit of meters. It is best understood as an effect-size summary accompanying the nonparametric hypothesis test. It does not mean that every individual patient experienced exactly a 40.8-meter difference.
Why is a hazard ratio used for clinical worsening?
The endpoint is defined by the time until death or the first clinical worsening event. A hazard ratio summarizes the relative instantaneous event rate between treatment groups under the Cox model. Unlike a simple proportion of participants with an event, it uses the timing of events and the available follow-up.
What does HR 0.163 mean?
Under the fitted model, the estimated instantaneous hazard in the sotatercept group is 0.163 times the hazard in the placebo group. The corresponding arithmetic interpretation is an approximately 83.7% lower estimated hazard. It is not an 83.7-percentage-point reduction in event probability and does not mean that every patient receives exactly the same relative reduction.
Why use a stratified log-rank test?
The registry analysis specifies WHO FC II/III and background PAH therapy as strata. A stratified log-rank test allows the time-to-event comparison to account for that stratified structure rather than ignoring it. The resulting P-value is therefore tied to the stratified comparison described in the registry.
Why use the CMH test for binary outcomes?
The CMH test is designed for categorical outcomes when the analysis includes strata. In STELLAR, the ClinicalTrials.gov record identifies WHO FC II/III and background PAH therapy as the strata. The method combines evidence across those strata while preserving the stratified design.
Why does missing-data handling matter so much here?
For the ARSW analyses, deaths and certain clinical-worsening-related missing observations were assigned explicitly poor ranks. That means the analysis encodes the clinical significance of these events rather than treating all missing measurements as equivalent. The resulting estimate therefore depends not only on observed 6MWD or other continuous values but also on the prespecified rank rules.
17. Confidence Intervals and P-values
The reported STELLAR analyses illustrate why confidence intervals and P-values should be read together but should not be treated as interchangeable statistics.
| Endpoint | Estimate | 95% CI | P-value |
|---|---|---|---|
| 6MWD treatment difference | 40.8 | 27.53–54.14 | <0.001 |
| PVR treatment difference | −234.6 | −288.37 to −180.75 | <0.001 |
| NT-proBNP treatment difference | −441.6 | −573.54 to −309.61 | <.001 |
| Death / first clinical worsening HR | 0.163 | 0.076–0.347 | <0.001 |
| PAH-SYMPACT physical impacts | −0.26 | −0.490 to −0.040 | 0.010 |
| PAH-SYMPACT cardiopulmonary symptoms | −0.13 | −0.256 to −0.014 | 0.028 |
| PAH-SYMPACT cognitive/emotional impacts | −0.16 | −0.399 to 0.084 | 0.156 |
A confidence interval gives an uncertainty range around the reported effect estimate under the relevant statistical framework. Its interpretation depends on the estimator and assumptions used to construct it. It should not be interpreted as a prediction interval for future individual patients.
A P-value instead measures how compatible the observed data are with a null hypothesis under the specified test. A very small P-value can accompany a modest effect when the data are sufficiently informative, while a larger P-value can occur with a potentially important point estimate when uncertainty is substantial.
First: identify the outcome and effect measure. Second: determine the direction of the estimate. Third: examine the confidence interval for precision and whether it includes the null value. Fourth: read the P-value as evidence from the specified hypothesis test rather than as a measure of effect magnitude.
18. Comparing the Reported Effect Measures
The numerical estimates in STELLAR live on different scales. Treating them as if they were directly comparable would be a statistical error.
| Measure | Meaning | Reported STELLAR example |
|---|---|---|
| Treatment difference | Difference in the location of the outcome distribution under the reported rank-based analysis | 6MWD: 40.8 meters |
| Treatment difference | Location-shift estimate for a physiologic or biomarker outcome | PVR: −234.6 dynes*sec/cm5 |
| Treatment difference | Location-shift estimate for a biomarker outcome | NT-proBNP: −441.6 pg/mL |
| Hazard ratio | Relative instantaneous event rate under the Cox model | Death/clinical worsening: 0.163 |
| CMH P-value | Evidence from a stratified categorical comparison | Multicomponent improvement: <0.001 |
A difference of 40.8 meters cannot be compared with a hazard ratio of 0.163 by asking which number is “larger” or “smaller.” They represent fundamentally different quantities. Statistical interpretation begins with understanding the estimand before considering the numerical result.
19. Safety Results
The registered primary safety endpoints were the number of participants who experienced an adverse event and the number who discontinued study treatment due to an adverse event, both assessed up to approximately 24 weeks. The ClinicalTrials.gov record does not provide formal comparative estimates, confidence intervals, or P-values for these two primary safety endpoints.
| Primary safety endpoint | Time frame | Formal statistical comparison reported |
|---|---|---|
| Number of Participants Who Experienced an Adverse Event (AE) | Up to approximately 24 weeks | Not reported in the listed primary statistical analysis |
| Number of Participants Who Discontinued Study Treatment Due to an AE | Up to approximately 24 weeks | Not reported in the listed primary statistical analysis |
Serious adverse events by arm are reported in the ClinicalTrials.gov record. The registry data provide affected participants over the corresponding risk sets in two labeled periods.
| Period label in the ClinicalTrials.gov record | Treatment arm | Serious adverse events, affected / at risk |
|---|---|---|
| DBP | Sotatercept Plus Background PAH Therapy | 23/163 |
| DBP | Placebo Plus Background PAH Therapy | 36/160 |
| LTD | Sotatercept Plus Background PAH Therapy | 26/158 |
| LTD | Placebo Plus Background PAH Therapy | 14/142 |
These serious-adverse-event figures are descriptive affected/at-risk counts as reported in the registry by the registry. They should not be transformed here into a new comparative risk estimate because the ClinicalTrials.gov record does not specify a formal statistical analysis for these safety counts, nor do they establish that the two labeled periods are directly interchangeable exposure windows.
20. Design Topics That Are Not Supported by the Supplied Record
Several statistical topics commonly appear in clinical-trial analyses but are not documented sufficiently in the registry-reported STELLAR data to support a detailed trial-specific reconstruction.
| Topic | What can be stated from the ClinicalTrials.gov record |
|---|---|
| Non-inferiority margin | No non-inferiority margin is reported in the ClinicalTrials.gov record. |
| Crossover | No crossover rule or crossover analysis is reported in the ClinicalTrials.gov record. |
| Factorial design | The design is described as parallel; no factorial structure is reported. |
| Bayesian methods | No Bayesian analysis is identified in the registry-reported methods. |
| Interim analysis | No interim-analysis boundary, alpha-spending method, or interim decision rule is reported. |
| Multiplicity | The ClinicalTrials.gov record identifies multiple endpoints and P-values but do not provide a detailed multiplicity-control procedure. |
| Imputation model | Multiple imputation / missing data is identified as an analysis concept, but the ClinicalTrials.gov record does not specify an imputation model or number of imputations. |
This distinction is important for a statistical teaching page. The absence of a documented method in the ClinicalTrials.gov record should not be replaced by a generic assumption about what a phase 3 trial “must have” done. Trial-specific methodology should be reported only when the available record supports it.
21. Primary Endpoint Interpretation in Context
The primary 6MWD result is statistically notable because it combines four layers of information: a prespecified Week 24 functional endpoint, a rank-based analysis, an explicit treatment-difference estimate, and a missing-data rule that assigns poor ranks to clinically serious intercurrent events.
Endpoint scale
The effect is expressed in meters, making the treatment contrast directly interpretable on the 6MWD scale.
Nonparametric method
The ARSW analysis uses ranks rather than relying solely on a conventional parametric comparison of means.
Precision
The 95% CI of 27.53–54.14 provides information about uncertainty around the 40.8 estimate.
Clinical events
Death and non-fatal clinical worsening are explicitly incorporated into the ranking rules for the analysis.
The result should therefore be read as a statistical treatment contrast under a particular estimand and analysis rule, rather than simply as “the average improvement in walking distance.” That distinction is especially important when outcomes can be affected by death, clinical worsening, or missing follow-up measurements.
22. Interpreting the Clinical Worsening Hazard Ratio
The time-to-event result adds a different dimension to the evidence. The 6MWD analysis asks about functional capacity at Week 24, while the clinical-worsening analysis asks about when a participant first experienced death or a clinical worsening event.
A hazard ratio below 1 indicates a lower estimated instantaneous event hazard in the sotatercept group relative to the placebo group under the fitted model. The numerical complement, 1 − HR, should not be confused with an absolute event-rate reduction.
The reported confidence interval, 0.076 to 0.347, is important because it shows that the point estimate is not the only plausible value supported by the statistical analysis. The interval is relatively broad on the multiplicative scale, even though it remains below 1.
Another important distinction is between hazard and
23. What the Secondary Results Add
The secondary analyses span several domains: pulmonary vascular physiology, biomarkers, time-to-event outcomes, patient-reported symptoms and impacts, functional class, multicomponent improvement, and risk status. This breadth is statistically useful because different endpoints capture different dimensions of the disease process.
| Domain | Endpoint examples | Statistical method |
|---|---|---|
| Functional capacity | 6MWD | ARSW / Hodges-Lehmann |
| Pulmonary vascular physiology | PVR | ARSW / Hodges-Lehmann |
| Biomarker | NT-proBNP | ARSW / Hodges-Lehmann |
| Clinical events | Death or first clinical worsening | Log-rank / Cox HR |
| Patient-reported outcomes | PAH-SYMPACT domains | ARSW / Hodges-Lehmann |
| Categorical improvement | Multicomponent improvement, WHO FC improvement, low risk score | CMH |
The statistical lesson is that consistency across different outcome domains can be informative, but it does not turn every secondary endpoint into an independent confirmatory conclusion. Each endpoint has its own measurement properties, analysis population, estimand, uncertainty, and potential multiplicity considerations.
24. Limitations and Interpretation Issues
- Registry-level detail: This analysis is constrained to the ClinicalTrials.gov record. It does not reconstruct statistical procedures that are not documented in that record.
- Multiple endpoints: The trial has three registered primary endpoints and multiple secondary outcomes. The ClinicalTrials.gov record does not specify a complete multiplicity-control framework, so nominal P-values should not automatically be treated as independent confirmatory evidence.
- Missing-data assumptions: The ARSW analyses explicitly assign poor ranks to deaths and non-fatal clinical-worsening-related missing observations, but the ClinicalTrials.gov record does not provide a complete sensitivity-analysis framework.
- Analysis populations: The listed analyses generally require at least one treatment dose and a relevant baseline measurement. This differs operationally from simply analyzing every randomized participant regardless of available baseline data.
- Hazard-ratio assumptions: The clinical-worsening HR comes from a Cox proportional-hazards model. A single HR is most straightforward to interpret when the proportional-hazards structure is appropriate; the ClinicalTrials.gov record does not provide a formal diagnostic for that assumption.
- Composite endpoint: Death or first clinical worsening combines more than one type of event. The HR therefore describes the composite endpoint rather than establishing identical effects on every component.
- Safety interpretation: The registry-reported serious-adverse-event counts are descriptive affected/at-risk counts. They do not constitute a formal comparative safety analysis by themselves.
- Generalizability: The ClinicalTrials.gov record establishes the trial population and condition but does not provide a detailed baseline-characteristic table from which broader population-representativeness conclusions could be drawn.
- Unreported methodology: Non-inferiority margins, Bayesian methods, crossover, interim boundaries, and detailed imputation procedures are not documented in the ClinicalTrials.gov record and therefore are not inferred here.
25. Why This Trial Matters Statistically
STELLAR is a useful teaching case because it demonstrates how a single randomized clinical trial can require several distinct statistical frameworks depending on the endpoint. Its the ClinicalTrials.gov record includes rank-based continuous-outcome analyses, a stratified survival analysis, a Cox model, and stratified categorical testing.
| Concept | How it appears in STELLAR |
|---|---|
| Randomization | Randomized parallel-group phase 3 design with 324 enrolled participants. |
| Masking | Quadruple-masked design. |
| Nonparametric analysis | ARSW analyses for 6MWD and several continuous secondary outcomes. |
| Hodges-Lehmann estimate | Used to report treatment differences for rank-based continuous outcomes. |
| Missing-data handling | Deaths and non-fatal clinical-worsening-related missing values assigned poor ranks in the ARSW analyses. |
| Confidence intervals | Reported for the primary 6MWD estimate and several secondary continuous and time-to-event estimates. |
| Log-rank test | Used for time to death or first clinical worsening. |
| Hazard ratio | Cox model estimate of the relative hazard for death or clinical worsening. |
| Stratified survival analysis | WHO FC II/III and background PAH therapy used as strata. |
| Cochran-Mantel-Haenszel test | Used for reported binary secondary outcomes with the same strata. |
| Multiplicity | Multiple primary and secondary endpoints require careful interpretation of nominal P-values. |
26. Trial Timeline
Study start
The registry lists January 25, 2021 as the STELLAR study start date.
Primary completion
The registry lists August 26, 2022 as the primary completion date.
Registry status
The ClinicalTrials.gov record identifies the study status as completed.
the ClinicalTrials.gov record therefore provides a defined enrollment and completion history, while the statistical analyses cover both the Week 24 primary/secondary assessments and a clinical-worsening endpoint with follow-up extending to approximately 18 months.
27. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The primary 6MWD analysis produced a Hodges-Lehmann treatment difference of 40.8 meters with a two-sided 95% CI of 27.53–54.14 and P <0.001. The reported clinical-worsening analysis produced HR 0.163 with a 95% CI of 0.076–0.347 and P <0.001.
Clinical interpretation
The statistical results describe differences between randomized treatment groups across functional capacity and time to death or clinical worsening. Determining the clinical importance of any particular magnitude requires consideration of the endpoint's clinical context rather than relying on statistical significance alone.
The distinction is particularly important for the PAH-SYMPACT outcomes. The physical impacts and cardiopulmonary symptoms domains have reported P-values of 0.010 and 0.028, while the cognitive/emotional impacts domain has a P-value of 0.156. These results should be interpreted as separate endpoint-specific estimates, not collapsed into a single overall conclusion about patient-reported outcomes.
28. A Statistical Reading of the Complete Evidence Set
Viewed as a statistical system, the registry-reported STELLAR results have three complementary layers.
Functional outcome
6MWD provides a continuous measure at Week 24. The ARSW analysis reports the treatment contrast in meters through a Hodges-Lehmann estimate.
Physiologic and biomarker outcomes
PVR and NT-proBNP provide additional continuous endpoints, analyzed with the same general rank-based framework.
Clinical events
Time to death or clinical worsening captures event timing and is summarized through the log-rank test and Cox hazard ratio.
Categorical outcomes
CMH analyses evaluate reported binary secondary outcomes while accounting for WHO FC II/III and background PAH therapy strata.
The resulting statistical narrative is therefore broader than any one P-value. The primary endpoint supplies a treatment difference on a functional scale; secondary continuous endpoints provide additional treatment contrasts; the time-to-event endpoint provides a relative hazard measure; and the categorical analyses provide stratified evidence for binary outcomes.
29. Related Tutorials
Learn more about the methods used in this trial:
30. Related Statistical Calculators
31. Sources
- ClinicalTrials.gov: NCT04576988 — STELLAR.
- Linked publication: PubMed PMID 36877098.
- Linked publication: PubMed PMID 42246435.
- Linked publication: PubMed PMID 40526255.
- Linked publication: PubMed PMID 37851297.
- Linked publication: PubMed PMID 37696565.
The trial-specific numerical results and methodological descriptions on this page are restricted to the ClinicalTrials.gov record. The PubMed links are provided as navigation to the associated publications and are not used here to introduce additional numerical results.
Continue through the Clinical Biostats statistical learning pathway
Use the trial's endpoints as a starting point for deeper study of rank-based methods, stratified categorical analysis, survival analysis, hazard ratios, confidence intervals, and missing-data methods.
32. Record Summary
STELLAR provides a useful example of how a randomized phase 3 trial can use different statistical methods for different endpoint structures. The primary 6MWD endpoint was analyzed using an Aligned Rank Stratified Wilcoxon procedure with a Hodges-Lehmann treatment difference of 40.8 meters and a two-sided 95% CI of 27.53–54.14. Secondary continuous outcomes used the same general rank-based framework, while the time-to-death-or-clinical-worsening endpoint used a stratified log-rank test and Cox proportional-hazards model, producing an HR of 0.163 with a 95% CI of 0.076–0.347. Binary secondary outcomes used the Cochran-Mantel-Haenszel method with WHO FC II/III and background PAH therapy as strata.
The most important statistical lesson is that the numbers cannot be interpreted independently of their methods. A treatment difference in meters, a location shift in a biomarker, a hazard ratio, and a CMH P-value answer different questions. Confidence intervals describe precision around their respective estimates, while P-values provide hypothesis-testing evidence rather than measures of effect magnitude.