This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information reported in the ClinicalTrials.gov record. The registry record is the official source for the trial design and posted analyses.
1. Trial at a Glance
STEP-HFpEF was a randomized, double-blind, parallel phase 3 trial evaluating semaglutide versus placebo in people living with heart failure and obesity. The registry reports 529 participants, two treatment arms, two registered primary endpoints, and formal ANCOVA analyses for both primary endpoints.
| Feature | STEP-HFpEF |
|---|---|
| Trial name | STEP-HFpEF |
| ClinicalTrials.gov identifier | NCT04788511 |
| Phase | Phase 3 |
| Status | Completed |
| Start date | 2021-03-16 |
| Primary completion | 2023-04-18 |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
| Enrollment | 529 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Interventions | Semaglutide; placebo |
| Registered primary endpoints | 2 |
| Outcome measures posted | 15 |
| Statistical analyses posted | 2 |
2. Clinical Question
The central question was whether semaglutide differed from placebo with respect to change in Kansas City Cardiomyopathy Questionnaire Clinical Summary Score (KCCQ-CSS) and change in body weight from baseline to the end of treatment at week 52.
Population
People living with heart failure and obesity, as described by the trial's brief title and registry condition.
Intervention
Semaglutide. The registry's serious-adverse-event summary identifies the semaglutide arm as semaglutide 2.4 mg.
Comparator
Placebo.
Primary question
How does semaglutide compare with placebo for the registered KCCQ-CSS and body-weight changes through week 52?
3. Trial Design
Semaglutide 2.4 mg
- Semaglutide
- Registry safety summary identifies the arm as semaglutide 2.4 mg
- Compared with placebo in the randomized parallel design
Placebo
- Placebo
- Randomized comparator arm
- Used as the reference group for the primary treatment comparisons
4. Randomization, Stratification, and Analysis Population
The registry reports a randomized allocation and specifies that the primary ANCOVA models included randomized treatment arm. and BMI stratification as factors. The two BMI strata were BMI <35.0 kg/m2 and BMI ≥35.0 kg/m2.
| Feature | Registry-supported description |
|---|---|
| Allocation | Randomized |
| Stratification | BMI <35.0 kg/m2 versus BMI ≥35.0 kg/m2 |
| Primary analysis population | Full Analysis Set (FAS) |
| FAS definition | All randomized participants |
| Week-52 analysis count | The registry states that the overall number analyzed signifies participants who had an observed value at week 52 |
| Primary model | ANCOVA |
This distinction between the FAS definition and the participants with an observed week-52 value is important. The registry describes the FAS as all randomized participants, while the reported analysis count is tied to those with an observed value at week 52. The analysis notes also state that missing week-52 observations were multiply imputed.
5. Primary Endpoints
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Change in Kansas City Cardiomyopathy Questionnaire Clinical Summary Score (KCCQ-CSS) | From baseline (week 0) to end of treatment (week 52) | ANCOVA; estimated treatment difference |
| Change in Body Weight | From baseline (week 0) to end of treatment (week 52) | ANCOVA; estimated treatment difference |
The registry describes KCCQ-CSS as a score derived from the Kansas City Cardiomyopathy Questionnaire, a standardized 23-item, self-administered instrument addressing heart failure symptoms, physical limitation, quality of life, and social limitation. The registry definition specifically identifies KCCQ-CSS as including the symptom and physical-limitation components.
6. Statistical Methodology
Analysis of covariance
Both registered primary endpoints were analyzed using analysis of covariance (ANCOVA). The models included randomized treatment and BMI stratification as factors and the relevant baseline measure as a covariate.
The registry describes the response at week 52 as being analyzed with randomized treatment and BMI stratification as factors and baseline KCCQ-CSS as a covariate.
For body weight, the registry describes baseline body weight in kilograms as the covariate.
Why baseline adjustment matters
ANCOVA uses the baseline value to account for differences in the outcome that were already present before treatment. In a randomized trial, baseline adjustment is not needed to make randomization valid, but it can improve precision by explaining some of the variation in the week-52 outcome.
Why BMI stratification matters
The analysis model retained the BMI strata used in the randomized analysis. Including the stratification factor helps ensure that the treatment comparison reflects the structure used to allocate participants and avoids treating the randomized strata as irrelevant to the primary model.
Estimated treatment difference
The primary effect measure was an estimated treatment difference. For KCCQ-CSS, the estimate is expressed in score units. For body weight, the registry reports the outcome unit as percentage of body weight, so the estimated treatment difference is expressed on that percentage scale.
Missing week-52 observations
The registry analysis notes state that missing observations at week 52 were multiple (x1000) imputed from retrieved participants of the same randomized treatment arm.. This is materially different from simply deleting every participant without an observed week-52 value. The imputation procedure attempts to preserve information from participants who were retrieved while incorporating uncertainty about missing observations.
7. Primary Results: KCCQ-CSS
The first primary endpoint was change in KCCQ-CSS from baseline at week 0 to the end of treatment at week 52. The registry reports an ANCOVA-based estimated treatment difference comparing semaglutide 2.4 mg with placebo.
Estimated treatment difference in KCCQ-CSS
95% CI: 4.8–10.9 · P < 0.0001
Superiority hypothesis; two-sided 95% confidence interval.
| Element | Reported result |
|---|---|
| Endpoint | Change in Kansas City Cardiomyopathy Questionnaire Clinical Summary Score (KCCQ-CSS) |
| Time frame | From baseline (week 0) to end of treatment (week 52) |
| Comparison | Semaglutide 2.4 mg vs placebo |
| Method | ANCOVA |
| Effect measure | Estimated Treatment Difference |
| Estimate | 7.8 |
| 95% CI | 4.8–10.9 |
| P-value | <0.0001 |
| Hypothesis type | Superiority |
The estimated treatment difference of 7.8 means that the model-estimated week-52 change in KCCQ-CSS was 7.8 score units higher for the semaglutide group than for the placebo group after accounting for the prespecified baseline KCCQ-CSS covariate and BMI stratification.
The estimate is not a statement that every participant improved by 7.8 points. It is a between-group adjusted estimate. It also does not describe the distribution of individual patient responses.
The 95% confidence interval of 4.8–10.9 describes the statistical uncertainty around the estimated treatment difference under the specified analysis framework. A narrower interval would indicate greater precision; the interval itself is not a range containing the treatment effect for 95% of individual patients.
The P < 0.0001 result addresses evidence against the null hypothesis under the prespecified statistical comparison. A p-value does not measure the size, importance, or clinical usefulness of an effect. The magnitude of 7.8 and its confidence interval provide the effect-size information.
The analysis also depends on the FAS framework, the week-52 outcome, the BMI-stratified ANCOVA model, and the multiple-imputation approach for missing observations. The ClinicalTrials.gov record does not report a separate multiplicity procedure beyond identifying both endpoints as primary and the hypothesis as superiority.
8. Primary Results: Body Weight
The second primary endpoint was change in body weight from baseline at week 0 to the end of treatment at week 52. The registry reports an ANCOVA estimated treatment difference for semaglutide 2.4 mg versus placebo.
Estimated treatment difference in body weight
95% CI: −11.9–−9.4 · P < 0.0001
Superiority hypothesis; two-sided 95% confidence interval.
| Element | Reported result |
|---|---|
| Endpoint | Change in Body Weight |
| Time frame | From baseline (week 0) to end of treatment (week 52) |
| Comparison | Semaglutide 2.4 mg vs placebo |
| Method | ANCOVA |
| Effect measure | Estimated Treatment Difference |
| Estimate | −10.7 |
| 95% CI | −11.9–−9.4 |
| P-value | <0.0001 |
| Hypothesis type | Superiority |
The estimated treatment difference of −10.7 indicates that the adjusted change in body weight was estimated to be 10.7 percentage points lower in the semaglutide group than in the placebo group on the registry's reported percentage-of-body-weight scale.
The negative sign identifies the direction of the estimated between-group difference. It does not mean that every participant lost exactly 10.7% of body weight, nor does it provide the individual-level distribution of weight changes.
The 95% confidence interval of −11.9–−9.4 describes uncertainty around the adjusted treatment difference. Because the entire interval is below zero, the reported estimate is consistently in the semaglutide-favoring direction within this confidence interval.
The P < 0.0001 result provides evidence against the null hypothesis of no treatment difference under the reported superiority analysis. It is not a measure of how large the treatment effect is; the estimate and confidence interval describe magnitude and precision.
The same interpretation cautions apply here as for KCCQ-CSS: the estimate depends on the FAS analysis, BMI stratification, baseline adjustment, week-52 assessment, and the stated multiple-imputation procedure for missing observations.
9. Reading the Two Primary Results Together
| Primary endpoint | Estimated treatment difference | 95% CI | P-value | Analysis |
|---|---|---|---|---|
| Change in KCCQ-CSS | 7.8 | 4.8–10.9 | <0.0001 | ANCOVA |
| Change in body weight | −10.7 | −11.9–−9.4 | <0.0001 | ANCOVA |
These endpoints measure different constructs. KCCQ-CSS is a patient-reported score covering heart failure symptoms and physical limitations within the registry's definition, whereas body weight is a physical measurement expressed in the posted analysis as a percentage-of-body-weight outcome.
Consequently, the two estimates should not be combined into a single effect measure. A treatment difference of 7.8 KCCQ-CSS units and a treatment difference of −10.7 percentage points are on different scales and answer different questions.
10. Statistical Methods Explained
Why was ANCOVA used?
ANCOVA is appropriate for a continuous outcome measured at a specified follow-up time when the analysis can benefit from adjustment for the corresponding baseline measurement. Here, the week-52 response was modeled while adjusting for baseline KCCQ-CSS for the KCCQ endpoint and baseline body weight for the body-weight endpoint.
Why include baseline KCCQ-CSS or baseline body weight?
Participants enter a randomized trial with different baseline measurements. Including the baseline value as a covariate can explain part of the variability in the week-52 outcome. The treatment comparison then reflects an adjusted difference rather than relying only on the raw difference in observed week-52 values.
Why was BMI stratification included in the model?
The registry states that the analysis included BMI <35.0 kg/m2 and BMI ≥35.0 kg/m2 as stratification factors. Including the randomized stratification factor in the analysis is consistent with the design and accounts for systematic structure introduced by the stratified randomization.
What does the KCCQ-CSS estimate of 7.8 mean?
It is the estimated adjusted difference between semaglutide 2.4 mg and placebo for the change in KCCQ-CSS from week 0 to week 52. A positive value indicates a higher estimated change in the semaglutide group on this score scale. It does not mean that every patient experienced a 7.8-point change.
What does the body-weight estimate of −10.7 mean?
It is the estimated adjusted treatment difference for the percentage-of-body-weight outcome reported by the registry. The negative sign means the estimated change was lower in the semaglutide group than in the placebo group by 10.7 percentage points on that analysis scale.
What does the 95% confidence interval tell us?
A 95% confidence interval describes uncertainty around the estimated treatment difference under the statistical model and sampling framework. For KCCQ-CSS, the interval is 4.8–10.9; for body weight, it is −11.9–−9.4. Neither interval describes the range of individual participant responses.
Why doesn't the p-value measure effect size?
A p-value summarizes evidence against a null hypothesis under a specified statistical model. It depends on both the estimated effect and the amount of information available. Therefore, two studies can have similarly small p-values despite different effect sizes, and a p-value should be interpreted alongside the estimate and confidence interval.
11. Missing Data and Imputation
Missing week-52 observations are explicitly addressed in both primary analysis descriptions. The registry states that missing observations at week 52 were multiple (x1000) imputed from retrieved participants of the same randomized treatment arm..
Why impute?
Simply analyzing only participants with observed week-52 values can change the analysis population and potentially introduce bias if missingness is related to participant characteristics or outcomes.
Why multiple imputation?
Multiple imputation is designed to reflect uncertainty about missing values by creating multiple plausible completed datasets rather than treating a single filled-in value as known with certainty.
Same-treatment retrieval
The registry-reported analysis notes state that imputation used retrieved participants from the same randomized treatment.
What is not available here?
The ClinicalTrials.gov record does not provide enough detail to reconstruct the complete imputation model, auxiliary variables, or sensitivity analyses.
The presence of imputation is therefore part of the interpretation of both primary estimates. The reported treatment differences are not simply unadjusted comparisons among participants who happened to have an observed week-52 value.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The reported counts are presented as affected participants over participants at risk.
| Safety measure | Semaglutide 2.4 mg | Placebo |
|---|---|---|
| Serious adverse events | 35/263 | 71/266 |
These figures provide the number of participants affected and the corresponding number at risk in each arm. They are a safety count summary rather than a formal statistical comparison reported in the ClinicalTrials.gov record.
13. What the Estimated Treatment Difference Does — and Does Not — Mean
An estimated treatment difference of 7.8 means the ANCOVA model estimated a 7.8-unit difference between semaglutide 2.4 mg and placebo in the change in KCCQ-CSS from baseline to week 52, after accounting for baseline KCCQ-CSS and BMI stratification.
It does not mean that every participant experienced a 7.8-unit improvement, that every participant benefited, or that the effect was identical across individuals.
An estimated treatment difference of −10.7 means the adjusted difference in the reported percentage-of-body-weight outcome favored the semaglutide group by 10.7 percentage points on the analysis scale.
It does not mean that every participant lost exactly 10.7% of body weight or that the same effect occurred in every subgroup or individual.
The 95% confidence interval of 4.8–10.9 for KCCQ-CSS and −11.9–−9.4 for body weight quantify statistical uncertainty around their respective model-based estimates. They are not prediction intervals for individual participants.
Both reported primary comparisons have P < 0.0001. This indicates strong evidence against the corresponding null hypothesis under the reported analysis, but it does not by itself establish the magnitude, practical importance, or individual-level consistency of the treatment effect.
14. Primary Endpoint Analysis in Statistical Detail
| Model component | KCCQ-CSS | Body weight |
|---|---|---|
| Response | Week-52 KCCQ-CSS response | Week-52 body-weight response |
| Treatment factor | Randomized treatment arm. | Randomized treatment arm. |
| Stratification factor | BMI <35.0 vs BMI ≥35.0 kg/m2 | BMI <35.0 vs BMI ≥35.0 kg/m2 |
| Baseline covariate | Baseline KCCQ-CSS | Baseline body weight (kg) |
| Analysis population | FAS | FAS |
| Missing week-52 observations | Multiple (x1000) imputation | Multiple (x1000) imputation |
| Effect measure | Estimated Treatment Difference | Estimated Treatment Difference |
| Hypothesis | Superiority | Superiority |
This structure illustrates an important principle in clinical-trial statistics: the reported treatment effect is the output of a prespecified model, not merely the arithmetic difference between two unadjusted group means.
15. Endpoint Type and Model Interpretation
The ClinicalTrials.gov record identifies KCCQ-CSS as a continuous endpoint and body weight as a binary endpoint in the statistical-analysis metadata, while the body-weight outcome unit is reported as percentage of body weight and the method is ANCOVA. That combination should be read carefully rather than silently replacing the registry's classification.
16. Multiplicity and the Two Primary Endpoints
The registry identifies two primary endpoints and reports a superiority hypothesis for each formal primary analysis. The ClinicalTrials.gov record does not specify an alpha-allocation scheme, hierarchical testing sequence, gatekeeping procedure, or other multiplicity adjustment connecting the two endpoints.
| Endpoint | Role | Formal analysis posted | Multiplicity procedure reported? |
|---|---|---|---|
| Change in KCCQ-CSS | Primary | Yes | Not specified in the ClinicalTrials.gov record |
| Change in Body Weight | Primary | Yes | Not specified in the ClinicalTrials.gov record |
This matters because multiple primary endpoints can create a family of hypotheses. The fact that both reported p-values are below 0.0001 is an important descriptive result, but the ClinicalTrials.gov record does not establish how type I error across the two primary endpoints was controlled.
17. Blinding and Its Statistical Role
The trial is registered as double masked. Blinding is not itself a statistical test, but it is an important design feature because knowledge of treatment assignment can influence behavior, reporting, assessment, and other aspects of trial conduct.
Patient-reported endpoint
KCCQ-CSS is a patient-reported measure, making the protection provided by masking particularly relevant to the interpretation of observed responses.
Objective measurement
Body weight is measured differently from a patient-reported score, but blinded allocation remains part of the overall randomized trial design.
Blinding does not eliminate all sources of bias, and it does not change the mathematical interpretation of an ANCOVA estimate. Its principal role is to strengthen the design from which the statistical comparison is generated.
18. Trial Timeline
Trial start
The registry lists 2021-03-16 as the study start date.
Randomized parallel design
The trial used randomized allocation, a parallel design, double masking, and two treatment arms.
Primary endpoint assessment window
Both registered primary endpoints were defined as changes from baseline at week 0 to the end of treatment at week 52.
Primary completion
The registry lists 2023-04-18 as the primary completion date.
19. Limitations
- Registry-level detail: the ClinicalTrials.gov record contains the posted endpoint definitions and statistical analyses, but not the complete statistical analysis plan.
- Multiplicity: two primary endpoints are reported, but the ClinicalTrials.gov record does not specify how type I error was allocated across them.
- Missing-data assumptions: multiple imputation was used for missing week-52 observations, but the ClinicalTrials.gov record does not provide the full imputation model or sensitivity-analysis framework.
- Analysis population: the FAS is defined as all randomized participants, while the posted overall number analyzed signifies participants with an observed week-52 value. The distinction should be retained when interpreting analysis counts.
- Endpoint metadata: the registry metadata label body weight as a binary endpoint while reporting an ANCOVA analysis and a percentage-of-body-weight outcome unit. The posted analysis method is therefore more informative for interpreting the reported estimate than the normalized endpoint label alone.
- No subgroup estimates reported: the ClinicalTrials.gov record does not contain formal subgroup treatment estimates, so subgroup consistency cannot be evaluated from this record.
- No formal safety comparison reported: serious adverse-event counts by arm are available, but no comparative statistical analysis is posted on ClinicalTrials.gov for those safety data.
- No non-inferiority framework: the registry identifies superiority as the hypothesis type; there is no non-inferiority margin in the ClinicalTrials.gov record.
- No crossover analysis reported: the registry-reported design data do not identify a crossover component or crossover analysis.
- No Bayesian analysis reported: the posted primary analyses are ANCOVA models; the ClinicalTrials.gov record does not identify a Bayesian component.
20. Why This Trial Matters Statistically
STEP-HFpEF is a useful teaching case because its two primary endpoints demonstrate how a randomized clinical trial can analyze different outcome domains using the same general modeling framework while tailoring the covariate adjustment to the endpoint.
| Concept | How it appears in STEP-HFpEF |
|---|---|
| Randomization | Participants were randomized to semaglutide or placebo in a parallel phase 3 design. |
| Blinding | The registry identifies the trial as double masked. |
| Stratified analysis | BMI <35.0 kg/m2 versus BMI ≥35.0 kg/m2 was included as a factor. |
| ANCOVA | Used for both posted primary endpoint analyses. |
| Covariate adjustment | Baseline KCCQ-CSS and baseline body weight were included for their respective endpoints. |
| Confidence intervals | Both primary treatment differences include two-sided 95% confidence intervals. |
| P-values | Both primary analyses report P < 0.0001. |
| Missing-data handling | Missing week-52 observations were multiply imputed from retrieved participants of the same randomized treatment. |
| Superiority testing | Both primary analyses are identified as superiority hypotheses. |
| Safety summaries | Serious adverse events are reported by randomized arm. |
21. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: NCT04788511 — STEP-HFpEF.
- PubMed: PMID 37622681.
- PubMed: PMID 39222642.
- PubMed: PMID 38739118.
- PubMed: PMID 38599221.
- PubMed: PMID 37993201.
Continue with the statistical methods behind this trial
Explore the Clinical Biostats tutorials and statistical calculators related to ANCOVA, confidence intervals, covariate adjustment, randomization, blinding, p-values, and stratified analysis.
24. Record Summary
STEP-HFpEF provides a clear example of an adjusted randomized-trial analysis in which two primary endpoints were evaluated at the same week-52 time point using ANCOVA. The KCCQ-CSS analysis produced an estimated treatment difference of 7.8 with a two-sided 95% CI of 4.8–10.9 and P < 0.0001. The body-weight analysis produced an estimated treatment difference of −10.7 with a two-sided 95% CI of −11.9–−9.4 and P < 0.0001.
The statistical story is more than the two p-values. The estimates arise from ANCOVA models incorporating randomized treatment, BMI stratification, and the relevant baseline covariate, with multiple imputation used for missing week-52 observations. Interpreting the results therefore requires attention to the effect estimate, confidence interval, p-value, analysis population, covariate adjustment, stratification, and missing-data approach.