← Clinical Trials
Heart Failure & Obesity Phase 3 Completed NCT04788511

STEP-HFpEF: Complete Statistical Analysis of Semaglutide in Heart Failure and Obesity

An independent statistical review of the randomized, double-blind phase 3 STEP-HFpEF trial evaluating semaglutide versus placebo in people living with heart failure and obesity, with emphasis on its two registered primary endpoints and the ANCOVA framework used for their analysis.

Phase 3  ·  Randomized  ·  Double-blind  ·  529 participants
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information reported in the ClinicalTrials.gov record. The registry record is the official source for the trial design and posted analyses.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

STEP-HFpEF was a randomized, double-blind, parallel phase 3 trial evaluating semaglutide versus placebo in people living with heart failure and obesity. The registry reports 529 participants, two treatment arms, two registered primary endpoints, and formal ANCOVA analyses for both primary endpoints.

529
Enrolled
2 treatment arms
3
Phase
Phase 3
7.8
KCCQ-CSS Difference
95% CI 4.8–10.9
−10.7
Body Weight Difference
95% CI −11.9–−9.4
FeatureSTEP-HFpEF
Trial nameSTEP-HFpEF
ClinicalTrials.gov identifierNCT04788511
PhasePhase 3
StatusCompleted
Start date2021-03-16
Primary completion2023-04-18
Lead sponsorNovo Nordisk A/S
Sponsor typeIndustry
Enrollment529
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
InterventionsSemaglutide; placebo
Registered primary endpoints2
Outcome measures posted15
Statistical analyses posted2

2. Clinical Question

The central question was whether semaglutide differed from placebo with respect to change in Kansas City Cardiomyopathy Questionnaire Clinical Summary Score (KCCQ-CSS) and change in body weight from baseline to the end of treatment at week 52.

Population

People living with heart failure and obesity, as described by the trial's brief title and registry condition.

Intervention

Semaglutide. The registry's serious-adverse-event summary identifies the semaglutide arm as semaglutide 2.4 mg.

Comparator

Placebo.

Primary question

How does semaglutide compare with placebo for the registered KCCQ-CSS and body-weight changes through week 52?

3. Trial Design

01
Randomize529 participants
02
2 ArmsSemaglutide vs placebo
03
Double-blindMasked treatment assignment
04
Week 52End of treatment
05
AnalyzeANCOVA in FAS
ARM A · SEMAGLUTIDE

Semaglutide 2.4 mg

  • Semaglutide
  • Registry safety summary identifies the arm as semaglutide 2.4 mg
  • Compared with placebo in the randomized parallel design
ARM B · PLACEBO

Placebo

  • Placebo
  • Randomized comparator arm
  • Used as the reference group for the primary treatment comparisons
What the design contributes statistically. Randomization creates the framework for comparing treatment groups, while double masking reduces the opportunity for knowledge of treatment assignment to influence participant or study-team behavior. The parallel design means the randomized treatment groups are compared across the same trial period rather than through a crossover sequence.

4. Randomization, Stratification, and Analysis Population

The registry reports a randomized allocation and specifies that the primary ANCOVA models included randomized treatment arm. and BMI stratification as factors. The two BMI strata were BMI <35.0 kg/m2 and BMI ≥35.0 kg/m2.

FeatureRegistry-supported description
AllocationRandomized
StratificationBMI <35.0 kg/m2 versus BMI ≥35.0 kg/m2
Primary analysis populationFull Analysis Set (FAS)
FAS definitionAll randomized participants
Week-52 analysis countThe registry states that the overall number analyzed signifies participants who had an observed value at week 52
Primary modelANCOVA

This distinction between the FAS definition and the participants with an observed week-52 value is important. The registry describes the FAS as all randomized participants, while the reported analysis count is tied to those with an observed value at week 52. The analysis notes also state that missing week-52 observations were multiply imputed.

5. Primary Endpoints

EndpointRegistry definition / time frameAnalysis
Change in Kansas City Cardiomyopathy Questionnaire Clinical Summary Score (KCCQ-CSS) From baseline (week 0) to end of treatment (week 52) ANCOVA; estimated treatment difference
Change in Body Weight From baseline (week 0) to end of treatment (week 52) ANCOVA; estimated treatment difference

The registry describes KCCQ-CSS as a score derived from the Kansas City Cardiomyopathy Questionnaire, a standardized 23-item, self-administered instrument addressing heart failure symptoms, physical limitation, quality of life, and social limitation. The registry definition specifically identifies KCCQ-CSS as including the symptom and physical-limitation components.

6. Statistical Methodology

Analysis of covariance

Both registered primary endpoints were analyzed using analysis of covariance (ANCOVA). The models included randomized treatment and BMI stratification as factors and the relevant baseline measure as a covariate.

KCCQ-CSS model
Week-52 KCCQ-CSS response ~ randomized treatment + BMI stratum + baseline KCCQ-CSS

The registry describes the response at week 52 as being analyzed with randomized treatment and BMI stratification as factors and baseline KCCQ-CSS as a covariate.

Body-weight model
Week-52 body-weight response ~ randomized treatment + BMI stratum + baseline body weight

For body weight, the registry describes baseline body weight in kilograms as the covariate.

Why baseline adjustment matters

ANCOVA uses the baseline value to account for differences in the outcome that were already present before treatment. In a randomized trial, baseline adjustment is not needed to make randomization valid, but it can improve precision by explaining some of the variation in the week-52 outcome.

Why BMI stratification matters

The analysis model retained the BMI strata used in the randomized analysis. Including the stratification factor helps ensure that the treatment comparison reflects the structure used to allocate participants and avoids treating the randomized strata as irrelevant to the primary model.

Estimated treatment difference

The primary effect measure was an estimated treatment difference. For KCCQ-CSS, the estimate is expressed in score units. For body weight, the registry reports the outcome unit as percentage of body weight, so the estimated treatment difference is expressed on that percentage scale.

Missing week-52 observations

The registry analysis notes state that missing observations at week 52 were multiple (x1000) imputed from retrieved participants of the same randomized treatment arm.. This is materially different from simply deleting every participant without an observed week-52 value. The imputation procedure attempts to preserve information from participants who were retrieved while incorporating uncertainty about missing observations.

Interpretation caution: the statistical conclusion from an imputed analysis depends on the assumptions underlying the imputation strategy. The ClinicalTrials.gov record identifies the multiple-imputation approach, but do not provide enough detail here to reconstruct every component of the imputation model or evaluate its assumptions independently.

7. Primary Results: KCCQ-CSS

The first primary endpoint was change in KCCQ-CSS from baseline at week 0 to the end of treatment at week 52. The registry reports an ANCOVA-based estimated treatment difference comparing semaglutide 2.4 mg with placebo.

Estimated treatment difference in KCCQ-CSS

7.8

95% CI: 4.8–10.9   ·   P < 0.0001

Superiority hypothesis; two-sided 95% confidence interval.

ElementReported result
EndpointChange in Kansas City Cardiomyopathy Questionnaire Clinical Summary Score (KCCQ-CSS)
Time frameFrom baseline (week 0) to end of treatment (week 52)
ComparisonSemaglutide 2.4 mg vs placebo
MethodANCOVA
Effect measureEstimated Treatment Difference
Estimate7.8
95% CI4.8–10.9
P-value<0.0001
Hypothesis typeSuperiority
Clinical Biostats interpretation

The estimated treatment difference of 7.8 means that the model-estimated week-52 change in KCCQ-CSS was 7.8 score units higher for the semaglutide group than for the placebo group after accounting for the prespecified baseline KCCQ-CSS covariate and BMI stratification.

The estimate is not a statement that every participant improved by 7.8 points. It is a between-group adjusted estimate. It also does not describe the distribution of individual patient responses.

The 95% confidence interval of 4.8–10.9 describes the statistical uncertainty around the estimated treatment difference under the specified analysis framework. A narrower interval would indicate greater precision; the interval itself is not a range containing the treatment effect for 95% of individual patients.

The P < 0.0001 result addresses evidence against the null hypothesis under the prespecified statistical comparison. A p-value does not measure the size, importance, or clinical usefulness of an effect. The magnitude of 7.8 and its confidence interval provide the effect-size information.

The analysis also depends on the FAS framework, the week-52 outcome, the BMI-stratified ANCOVA model, and the multiple-imputation approach for missing observations. The ClinicalTrials.gov record does not report a separate multiplicity procedure beyond identifying both endpoints as primary and the hypothesis as superiority.

8. Primary Results: Body Weight

The second primary endpoint was change in body weight from baseline at week 0 to the end of treatment at week 52. The registry reports an ANCOVA estimated treatment difference for semaglutide 2.4 mg versus placebo.

Estimated treatment difference in body weight

−10.7

95% CI: −11.9–−9.4   ·   P < 0.0001

Superiority hypothesis; two-sided 95% confidence interval.

ElementReported result
EndpointChange in Body Weight
Time frameFrom baseline (week 0) to end of treatment (week 52)
ComparisonSemaglutide 2.4 mg vs placebo
MethodANCOVA
Effect measureEstimated Treatment Difference
Estimate−10.7
95% CI−11.9–−9.4
P-value<0.0001
Hypothesis typeSuperiority
Clinical Biostats interpretation

The estimated treatment difference of −10.7 indicates that the adjusted change in body weight was estimated to be 10.7 percentage points lower in the semaglutide group than in the placebo group on the registry's reported percentage-of-body-weight scale.

The negative sign identifies the direction of the estimated between-group difference. It does not mean that every participant lost exactly 10.7% of body weight, nor does it provide the individual-level distribution of weight changes.

The 95% confidence interval of −11.9–−9.4 describes uncertainty around the adjusted treatment difference. Because the entire interval is below zero, the reported estimate is consistently in the semaglutide-favoring direction within this confidence interval.

The P < 0.0001 result provides evidence against the null hypothesis of no treatment difference under the reported superiority analysis. It is not a measure of how large the treatment effect is; the estimate and confidence interval describe magnitude and precision.

The same interpretation cautions apply here as for KCCQ-CSS: the estimate depends on the FAS analysis, BMI stratification, baseline adjustment, week-52 assessment, and the stated multiple-imputation procedure for missing observations.

9. Reading the Two Primary Results Together

Primary endpointEstimated treatment difference95% CIP-valueAnalysis
Change in KCCQ-CSS 7.8 4.8–10.9 <0.0001 ANCOVA
Change in body weight −10.7 −11.9–−9.4 <0.0001 ANCOVA

These endpoints measure different constructs. KCCQ-CSS is a patient-reported score covering heart failure symptoms and physical limitations within the registry's definition, whereas body weight is a physical measurement expressed in the posted analysis as a percentage-of-body-weight outcome.

Consequently, the two estimates should not be combined into a single effect measure. A treatment difference of 7.8 KCCQ-CSS units and a treatment difference of −10.7 percentage points are on different scales and answer different questions.

Important statistical distinction: both endpoints have P-values below 0.0001, but that does not mean the two treatment effects have equal evidentiary strength or that their magnitudes can be compared numerically. The effect measures, confidence intervals, and clinical constructs are different.

10. Statistical Methods Explained

Why was ANCOVA used?

ANCOVA is appropriate for a continuous outcome measured at a specified follow-up time when the analysis can benefit from adjustment for the corresponding baseline measurement. Here, the week-52 response was modeled while adjusting for baseline KCCQ-CSS for the KCCQ endpoint and baseline body weight for the body-weight endpoint.

Why include baseline KCCQ-CSS or baseline body weight?

Participants enter a randomized trial with different baseline measurements. Including the baseline value as a covariate can explain part of the variability in the week-52 outcome. The treatment comparison then reflects an adjusted difference rather than relying only on the raw difference in observed week-52 values.

Why was BMI stratification included in the model?

The registry states that the analysis included BMI <35.0 kg/m2 and BMI ≥35.0 kg/m2 as stratification factors. Including the randomized stratification factor in the analysis is consistent with the design and accounts for systematic structure introduced by the stratified randomization.

What does the KCCQ-CSS estimate of 7.8 mean?

It is the estimated adjusted difference between semaglutide 2.4 mg and placebo for the change in KCCQ-CSS from week 0 to week 52. A positive value indicates a higher estimated change in the semaglutide group on this score scale. It does not mean that every patient experienced a 7.8-point change.

What does the body-weight estimate of −10.7 mean?

It is the estimated adjusted treatment difference for the percentage-of-body-weight outcome reported by the registry. The negative sign means the estimated change was lower in the semaglutide group than in the placebo group by 10.7 percentage points on that analysis scale.

What does the 95% confidence interval tell us?

A 95% confidence interval describes uncertainty around the estimated treatment difference under the statistical model and sampling framework. For KCCQ-CSS, the interval is 4.8–10.9; for body weight, it is −11.9–−9.4. Neither interval describes the range of individual participant responses.

Why doesn't the p-value measure effect size?

A p-value summarizes evidence against a null hypothesis under a specified statistical model. It depends on both the estimated effect and the amount of information available. Therefore, two studies can have similarly small p-values despite different effect sizes, and a p-value should be interpreted alongside the estimate and confidence interval.

11. Missing Data and Imputation

Missing week-52 observations are explicitly addressed in both primary analysis descriptions. The registry states that missing observations at week 52 were multiple (x1000) imputed from retrieved participants of the same randomized treatment arm..

Why impute?

Simply analyzing only participants with observed week-52 values can change the analysis population and potentially introduce bias if missingness is related to participant characteristics or outcomes.

Why multiple imputation?

Multiple imputation is designed to reflect uncertainty about missing values by creating multiple plausible completed datasets rather than treating a single filled-in value as known with certainty.

Same-treatment retrieval

The registry-reported analysis notes state that imputation used retrieved participants from the same randomized treatment.

What is not available here?

The ClinicalTrials.gov record does not provide enough detail to reconstruct the complete imputation model, auxiliary variables, or sensitivity analyses.

The presence of imputation is therefore part of the interpretation of both primary estimates. The reported treatment differences are not simply unadjusted comparisons among participants who happened to have an observed week-52 value.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The reported counts are presented as affected participants over participants at risk.

Safety measureSemaglutide 2.4 mgPlacebo
Serious adverse events35/26371/266

These figures provide the number of participants affected and the corresponding number at risk in each arm. They are a safety count summary rather than a formal statistical comparison reported in the ClinicalTrials.gov record.

Safety interpretation: the ClinicalTrials.gov record does not provide a statistical analysis, confidence interval, or p-value for the serious-adverse-event comparison. The two arm-level counts therefore should not be converted into a formal comparative inference that the registry did not report.

13. What the Estimated Treatment Difference Does — and Does Not — Mean

KCCQ-CSS

An estimated treatment difference of 7.8 means the ANCOVA model estimated a 7.8-unit difference between semaglutide 2.4 mg and placebo in the change in KCCQ-CSS from baseline to week 52, after accounting for baseline KCCQ-CSS and BMI stratification.

It does not mean that every participant experienced a 7.8-unit improvement, that every participant benefited, or that the effect was identical across individuals.

Body weight

An estimated treatment difference of −10.7 means the adjusted difference in the reported percentage-of-body-weight outcome favored the semaglutide group by 10.7 percentage points on the analysis scale.

It does not mean that every participant lost exactly 10.7% of body weight or that the same effect occurred in every subgroup or individual.

Confidence intervals

The 95% confidence interval of 4.8–10.9 for KCCQ-CSS and −11.9–−9.4 for body weight quantify statistical uncertainty around their respective model-based estimates. They are not prediction intervals for individual participants.

P-values

Both reported primary comparisons have P < 0.0001. This indicates strong evidence against the corresponding null hypothesis under the reported analysis, but it does not by itself establish the magnitude, practical importance, or individual-level consistency of the treatment effect.

14. Primary Endpoint Analysis in Statistical Detail

Model componentKCCQ-CSSBody weight
ResponseWeek-52 KCCQ-CSS responseWeek-52 body-weight response
Treatment factorRandomized treatment arm.Randomized treatment arm.
Stratification factorBMI <35.0 vs BMI ≥35.0 kg/m2BMI <35.0 vs BMI ≥35.0 kg/m2
Baseline covariateBaseline KCCQ-CSSBaseline body weight (kg)
Analysis populationFASFAS
Missing week-52 observationsMultiple (x1000) imputationMultiple (x1000) imputation
Effect measureEstimated Treatment DifferenceEstimated Treatment Difference
HypothesisSuperioritySuperiority

This structure illustrates an important principle in clinical-trial statistics: the reported treatment effect is the output of a prespecified model, not merely the arithmetic difference between two unadjusted group means.

15. Endpoint Type and Model Interpretation

The ClinicalTrials.gov record identifies KCCQ-CSS as a continuous endpoint and body weight as a binary endpoint in the statistical-analysis metadata, while the body-weight outcome unit is reported as percentage of body weight and the method is ANCOVA. That combination should be read carefully rather than silently replacing the registry's classification.

Registry-method note: the page preserves the posted analysis metadata. Because the formal analysis is reported as ANCOVA with an estimated treatment difference, the interpretation here follows the posted model and effect measure rather than substituting a different analysis based solely on the normalized endpoint-type label.

16. Multiplicity and the Two Primary Endpoints

The registry identifies two primary endpoints and reports a superiority hypothesis for each formal primary analysis. The ClinicalTrials.gov record does not specify an alpha-allocation scheme, hierarchical testing sequence, gatekeeping procedure, or other multiplicity adjustment connecting the two endpoints.

EndpointRoleFormal analysis postedMultiplicity procedure reported?
Change in KCCQ-CSS Primary Yes Not specified in the ClinicalTrials.gov record
Change in Body Weight Primary Yes Not specified in the ClinicalTrials.gov record

This matters because multiple primary endpoints can create a family of hypotheses. The fact that both reported p-values are below 0.0001 is an important descriptive result, but the ClinicalTrials.gov record does not establish how type I error across the two primary endpoints was controlled.

17. Blinding and Its Statistical Role

The trial is registered as double masked. Blinding is not itself a statistical test, but it is an important design feature because knowledge of treatment assignment can influence behavior, reporting, assessment, and other aspects of trial conduct.

Patient-reported endpoint

KCCQ-CSS is a patient-reported measure, making the protection provided by masking particularly relevant to the interpretation of observed responses.

Objective measurement

Body weight is measured differently from a patient-reported score, but blinded allocation remains part of the overall randomized trial design.

Blinding does not eliminate all sources of bias, and it does not change the mathematical interpretation of an ANCOVA estimate. Its principal role is to strengthen the design from which the statistical comparison is generated.

18. Trial Timeline

2021-03-16

Trial start

The registry lists 2021-03-16 as the study start date.

Phase 3

Randomized parallel design

The trial used randomized allocation, a parallel design, double masking, and two treatment arms.

Week 0 → Week 52

Primary endpoint assessment window

Both registered primary endpoints were defined as changes from baseline at week 0 to the end of treatment at week 52.

2023-04-18

Primary completion

The registry lists 2023-04-18 as the primary completion date.

19. Limitations

20. Why This Trial Matters Statistically

STEP-HFpEF is a useful teaching case because its two primary endpoints demonstrate how a randomized clinical trial can analyze different outcome domains using the same general modeling framework while tailoring the covariate adjustment to the endpoint.

ConceptHow it appears in STEP-HFpEF
RandomizationParticipants were randomized to semaglutide or placebo in a parallel phase 3 design.
BlindingThe registry identifies the trial as double masked.
Stratified analysisBMI <35.0 kg/m2 versus BMI ≥35.0 kg/m2 was included as a factor.
ANCOVAUsed for both posted primary endpoint analyses.
Covariate adjustmentBaseline KCCQ-CSS and baseline body weight were included for their respective endpoints.
Confidence intervalsBoth primary treatment differences include two-sided 95% confidence intervals.
P-valuesBoth primary analyses report P < 0.0001.
Missing-data handlingMissing week-52 observations were multiply imputed from retrieved participants of the same randomized treatment.
Superiority testingBoth primary analyses are identified as superiority hypotheses.
Safety summariesSerious adverse events are reported by randomized arm.

21. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

Continue with the statistical methods behind this trial

Explore the Clinical Biostats tutorials and statistical calculators related to ANCOVA, confidence intervals, covariate adjustment, randomization, blinding, p-values, and stratified analysis.

24. Record Summary

STEP-HFpEF provides a clear example of an adjusted randomized-trial analysis in which two primary endpoints were evaluated at the same week-52 time point using ANCOVA. The KCCQ-CSS analysis produced an estimated treatment difference of 7.8 with a two-sided 95% CI of 4.8–10.9 and P < 0.0001. The body-weight analysis produced an estimated treatment difference of −10.7 with a two-sided 95% CI of −11.9–−9.4 and P < 0.0001.

The statistical story is more than the two p-values. The estimates arise from ANCOVA models incorporating randomized treatment, BMI stratification, and the relevant baseline covariate, with multiple imputation used for missing week-52 observations. Interpreting the results therefore requires attention to the effect estimate, confidence interval, p-value, analysis population, covariate adjustment, stratification, and missing-data approach.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and preserving the limits of the available analysis data.