This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
STEP 8 was a completed, randomized, parallel phase 3 trial in people living with overweight or obesity. The registered primary comparison evaluated semaglutide 2.4 mg versus liraglutide 3.0 mg for change from baseline (week 0) to week 68 in body weight (%).
| Feature | STEP 8 |
|---|---|
| Trial | STEP 8 |
| NCT ID | NCT04074161 |
| Therapeutic area | Endocrinology |
| Condition | Overweight; Obesity |
| Phase | Phase 3 |
| Status | COMPLETED |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | QUADRUPLE |
| Primary purpose | TREATMENT |
| Enrollment | 338.0 |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | INDUSTRY |
2. Clinical Question
The central statistical question was whether semaglutide 2.4 mg produced a different change in body weight from baseline (week 0) to week 68 compared with liraglutide 3.0 mg. The registered hypothesis type was superiority, so the comparison was framed around whether the two randomized treatment groups differed rather than around a non-inferiority margin.
Population
People living with overweight or obesity, as described by the trial's registered conditions.
Intervention
Semaglutide, with the statistical analysis specifically comparing the semaglutide 2.4 mg group with the liraglutide 3.0 mg group.
Comparator
Liraglutide 3.0 mg for the primary treatment comparison.
Primary question
What is the treatment difference in change from baseline (week 0) to week 68 in body weight (%) between semaglutide 2.4 mg and liraglutide 3.0 mg?
3. Trial Design
Semaglutide
- Semaglutide (drug)
- Primary comparison group: semaglutide 2.4 mg
Placebo (semaglutide)
- Placebo (semaglutide) (drug)
Liraglutide
- Liraglutide (drug)
- Primary comparison group: liraglutide 3.0 mg
Placebo (liraglutide)
- Placebo (liraglutide) (drug)
4. Endpoints
The registry identifies one primary endpoint. Its registered definition specifies change from baseline (week 0) to week 68 in body weight (%) for the semaglutide 2.4 mg versus liraglutide 3.0 mg comparison.
| Endpoint | Definition / time frame | Analysis |
|---|---|---|
| Primary endpoint | Change From Baseline (Week 0) to Week 68 in Body Weight (%) (Semaglutide 2.4 mg Versus Liraglutide 3.0 mg) | ANCOVA |
| Time frame | Baseline (week 0), week 68 | Two-sided 95% CI |
| Outcome unit | Percentage of body weight | Treatment difference |
| Endpoint type | Binary, as recorded in the registry analysis data | Superiority hypothesis |
The registry states that change from baseline (week 0) to week 68 in body weight (%) is presented. Data are reported for the in-trial period, defined as the uninterrupted time interval from the date of randomization to the date of last contact with the trial site.
5. Primary Result: Change in Body Weight
The posted formal analysis compares semaglutide 2.4 mg with liraglutide 3.0 mg using an analysis of covariance model. The full analysis set included all randomized participants, while the number analyzed for the outcome was defined as participants with available data for that outcome measure.
Semaglutide 2.4 mg vs liraglutide 3.0 mg
Treatment difference in change from baseline to week 68 in body weight (%)
95% CI: -11.97 to -6.80 · P < 0.0001
| Analysis element | Reported result |
|---|---|
| Outcome measure | Change From Baseline (Week 0) to Week 68 in Body Weight (%) |
| Groups compared | Semaglutide 2.4 mg vs Liraglutide 3.0 mg |
| Analysis population | The full analysis set (FAS) included all randomized participants; participants with available data for this outcome measure contributed to the reported analysis. |
| Method | ANCOVA |
| Effect measure | Treatment difference |
| Estimate | -9.38 |
| Confidence interval | 95% two-sided CI: -11.97 to -6.80 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The estimated treatment difference was -9.38 percentage points for the semaglutide 2.4 mg group relative to the liraglutide 3.0 mg group under the registered analysis. Because the estimate is negative, the modeled change in body weight (%) was lower in the semaglutide group relative to the liraglutide group according to the direction used for the reported treatment difference.
The estimate does not mean that every participant experienced exactly a -9.38 change, nor does it describe the response of an individual participant. It is a group-level treatment contrast produced by the specified ANCOVA analysis.
The 95% two-sided confidence interval extends from -11.97 to -6.80. This interval describes statistical uncertainty around the estimated treatment difference under the analysis framework; it is not a range containing the individual treatment effects of all participants.
The P-value <0.0001 addresses evidence against the null hypothesis of no treatment difference within the specified testing framework. A p-value is not a measure of the magnitude of the effect. The size and direction of the estimated treatment difference are communicated by -9.38, while its statistical uncertainty is communicated by the confidence interval.
The analysis was specified as a superiority comparison, not a non-inferiority analysis. Consequently, there is no non-inferiority margin to interpret here. The analysis is also not a time-to-event model, so proportional-hazards assumptions and censoring rules associated with Cox regression are not the central interpretive issues for this endpoint.
6. Statistical Methodology
Analysis of covariance
The registry reports an analysis of covariance (ANCOVA) model with randomized treatment as a factor and baseline body weight as a covariate. This combines information about treatment assignment with baseline body weight when estimating the treatment difference at the endpoint.
For STEP 8, the registered analysis describes randomized treatment as the factor and baseline body weight as the covariate. The resulting treatment contrast is reported as a treatment difference.
Adjusting for baseline body weight can make the treatment comparison more statistically efficient because baseline measurements contain information about where participants started. The key distinction is that baseline body weight is used as a covariate rather than becoming a second treatment group or a post-randomization outcome.
Randomized treatment as a factor
The registry states that randomized treatment was included as a factor in the ANCOVA model. This preserves the randomized comparison at the center of the analysis: treatment assignment defines the groups being contrasted, while baseline body weight provides covariate information.
Baseline body weight as a covariate
Baseline body weight is measured before the week 68 outcome and is therefore available as a baseline characteristic for adjustment. In an ANCOVA framework, this allows the estimated treatment contrast to account statistically for baseline body-weight differences rather than relying solely on an unadjusted comparison of endpoint values.
Full analysis set and available outcome data
The analysis record states that the full analysis set (FAS) included all randomized participants. It also states that the overall number of participants analyzed for the outcome measure consisted of participants with available data for that outcome measure. These two statements are important to keep separate: the FAS is defined from randomization, while the reported outcome analysis depends on availability of data for the particular endpoint.
7. Statistical Methods Explained
Why was ANCOVA used?
ANCOVA is appropriate for comparing a quantitative outcome between randomized treatment groups while incorporating a prespecified baseline covariate. In STEP 8, the registry explicitly reports randomized treatment as the factor and baseline body weight as the covariate. This makes the treatment difference the principal parameter while using baseline information to improve the statistical comparison.
What does a treatment difference of -9.38 mean?
The reported treatment difference is the estimated contrast between semaglutide 2.4 mg and liraglutide 3.0 mg for change from baseline (week 0) to week 68 in body weight (%). The negative sign identifies the direction of the contrast as reported. It should not be converted into a percentage reduction in the probability of an event or interpreted as an individual-level effect.
Why does baseline body weight enter the model?
Participants begin the study with different body weights. Including baseline body weight as a covariate allows the model to account for that starting-point information when estimating the treatment contrast. This is different from simply comparing the raw baseline values or treating baseline weight as an outcome.
What does the 95% confidence interval tell us?
The 95% two-sided confidence interval is -11.97 to -6.80. It expresses uncertainty around the estimated treatment difference under the statistical model and sampling framework. A confidence interval is more informative than the point estimate alone because it shows the range of treatment contrasts compatible with the observed data under that framework.
Why does the p-value not measure effect size?
The reported P <0.0001 indicates the strength of evidence against the null hypothesis within the specified superiority analysis. It does not say that the treatment effect is "less than 0.0001" or that the treatment difference is a probability. The effect size is the treatment difference of -9.38, while the confidence interval communicates its precision.
How does randomization support the comparison?
Randomization determines treatment assignment independently of participants' subsequent outcomes. This is the foundation for interpreting differences between randomized groups as treatment contrasts rather than simply as associations between treatment exposure and outcome. The strength of that interpretation depends on the actual conduct and analysis of the randomized trial.
8. Safety Results
The trial data provide serious adverse events by randomized treatment category as affected participants divided by participants at risk. These figures are reported separately from the efficacy analysis and should not be combined into a single numerical benefit-risk measure.
| Safety group | Serious adverse events | Interpretation of reported format |
|---|---|---|
| Semaglutide 2.4 mg | 10/126 | 10 affected among 126 at risk |
| Liraglutide 3.0 mg | 14/127 | 14 affected among 127 at risk |
| Pooled Placebo | 6/85 | 6 affected among 85 at risk |
The serious-adverse-event figures provide counts of affected participants relative to those at risk for the three reported safety groups. They do not by themselves establish causality, severity distributions beyond the serious-adverse-event classification, timing, or comparative statistical significance.
Efficacy and safety are distinct analyses
The primary efficacy estimate concerns change in body weight, whereas serious adverse events describe a separate safety domain. The two should be interpreted independently.
Do not infer an unreported test
The ClinicalTrials.gov record reports affected/at-risk counts for serious adverse events but do not provide a formal statistical comparison, confidence interval, or p-value for these safety counts.
9. What the Primary Estimate Does — and Does Not — Mean
The treatment difference of -9.38 means that the estimated change in body weight (%) was lower for semaglutide 2.4 mg than for liraglutide 3.0 mg according to the direction of the reported treatment contrast.
The 95% two-sided confidence interval of -11.97 to -6.80 shows that the point estimate should not be viewed as exact. The interval provides a measure of uncertainty around the estimated treatment contrast.
P <0.0001 is evidence against a null treatment difference under the specified superiority analysis. It does not quantify how large the treatment effect is and does not replace the effect estimate or its confidence interval.
This result is not a hazard ratio, risk ratio, odds ratio, or survival probability. There is therefore no reason to apply time-to-event interpretations such as proportional-hazards assumptions to this primary analysis.
10. Superiority Rather Than Non-Inferiority
The registry identifies the primary hypothesis type as superiority. That distinction determines how the treatment comparison should be interpreted.
| Feature | STEP 8 primary analysis |
|---|---|
| Hypothesis type | Superiority |
| Effect measure | Treatment difference |
| Estimate | -9.38 |
| Confidence interval | 95% two-sided: -11.97 to -6.80 |
| Non-inferiority margin | Not part of the registry-reported primary analysis data |
| Formal method | ANCOVA |
In a superiority analysis, the question is whether the randomized treatment groups differ. A non-inferiority analysis instead asks whether a new treatment is not unacceptably worse than a comparator by more than a prespecified margin. Because STEP 8's registry-reported analysis is identified as superiority, the interpretation should remain centered on the treatment difference and its two-sided confidence interval rather than on a non-inferiority margin.
11. Trial Timeline
Trial start
The registered trial start date was 2019-09-11.
Primary completion
The registered primary completion date was 2021-03-27.
Results posted
The ClinicalTrials.gov record identifies the trial as COMPLETED and report that results were posted, with 32 outcome measures and 1 statistical analysis posted.
12. Important Limitations and Interpretation Issues
- Endpoint classification: the registry analysis record classifies the primary endpoint as binary even though its wording describes change in body weight (%). This page preserves the registry classification rather than substituting a different classification.
- Available outcome data: the analysis record defines the FAS as all randomized participants but states that the overall number analyzed for the outcome consists of participants with available data for that outcome measure. This distinction matters when interpreting the relationship between randomization and the analyzed observations.
- Point estimate versus individual response: the treatment difference of -9.38 is a group-level model estimate and should not be interpreted as the response experienced by every participant.
- Confidence interval: the 95% CI describes uncertainty around the estimated treatment contrast; it is not a prediction interval for individual participant outcomes.
- P-value interpretation: P <0.0001 measures evidence against the null hypothesis under the specified framework, not the magnitude or clinical importance of the treatment difference.
- Safety information: the ClinicalTrials.gov record reports serious adverse events by arm but do not provide a formal comparative analysis for those safety counts.
- Limited statistical-analysis record: the registry extraction reports 1 statistical analysis and 1 primary-endpoint analysis. Other posted outcome measures are not accompanied by additional statistical analyses in the ClinicalTrials.gov record.
- Scope of inference: the statistical interpretation is tied to the randomized comparison and the registered endpoint, analysis population, and ANCOVA specification reported in the ClinicalTrials.gov record.
13. Why This Trial Matters Statistically
STEP 8 is a useful teaching case because it illustrates a different statistical structure from trials dominated by survival analysis. Its primary comparison uses an ANCOVA model for a body-weight change endpoint, with randomized treatment as a factor and baseline body weight as a covariate. The result therefore provides a compact example of how randomized clinical-trial evidence can be expressed as an adjusted treatment difference rather than a hazard ratio.
| Concept | How it appears in STEP 8 |
|---|---|
| Randomization | The trial uses RANDOMIZED allocation in a phase 3 parallel design. |
| Blinding | The registered masking level is QUADRUPLE. |
| ANCOVA | The primary statistical method is ANCOVA with randomized treatment as factor and baseline body weight as covariate. |
| Baseline adjustment | Baseline body weight is explicitly incorporated as a covariate. |
| Intention-to-treat analysis | Intention-to-treat analysis is identified among the concepts in the analysis text, with the FAS defined as all randomized participants. |
| Treatment difference | The primary effect measure is reported as a treatment difference of -9.38. |
| Confidence interval | The treatment difference has a 95% two-sided CI of -11.97 to -6.80. |
| P-value | The superiority comparison reports P <0.0001. |
| Safety analysis | Serious adverse events are reported as affected/at-risk counts by treatment group. |
| Endpoint timing | The registered primary endpoint compares baseline (week 0) with week 68. |
14. Statistical Concepts in This Trial
This trial provides a focused pathway into the statistical concepts used to design and interpret randomized clinical trials with adjusted continuous outcomes:
15. Related Tutorials
Learn more about the methods used in this trial:
16. Related Calculators
Practice the core quantitative methods represented in the STEP 8 analysis:
17. Sources
- ClinicalTrials.gov: STEP 8 (NCT04074161).
- Linked publication: PubMed record: PMID 34775881.
- Linked publication: PubMed record: PMID 35015037.
- Linked publication: PubMed record: PMID 35791625.
Continue through the Clinical Biostats statistical methods library
Use the related tutorials and calculators to examine the core concepts behind randomized treatment comparisons, covariate adjustment, confidence intervals, and hypothesis testing.
18. Record Summary
STEP 8 provides a clear example of a randomized phase 3 superiority comparison analyzed with ANCOVA. The registered primary endpoint was change from baseline (week 0) to week 68 in body weight (%), comparing semaglutide 2.4 mg with liraglutide 3.0 mg. The reported treatment difference was -9.38, with a 95% two-sided confidence interval of -11.97 to -6.80 and P <0.0001. The analysis used randomized treatment as a factor and baseline body weight as a covariate, with the full analysis set defined as all randomized participants and the outcome analysis based on participants with available data for that measure.
Statistically, the most important lesson is the distinction between an effect estimate, its uncertainty, and the evidence against a null hypothesis. The treatment difference communicates the estimated magnitude and direction, the confidence interval communicates precision, and the p-value addresses the hypothesis-testing question. None of these quantities should be interpreted as an individual patient's expected response or as a complete description of the trial's safety profile.