← Clinical Trials
Overweight / Obesity Phase 3 Superiority NCT04074161

STEP 8: Complete Statistical Analysis of Semaglutide in Overweight or Obesity

An independent statistical review of the randomized phase 3 STEP 8 trial comparing semaglutide 2.4 mg with liraglutide 3.0 mg in people living with overweight or obesity, with the primary analysis focused on change from baseline (week 0) to week 68 in body weight (%).

Trial start: 2019-09-11  ·  Primary completion: 2021-03-27  ·  Enrollment: 338.0
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

STEP 8 was a completed, randomized, parallel phase 3 trial in people living with overweight or obesity. The registered primary comparison evaluated semaglutide 2.4 mg versus liraglutide 3.0 mg for change from baseline (week 0) to week 68 in body weight (%).

338.0
Enrollment
Randomized trial
4
Arms
Parallel design
-9.38
Treatment difference
Semaglutide vs liraglutide
<0.0001
P-value
Superiority analysis
FeatureSTEP 8
TrialSTEP 8
NCT IDNCT04074161
Therapeutic areaEndocrinology
ConditionOverweight; Obesity
PhasePhase 3
StatusCOMPLETED
AllocationRANDOMIZED
Design modelPARALLEL
MaskingQUADRUPLE
Primary purposeTREATMENT
Enrollment338.0
Lead sponsorNovo Nordisk A/S
Sponsor typeINDUSTRY

2. Clinical Question

The central statistical question was whether semaglutide 2.4 mg produced a different change in body weight from baseline (week 0) to week 68 compared with liraglutide 3.0 mg. The registered hypothesis type was superiority, so the comparison was framed around whether the two randomized treatment groups differed rather than around a non-inferiority margin.

Population

People living with overweight or obesity, as described by the trial's registered conditions.

Intervention

Semaglutide, with the statistical analysis specifically comparing the semaglutide 2.4 mg group with the liraglutide 3.0 mg group.

Comparator

Liraglutide 3.0 mg for the primary treatment comparison.

Primary question

What is the treatment difference in change from baseline (week 0) to week 68 in body weight (%) between semaglutide 2.4 mg and liraglutide 3.0 mg?

3. Trial Design

01
Randomize338.0 participants
02
Parallel groups4 registered arms
03
Quadruple maskingRegistered masking
04
Week 0Baseline body weight
05
Week 68Primary endpoint
REGISTERED ARM

Semaglutide

  • Semaglutide (drug)
  • Primary comparison group: semaglutide 2.4 mg
REGISTERED ARM

Placebo (semaglutide)

  • Placebo (semaglutide) (drug)
REGISTERED ARM

Liraglutide

  • Liraglutide (drug)
  • Primary comparison group: liraglutide 3.0 mg
REGISTERED ARM

Placebo (liraglutide)

  • Placebo (liraglutide) (drug)
Why the design matters. Randomization creates the basis for comparing treatment groups without assigning participants according to observed outcomes. The parallel structure means the randomized groups are followed as distinct treatment assignments rather than being switched between treatments as part of the registered design. Quadruple masking indicates that masking was part of the registered trial design, reducing opportunities for knowledge of treatment assignment to influence trial conduct or assessment.

4. Endpoints

The registry identifies one primary endpoint. Its registered definition specifies change from baseline (week 0) to week 68 in body weight (%) for the semaglutide 2.4 mg versus liraglutide 3.0 mg comparison.

EndpointDefinition / time frameAnalysis
Primary endpointChange From Baseline (Week 0) to Week 68 in Body Weight (%) (Semaglutide 2.4 mg Versus Liraglutide 3.0 mg)ANCOVA
Time frameBaseline (week 0), week 68Two-sided 95% CI
Outcome unitPercentage of body weightTreatment difference
Endpoint typeBinary, as recorded in the registry analysis dataSuperiority hypothesis

The registry states that change from baseline (week 0) to week 68 in body weight (%) is presented. Data are reported for the in-trial period, defined as the uninterrupted time interval from the date of randomization to the date of last contact with the trial site.

Registry terminology matters. The registry's analysis record classifies the primary endpoint as binary even though the endpoint description concerns change in body weight expressed as a percentage. This page preserves that registry classification rather than silently replacing it with a different endpoint type.

5. Primary Result: Change in Body Weight

The posted formal analysis compares semaglutide 2.4 mg with liraglutide 3.0 mg using an analysis of covariance model. The full analysis set included all randomized participants, while the number analyzed for the outcome was defined as participants with available data for that outcome measure.

Semaglutide 2.4 mg vs liraglutide 3.0 mg

-9.38

Treatment difference in change from baseline to week 68 in body weight (%)

95% CI: -11.97 to -6.80   ·   P < 0.0001

Analysis elementReported result
Outcome measureChange From Baseline (Week 0) to Week 68 in Body Weight (%)
Groups comparedSemaglutide 2.4 mg vs Liraglutide 3.0 mg
Analysis populationThe full analysis set (FAS) included all randomized participants; participants with available data for this outcome measure contributed to the reported analysis.
MethodANCOVA
Effect measureTreatment difference
Estimate-9.38
Confidence interval95% two-sided CI: -11.97 to -6.80
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The estimated treatment difference was -9.38 percentage points for the semaglutide 2.4 mg group relative to the liraglutide 3.0 mg group under the registered analysis. Because the estimate is negative, the modeled change in body weight (%) was lower in the semaglutide group relative to the liraglutide group according to the direction used for the reported treatment difference.

The estimate does not mean that every participant experienced exactly a -9.38 change, nor does it describe the response of an individual participant. It is a group-level treatment contrast produced by the specified ANCOVA analysis.

The 95% two-sided confidence interval extends from -11.97 to -6.80. This interval describes statistical uncertainty around the estimated treatment difference under the analysis framework; it is not a range containing the individual treatment effects of all participants.

The P-value <0.0001 addresses evidence against the null hypothesis of no treatment difference within the specified testing framework. A p-value is not a measure of the magnitude of the effect. The size and direction of the estimated treatment difference are communicated by -9.38, while its statistical uncertainty is communicated by the confidence interval.

The analysis was specified as a superiority comparison, not a non-inferiority analysis. Consequently, there is no non-inferiority margin to interpret here. The analysis is also not a time-to-event model, so proportional-hazards assumptions and censoring rules associated with Cox regression are not the central interpretive issues for this endpoint.

6. Statistical Methodology

Analysis of covariance

The registry reports an analysis of covariance (ANCOVA) model with randomized treatment as a factor and baseline body weight as a covariate. This combines information about treatment assignment with baseline body weight when estimating the treatment difference at the endpoint.

Conceptual ANCOVA structure
Outcome = treatment effect + baseline covariate effect + residual variation

For STEP 8, the registered analysis describes randomized treatment as the factor and baseline body weight as the covariate. The resulting treatment contrast is reported as a treatment difference.

Adjusting for baseline body weight can make the treatment comparison more statistically efficient because baseline measurements contain information about where participants started. The key distinction is that baseline body weight is used as a covariate rather than becoming a second treatment group or a post-randomization outcome.

Randomized treatment as a factor

The registry states that randomized treatment was included as a factor in the ANCOVA model. This preserves the randomized comparison at the center of the analysis: treatment assignment defines the groups being contrasted, while baseline body weight provides covariate information.

Baseline body weight as a covariate

Baseline body weight is measured before the week 68 outcome and is therefore available as a baseline characteristic for adjustment. In an ANCOVA framework, this allows the estimated treatment contrast to account statistically for baseline body-weight differences rather than relying solely on an unadjusted comparison of endpoint values.

Full analysis set and available outcome data

The analysis record states that the full analysis set (FAS) included all randomized participants. It also states that the overall number of participants analyzed for the outcome measure consisted of participants with available data for that outcome measure. These two statements are important to keep separate: the FAS is defined from randomization, while the reported outcome analysis depends on availability of data for the particular endpoint.

Intention-to-treat principle. The analysis text identifies intention-to-treat analysis as an analysis concept. The core idea is to preserve the randomized treatment assignment when evaluating efficacy. The registry's specific wording should be distinguished from a claim that every randomized participant necessarily contributed an observed week 68 value; the analysis record explicitly describes the outcome analysis as involving participants with available data for the outcome measure.

7. Statistical Methods Explained

Why was ANCOVA used?

ANCOVA is appropriate for comparing a quantitative outcome between randomized treatment groups while incorporating a prespecified baseline covariate. In STEP 8, the registry explicitly reports randomized treatment as the factor and baseline body weight as the covariate. This makes the treatment difference the principal parameter while using baseline information to improve the statistical comparison.

What does a treatment difference of -9.38 mean?

The reported treatment difference is the estimated contrast between semaglutide 2.4 mg and liraglutide 3.0 mg for change from baseline (week 0) to week 68 in body weight (%). The negative sign identifies the direction of the contrast as reported. It should not be converted into a percentage reduction in the probability of an event or interpreted as an individual-level effect.

Why does baseline body weight enter the model?

Participants begin the study with different body weights. Including baseline body weight as a covariate allows the model to account for that starting-point information when estimating the treatment contrast. This is different from simply comparing the raw baseline values or treating baseline weight as an outcome.

What does the 95% confidence interval tell us?

The 95% two-sided confidence interval is -11.97 to -6.80. It expresses uncertainty around the estimated treatment difference under the statistical model and sampling framework. A confidence interval is more informative than the point estimate alone because it shows the range of treatment contrasts compatible with the observed data under that framework.

Why does the p-value not measure effect size?

The reported P <0.0001 indicates the strength of evidence against the null hypothesis within the specified superiority analysis. It does not say that the treatment effect is "less than 0.0001" or that the treatment difference is a probability. The effect size is the treatment difference of -9.38, while the confidence interval communicates its precision.

How does randomization support the comparison?

Randomization determines treatment assignment independently of participants' subsequent outcomes. This is the foundation for interpreting differences between randomized groups as treatment contrasts rather than simply as associations between treatment exposure and outcome. The strength of that interpretation depends on the actual conduct and analysis of the randomized trial.

8. Safety Results

The trial data provide serious adverse events by randomized treatment category as affected participants divided by participants at risk. These figures are reported separately from the efficacy analysis and should not be combined into a single numerical benefit-risk measure.

Safety groupSerious adverse eventsInterpretation of reported format
Semaglutide 2.4 mg10/12610 affected among 126 at risk
Liraglutide 3.0 mg14/12714 affected among 127 at risk
Pooled Placebo6/856 affected among 85 at risk

The serious-adverse-event figures provide counts of affected participants relative to those at risk for the three reported safety groups. They do not by themselves establish causality, severity distributions beyond the serious-adverse-event classification, timing, or comparative statistical significance.

Efficacy and safety are distinct analyses

The primary efficacy estimate concerns change in body weight, whereas serious adverse events describe a separate safety domain. The two should be interpreted independently.

Do not infer an unreported test

The ClinicalTrials.gov record reports affected/at-risk counts for serious adverse events but do not provide a formal statistical comparison, confidence interval, or p-value for these safety counts.

9. What the Primary Estimate Does — and Does Not — Mean

Direction of the effect

The treatment difference of -9.38 means that the estimated change in body weight (%) was lower for semaglutide 2.4 mg than for liraglutide 3.0 mg according to the direction of the reported treatment contrast.

Precision of the estimate

The 95% two-sided confidence interval of -11.97 to -6.80 shows that the point estimate should not be viewed as exact. The interval provides a measure of uncertainty around the estimated treatment contrast.

What the p-value means

P <0.0001 is evidence against a null treatment difference under the specified superiority analysis. It does not quantify how large the treatment effect is and does not replace the effect estimate or its confidence interval.

What is not being estimated

This result is not a hazard ratio, risk ratio, odds ratio, or survival probability. There is therefore no reason to apply time-to-event interpretations such as proportional-hazards assumptions to this primary analysis.

10. Superiority Rather Than Non-Inferiority

The registry identifies the primary hypothesis type as superiority. That distinction determines how the treatment comparison should be interpreted.

FeatureSTEP 8 primary analysis
Hypothesis typeSuperiority
Effect measureTreatment difference
Estimate-9.38
Confidence interval95% two-sided: -11.97 to -6.80
Non-inferiority marginNot part of the registry-reported primary analysis data
Formal methodANCOVA

In a superiority analysis, the question is whether the randomized treatment groups differ. A non-inferiority analysis instead asks whether a new treatment is not unacceptably worse than a comparator by more than a prespecified margin. Because STEP 8's registry-reported analysis is identified as superiority, the interpretation should remain centered on the treatment difference and its two-sided confidence interval rather than on a non-inferiority margin.

11. Trial Timeline

2019-09-11

Trial start

The registered trial start date was 2019-09-11.

2021-03-27

Primary completion

The registered primary completion date was 2021-03-27.

Completed registry record

Results posted

The ClinicalTrials.gov record identifies the trial as COMPLETED and report that results were posted, with 32 outcome measures and 1 statistical analysis posted.

12. Important Limitations and Interpretation Issues

13. Why This Trial Matters Statistically

STEP 8 is a useful teaching case because it illustrates a different statistical structure from trials dominated by survival analysis. Its primary comparison uses an ANCOVA model for a body-weight change endpoint, with randomized treatment as a factor and baseline body weight as a covariate. The result therefore provides a compact example of how randomized clinical-trial evidence can be expressed as an adjusted treatment difference rather than a hazard ratio.

ConceptHow it appears in STEP 8
RandomizationThe trial uses RANDOMIZED allocation in a phase 3 parallel design.
BlindingThe registered masking level is QUADRUPLE.
ANCOVAThe primary statistical method is ANCOVA with randomized treatment as factor and baseline body weight as covariate.
Baseline adjustmentBaseline body weight is explicitly incorporated as a covariate.
Intention-to-treat analysisIntention-to-treat analysis is identified among the concepts in the analysis text, with the FAS defined as all randomized participants.
Treatment differenceThe primary effect measure is reported as a treatment difference of -9.38.
Confidence intervalThe treatment difference has a 95% two-sided CI of -11.97 to -6.80.
P-valueThe superiority comparison reports P <0.0001.
Safety analysisSerious adverse events are reported as affected/at-risk counts by treatment group.
Endpoint timingThe registered primary endpoint compares baseline (week 0) with week 68.

14. Statistical Concepts in This Trial

This trial provides a focused pathway into the statistical concepts used to design and interpret randomized clinical trials with adjusted continuous outcomes:

15. Related Tutorials

Learn more about the methods used in this trial:

16. Related Calculators

Practice the core quantitative methods represented in the STEP 8 analysis:

17. Sources

Continue through the Clinical Biostats statistical methods library

Use the related tutorials and calculators to examine the core concepts behind randomized treatment comparisons, covariate adjustment, confidence intervals, and hypothesis testing.

18. Record Summary

STEP 8 provides a clear example of a randomized phase 3 superiority comparison analyzed with ANCOVA. The registered primary endpoint was change from baseline (week 0) to week 68 in body weight (%), comparing semaglutide 2.4 mg with liraglutide 3.0 mg. The reported treatment difference was -9.38, with a 95% two-sided confidence interval of -11.97 to -6.80 and P <0.0001. The analysis used randomized treatment as a factor and baseline body weight as a covariate, with the full analysis set defined as all randomized participants and the outcome analysis based on participants with available data for that measure.

Statistically, the most important lesson is the distinction between an effect estimate, its uncertainty, and the evidence against a null hypothesis. The treatment difference communicates the estimated magnitude and direction, the confidence interval communicates precision, and the p-value addresses the hypothesis-testing question. None of these quantities should be interpreted as an individual patient's expected response or as a complete description of the trial's safety profile.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry record. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and preserving the definitions, analysis population, effect measure, confidence interval, and hypothesis-testing framework actually reported by the source data.