← Clinical Trials
Overweight / Obesity Phase 3 Completed NCT03548987

STEP 4: Complete Statistical Analysis of Semaglutide in Overweight or Obesity

An independent statistical review of the randomized phase 3 STEP 4 trial evaluating change in body weight from randomisation to week 68 with semaglutide 2.4 mg versus placebo in people with overweight or obesity.

Trial start: June 4, 2018  ·  Primary completion: February 22, 2020  ·  Lead sponsor: Novo Nordisk A/S
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

STEP 4 was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating semaglutide versus placebo in people with overweight or obesity. The registry reports 902 enrolled participants and a primary endpoint concerning change in body weight from randomisation at week 20 to week 68.

902
Enrolled
Phase 3 trial
2
Arms
Semaglutide / placebo
-14.75
Treatment difference
95% CI -16.00 to -13.50
<0.0001
P-value
ANCOVA primary analysis
FeatureSTEP 4
Trial nameSTEP 4
NCT identifierNCT03548987
PhasePhase 3
StatusCompleted
Therapeutic areaEndocrinology
ConditionsMetabolism and Nutrition Disorder; Obesity
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment902
InterventionsSemaglutide; placebo
Results postedYes
Outcome measures posted37
Statistical analyses posted2
Lead sponsorNovo Nordisk A/S
Sponsor typeIndustry

2. Clinical Question

The primary statistical question was whether the change in body weight from randomisation at week 20 to week 68 differed between participants assigned to semaglutide 2.4 mg and those assigned to placebo.

Population

The registry describes the study as investigating how well semaglutide works in people suffering from overweight or obesity. The listed conditions are metabolism and nutrition disorder and obesity.

Intervention

Semaglutide. The posted primary statistical analyses identify the treatment comparison specifically as semaglutide 2.4 mg versus placebo.

Comparator

Placebo.

Primary question

What is the treatment difference in change from randomisation to week 68 in body weight, comparing semaglutide 2.4 mg with placebo?

3. Trial Design

01
Enroll902 participants
02
RandomizeTwo parallel arms
03
MaskQuadruple masking
04
AssessBody weight through week 68
05
AnalyzeANCOVA and MMRM
Allocation
Randomized
Participants were assigned to parallel treatment groups.
Masking
Quadruple
The registry identifies the study as quadruple-masked.
Primary purpose
Treatment
The study was designed to evaluate treatment effects.
Study period
2018–2020
Start: June 4, 2018. Primary completion: February 22, 2020.
TREATMENT

Semaglutide 2.4 mg

  • Semaglutide
  • Included in the primary comparison of change in body weight
  • Primary analyses used the full analysis set framework
COMPARATOR

Placebo

  • Placebo
  • Compared with semaglutide 2.4 mg for the primary endpoint
  • Primary analyses used the full analysis set framework

4. Endpoint Framework

The registry identifies one registered primary endpoint and posts two formal statistical analyses for that endpoint. The two analyses use different statistical frameworks and are associated with different estimand descriptions.

EndpointRegistry time frameTypeAnalyses posted
Change From Randomisation to Week 68 in Body Weight (%) Randomisation (week 20) to week 68 Binary, as classified in the ClinicalTrials.gov record ANCOVA; MMRM

Registered endpoint definition

The registry states that change in body weight from baseline at week 20 to week 68 is presented. The endpoint was evaluated based on data from both in-trial and on-treatment observation periods.

The registry defines the in-trial observation period as the uninterrupted time interval from the start of run-in at week 0 to the last trial-related subject-site contact at week 75. The registry also defines an on-treatment observation period, described as including all time intervals in which participants were receiving treatment; the registry-reported endpoint definition is truncated after that point, so no additional wording is inferred here.

Important registry detail: the ClinicalTrials.gov record classifies this endpoint as “Binary,” while the posted outcome unit is “Percentage point” and the reported treatment effect is a treatment difference. This page preserves the registry classification rather than silently replacing it with a different endpoint type.

5. Statistical Analysis Populations

Both posted primary analyses identify the full analysis set (FAS) as the overall analysis population. The registry describes the FAS as comprising all randomised participants, with the number analyzed defined as participants with available data.

PopulationRegistry descriptionRole in posted analysis
Full analysis set All randomised participants Primary efficacy analysis population
Participants with available data Number analyzed is defined as participants with available data Contributes observations to the posted analysis

This distinction is important statistically. “All randomised participants” describes the FAS definition, whereas the posted analysis record separately states that the number analyzed is participants with available data. The registry therefore provides the population framework but does not, in the ClinicalTrials.gov record, provide a complete accounting of every missing observation.

6. Primary Results

ClinicalTrials.gov posts two formal statistical analyses for the same primary endpoint: one using ANCOVA and one using an MMRM. Both compare semaglutide 2.4 mg with placebo, but the analysis notes associate the ANCOVA result with a treatment policy estimand and the MMRM result with a hypothetical estimand.

ANCOVA Analysis

Treatment difference in change in body weight

-14.75 percentage points

95% CI: -16.00 to -13.50   ·   P < 0.0001

Semaglutide 2.4 mg vs placebo · Superiority hypothesis · Treatment policy estimand

CharacteristicPosted result
EndpointChange From Randomisation to Week 68 in Body Weight (%)
Time frameRandomisation (week 20) to week 68
ComparisonSemaglutide 2.4 mg vs placebo
MethodANCOVA
Effect measureTreatment difference
Estimate-14.75 percentage points
95% CI-16.00 to -13.50
P-value<0.0001
HypothesisSuperiority
Estimand noteTreatment policy estimand
Clinical Biostats interpretation

The reported treatment difference of -14.75 percentage points means that the estimated change in body weight favored semaglutide 2.4 mg relative to placebo by 14.75 percentage points under the posted ANCOVA analysis and its treatment-policy estimand.

The negative sign is important because the effect is expressed as the treatment difference for change in body weight. It does not mean that every individual participant experienced exactly a 14.75-percentage-point difference, nor does it describe the response of an individual patient.

The 95% confidence interval, -16.00 to -13.50, describes the statistical uncertainty around the estimated treatment difference under the analysis framework. It is an interval for the estimated population-level treatment effect, not a range containing the individual treatment effects of 95% of participants.

The p-value of <0.0001 addresses evidence against the null hypothesis within the posted statistical framework. It does not measure the size of the treatment effect, the probability that the treatment is effective, or the clinical importance of the result. Effect size and uncertainty are better conveyed by the estimate and its confidence interval.

Because the analysis is identified as a treatment-policy estimand, its interpretation is tied to the specified treatment-policy question rather than simply describing outcomes while participants remain continuously on treatment. The ClinicalTrials.gov record does not provide the complete estimand definition beyond this classification.

MMRM Analysis

Treatment difference in change in body weight

-15.33 percentage points

95% CI: -16.52 to -14.13   ·   P < 0.0001

Semaglutide 2.4 mg vs placebo · Hypothetical estimand

CharacteristicPosted result
EndpointChange From Randomisation to Week 68 in Body Weight (%)
Time frameRandomisation (week 20) to week 68
ComparisonSemaglutide 2.4 mg vs placebo
MethodMMRM (mixed model repeated measurement)
Effect measureTreatment difference
Estimate-15.33 percentage points
95% CI-16.52 to -14.13
P-value<0.0001
Hypothesis typeOther / not stated
Estimand noteHypothetical estimand
Clinical Biostats interpretation

The MMRM estimate of -15.33 percentage points is the estimated treatment difference from randomisation to week 68 under the posted longitudinal model and hypothetical estimand.

MMRM uses the repeated-measure structure of longitudinal data rather than reducing the entire follow-up history to a single observation without regard to intermediate measurements. That makes the method particularly relevant when an endpoint is measured repeatedly over time.

The 95% confidence interval of -16.52 to -14.13 quantifies uncertainty around the model-based treatment-difference estimate. It does not mean that 95% of participants had changes within that interval.

The p-value of <0.0001 indicates strong statistical evidence under the specified model and hypothesis-testing framework, but it does not quantify the magnitude or practical importance of the treatment effect.

The MMRM analysis is explicitly labeled as a hypothetical estimand. Consequently, its interpretation depends on the hypothetical treatment-policy scenario represented by that estimand. The ClinicalTrials.gov record does not provide enough detail to reconstruct the full hypothetical scenario, so no additional assumptions are imposed here.

7. Comparing the Two Primary Analyses

The two posted estimates are not identical, but they are directionally consistent. The ANCOVA treatment difference is -14.75, while the MMRM treatment difference is -15.33. Both confidence intervals lie entirely below zero and both reported p-values are <0.0001.

FeatureANCOVAMMRM
Estimate-14.75-15.33
95% CI-16.00 to -13.50-16.52 to -14.13
P-value<0.0001<0.0001
Effect measureTreatment differenceTreatment difference
Analysis familyLinear modelLongitudinal / mixed model
Estimand noteTreatment policyHypothetical
Analysis populationFAS; participants with available dataFAS; participants with available data

The difference between the point estimates should not be interpreted as evidence that one method is “correct” and the other is “incorrect.” They answer related but not necessarily identical statistical questions because the methods and estimands differ. ANCOVA is a cross-sectional linear-model approach to the specified endpoint, whereas MMRM is designed for repeated measurements and can model the longitudinal trajectory.

What can and cannot be concluded from the comparison: the ClinicalTrials.gov record supports the statement that both formal analyses report a negative treatment difference with a 95% confidence interval entirely below zero and a p-value <0.0001. They do not provide enough information to determine which model assumptions were most influential or to reconstruct the complete statistical analysis plan.

8. Statistical Methodology

ANCOVA

ANCOVA, or analysis of covariance, is a linear-model framework commonly used when the outcome is continuous and the analysis needs to compare treatment groups while accounting for one or more covariates. In this trial, the registry explicitly identifies ANCOVA as the method for the primary treatment-policy analysis.

Conceptual model
Y = β0 + β1(Treatment) + β2(Covariates) + ε

The treatment coefficient represents an adjusted between-group difference under the fitted model. The exact covariates and model specification are not reported in the ClinicalTrials.gov record and therefore are not inferred.

MMRM: Mixed Model for Repeated Measures

MMRM is a longitudinal modeling framework for repeated outcome measurements. Rather than treating each measurement as an isolated observation, the model accounts for the fact that measurements from the same participant are related.

Longitudinal structure
Yij = treatment/time effects + participant-level dependence + error

The exact covariance structure, fixed effects, time points, and estimation details are not provided in the ClinicalTrials.gov record. The important point for interpretation is that the posted MMRM uses repeated measurements rather than only a single cross-sectional comparison.

Treatment differences

The reported effect measure for both analyses is treatment difference. This is a difference-scale measure rather than a ratio or hazard ratio. The reported outcome unit is “percentage point,” so the estimate is interpreted on that scale.

Confidence intervals

A 95% confidence interval describes uncertainty associated with the estimated treatment difference under the specified statistical model and sampling framework. Narrower intervals generally indicate greater statistical precision than wider intervals, although precision and clinical importance are separate concepts.

P-values

The p-value evaluates compatibility of the observed data with a specified null hypothesis under the model and testing framework. It is not an effect-size measure. A very small p-value can accompany either a relatively small or a large effect, depending on sample size and variability.

Full analysis set

The registry defines the FAS as all randomised participants. This preserves the randomized assignment as the foundation for the efficacy analysis. The posted records also state that the number analyzed consists of participants with available data, which makes the handling of unavailable measurements an important part of interpreting the results.

9. Estimands: Treatment Policy vs Hypothetical

One of the most statistically informative features of the STEP 4 registry results is that the two primary analyses are associated with different estimand descriptions.

Treatment policy estimand

The ANCOVA analysis is labeled as a treatment policy estimand. Conceptually, a treatment-policy estimand asks about the treatment effect according to a policy that does not simply remove participants from the treatment comparison when an intercurrent event occurs.

Hypothetical estimand

The MMRM analysis is labeled as a hypothetical estimand. Conceptually, this asks what the treatment effect would be under a specified hypothetical scenario after an intercurrent event.

The exact intercurrent-event strategy and complete estimand definitions are not reported in the ClinicalTrials.gov record. Therefore, these concepts should be used to understand why the analyses can differ without assigning an unstated clinical scenario to either analysis.

10. Statistical Methods Explained

Why was ANCOVA used?

The registry explicitly reports ANCOVA for the primary endpoint. ANCOVA is a linear-model approach that estimates a treatment difference while allowing the model to account for specified covariates. The ClinicalTrials.gov record does not identify the individual covariates, so the analysis should not be expanded beyond the documented method.

Why was MMRM used?

MMRM is designed for repeated measurements. When body weight is observed longitudinally, measurements from the same participant are statistically correlated. A mixed model can represent this within-participant structure while estimating treatment differences over time.

What does a treatment difference of -14.75 mean?

It means that the estimated difference between the semaglutide 2.4 mg and placebo groups for the specified change-from-randomisation endpoint was -14.75 percentage points under the ANCOVA analysis. The negative direction indicates that the semaglutide group had the lower estimated value for the defined change measure.

What does the confidence interval tell us?

The ANCOVA 95% CI is -16.00 to -13.50. The MMRM 95% CI is -16.52 to -14.13. These intervals describe uncertainty around their respective model-based treatment estimates. They are not individual-patient prediction intervals.

Why is the p-value not the same thing as effect size?

A p-value answers a hypothesis-testing question. The treatment difference answers an effect-size question. The confidence interval adds information about precision. A complete statistical interpretation therefore considers all three rather than treating the p-value as a measure of treatment magnitude.

Why do the ANCOVA and MMRM estimates differ?

The models use different statistical structures, and the registry assigns them different estimand descriptions. ANCOVA is a linear-model analysis, while MMRM explicitly handles repeated measurements. A difference between their estimates is therefore not, by itself, evidence of a contradiction.

What does the FAS contribute to interpretation?

The FAS is defined as all randomised participants, preserving the randomized trial framework. However, the posted analysis records also specify that the number analyzed is participants with available data. This makes the distinction between randomized population and observed data important when considering missing measurements.

11. Primary Endpoint Interpretation in Statistical Context

Effect estimate

The ANCOVA treatment difference was -14.75 percentage points. The MMRM treatment difference was -15.33 percentage points. These are absolute difference-scale estimates for the registered body-weight-change endpoint, not relative ratios.

Precision

The ANCOVA 95% CI spans -16.00 to -13.50, while the MMRM 95% CI spans -16.52 to -14.13. Each interval gives a range of values compatible with the corresponding estimate under its statistical framework.

Statistical evidence

Both analyses report P < 0.0001. This provides strong evidence against the relevant null hypothesis within the posted testing framework, but the p-value itself does not establish how large or clinically important the effect is.

Model dependence

The results are model-based. ANCOVA relies on its linear-model assumptions, while MMRM relies on assumptions concerning the longitudinal mean structure, within-participant dependence, and missing-data framework. The ClinicalTrials.gov record does not provide enough detail to evaluate every assumption empirically.

12. Missing Data and Estimand Considerations

Missing data are particularly important in longitudinal trials because participants may have different numbers of observations available by week 68. The registry's posted analysis records state that the FAS comprises all randomised participants but define the number analyzed as participants with available data.

For an ANCOVA analysis, the treatment estimate depends on the observations available for the modeled endpoint and on the assumptions used for handling unavailable data. For an MMRM, repeated observations can contribute information even when a participant does not have a complete sequence of measurements, depending on the model and missing-data assumptions.

Do not overinterpret the missing-data mechanism: the ClinicalTrials.gov record does not state the exact missing-data assumptions, imputation method, dropout pattern, or sensitivity-analysis framework. Those details should not be reconstructed from the posted point estimates alone.

13. Multiplicity and Hypothesis Testing

The registry identifies one registered primary endpoint and two formal statistical analyses for it. The posted ANCOVA analysis is explicitly labeled with a superiority hypothesis. The MMRM analysis has its hypothesis type recorded as Other / not stated.

ComponentRegistry-supported interpretation
Registered primary endpoints1
Primary endpoint analyses2
ANCOVA hypothesisSuperiority
MMRM hypothesisOther / not stated
Posted p-values<0.0001 for both analyses

The ClinicalTrials.gov record does not describe a multiplicity-adjustment procedure, alpha allocation, gatekeeping strategy, or hierarchical testing procedure. Accordingly, no additional multiplicity framework is inferred.

Statistical caution: two analyses of the same endpoint should not automatically be interpreted as two independent confirmatory tests. Their role depends on the prespecified statistical analysis plan, estimand framework, and multiplicity strategy. Those details are not contained in the ClinicalTrials.gov record.

14. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm and observation period. These are presented as affected participants divided by the corresponding at-risk population.

Arm / periodSerious adverse eventsInterpretive note
Semaglutide — run-in period21/902Affected / at risk
Semaglutide 2.4 — treatment period41/535Affected / at risk
Placebo — treatment period15/268Affected / at risk

These figures should be kept tied to their stated observation periods. In particular, the 21/902 figure refers to the semaglutide run-in period, whereas 41/535 and 15/268 refer to treatment-period observations for semaglutide 2.4 and placebo, respectively.

Safety interpretation: the serious-adverse-event figures are not directly interchangeable with the primary efficacy analysis. They have different observation-period definitions and denominators, and the ClinicalTrials.gov record does not provide a formal between-arm statistical test for serious adverse events.

15. What the Treatment Difference Does — and Does Not — Mean

Difference-scale interpretation

A treatment difference of -14.75 means the estimated difference between semaglutide 2.4 mg and placebo for the specified body-weight-change endpoint is 14.75 percentage points in the negative direction. It is a population-level statistical estimate, not a statement that every participant experienced the same effect.

Confidence interval interpretation

The 95% CI of -16.00 to -13.50 indicates the statistical uncertainty around the ANCOVA estimate. The corresponding MMRM interval is -16.52 to -14.13. Neither interval describes the distribution of individual patient responses.

P-value interpretation

A p-value of <0.0001 is evidence against the relevant null hypothesis under the analysis framework. It does not mean there is a <0.0001 probability that the observed result occurred by chance, and it does not quantify treatment benefit.

Estimand interpretation

The ANCOVA result is labeled a treatment policy estimand and the MMRM result a hypothetical estimand. These labels matter because an estimand defines the precise treatment-effect question being estimated. The ClinicalTrials.gov record does not provide enough information to specify the full intercurrent-event strategies beyond those labels.

16. Trial Timeline

June 4, 2018

Trial start

The registry lists June 4, 2018 as the study start date.

Phase 3

Randomized parallel-group evaluation

The study is classified as a phase 3, randomized, parallel trial with quadruple masking and treatment as its primary purpose.

Week 20 → Week 68

Primary endpoint window

The registered primary endpoint measures change from randomisation at week 20 to week 68 in body weight.

February 22, 2020

Primary completion

The registry lists February 22, 2020 as the primary completion date.

17. Design Features That Matter Statistically

ConceptHow it appears in STEP 4Why it matters
Randomization Allocation is randomized Creates the foundation for comparing treatment groups while reducing systematic allocation differences.
Parallel design Two parallel arms Participants remain associated with their randomized treatment comparison rather than switching through a crossover design.
Quadruple masking Masking is recorded as quadruple Masking can reduce the potential influence of treatment knowledge on trial conduct and assessment.
FAS All randomised participants Maintains randomization as the basis for the primary efficacy population.
ANCOVA Primary treatment-policy analysis Provides a linear-model estimate of the treatment difference for the endpoint.
MMRM Primary hypothetical-estimand analysis Uses a repeated-measures framework suited to longitudinal observations.
Confidence intervals 95% two-sided intervals for both posted estimates Shows uncertainty around the point estimates rather than reporting only p-values.
Estimands Treatment policy and hypothetical Clarifies that apparently similar analyses can answer different treatment-effect questions.

18. Limitations and Interpretation Issues

19. Why This Trial Matters Statistically

STEP 4 is a useful teaching example because its registry results illustrate an issue that is central to modern clinical-trial statistics: the same clinical endpoint can be analyzed under different statistical models and different estimand frameworks.

Statistical conceptSTEP 4 example
RandomizationRandomized allocation in a phase 3 parallel-group design
BlindingQuadruple masking
Continuous differenceTreatment difference reported in percentage-point units
ANCOVAPosted analysis using a linear-model framework
MMRMPosted longitudinal mixed-model analysis
Confidence intervals95% two-sided CIs for both primary analyses
P-values<0.0001 for both posted primary analyses
EstimandsTreatment policy versus hypothetical
Analysis populationFull analysis set of all randomised participants, with participants with available data analyzed
Missing dataRelevant because the posted analysis specifies participants with available data
Safety analysisSerious adverse events reported separately by arm and observation period

20. A Practical Reading Strategy for the Results

A statistically disciplined reading of the STEP 4 results can proceed in a fixed order.

  1. Identify the estimand: determine whether the reported estimate corresponds to the treatment-policy or hypothetical analysis.
  2. Identify the effect measure: here the registry reports a treatment difference rather than a hazard ratio or ratio measure.
  3. Read the point estimate: the ANCOVA estimate is -14.75 and the MMRM estimate is -15.33.
  4. Read the confidence interval: assess the uncertainty surrounding each estimate rather than relying on the p-value alone.
  5. Read the p-value: interpret it as evidence against the relevant null hypothesis, not as a measure of effect size.
  6. Check the analysis population: both records identify the FAS and participants with available data.
  7. Check model assumptions and missing-data strategy: recognize that the registry summary does not provide enough detail to fully evaluate them.
The central statistical lesson: a trial result is not just a number. The estimate, confidence interval, p-value, analysis population, model, and estimand together define what the reported result actually means.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

Continue through the Clinical Biostats statistical tutorials

Explore the statistical concepts that connect trial design, repeated-measures analysis, confidence intervals, hypothesis testing, and randomized comparisons.

24. Record Summary

STEP 4 provides a useful example of how a randomized phase 3 trial can produce complementary statistical analyses for the same primary endpoint. The registry reports a treatment difference of -14.75 percentage points with a 95% CI of -16.00 to -13.50 and P < 0.0001 using ANCOVA, and a treatment difference of -15.33 percentage points with a 95% CI of -16.52 to -14.13 and P < 0.0001 using MMRM.

The statistical distinction between the analyses is important. The ANCOVA result is identified as a treatment policy estimand, whereas the MMRM result is identified as a hypothetical estimand. The FAS comprises all randomised participants, while the posted analysis records specify that participants with available data were analyzed. These details help define the scope of the reported estimates and caution against interpreting either number independently of its statistical framework.

Clinical Biostats methodology: A trial-results page should not merely repeat an effect estimate. The objective is to explain the statistical question, analysis population, effect measure, confidence interval, p-value, model, estimand, and limitations so that readers can understand exactly what the reported result does and does not establish.