This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
STEP 4 was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating semaglutide versus placebo in people with overweight or obesity. The registry reports 902 enrolled participants and a primary endpoint concerning change in body weight from randomisation at week 20 to week 68.
| Feature | STEP 4 |
|---|---|
| Trial name | STEP 4 |
| NCT identifier | NCT03548987 |
| Phase | Phase 3 |
| Status | Completed |
| Therapeutic area | Endocrinology |
| Conditions | Metabolism and Nutrition Disorder; Obesity |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 902 |
| Interventions | Semaglutide; placebo |
| Results posted | Yes |
| Outcome measures posted | 37 |
| Statistical analyses posted | 2 |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
2. Clinical Question
The primary statistical question was whether the change in body weight from randomisation at week 20 to week 68 differed between participants assigned to semaglutide 2.4 mg and those assigned to placebo.
Population
The registry describes the study as investigating how well semaglutide works in people suffering from overweight or obesity. The listed conditions are metabolism and nutrition disorder and obesity.
Intervention
Semaglutide. The posted primary statistical analyses identify the treatment comparison specifically as semaglutide 2.4 mg versus placebo.
Comparator
Placebo.
Primary question
What is the treatment difference in change from randomisation to week 68 in body weight, comparing semaglutide 2.4 mg with placebo?
3. Trial Design
Participants were assigned to parallel treatment groups.
The registry identifies the study as quadruple-masked.
The study was designed to evaluate treatment effects.
Start: June 4, 2018. Primary completion: February 22, 2020.
Semaglutide 2.4 mg
- Semaglutide
- Included in the primary comparison of change in body weight
- Primary analyses used the full analysis set framework
Placebo
- Placebo
- Compared with semaglutide 2.4 mg for the primary endpoint
- Primary analyses used the full analysis set framework
4. Endpoint Framework
The registry identifies one registered primary endpoint and posts two formal statistical analyses for that endpoint. The two analyses use different statistical frameworks and are associated with different estimand descriptions.
| Endpoint | Registry time frame | Type | Analyses posted |
|---|---|---|---|
| Change From Randomisation to Week 68 in Body Weight (%) | Randomisation (week 20) to week 68 | Binary, as classified in the ClinicalTrials.gov record | ANCOVA; MMRM |
Registered endpoint definition
The registry states that change in body weight from baseline at week 20 to week 68 is presented. The endpoint was evaluated based on data from both in-trial and on-treatment observation periods.
The registry defines the in-trial observation period as the uninterrupted time interval from the start of run-in at week 0 to the last trial-related subject-site contact at week 75. The registry also defines an on-treatment observation period, described as including all time intervals in which participants were receiving treatment; the registry-reported endpoint definition is truncated after that point, so no additional wording is inferred here.
5. Statistical Analysis Populations
Both posted primary analyses identify the full analysis set (FAS) as the overall analysis population. The registry describes the FAS as comprising all randomised participants, with the number analyzed defined as participants with available data.
| Population | Registry description | Role in posted analysis |
|---|---|---|
| Full analysis set | All randomised participants | Primary efficacy analysis population |
| Participants with available data | Number analyzed is defined as participants with available data | Contributes observations to the posted analysis |
This distinction is important statistically. “All randomised participants” describes the FAS definition, whereas the posted analysis record separately states that the number analyzed is participants with available data. The registry therefore provides the population framework but does not, in the ClinicalTrials.gov record, provide a complete accounting of every missing observation.
6. Primary Results
ClinicalTrials.gov posts two formal statistical analyses for the same primary endpoint: one using ANCOVA and one using an MMRM. Both compare semaglutide 2.4 mg with placebo, but the analysis notes associate the ANCOVA result with a treatment policy estimand and the MMRM result with a hypothetical estimand.
ANCOVA Analysis
Treatment difference in change in body weight
95% CI: -16.00 to -13.50 · P < 0.0001
Semaglutide 2.4 mg vs placebo · Superiority hypothesis · Treatment policy estimand
| Characteristic | Posted result |
|---|---|
| Endpoint | Change From Randomisation to Week 68 in Body Weight (%) |
| Time frame | Randomisation (week 20) to week 68 |
| Comparison | Semaglutide 2.4 mg vs placebo |
| Method | ANCOVA |
| Effect measure | Treatment difference |
| Estimate | -14.75 percentage points |
| 95% CI | -16.00 to -13.50 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Estimand note | Treatment policy estimand |
The reported treatment difference of -14.75 percentage points means that the estimated change in body weight favored semaglutide 2.4 mg relative to placebo by 14.75 percentage points under the posted ANCOVA analysis and its treatment-policy estimand.
The negative sign is important because the effect is expressed as the treatment difference for change in body weight. It does not mean that every individual participant experienced exactly a 14.75-percentage-point difference, nor does it describe the response of an individual patient.
The 95% confidence interval, -16.00 to -13.50, describes the statistical uncertainty around the estimated treatment difference under the analysis framework. It is an interval for the estimated population-level treatment effect, not a range containing the individual treatment effects of 95% of participants.
The p-value of <0.0001 addresses evidence against the null hypothesis within the posted statistical framework. It does not measure the size of the treatment effect, the probability that the treatment is effective, or the clinical importance of the result. Effect size and uncertainty are better conveyed by the estimate and its confidence interval.
Because the analysis is identified as a treatment-policy estimand, its interpretation is tied to the specified treatment-policy question rather than simply describing outcomes while participants remain continuously on treatment. The ClinicalTrials.gov record does not provide the complete estimand definition beyond this classification.
MMRM Analysis
Treatment difference in change in body weight
95% CI: -16.52 to -14.13 · P < 0.0001
Semaglutide 2.4 mg vs placebo · Hypothetical estimand
| Characteristic | Posted result |
|---|---|
| Endpoint | Change From Randomisation to Week 68 in Body Weight (%) |
| Time frame | Randomisation (week 20) to week 68 |
| Comparison | Semaglutide 2.4 mg vs placebo |
| Method | MMRM (mixed model repeated measurement) |
| Effect measure | Treatment difference |
| Estimate | -15.33 percentage points |
| 95% CI | -16.52 to -14.13 |
| P-value | <0.0001 |
| Hypothesis type | Other / not stated |
| Estimand note | Hypothetical estimand |
The MMRM estimate of -15.33 percentage points is the estimated treatment difference from randomisation to week 68 under the posted longitudinal model and hypothetical estimand.
MMRM uses the repeated-measure structure of longitudinal data rather than reducing the entire follow-up history to a single observation without regard to intermediate measurements. That makes the method particularly relevant when an endpoint is measured repeatedly over time.
The 95% confidence interval of -16.52 to -14.13 quantifies uncertainty around the model-based treatment-difference estimate. It does not mean that 95% of participants had changes within that interval.
The p-value of <0.0001 indicates strong statistical evidence under the specified model and hypothesis-testing framework, but it does not quantify the magnitude or practical importance of the treatment effect.
The MMRM analysis is explicitly labeled as a hypothetical estimand. Consequently, its interpretation depends on the hypothetical treatment-policy scenario represented by that estimand. The ClinicalTrials.gov record does not provide enough detail to reconstruct the full hypothetical scenario, so no additional assumptions are imposed here.
7. Comparing the Two Primary Analyses
The two posted estimates are not identical, but they are directionally consistent. The ANCOVA treatment difference is -14.75, while the MMRM treatment difference is -15.33. Both confidence intervals lie entirely below zero and both reported p-values are <0.0001.
| Feature | ANCOVA | MMRM |
|---|---|---|
| Estimate | -14.75 | -15.33 |
| 95% CI | -16.00 to -13.50 | -16.52 to -14.13 |
| P-value | <0.0001 | <0.0001 |
| Effect measure | Treatment difference | Treatment difference |
| Analysis family | Linear model | Longitudinal / mixed model |
| Estimand note | Treatment policy | Hypothetical |
| Analysis population | FAS; participants with available data | FAS; participants with available data |
The difference between the point estimates should not be interpreted as evidence that one method is “correct” and the other is “incorrect.” They answer related but not necessarily identical statistical questions because the methods and estimands differ. ANCOVA is a cross-sectional linear-model approach to the specified endpoint, whereas MMRM is designed for repeated measurements and can model the longitudinal trajectory.
8. Statistical Methodology
ANCOVA
ANCOVA, or analysis of covariance, is a linear-model framework commonly used when the outcome is continuous and the analysis needs to compare treatment groups while accounting for one or more covariates. In this trial, the registry explicitly identifies ANCOVA as the method for the primary treatment-policy analysis.
The treatment coefficient represents an adjusted between-group difference under the fitted model. The exact covariates and model specification are not reported in the ClinicalTrials.gov record and therefore are not inferred.
MMRM: Mixed Model for Repeated Measures
MMRM is a longitudinal modeling framework for repeated outcome measurements. Rather than treating each measurement as an isolated observation, the model accounts for the fact that measurements from the same participant are related.
The exact covariance structure, fixed effects, time points, and estimation details are not provided in the ClinicalTrials.gov record. The important point for interpretation is that the posted MMRM uses repeated measurements rather than only a single cross-sectional comparison.
Treatment differences
The reported effect measure for both analyses is treatment difference. This is a difference-scale measure rather than a ratio or hazard ratio. The reported outcome unit is “percentage point,” so the estimate is interpreted on that scale.
Confidence intervals
A 95% confidence interval describes uncertainty associated with the estimated treatment difference under the specified statistical model and sampling framework. Narrower intervals generally indicate greater statistical precision than wider intervals, although precision and clinical importance are separate concepts.
P-values
The p-value evaluates compatibility of the observed data with a specified null hypothesis under the model and testing framework. It is not an effect-size measure. A very small p-value can accompany either a relatively small or a large effect, depending on sample size and variability.
Full analysis set
The registry defines the FAS as all randomised participants. This preserves the randomized assignment as the foundation for the efficacy analysis. The posted records also state that the number analyzed consists of participants with available data, which makes the handling of unavailable measurements an important part of interpreting the results.
9. Estimands: Treatment Policy vs Hypothetical
One of the most statistically informative features of the STEP 4 registry results is that the two primary analyses are associated with different estimand descriptions.
Treatment policy estimand
The ANCOVA analysis is labeled as a treatment policy estimand. Conceptually, a treatment-policy estimand asks about the treatment effect according to a policy that does not simply remove participants from the treatment comparison when an intercurrent event occurs.
Hypothetical estimand
The MMRM analysis is labeled as a hypothetical estimand. Conceptually, this asks what the treatment effect would be under a specified hypothetical scenario after an intercurrent event.
The exact intercurrent-event strategy and complete estimand definitions are not reported in the ClinicalTrials.gov record. Therefore, these concepts should be used to understand why the analyses can differ without assigning an unstated clinical scenario to either analysis.
10. Statistical Methods Explained
Why was ANCOVA used?
The registry explicitly reports ANCOVA for the primary endpoint. ANCOVA is a linear-model approach that estimates a treatment difference while allowing the model to account for specified covariates. The ClinicalTrials.gov record does not identify the individual covariates, so the analysis should not be expanded beyond the documented method.
Why was MMRM used?
MMRM is designed for repeated measurements. When body weight is observed longitudinally, measurements from the same participant are statistically correlated. A mixed model can represent this within-participant structure while estimating treatment differences over time.
What does a treatment difference of -14.75 mean?
It means that the estimated difference between the semaglutide 2.4 mg and placebo groups for the specified change-from-randomisation endpoint was -14.75 percentage points under the ANCOVA analysis. The negative direction indicates that the semaglutide group had the lower estimated value for the defined change measure.
What does the confidence interval tell us?
The ANCOVA 95% CI is -16.00 to -13.50. The MMRM 95% CI is -16.52 to -14.13. These intervals describe uncertainty around their respective model-based treatment estimates. They are not individual-patient prediction intervals.
Why is the p-value not the same thing as effect size?
A p-value answers a hypothesis-testing question. The treatment difference answers an effect-size question. The confidence interval adds information about precision. A complete statistical interpretation therefore considers all three rather than treating the p-value as a measure of treatment magnitude.
Why do the ANCOVA and MMRM estimates differ?
The models use different statistical structures, and the registry assigns them different estimand descriptions. ANCOVA is a linear-model analysis, while MMRM explicitly handles repeated measurements. A difference between their estimates is therefore not, by itself, evidence of a contradiction.
What does the FAS contribute to interpretation?
The FAS is defined as all randomised participants, preserving the randomized trial framework. However, the posted analysis records also specify that the number analyzed is participants with available data. This makes the distinction between randomized population and observed data important when considering missing measurements.
11. Primary Endpoint Interpretation in Statistical Context
The ANCOVA treatment difference was -14.75 percentage points. The MMRM treatment difference was -15.33 percentage points. These are absolute difference-scale estimates for the registered body-weight-change endpoint, not relative ratios.
The ANCOVA 95% CI spans -16.00 to -13.50, while the MMRM 95% CI spans -16.52 to -14.13. Each interval gives a range of values compatible with the corresponding estimate under its statistical framework.
Both analyses report P < 0.0001. This provides strong evidence against the relevant null hypothesis within the posted testing framework, but the p-value itself does not establish how large or clinically important the effect is.
The results are model-based. ANCOVA relies on its linear-model assumptions, while MMRM relies on assumptions concerning the longitudinal mean structure, within-participant dependence, and missing-data framework. The ClinicalTrials.gov record does not provide enough detail to evaluate every assumption empirically.
12. Missing Data and Estimand Considerations
Missing data are particularly important in longitudinal trials because participants may have different numbers of observations available by week 68. The registry's posted analysis records state that the FAS comprises all randomised participants but define the number analyzed as participants with available data.
For an ANCOVA analysis, the treatment estimate depends on the observations available for the modeled endpoint and on the assumptions used for handling unavailable data. For an MMRM, repeated observations can contribute information even when a participant does not have a complete sequence of measurements, depending on the model and missing-data assumptions.
13. Multiplicity and Hypothesis Testing
The registry identifies one registered primary endpoint and two formal statistical analyses for it. The posted ANCOVA analysis is explicitly labeled with a superiority hypothesis. The MMRM analysis has its hypothesis type recorded as Other / not stated.
| Component | Registry-supported interpretation |
|---|---|
| Registered primary endpoints | 1 |
| Primary endpoint analyses | 2 |
| ANCOVA hypothesis | Superiority |
| MMRM hypothesis | Other / not stated |
| Posted p-values | <0.0001 for both analyses |
The ClinicalTrials.gov record does not describe a multiplicity-adjustment procedure, alpha allocation, gatekeeping strategy, or hierarchical testing procedure. Accordingly, no additional multiplicity framework is inferred.
14. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm and observation period. These are presented as affected participants divided by the corresponding at-risk population.
| Arm / period | Serious adverse events | Interpretive note |
|---|---|---|
| Semaglutide — run-in period | 21/902 | Affected / at risk |
| Semaglutide 2.4 — treatment period | 41/535 | Affected / at risk |
| Placebo — treatment period | 15/268 | Affected / at risk |
These figures should be kept tied to their stated observation periods. In particular, the 21/902 figure refers to the semaglutide run-in period, whereas 41/535 and 15/268 refer to treatment-period observations for semaglutide 2.4 and placebo, respectively.
15. What the Treatment Difference Does — and Does Not — Mean
A treatment difference of -14.75 means the estimated difference between semaglutide 2.4 mg and placebo for the specified body-weight-change endpoint is 14.75 percentage points in the negative direction. It is a population-level statistical estimate, not a statement that every participant experienced the same effect.
The 95% CI of -16.00 to -13.50 indicates the statistical uncertainty around the ANCOVA estimate. The corresponding MMRM interval is -16.52 to -14.13. Neither interval describes the distribution of individual patient responses.
A p-value of <0.0001 is evidence against the relevant null hypothesis under the analysis framework. It does not mean there is a <0.0001 probability that the observed result occurred by chance, and it does not quantify treatment benefit.
The ANCOVA result is labeled a treatment policy estimand and the MMRM result a hypothetical estimand. These labels matter because an estimand defines the precise treatment-effect question being estimated. The ClinicalTrials.gov record does not provide enough information to specify the full intercurrent-event strategies beyond those labels.
16. Trial Timeline
Trial start
The registry lists June 4, 2018 as the study start date.
Randomized parallel-group evaluation
The study is classified as a phase 3, randomized, parallel trial with quadruple masking and treatment as its primary purpose.
Primary endpoint window
The registered primary endpoint measures change from randomisation at week 20 to week 68 in body weight.
Primary completion
The registry lists February 22, 2020 as the primary completion date.
17. Design Features That Matter Statistically
| Concept | How it appears in STEP 4 | Why it matters |
|---|---|---|
| Randomization | Allocation is randomized | Creates the foundation for comparing treatment groups while reducing systematic allocation differences. |
| Parallel design | Two parallel arms | Participants remain associated with their randomized treatment comparison rather than switching through a crossover design. |
| Quadruple masking | Masking is recorded as quadruple | Masking can reduce the potential influence of treatment knowledge on trial conduct and assessment. |
| FAS | All randomised participants | Maintains randomization as the basis for the primary efficacy population. |
| ANCOVA | Primary treatment-policy analysis | Provides a linear-model estimate of the treatment difference for the endpoint. |
| MMRM | Primary hypothetical-estimand analysis | Uses a repeated-measures framework suited to longitudinal observations. |
| Confidence intervals | 95% two-sided intervals for both posted estimates | Shows uncertainty around the point estimates rather than reporting only p-values. |
| Estimands | Treatment policy and hypothetical | Clarifies that apparently similar analyses can answer different treatment-effect questions. |
18. Limitations and Interpretation Issues
- Limited registry detail: the ClinicalTrials.gov record provides the principal estimates and methods but do not contain the complete statistical analysis plan.
- Endpoint classification: the registry data classify the primary endpoint as binary even though the outcome unit is percentage point and the effect measure is a treatment difference. This page preserves the registry-reported classification rather than silently changing it.
- Incomplete estimand definitions: the ANCOVA result is labeled treatment policy and the MMRM result hypothetical, but the complete intercurrent-event strategies are not reported.
- Missing-data assumptions: the registry states that participants with available data were analyzed but does not provide enough information here to identify the complete missing-data mechanism or sensitivity-analysis strategy.
- Model specification: the ClinicalTrials.gov record does not specify the covariates for ANCOVA or the covariance structure and fixed-effect specification for MMRM.
- Two analyses of one endpoint: the presence of two formal analyses does not by itself establish that they represent independent confirmatory hypothesis tests. The multiplicity strategy is not provided.
- Safety denominators differ by period: serious adverse events are reported for a run-in period and treatment periods with different at-risk populations, so the figures should not be compared without respecting their observation periods.
- No individual-response distribution: the ClinicalTrials.gov record provides treatment-level estimates and confidence intervals, not a distribution of individual treatment responses.
- No subgroup results in the ClinicalTrials.gov record: subgroup estimates, forest plots, or subgroup interaction tests are not included and therefore are not presented here.
19. Why This Trial Matters Statistically
STEP 4 is a useful teaching example because its registry results illustrate an issue that is central to modern clinical-trial statistics: the same clinical endpoint can be analyzed under different statistical models and different estimand frameworks.
| Statistical concept | STEP 4 example |
|---|---|
| Randomization | Randomized allocation in a phase 3 parallel-group design |
| Blinding | Quadruple masking |
| Continuous difference | Treatment difference reported in percentage-point units |
| ANCOVA | Posted analysis using a linear-model framework |
| MMRM | Posted longitudinal mixed-model analysis |
| Confidence intervals | 95% two-sided CIs for both primary analyses |
| P-values | <0.0001 for both posted primary analyses |
| Estimands | Treatment policy versus hypothetical |
| Analysis population | Full analysis set of all randomised participants, with participants with available data analyzed |
| Missing data | Relevant because the posted analysis specifies participants with available data |
| Safety analysis | Serious adverse events reported separately by arm and observation period |
20. A Practical Reading Strategy for the Results
A statistically disciplined reading of the STEP 4 results can proceed in a fixed order.
- Identify the estimand: determine whether the reported estimate corresponds to the treatment-policy or hypothetical analysis.
- Identify the effect measure: here the registry reports a treatment difference rather than a hazard ratio or ratio measure.
- Read the point estimate: the ANCOVA estimate is -14.75 and the MMRM estimate is -15.33.
- Read the confidence interval: assess the uncertainty surrounding each estimate rather than relying on the p-value alone.
- Read the p-value: interpret it as evidence against the relevant null hypothesis, not as a measure of effect size.
- Check the analysis population: both records identify the FAS and participants with available data.
- Check model assumptions and missing-data strategy: recognize that the registry summary does not provide enough detail to fully evaluate them.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: NCT03548987 — STEP 4. Official trial registry record and source for the trial characteristics and statistical analyses summarized on this page.
- PubMed: PMID 33755728.
- PubMed: PMID 38698650.
- PubMed: PMID 37605636.
- PubMed: PMID 36200477.
- PubMed: PMID 35724304.
Continue through the Clinical Biostats statistical tutorials
Explore the statistical concepts that connect trial design, repeated-measures analysis, confidence intervals, hypothesis testing, and randomized comparisons.
24. Record Summary
STEP 4 provides a useful example of how a randomized phase 3 trial can produce complementary statistical analyses for the same primary endpoint. The registry reports a treatment difference of -14.75 percentage points with a 95% CI of -16.00 to -13.50 and P < 0.0001 using ANCOVA, and a treatment difference of -15.33 percentage points with a 95% CI of -16.52 to -14.13 and P < 0.0001 using MMRM.
The statistical distinction between the analyses is important. The ANCOVA result is identified as a treatment policy estimand, whereas the MMRM result is identified as a hypothetical estimand. The FAS comprises all randomised participants, while the posted analysis records specify that participants with available data were analyzed. These details help define the scope of the reported estimates and caution against interpreting either number independently of its statistical framework.