This page separates reported trial results from statistical interpretation. The numerical results and trial characteristics presented here are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
STEP 5 was a randomized, parallel-group, quadruple-masked phase 3 trial comparing semaglutide with placebo in people with overweight or obesity. The trial enrolled 304 participants, with 152 participants assigned to each arm, and evaluated two registered primary endpoints through week 104.
| Feature | STEP 5 |
|---|---|
| Trial name | STEP 5 |
| Phase | Phase 3 |
| Status | COMPLETED |
| Therapeutic area | Endocrinology |
| Conditions | Overweight; Obesity |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 304 |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03693430 |
2. Clinical Question
The primary statistical question was whether semaglutide differed from placebo with respect to the two registered primary outcomes: percentage change from baseline at week 104 in body weight and achievement of body-weight reduction greater than or equal to 5% at week 104.
Population
People with overweight or obesity enrolled in the STEP 5 phase 3 trial.
Intervention
Semaglutide, with the reported serious-adverse-event analysis identifying the treatment arm as semaglutide 2.4 mg.
Comparator
Placebo, identified in the registry results as the placebo arm.
Primary question
How did semaglutide compare with placebo at week 104 for percentage change in body weight and achievement of at least 5% body-weight reduction?
3. Trial Design
Semaglutide
- Semaglutide 2.4 mg as identified in the serious-adverse-event results
- Primary assessment at week 104
- Primary analyses included ANCOVA, MMRM, and logistic regression
Placebo
- Placebo comparator
- Primary assessment at week 104
- Primary analyses included ANCOVA, MMRM, and logistic-regression comparisons
4. Endpoints
The registry lists two primary endpoints. Both had posted formal statistical analyses and both had an estimate, a two-sided 95% confidence interval, and a p-value in the ClinicalTrials.gov record.
| Primary endpoint | Registry time frame | Endpoint type | Principal analysis methods reported |
|---|---|---|---|
| Percentage Change From Baseline (Week 0) to Week 104 in Body Weight | From Baseline (Week 0) to Week 104 | Binary in the registry inference | ANCOVA; MMRM |
| Number of Participants Who Achieved (Yes/no): Body Weight Reduction More Than or Equal to 5% | At Week 104 | Binary | Logistic regression; MMRM |
Endpoint definitions
For the first endpoint, the registry states that percentage change in body weight for both the in-trial and on-treatment observation periods from baseline (week 0) to week 104 is presented. The outcome measure was evaluated based on data from both in-trial and on-treatment periods.
For the second endpoint, the registry states that the number of participants who achieved greater than or equal to 5% weight loss at 104 weeks is presented. In the reported data, “Yes” infers participants who achieved greater than or equal to 5% weight loss, whereas “No” infers participants who did not achieve greater than or equal to 5% weight loss.
5. Analysis Populations and Estimands
The ClinicalTrials.gov record identifies the full analysis set (FAS) as including all randomized participants according to the intention-to-treat principle. This is important because the treatment comparison remains anchored to randomized assignment rather than being restricted to participants who completed treatment exactly as planned.
The distinction between a treatment-policy and hypothetical estimand is statistically important. A treatment-policy estimand asks about the treatment strategy while retaining the consequences of treatment discontinuation or other post-randomization events according to the specified strategy. A hypothetical estimand instead targets what the outcome would look like under a hypothetical scenario in which specified post-randomization events did not occur.
6. Results: Percentage Change in Body Weight
The first primary endpoint was the percentage change from baseline (week 0) to week 104 in body weight. Two formal analyses were posted for this endpoint: ANCOVA under a treatment-policy estimand and MMRM under a hypothetical estimand.
ANCOVA: Treatment-Policy Estimand
Estimated treatment difference
95% CI: -15.33 to -9.77 · P < .0001
Semaglutide 2.4 mg vs Placebo
| Feature | Reported result |
|---|---|
| Outcome | Percentage Change From Baseline (Week 0) to Week 104 in Body Weight |
| Method | ANCOVA |
| Effect measure | Treatment difference |
| Estimate | -12.55 |
| 95% CI | -15.33 to -9.77 |
| P-value | <.0001 |
| Hypothesis | Superiority |
| Estimand | Treatment policy |
The registry states that week 104 responses were analyzed using an analysis of covariance model with randomized treatment as a factor and baseline body weight as a covariate. Missing observations were multiple (x1000) imputed from retrieved subjects of the same randomized treatment arm.
The estimate of -12.55 is the reported treatment difference in percentage change in body weight between semaglutide 2.4 mg and placebo under the treatment-policy analysis. Its negative direction indicates a lower percentage-change value for semaglutide relative to placebo.
The estimate does not mean that every participant experienced exactly a 12.55 percentage-point difference. It is a group-level adjusted treatment contrast from the ANCOVA model.
The two-sided 95% confidence interval of -15.33 to -9.77 describes the statistical uncertainty around the estimated treatment difference under the analysis framework. Because the interval remains below zero, it is consistent with a treatment difference favoring the semaglutide direction for this endpoint.
The p-value of <.0001 addresses the evidence against the null hypothesis under the specified test. It does not measure the size of the treatment effect, clinical importance, or probability that the treatment is effective.
The interpretation also depends on the treatment-policy estimand and the multiple-imputation approach specified in the registry analysis description. It should not be silently substituted for the separate hypothetical-estimand MMRM result.
MMRM: Hypothetical Estimand
Estimated treatment difference
95% CI: -18.64 to -13.45 · P <0.0001
Semaglutide 2.4 mg vs Placebo
| Feature | Reported result |
|---|---|
| Outcome | Percentage Change From Baseline (Week 0) to Week 104 in Body Weight |
| Method | MMRM (mixed model for repeated measures) |
| Effect measure | Treatment difference |
| Estimate | -16.05 |
| 95% CI | -18.64 to -13.45 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Estimand | Hypothetical |
The registry states that all responses prior to first discontinuation of treatment, or initiation of other anti-obesity medication or bariatric surgery, were included in a mixed model for repeated measurements with randomized treatment as factor and baseline body weight as covariate, all nested within visit.
The reported -16.05 is the treatment difference from the MMRM analysis under a hypothetical estimand. It therefore answers a different statistical question from the treatment-policy ANCOVA estimate of -12.55.
The 95% confidence interval of -18.64 to -13.45 quantifies uncertainty around this model-based longitudinal treatment contrast. The entire interval is below zero.
The p-value of <0.0001 indicates strong evidence against the corresponding null hypothesis within this analysis. It is not an effect-size metric and should not be used as a substitute for examining the estimated difference and its confidence interval.
Because this is an MMRM, interpretation also depends on the longitudinal model and its assumptions about the repeated measurements and the hypothetical treatment scenario. The MMRM result should therefore be understood as an estimand-specific model result rather than as a universal estimate of what every participant would have experienced.
7. Results: Achievement of at Least 5% Weight Reduction
The second primary endpoint was binary: whether a participant achieved body-weight reduction greater than or equal to 5% at week 104. Two formal analyses were posted: logistic regression under a treatment-policy estimand and MMRM under a hypothetical estimand.
Logistic Regression: Treatment-Policy Estimand
Odds ratio
95% CI: 2.95 to 8.42 · P <0.0001
Semaglutide 2.4 mg vs Placebo
| Feature | Reported result |
|---|---|
| Outcome | Body Weight Reduction More Than or Equal to 5% at Week 104 |
| Method | Logistic regression |
| Effect measure | Odds Ratio (OR) |
| Estimate | 4.99 |
| 95% CI | 2.95 to 8.42 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Estimand | Treatment policy |
An odds ratio of 4.99 means that the modeled odds of achieving at least 5% body-weight reduction were estimated to be 4.99 times as high with semaglutide 2.4 mg as with placebo under this analysis.
An odds ratio is not a risk ratio. It does not mean that the probability of achieving the endpoint was 4.99 times as high. The distinction becomes especially important when the outcome is not rare.
The 95% confidence interval of 2.95 to 8.42 expresses uncertainty around the estimated odds ratio. It remains above 1, the null value for an odds ratio.
The p-value of <0.0001 describes the evidence against the null hypothesis under the specified model. It does not describe the magnitude of the odds ratio or the probability that a particular participant will achieve the endpoint.
The analysis is also explicitly tied to a treatment-policy estimand, so its interpretation should not be merged with the separate MMRM result that uses a hypothetical estimand.
MMRM: Hypothetical Estimand
Odds ratio
95% CI: 10.04 to 32.49 · P <0.0001
Semaglutide 2.4 mg vs Placebo
| Feature | Reported result |
|---|---|
| Outcome | Body Weight Reduction More Than or Equal to 5% at Week 104 |
| Method | MMRM (mixed model for repeated measures) |
| Effect measure | Odds Ratio (OR) |
| Estimate | 18.06 |
| 95% CI | 10.04 to 32.49 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Estimand | Hypothetical |
The reported odds ratio of 18.06 indicates substantially higher modeled odds of achieving at least 5% body-weight reduction with semaglutide than placebo under the hypothetical-estimand MMRM analysis.
This is an odds ratio, not a probability ratio. An odds ratio of 18.06 cannot be read as saying that 18.06 times as many participants achieved the endpoint.
The 95% confidence interval of 10.04 to 32.49 indicates uncertainty around the estimate while remaining above the null value of 1.
The p-value of <0.0001 indicates strong evidence against the null hypothesis within this model. It does not quantify clinical magnitude or individual-level benefit.
The hypothetical estimand also matters: this analysis is not answering precisely the same question as the treatment-policy logistic-regression analysis, even though both concern the same week-104 binary endpoint.
8. Comparing the Two Primary Endpoints
| Primary endpoint | Method | Estimand | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Percentage change in body weight | ANCOVA | Treatment policy | Treatment difference | -12.55 | -15.33 to -9.77 | <.0001 |
| Percentage change in body weight | MMRM | Hypothetical | Treatment difference | -16.05 | -18.64 to -13.45 | <0.0001 |
| At least 5% weight reduction | Logistic regression | Treatment policy | Odds ratio | 4.99 | 2.95 to 8.42 | <0.0001 |
| At least 5% weight reduction | MMRM | Hypothetical | Odds ratio | 18.06 | 10.04 to 32.49 | <0.0001 |
The most important statistical distinction is not simply the numerical size of the four estimates. The analyses answer different questions. The two ANCOVA/logistic-regression results use a treatment-policy estimand, while the two MMRM results use a hypothetical estimand. Comparing those pairs without recognizing the estimand difference would make the statistical interpretation less precise.
9. Safety Results
The ClinicalTrials.gov record includes serious adverse events by randomized treatment arm. The ClinicalTrials.gov record does not supply a broader adverse-event table, so safety interpretation is limited to the reported serious-adverse-event counts.
| Safety measure | Semaglutide 2.4 mg | Placebo |
|---|---|---|
| Participants affected / at risk | 12 / 152 | 18 / 152 |
These counts describe serious adverse events in each randomized arm as reported in the registry. They do not establish that every observed event was caused by the assigned intervention, and they should not be converted into a comparative risk measure that was not reported in the ClinicalTrials.gov record.
10. Statistical Methodology
ANCOVA
ANCOVA, or analysis of covariance, was used for the treatment-policy analysis of percentage change in body weight. The registry specifies randomized treatment as a factor and baseline body weight as a covariate.
The practical purpose of the covariate is to account for baseline body weight when estimating the treatment contrast at the specified analysis time point.
For a continuous outcome such as percentage change in body weight, the treatment difference is naturally interpreted on the outcome scale. A negative treatment difference here means the modeled percentage-change outcome was lower in the semaglutide group than in the placebo group.
MMRM
MMRM, or mixed model for repeated measures, was used for both primary endpoints in the hypothetical-estimand analyses. The registry states that responses were modeled with randomized treatment as a factor and baseline body weight as a covariate, with terms nested within visit.
The principal advantage of a repeated-measures model is that it can use longitudinal information rather than treating each participant's week-104 observation as an isolated data point. The interpretation, however, depends on the model specification and the estimand it is intended to estimate.
Logistic regression
The treatment-policy analysis of achieving at least 5% weight reduction used logistic regression. Logistic regression is appropriate for a binary outcome such as “Yes” or “No.” Its natural effect measure is the odds ratio.
An OR above 1 indicates higher odds of the defined binary outcome in the semaglutide group under the fitted model. It should not be interpreted as a risk ratio without additional assumptions or calculations.
Intention-to-treat analysis
The ClinicalTrials.gov record states that the full analysis set included all randomized participants according to the intention-to-treat principle. This preserves the treatment assignment created by randomization and avoids redefining the primary comparison based solely on subsequent treatment behavior.
Superiority testing
All four registry-reported statistical analyses identify the hypothesis type as superiority. In a superiority framework, the statistical question is whether the treatment groups differ in the prespecified direction under the relevant model, rather than whether a treatment is merely no worse than a comparator by a predefined non-inferiority margin.
11. Statistical Methods Explained
Why was ANCOVA used for percentage change in body weight?
Percentage change in body weight is a continuous outcome, making a linear-model approach appropriate. ANCOVA allows the analysis to include randomized treatment as the principal factor while also accounting for baseline body weight as a covariate. The resulting treatment difference is adjusted for baseline body weight rather than being a simple unadjusted comparison of observed week-104 means.
Why was MMRM also used?
The MMRM analysis addresses repeated measurements over visits rather than focusing solely on a single endpoint measurement. In STEP 5, the registry text specifies a longitudinal model with randomized treatment, baseline body weight, and visit structure. The MMRM results therefore provide a different estimand-specific analysis of the same primary endpoint concepts.
What does an odds ratio of 4.99 mean?
An odds ratio of 4.99 means the modeled odds of achieving at least 5% weight reduction were 4.99 times as high in the semaglutide group as in the placebo group under that logistic-regression analysis. It does not mean that the probability was 4.99 times as high, because odds and probability are different quantities.
Why does the confidence interval matter?
A point estimate alone does not describe statistical uncertainty. The 95% confidence interval gives a range of values compatible with the model and sampling framework at the stated confidence level. For the ANCOVA estimate, the interval is -15.33 to -9.77; for the treatment-policy odds ratio, it is 2.95 to 8.42. These intervals communicate both direction and precision more fully than the p-values alone.
Why does the p-value not measure effect size?
The p-value measures how inconsistent the observed data are with a specified null hypothesis under the statistical model. It is affected by the size and variability of the dataset and does not tell the reader how large the treatment effect is. That is why the estimate and confidence interval should be read before considering the p-value.
Why does intention-to-treat analysis matter?
Analyzing randomized participants according to randomized assignment preserves the comparison created by randomization. It avoids selectively removing participants after randomization based on treatment discontinuation or other subsequent events. The registry analysis explicitly identifies the full analysis set with the intention-to-treat principle.
Why are there different estimands for the same primary endpoints?
The treatment-policy and hypothetical estimands describe different scientific questions. A treatment-policy analysis incorporates the consequences of post-randomization events according to the treatment strategy being evaluated, while a hypothetical analysis asks what would be expected under a specified hypothetical scenario in which particular events did not occur. The numerical estimates should therefore not be treated as interchangeable.
12. Missing Data and Imputation
The ANCOVA analysis description explicitly reports a multiple-imputation approach. Missing observations were multiple (x1000) imputed from retrieved subjects of the same randomized treatment arm.
Why imputation matters
A week-104 analysis can be affected when participants do not have an observed outcome at the target time. Multiple imputation replaces a single missing value with multiple plausible values and combines the resulting analyses.
Why assumptions matter
Imputation does not make missingness disappear. The validity of the resulting estimate depends on the assumptions and model used to generate the imputed values.
The registry-reported MMRM analysis instead uses longitudinal observations prior to the specified discontinuation or other events for its hypothetical estimand. The two approaches therefore address missing and post-randomization information differently.
13. Multiplicity and Multiple Primary Endpoints
The trial has two registered primary endpoints, and four primary-endpoint statistical analyses were posted: two analyses for percentage change in body weight and two analyses for achievement of at least 5% weight reduction.
| Multiplicity issue | What the ClinicalTrials.gov record establishes |
|---|---|
| Number of registered primary endpoints | 2 |
| Formal primary analyses posted | 4 |
| Hypothesis type | Superiority |
| Multiplicity adjustment reported in the ClinicalTrials.gov record | Not specified |
Two primary endpoints create a multiplicity question because multiple confirmatory questions can increase the chance of observing at least one apparently positive result by chance if each is tested independently at the same nominal level. However, the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure or an alpha-allocation strategy. Accordingly, this page does not assign a particular familywise-error interpretation to the four reported p-values.
14. Blinding and Randomization
The trial was randomized and quadruple-masked. These are design features rather than statistical tests, but they are important to the credibility of the comparison.
Randomization is especially important for causal interpretation because it creates the basis for comparing outcomes between treatment assignments. Masking complements randomization by reducing the potential influence of knowledge of treatment assignment on participants, investigators, assessors, or other trial personnel covered by the masking designation.
15. Why the Treatment Difference and Odds Ratio Are Different Measures
| Endpoint | Effect measure | Null value | Interpretation |
|---|---|---|---|
| Percentage change in body weight | Treatment difference | 0 | Difference between treatment-group means or modeled outcome contrasts on the percentage-change scale. |
| Achievement of ≥5% weight reduction | Odds ratio | 1 | Ratio of the odds of achieving the binary endpoint between treatment groups. |
This distinction is essential when reading the four primary results. The value -12.55 is not directly comparable in meaning to the odds ratio 4.99. One is a treatment difference on a percentage-change scale; the other is a ratio of odds for a binary outcome.
16. Understanding the Two Primary Result Frameworks
Continuous outcome
The percentage-change endpoint preserves information about the magnitude of weight change and is summarized through a treatment difference. ANCOVA and MMRM are model-based approaches suited to this type of outcome.
Binary outcome
The ≥5% endpoint reduces the outcome to achievement versus non-achievement and is summarized through an odds ratio in the reported analyses.
Treatment policy
The treatment-policy analyses address the outcome under the treatment strategy and are represented by the ANCOVA and logistic-regression results.
Hypothetical
The hypothetical analyses address a specified scenario without the listed post-randomization events and are represented by the two MMRM results.
17. Clinical Biostats Interpretation of the Primary Results
The treatment-policy ANCOVA estimated a treatment difference of -12.55, with a 95% CI of -15.33 to -9.77. The hypothetical-estimand MMRM estimated a treatment difference of -16.05, with a 95% CI of -18.64 to -13.45. Both estimates are negative, and both registry-reported confidence intervals remain below zero.
The difference between the two estimates should not automatically be described as a discrepancy. The analyses target different estimands and use different statistical frameworks.
The treatment-policy logistic-regression analysis produced an odds ratio of 4.99, with a 95% CI of 2.95 to 8.42. The hypothetical-estimand MMRM analysis produced an odds ratio of 18.06, with a 95% CI of 10.04 to 32.49.
Both estimates are above the null value of 1. Again, the different estimands and analysis methods mean that the two odds ratios should not be interpreted as competing estimates of exactly the same statistical quantity.
All four analyses report p-values below 0.0001 or, for the ANCOVA treatment-policy analysis, <.0001. These results provide strong evidence against their respective null hypotheses within the reported models. They do not tell the reader how large the treatment effect is, whether the effect is clinically important, or how an individual participant will respond.
18. Important Limitations and Interpretation Issues
- Different estimands: treatment-policy and hypothetical analyses answer different questions and should not be combined into a single pooled estimate.
- Different statistical scales: treatment differences and odds ratios have different interpretations and null values.
- Missing-data assumptions: the ANCOVA analysis uses multiple imputation, so the resulting estimate depends in part on the imputation framework and assumptions.
- MMRM assumptions: longitudinal mixed-model results depend on the specified repeated-measures model and its assumptions.
- Multiplicity: two registered primary endpoints are present, but the ClinicalTrials.gov record does not state an alpha-allocation or multiplicity-adjustment procedure.
- Safety scope: the ClinicalTrials.gov record provides serious-adverse-event counts by arm but do not provide a complete safety table.
- Endpoint-type metadata: the registry extract labels the percentage-change endpoint as binary in its inferred endpoint-type field even though the outcome is expressed as a percentage change and analyzed with ANCOVA and MMRM. This page retains the registry's registry-reported classification without treating it as a substantive description of the measurement scale.
19. Why This Trial Matters Statistically
STEP 5 is a useful teaching case because the trial places several core statistical ideas side by side: randomized treatment assignment, masking, intention-to-treat analysis, continuous and binary endpoints, covariate adjustment, longitudinal modeling, logistic regression, odds ratios, confidence intervals, p-values, multiple imputation, and distinct estimands.
| Concept | How it appears in STEP 5 |
|---|---|
| Randomization | Randomized allocation in a two-arm parallel-group phase 3 design. |
| Blinding | Quadruple masking. |
| Intention-to-treat | The full analysis set included all randomized participants according to the intention-to-treat principle. |
| ANCOVA | Used for percentage change in body weight under the treatment-policy estimand. |
| MMRM | Used for both primary endpoints under the hypothetical estimand. |
| Logistic regression | Used for the binary ≥5% weight-reduction endpoint under the treatment-policy estimand. |
| Odds ratio | Reported for the binary primary endpoint. |
| Confidence intervals | All four posted primary analyses include two-sided 95% confidence intervals. |
| P-values | All four posted primary analyses report p-values below 0.0001 or <.0001. |
| Missing-data methods | The ANCOVA analysis reports multiple (x1000) imputation. |
| Estimands | Treatment-policy and hypothetical estimands are both represented. |
20. Related Statistical Learning Pathway
Learn more about the methods used in this trial:
21. Related Statistical Calculators
22. Sources
- ClinicalTrials.gov: NCT03693430.
- STEP 5 primary publication: PubMed PMID 36216945.
- STEP 5-related analysis: PubMed PMID 37605636.
- Post hoc psychiatric safety analysis including STEP 5: PubMed PMID 39226070.
Continue through the Clinical Biostats statistical pathway
Explore the statistical methods behind randomized trials, longitudinal outcomes, binary endpoints, effect measures, and confidence intervals.
23. Record Summary
STEP 5 provides a compact example of how a modern randomized trial can evaluate the same clinical question through complementary statistical frameworks. The trial enrolled 304 participants in two randomized, quadruple-masked parallel arms and followed two registered primary endpoints through week 104. The registry analyses include ANCOVA and MMRM for percentage change in body weight and logistic regression and MMRM for achievement of at least 5% weight reduction.
The statistical story is defined not only by the point estimates, but by the analysis framework surrounding them. The treatment-policy ANCOVA estimated a body-weight treatment difference of -12.55 with a 95% CI of -15.33 to -9.77, while the hypothetical-estimand MMRM estimated -16.05 with a 95% CI of -18.64 to -13.45. For the binary endpoint, the treatment-policy logistic-regression analysis reported an odds ratio of 4.99 with a 95% CI of 2.95 to 8.42, while the hypothetical-estimand MMRM reported an odds ratio of 18.06 with a 95% CI of 10.04 to 32.49. All four analyses reported p-values below 0.0001 or <.0001.