← Clinical Trial Results
Overweight / Obesity Phase 3 104-Week Analysis NCT03693430

STEP 5: Complete Statistical Analysis of Semaglutide in Overweight or Obesity

An independent statistical analysis of the randomized phase 3 STEP 5 trial evaluating semaglutide versus placebo in people with overweight or obesity, focusing on the two registered primary endpoints at week 104 and the ANCOVA, MMRM, and logistic-regression methods used to analyze them.

Trial status: COMPLETED  ·  Enrollment: 304  ·  Primary completion: January 29, 2021
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results and trial characteristics presented here are restricted to the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

STEP 5 was a randomized, parallel-group, quadruple-masked phase 3 trial comparing semaglutide with placebo in people with overweight or obesity. The trial enrolled 304 participants, with 152 participants assigned to each arm, and evaluated two registered primary endpoints through week 104.

304
Randomized
152 per arm
2
Treatment arms
Semaglutide vs placebo
104
Primary time point
Weeks
2
Primary endpoints
Both formally analyzed
FeatureSTEP 5
Trial nameSTEP 5
PhasePhase 3
StatusCOMPLETED
Therapeutic areaEndocrinology
ConditionsOverweight; Obesity
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment304
Lead sponsorNovo Nordisk A/S
Sponsor typeIndustry
ClinicalTrials.govNCT03693430

2. Clinical Question

The primary statistical question was whether semaglutide differed from placebo with respect to the two registered primary outcomes: percentage change from baseline at week 104 in body weight and achievement of body-weight reduction greater than or equal to 5% at week 104.

Population

People with overweight or obesity enrolled in the STEP 5 phase 3 trial.

Intervention

Semaglutide, with the reported serious-adverse-event analysis identifying the treatment arm as semaglutide 2.4 mg.

Comparator

Placebo, identified in the registry results as the placebo arm.

Primary question

How did semaglutide compare with placebo at week 104 for percentage change in body weight and achievement of at least 5% body-weight reduction?

3. Trial Design

01
Randomize 304 participants
02
Parallel arms Semaglutide or placebo
03
Follow Through week 104
04
Analyze Continuous and binary endpoints
05
Interpret Effect estimates and uncertainty
ARM A · n = 152

Semaglutide

  • Semaglutide 2.4 mg as identified in the serious-adverse-event results
  • Primary assessment at week 104
  • Primary analyses included ANCOVA, MMRM, and logistic regression
ARM B · n = 152

Placebo

  • Placebo comparator
  • Primary assessment at week 104
  • Primary analyses included ANCOVA, MMRM, and logistic-regression comparisons
Allocation
Randomized allocation in a parallel-group design.
Masking
Quadruple masking was registered for the trial.
Primary purpose
Treatment.
Hypothesis type
Superiority.

4. Endpoints

The registry lists two primary endpoints. Both had posted formal statistical analyses and both had an estimate, a two-sided 95% confidence interval, and a p-value in the ClinicalTrials.gov record.

Primary endpointRegistry time frameEndpoint typePrincipal analysis methods reported
Percentage Change From Baseline (Week 0) to Week 104 in Body Weight From Baseline (Week 0) to Week 104 Binary in the registry inference ANCOVA; MMRM
Number of Participants Who Achieved (Yes/no): Body Weight Reduction More Than or Equal to 5% At Week 104 Binary Logistic regression; MMRM

Endpoint definitions

For the first endpoint, the registry states that percentage change in body weight for both the in-trial and on-treatment observation periods from baseline (week 0) to week 104 is presented. The outcome measure was evaluated based on data from both in-trial and on-treatment periods.

For the second endpoint, the registry states that the number of participants who achieved greater than or equal to 5% weight loss at 104 weeks is presented. In the reported data, “Yes” infers participants who achieved greater than or equal to 5% weight loss, whereas “No” infers participants who did not achieve greater than or equal to 5% weight loss.

5. Analysis Populations and Estimands

The ClinicalTrials.gov record identifies the full analysis set (FAS) as including all randomized participants according to the intention-to-treat principle. This is important because the treatment comparison remains anchored to randomized assignment rather than being restricted to participants who completed treatment exactly as planned.

The distinction between a treatment-policy and hypothetical estimand is statistically important. A treatment-policy estimand asks about the treatment strategy while retaining the consequences of treatment discontinuation or other post-randomization events according to the specified strategy. A hypothetical estimand instead targets what the outcome would look like under a hypothetical scenario in which specified post-randomization events did not occur.

6. Results: Percentage Change in Body Weight

The first primary endpoint was the percentage change from baseline (week 0) to week 104 in body weight. Two formal analyses were posted for this endpoint: ANCOVA under a treatment-policy estimand and MMRM under a hypothetical estimand.

ANCOVA: Treatment-Policy Estimand

Estimated treatment difference

-12.55

95% CI: -15.33 to -9.77   ·   P < .0001

Semaglutide 2.4 mg vs Placebo

FeatureReported result
OutcomePercentage Change From Baseline (Week 0) to Week 104 in Body Weight
MethodANCOVA
Effect measureTreatment difference
Estimate-12.55
95% CI-15.33 to -9.77
P-value<.0001
HypothesisSuperiority
EstimandTreatment policy

The registry states that week 104 responses were analyzed using an analysis of covariance model with randomized treatment as a factor and baseline body weight as a covariate. Missing observations were multiple (x1000) imputed from retrieved subjects of the same randomized treatment arm.

Clinical Biostats interpretation

The estimate of -12.55 is the reported treatment difference in percentage change in body weight between semaglutide 2.4 mg and placebo under the treatment-policy analysis. Its negative direction indicates a lower percentage-change value for semaglutide relative to placebo.

The estimate does not mean that every participant experienced exactly a 12.55 percentage-point difference. It is a group-level adjusted treatment contrast from the ANCOVA model.

The two-sided 95% confidence interval of -15.33 to -9.77 describes the statistical uncertainty around the estimated treatment difference under the analysis framework. Because the interval remains below zero, it is consistent with a treatment difference favoring the semaglutide direction for this endpoint.

The p-value of <.0001 addresses the evidence against the null hypothesis under the specified test. It does not measure the size of the treatment effect, clinical importance, or probability that the treatment is effective.

The interpretation also depends on the treatment-policy estimand and the multiple-imputation approach specified in the registry analysis description. It should not be silently substituted for the separate hypothetical-estimand MMRM result.

MMRM: Hypothetical Estimand

Estimated treatment difference

-16.05

95% CI: -18.64 to -13.45   ·   P <0.0001

Semaglutide 2.4 mg vs Placebo

FeatureReported result
OutcomePercentage Change From Baseline (Week 0) to Week 104 in Body Weight
MethodMMRM (mixed model for repeated measures)
Effect measureTreatment difference
Estimate-16.05
95% CI-18.64 to -13.45
P-value<0.0001
HypothesisSuperiority
EstimandHypothetical

The registry states that all responses prior to first discontinuation of treatment, or initiation of other anti-obesity medication or bariatric surgery, were included in a mixed model for repeated measurements with randomized treatment as factor and baseline body weight as covariate, all nested within visit.

Clinical Biostats interpretation

The reported -16.05 is the treatment difference from the MMRM analysis under a hypothetical estimand. It therefore answers a different statistical question from the treatment-policy ANCOVA estimate of -12.55.

The 95% confidence interval of -18.64 to -13.45 quantifies uncertainty around this model-based longitudinal treatment contrast. The entire interval is below zero.

The p-value of <0.0001 indicates strong evidence against the corresponding null hypothesis within this analysis. It is not an effect-size metric and should not be used as a substitute for examining the estimated difference and its confidence interval.

Because this is an MMRM, interpretation also depends on the longitudinal model and its assumptions about the repeated measurements and the hypothetical treatment scenario. The MMRM result should therefore be understood as an estimand-specific model result rather than as a universal estimate of what every participant would have experienced.

7. Results: Achievement of at Least 5% Weight Reduction

The second primary endpoint was binary: whether a participant achieved body-weight reduction greater than or equal to 5% at week 104. Two formal analyses were posted: logistic regression under a treatment-policy estimand and MMRM under a hypothetical estimand.

Logistic Regression: Treatment-Policy Estimand

Odds ratio

4.99

95% CI: 2.95 to 8.42   ·   P <0.0001

Semaglutide 2.4 mg vs Placebo

FeatureReported result
OutcomeBody Weight Reduction More Than or Equal to 5% at Week 104
MethodLogistic regression
Effect measureOdds Ratio (OR)
Estimate4.99
95% CI2.95 to 8.42
P-value<0.0001
HypothesisSuperiority
EstimandTreatment policy
Clinical Biostats interpretation

An odds ratio of 4.99 means that the modeled odds of achieving at least 5% body-weight reduction were estimated to be 4.99 times as high with semaglutide 2.4 mg as with placebo under this analysis.

An odds ratio is not a risk ratio. It does not mean that the probability of achieving the endpoint was 4.99 times as high. The distinction becomes especially important when the outcome is not rare.

The 95% confidence interval of 2.95 to 8.42 expresses uncertainty around the estimated odds ratio. It remains above 1, the null value for an odds ratio.

The p-value of <0.0001 describes the evidence against the null hypothesis under the specified model. It does not describe the magnitude of the odds ratio or the probability that a particular participant will achieve the endpoint.

The analysis is also explicitly tied to a treatment-policy estimand, so its interpretation should not be merged with the separate MMRM result that uses a hypothetical estimand.

MMRM: Hypothetical Estimand

Odds ratio

18.06

95% CI: 10.04 to 32.49   ·   P <0.0001

Semaglutide 2.4 mg vs Placebo

FeatureReported result
OutcomeBody Weight Reduction More Than or Equal to 5% at Week 104
MethodMMRM (mixed model for repeated measures)
Effect measureOdds Ratio (OR)
Estimate18.06
95% CI10.04 to 32.49
P-value<0.0001
HypothesisSuperiority
EstimandHypothetical
Clinical Biostats interpretation

The reported odds ratio of 18.06 indicates substantially higher modeled odds of achieving at least 5% body-weight reduction with semaglutide than placebo under the hypothetical-estimand MMRM analysis.

This is an odds ratio, not a probability ratio. An odds ratio of 18.06 cannot be read as saying that 18.06 times as many participants achieved the endpoint.

The 95% confidence interval of 10.04 to 32.49 indicates uncertainty around the estimate while remaining above the null value of 1.

The p-value of <0.0001 indicates strong evidence against the null hypothesis within this model. It does not quantify clinical magnitude or individual-level benefit.

The hypothetical estimand also matters: this analysis is not answering precisely the same question as the treatment-policy logistic-regression analysis, even though both concern the same week-104 binary endpoint.

8. Comparing the Two Primary Endpoints

Primary endpointMethodEstimandEffect measureEstimate95% CIP-value
Percentage change in body weight ANCOVA Treatment policy Treatment difference -12.55 -15.33 to -9.77 <.0001
Percentage change in body weight MMRM Hypothetical Treatment difference -16.05 -18.64 to -13.45 <0.0001
At least 5% weight reduction Logistic regression Treatment policy Odds ratio 4.99 2.95 to 8.42 <0.0001
At least 5% weight reduction MMRM Hypothetical Odds ratio 18.06 10.04 to 32.49 <0.0001

The most important statistical distinction is not simply the numerical size of the four estimates. The analyses answer different questions. The two ANCOVA/logistic-regression results use a treatment-policy estimand, while the two MMRM results use a hypothetical estimand. Comparing those pairs without recognizing the estimand difference would make the statistical interpretation less precise.

9. Safety Results

The ClinicalTrials.gov record includes serious adverse events by randomized treatment arm. The ClinicalTrials.gov record does not supply a broader adverse-event table, so safety interpretation is limited to the reported serious-adverse-event counts.

Safety measureSemaglutide 2.4 mgPlacebo
Participants affected / at risk12 / 15218 / 152

These counts describe serious adverse events in each randomized arm as reported in the registry. They do not establish that every observed event was caused by the assigned intervention, and they should not be converted into a comparative risk measure that was not reported in the ClinicalTrials.gov record.

Safety interpretation: The serious-adverse-event counts are reported descriptively here. A causal safety conclusion would require the corresponding event definitions, timing, exposure information, and statistical safety framework; those additional details are not part of the ClinicalTrials.gov record.

10. Statistical Methodology

ANCOVA

ANCOVA, or analysis of covariance, was used for the treatment-policy analysis of percentage change in body weight. The registry specifies randomized treatment as a factor and baseline body weight as a covariate.

Conceptual model
Outcome at Week 104 = treatment effect + baseline body weight effect + residual variation

The practical purpose of the covariate is to account for baseline body weight when estimating the treatment contrast at the specified analysis time point.

For a continuous outcome such as percentage change in body weight, the treatment difference is naturally interpreted on the outcome scale. A negative treatment difference here means the modeled percentage-change outcome was lower in the semaglutide group than in the placebo group.

MMRM

MMRM, or mixed model for repeated measures, was used for both primary endpoints in the hypothetical-estimand analyses. The registry states that responses were modeled with randomized treatment as a factor and baseline body weight as a covariate, with terms nested within visit.

The principal advantage of a repeated-measures model is that it can use longitudinal information rather than treating each participant's week-104 observation as an isolated data point. The interpretation, however, depends on the model specification and the estimand it is intended to estimate.

Logistic regression

The treatment-policy analysis of achieving at least 5% weight reduction used logistic regression. Logistic regression is appropriate for a binary outcome such as “Yes” or “No.” Its natural effect measure is the odds ratio.

Odds-ratio interpretation
OR = odds of success in semaglutide ÷ odds of success in placebo

An OR above 1 indicates higher odds of the defined binary outcome in the semaglutide group under the fitted model. It should not be interpreted as a risk ratio without additional assumptions or calculations.

Intention-to-treat analysis

The ClinicalTrials.gov record states that the full analysis set included all randomized participants according to the intention-to-treat principle. This preserves the treatment assignment created by randomization and avoids redefining the primary comparison based solely on subsequent treatment behavior.

Superiority testing

All four registry-reported statistical analyses identify the hypothesis type as superiority. In a superiority framework, the statistical question is whether the treatment groups differ in the prespecified direction under the relevant model, rather than whether a treatment is merely no worse than a comparator by a predefined non-inferiority margin.

11. Statistical Methods Explained

Why was ANCOVA used for percentage change in body weight?

Percentage change in body weight is a continuous outcome, making a linear-model approach appropriate. ANCOVA allows the analysis to include randomized treatment as the principal factor while also accounting for baseline body weight as a covariate. The resulting treatment difference is adjusted for baseline body weight rather than being a simple unadjusted comparison of observed week-104 means.

Why was MMRM also used?

The MMRM analysis addresses repeated measurements over visits rather than focusing solely on a single endpoint measurement. In STEP 5, the registry text specifies a longitudinal model with randomized treatment, baseline body weight, and visit structure. The MMRM results therefore provide a different estimand-specific analysis of the same primary endpoint concepts.

What does an odds ratio of 4.99 mean?

An odds ratio of 4.99 means the modeled odds of achieving at least 5% weight reduction were 4.99 times as high in the semaglutide group as in the placebo group under that logistic-regression analysis. It does not mean that the probability was 4.99 times as high, because odds and probability are different quantities.

Why does the confidence interval matter?

A point estimate alone does not describe statistical uncertainty. The 95% confidence interval gives a range of values compatible with the model and sampling framework at the stated confidence level. For the ANCOVA estimate, the interval is -15.33 to -9.77; for the treatment-policy odds ratio, it is 2.95 to 8.42. These intervals communicate both direction and precision more fully than the p-values alone.

Why does the p-value not measure effect size?

The p-value measures how inconsistent the observed data are with a specified null hypothesis under the statistical model. It is affected by the size and variability of the dataset and does not tell the reader how large the treatment effect is. That is why the estimate and confidence interval should be read before considering the p-value.

Why does intention-to-treat analysis matter?

Analyzing randomized participants according to randomized assignment preserves the comparison created by randomization. It avoids selectively removing participants after randomization based on treatment discontinuation or other subsequent events. The registry analysis explicitly identifies the full analysis set with the intention-to-treat principle.

Why are there different estimands for the same primary endpoints?

The treatment-policy and hypothetical estimands describe different scientific questions. A treatment-policy analysis incorporates the consequences of post-randomization events according to the treatment strategy being evaluated, while a hypothetical analysis asks what would be expected under a specified hypothetical scenario in which particular events did not occur. The numerical estimates should therefore not be treated as interchangeable.

12. Missing Data and Imputation

The ANCOVA analysis description explicitly reports a multiple-imputation approach. Missing observations were multiple (x1000) imputed from retrieved subjects of the same randomized treatment arm.

Why imputation matters

A week-104 analysis can be affected when participants do not have an observed outcome at the target time. Multiple imputation replaces a single missing value with multiple plausible values and combines the resulting analyses.

Why assumptions matter

Imputation does not make missingness disappear. The validity of the resulting estimate depends on the assumptions and model used to generate the imputed values.

The registry-reported MMRM analysis instead uses longitudinal observations prior to the specified discontinuation or other events for its hypothetical estimand. The two approaches therefore address missing and post-randomization information differently.

13. Multiplicity and Multiple Primary Endpoints

The trial has two registered primary endpoints, and four primary-endpoint statistical analyses were posted: two analyses for percentage change in body weight and two analyses for achievement of at least 5% weight reduction.

Multiplicity issueWhat the ClinicalTrials.gov record establishes
Number of registered primary endpoints2
Formal primary analyses posted4
Hypothesis typeSuperiority
Multiplicity adjustment reported in the ClinicalTrials.gov recordNot specified

Two primary endpoints create a multiplicity question because multiple confirmatory questions can increase the chance of observing at least one apparently positive result by chance if each is tested independently at the same nominal level. However, the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure or an alpha-allocation strategy. Accordingly, this page does not assign a particular familywise-error interpretation to the four reported p-values.

Important distinction: the absence of a multiplicity procedure in the registry-reported extract does not establish that the trial had no prespecified multiplicity strategy. It means only that the ClinicalTrials.gov record does not provide enough information to describe such a strategy accurately.

14. Blinding and Randomization

The trial was randomized and quadruple-masked. These are design features rather than statistical tests, but they are important to the credibility of the comparison.

Randomization
Participants were allocated randomly to two parallel treatment arms, with 304 participants enrolled overall.
Quadruple masking
The registry identifies the trial as quadruple-masked, reducing opportunities for knowledge of treatment assignment to influence trial conduct or assessment.

Randomization is especially important for causal interpretation because it creates the basis for comparing outcomes between treatment assignments. Masking complements randomization by reducing the potential influence of knowledge of treatment assignment on participants, investigators, assessors, or other trial personnel covered by the masking designation.

15. Why the Treatment Difference and Odds Ratio Are Different Measures

EndpointEffect measureNull valueInterpretation
Percentage change in body weight Treatment difference 0 Difference between treatment-group means or modeled outcome contrasts on the percentage-change scale.
Achievement of ≥5% weight reduction Odds ratio 1 Ratio of the odds of achieving the binary endpoint between treatment groups.

This distinction is essential when reading the four primary results. The value -12.55 is not directly comparable in meaning to the odds ratio 4.99. One is a treatment difference on a percentage-change scale; the other is a ratio of odds for a binary outcome.

16. Understanding the Two Primary Result Frameworks

Continuous outcome

The percentage-change endpoint preserves information about the magnitude of weight change and is summarized through a treatment difference. ANCOVA and MMRM are model-based approaches suited to this type of outcome.

Binary outcome

The ≥5% endpoint reduces the outcome to achievement versus non-achievement and is summarized through an odds ratio in the reported analyses.

Treatment policy

The treatment-policy analyses address the outcome under the treatment strategy and are represented by the ANCOVA and logistic-regression results.

Hypothetical

The hypothetical analyses address a specified scenario without the listed post-randomization events and are represented by the two MMRM results.

17. Clinical Biostats Interpretation of the Primary Results

Reading the continuous endpoint

The treatment-policy ANCOVA estimated a treatment difference of -12.55, with a 95% CI of -15.33 to -9.77. The hypothetical-estimand MMRM estimated a treatment difference of -16.05, with a 95% CI of -18.64 to -13.45. Both estimates are negative, and both registry-reported confidence intervals remain below zero.

The difference between the two estimates should not automatically be described as a discrepancy. The analyses target different estimands and use different statistical frameworks.

Reading the binary endpoint

The treatment-policy logistic-regression analysis produced an odds ratio of 4.99, with a 95% CI of 2.95 to 8.42. The hypothetical-estimand MMRM analysis produced an odds ratio of 18.06, with a 95% CI of 10.04 to 32.49.

Both estimates are above the null value of 1. Again, the different estimands and analysis methods mean that the two odds ratios should not be interpreted as competing estimates of exactly the same statistical quantity.

Reading the p-values

All four analyses report p-values below 0.0001 or, for the ANCOVA treatment-policy analysis, <.0001. These results provide strong evidence against their respective null hypotheses within the reported models. They do not tell the reader how large the treatment effect is, whether the effect is clinically important, or how an individual participant will respond.

18. Important Limitations and Interpretation Issues

19. Why This Trial Matters Statistically

STEP 5 is a useful teaching case because the trial places several core statistical ideas side by side: randomized treatment assignment, masking, intention-to-treat analysis, continuous and binary endpoints, covariate adjustment, longitudinal modeling, logistic regression, odds ratios, confidence intervals, p-values, multiple imputation, and distinct estimands.

ConceptHow it appears in STEP 5
RandomizationRandomized allocation in a two-arm parallel-group phase 3 design.
BlindingQuadruple masking.
Intention-to-treatThe full analysis set included all randomized participants according to the intention-to-treat principle.
ANCOVAUsed for percentage change in body weight under the treatment-policy estimand.
MMRMUsed for both primary endpoints under the hypothetical estimand.
Logistic regressionUsed for the binary ≥5% weight-reduction endpoint under the treatment-policy estimand.
Odds ratioReported for the binary primary endpoint.
Confidence intervalsAll four posted primary analyses include two-sided 95% confidence intervals.
P-valuesAll four posted primary analyses report p-values below 0.0001 or <.0001.
Missing-data methodsThe ANCOVA analysis reports multiple (x1000) imputation.
EstimandsTreatment-policy and hypothetical estimands are both represented.

20. Related Statistical Learning Pathway

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods behind randomized trials, longitudinal outcomes, binary endpoints, effect measures, and confidence intervals.

23. Record Summary

STEP 5 provides a compact example of how a modern randomized trial can evaluate the same clinical question through complementary statistical frameworks. The trial enrolled 304 participants in two randomized, quadruple-masked parallel arms and followed two registered primary endpoints through week 104. The registry analyses include ANCOVA and MMRM for percentage change in body weight and logistic regression and MMRM for achievement of at least 5% weight reduction.

The statistical story is defined not only by the point estimates, but by the analysis framework surrounding them. The treatment-policy ANCOVA estimated a body-weight treatment difference of -12.55 with a 95% CI of -15.33 to -9.77, while the hypothetical-estimand MMRM estimated -16.05 with a 95% CI of -18.64 to -13.45. For the binary endpoint, the treatment-policy logistic-regression analysis reported an odds ratio of 4.99 with a 95% CI of 2.95 to 8.42, while the hypothetical-estimand MMRM reported an odds ratio of 18.06 with a 95% CI of 10.04 to 32.49. All four analyses reported p-values below 0.0001 or <.0001.

Clinical Biostats methodology: The purpose of a trial-results page is not merely to repeat numerical results. It is to explain what each estimate measures, which population and estimand it represents, how uncertainty is quantified, and which assumptions or design features affect interpretation.