← Clinical Trials
Overweight / Obesity Phase 3 Completed NCT03611582

STEP 3: Complete Statistical Analysis of Semaglutide in Overweight and Obesity

An independent statistical review of the randomized phase 3 STEP 3 trial evaluating semaglutide versus placebo for weight reduction in people with overweight or obesity, with emphasis on the registered primary endpoints and the reported ANCOVA, MMRM, and logistic-regression analyses.

Phase 3  ·  Randomized parallel design  ·  Enrollment 611  ·  Primary completion March 18, 2020
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the ClinicalTrials.gov record for NCT03611582.

Registry record: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

STEP 3 was a completed, randomized, quadruple-masked, parallel phase 3 trial in overweight and obesity. The registry reports 611 enrolled participants and two intervention groups: semaglutide and placebo.

611
Enrolled
Randomized trial
2
Arms
Semaglutide vs placebo
-10.27
ANCOVA difference
95% CI -11.97 to -8.57
6.11
5% weight-loss OR
95% CI 4.04 to 9.26
FeatureSTEP 3
Trial nameSTEP 3
ClinicalTrials.gov identifierNCT03611582
PhasePhase 3
Therapeutic areaEndocrinology
ConditionsOverweight; Obesity
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment611
InterventionsSemaglutide; placebo
Lead sponsorNovo Nordisk A/S
Sponsor typeIndustry
StatusCompleted

2. Clinical Question

The central statistical question was whether semaglutide differed from placebo with respect to weight-related outcomes in participants with overweight or obesity. The registry identifies two primary endpoints: change in body weight from baseline to week 68 and achievement of at least 5% body-weight reduction after 68 weeks.

Population

Participants enrolled in a phase 3 study for overweight or obesity. The ClinicalTrials.gov record reports 611 enrolled participants.

Intervention

Semaglutide. The serious-adverse-event data identify the semaglutide arm as semaglutide 2.4 mg.

Comparator

Placebo, described in the registry data as placebo for semaglutide.

Primary question

How does semaglutide compare with placebo for change in body weight and for achieving at least 5% weight reduction after 68 weeks?

3. Trial Design

01
Randomize611 enrolled
02
Parallel groupsSemaglutide vs placebo
03
Quadruple maskMasked trial design
04
AssessWeight outcomes
05
Week 68Primary endpoint time point
Allocation
Randomized allocation was used in a parallel-group phase 3 design.
Masking
The registry identifies the study as quadruple-masked.
Primary purpose
Treatment.
Observation
The primary continuous endpoint uses baseline (week 0) to week 68; the categorical endpoint is assessed after 68 weeks.
INTERVENTION

Semaglutide 2.4 mg

  • Semaglutide
  • Serious adverse events reported for 37 of 407 participants at risk
COMPARATOR

Placebo

  • Placebo for semaglutide
  • Serious adverse events reported for 6 of 204 participants at risk
Design implication: Randomization establishes the treatment comparison specified by the trial design, while quadruple masking is intended to reduce the influence of knowledge of treatment assignment on trial conduct and assessment. The ClinicalTrials.gov record does not provide further details about the masking assignments or procedures.

4. Endpoints

EndpointRegistered time frameEndpoint typeReported analysis
Change in Body Weight (%) Baseline (week 0) to week 68 Binary in the registry classification ANCOVA; MMRM
Participants Who Achieve (Yes/no): Body Weight Reduction More Than or Equal to 5% After 68 weeks Binary Logistic regression; MMRM

The registry states that the body-weight endpoint is presented as change in body weight from baseline (week 0) to week 68. For the ≥5% endpoint, “Yes” denotes participants who achieved greater than or equal to 5% weight loss and “No” denotes participants who did not achieve that threshold.

Registry classification note: The ClinicalTrials.gov record classifies “Change in Body Weight (%)” as a binary endpoint even though the registered endpoint name describes a continuous percentage change. This page preserves that registry classification rather than silently changing it. The reported analyses themselves identify ANCOVA and MMRM with a treatment-difference effect measure.

5. Statistical Methodology

Full analysis set and intention-to-treat principle

The reported efficacy analyses use the full analysis set (FAS), which the registry describes as comprising all randomized participants. For each reported analysis, the number analyzed is the number of participants with available data. The registry also identifies intention-to-treat analysis as a concept in the analysis text.

Analysis population
FAS = all randomized participants; analyzed number = participants with available data

This distinction matters because the randomized population defines the treatment comparison, while the number contributing observed data to a particular analysis can be smaller.

ANCOVA

ANCOVA, or analysis of covariance, is a linear-model approach commonly used when comparing a continuous outcome between randomized treatment groups while accounting for covariates specified in the model. In STEP 3, the registry reports ANCOVA for the baseline-to-week-68 body-weight endpoint and reports a treatment difference as the effect measure.

MMRM

The registry also reports a mixed model for repeated measures (MMRM). MMRM is designed for longitudinal data in which outcomes are measured repeatedly over time. Rather than reducing all observations to a single change score, the model can use the pattern of repeated observations to estimate treatment differences at the relevant time point.

Logistic regression

The ≥5% body-weight-reduction endpoint is binary: participants are classified as achieving the threshold or not achieving it. Logistic regression models the probability of the binary outcome and naturally expresses the treatment comparison through an odds ratio.

Treatment policy and hypothetical estimands

The registry distinguishes the estimand associated with the analyses. The ANCOVA and logistic-regression analyses are identified as using a treatment policy estimand, whereas the corresponding MMRM analyses are identified as using a hypothetical estimand. These are not merely different statistical algorithms; they represent different questions about how the treatment effect is defined in relation to post-randomization events and treatment experience.

Treatment policy

The treatment effect is defined under a strategy in which the outcome is considered regardless of intercurrent events in the way specified by the treatment-policy strategy.

Hypothetical

The treatment effect addresses a specified hypothetical scenario concerning what would have happened under the defined treatment condition.

6. Results: Change in Body Weight (%)

The first registered primary endpoint is change in body weight from baseline (week 0) to week 68. Two formal analyses are posted in the ClinicalTrials.gov record: ANCOVA using a treatment policy estimand and MMRM using a hypothetical estimand.

ANCOVA — treatment policy estimand

Semaglutide 2.4 mg vs placebo

-10.27

Treatment difference  ·  95% CI -11.97 to -8.57

Two-sided P < .0001  ·  Superiority

EndpointMethodEffect measureEstimate95% CIP-value
Change in Body Weight (%) ANCOVA Treatment difference -10.27 -11.97 to -8.57 <.0001
Clinical Biostats interpretation

The reported treatment difference of -10.27 indicates that the estimated change in the semaglutide group was 10.27 percentage points lower than the corresponding placebo-group value under the reported ANCOVA treatment-policy analysis.

The negative sign identifies the direction of the difference as reported; it does not by itself describe the absolute change within either treatment group.

The two-sided 95% confidence interval, -11.97 to -8.57, describes the statistical uncertainty around the estimated treatment difference under the analysis framework. It is an interval for the treatment-effect estimate, not a range containing individual participants' weight changes.

The P-value of <.0001 addresses the evidence against the null hypothesis specified for the superiority comparison. It does not measure the size of the treatment effect, clinical importance, or the probability that the null hypothesis is true.

The treatment-policy estimand is also important: this estimate answers the question defined by that estimand strategy, rather than automatically answering the hypothetical question represented by the separate MMRM analysis.

MMRM — hypothetical estimand

Semaglutide 2.4 mg vs placebo

-12.67

Treatment difference  ·  95% CI -14.34 to -11.00

Two-sided P < 0.0001  ·  Superiority

EndpointMethodEffect measureEstimate95% CIP-value
Change in Body Weight (%) MMRM Treatment difference -12.67 -14.34 to -11.00 <0.0001
Clinical Biostats interpretation

The MMRM estimate of -12.67 represents the reported treatment difference under the hypothetical estimand and repeated-measures model. As with the ANCOVA result, the negative value indicates a lower estimated body-weight change in the semaglutide group relative to placebo under the specified analysis.

The 95% CI of -14.34 to -11.00 quantifies uncertainty around this model-based estimate. Its width provides information about precision, while the interval itself does not describe the variability of individual weight changes.

The P-value of <0.0001 provides evidence against the null hypothesis used for the superiority comparison. A P-value should not be read as a measure of effect magnitude.

The estimate differs from the ANCOVA estimate because the analyses are not identical: the MMRM uses repeated-measures modeling and is identified in the registry as addressing a hypothetical estimand, while the ANCOVA result is identified as a treatment-policy analysis. The two estimates therefore should not simply be averaged or treated as independent replications of exactly the same estimand.

7. Results: Participants Achieving at Least 5% Weight Reduction

The second registered primary endpoint is the binary outcome “Participants Who Achieve (Yes/no): Body Weight Reduction More Than or Equal to 5%,” assessed after 68 weeks. The registry defines “Yes” as achieving ≥5% weight loss and “No” as not achieving ≥5% weight loss.

Logistic regression — treatment policy estimand

Odds ratio for achieving ≥5% weight reduction

6.11

95% CI: 4.04 to 9.26

Two-sided P < 0.0001  ·  Superiority

EndpointMethodEffect measureEstimate95% CIP-value
≥5% body-weight reduction after 68 weeks Logistic regression Odds ratio 6.11 4.04 to 9.26 <0.0001
Clinical Biostats interpretation

An odds ratio of 6.11 means that the reported odds of achieving at least 5% weight reduction were estimated to be 6.11 times as high with semaglutide as with placebo under the treatment-policy logistic-regression analysis.

An odds ratio is not a risk ratio and is not equivalent to saying that six times as many participants achieved the endpoint. The corresponding probabilities depend on the underlying event rates, which are not provided in the ClinicalTrials.gov record.

The 95% CI of 4.04 to 9.26 describes uncertainty around the odds-ratio estimate. It does not describe the range of individual treatment effects.

The two-sided P-value of <0.0001 addresses the hypothesis test; it does not quantify the magnitude or clinical importance of the odds ratio.

Because the analysis is based on a binary endpoint, the interpretation is specifically about achievement of the ≥5% threshold after 68 weeks, not about the continuous magnitude of weight change.

MMRM — hypothetical estimand

Odds ratio for achieving ≥5% weight reduction

11.67

95% CI: 7.64 to 17.81

Two-sided P < 0.0001  ·  Superiority

EndpointMethodEffect measureEstimate95% CIP-value
≥5% body-weight reduction after 68 weeks MMRM Odds ratio 11.67 7.64 to 17.81 <0.0001
Clinical Biostats interpretation

The reported odds ratio of 11.67 indicates substantially higher estimated odds of achieving the ≥5% threshold under the MMRM analysis and its hypothetical estimand.

The 95% CI of 7.64 to 17.81 provides the uncertainty interval for that estimate under the reported model. The interval is entirely above 1, consistent with the reported superiority test.

Again, this is an odds ratio rather than a probability ratio. Without the underlying event counts or probabilities from the ClinicalTrials.gov record, the odds ratio cannot be converted into an absolute percentage-point difference without additional information.

The difference between the reported OR of 6.11 from logistic regression and 11.67 from the MMRM analysis should be understood in the context of their different modeling and estimand specifications rather than treated as contradictory estimates of exactly the same statistical question.

8. Primary Results Side by Side

Primary endpointAnalysisEstimandEffect95% CIP-value
Change in Body Weight (%) ANCOVA Treatment policy -10.27 treatment difference -11.97 to -8.57 <.0001
Change in Body Weight (%) MMRM Hypothetical -12.67 treatment difference -14.34 to -11.00 <0.0001
≥5% weight reduction Logistic regression Treatment policy OR 6.11 4.04 to 9.26 <0.0001
≥5% weight reduction MMRM Hypothetical OR 11.67 7.64 to 17.81 <0.0001
How to read the table: The four reported analyses are not four independent primary endpoints. They are two registered primary endpoints analyzed under two statistical approaches/estimand specifications. The registry identifies all four analyses as superiority analyses with two-sided 95% confidence intervals.

9. Statistical Methods Explained

Why was ANCOVA used for change in body weight?

ANCOVA is a natural framework for comparing a continuous outcome between randomized groups while incorporating covariate information specified by the analysis model. The registry reports ANCOVA specifically for the baseline-to-week-68 body-weight endpoint and expresses the result as a treatment difference.

What does a treatment difference of -10.27 mean?

It is the reported estimated difference between the semaglutide and placebo groups under the ANCOVA treatment-policy analysis. The negative sign indicates the direction of the contrast as defined by the reported group comparison. It is not an odds ratio and should not be interpreted as a percentage probability.

Why use MMRM?

MMRM is designed for repeated measurements. Longitudinal trials can contain multiple observations per participant, and an MMRM can model those observations jointly rather than relying only on one observed change value. In STEP 3, the registry reports MMRM for both primary endpoints.

What does an odds ratio of 6.11 mean?

An OR of 6.11 means that the estimated odds of achieving the binary ≥5% weight-loss endpoint were 6.11 times as high in the semaglutide group as in the placebo group under the reported logistic-regression analysis. It does not mean that the probability was six times higher.

Why are there two different odds ratios for the same binary endpoint?

The registry reports logistic regression with a treatment-policy estimand and MMRM with a hypothetical estimand. Because these analyses address differently defined statistical questions and use different model structures, their estimates need not be identical.

What does the 95% confidence interval tell us?

The confidence interval describes uncertainty around the estimated treatment effect under the specified statistical model and sampling framework. For example, the ANCOVA estimate is -10.27 with a 95% CI from -11.97 to -8.57. The interval does not represent the range of responses among individual participants.

Why does the P-value not measure effect size?

A P-value quantifies evidence against a specified null hypothesis under the statistical model. It depends on both the estimated effect and the amount of information available. The effect estimate and confidence interval therefore remain necessary for understanding magnitude and precision.

10. Confidence Intervals and Precision

The four reported primary analyses all include two-sided 95% confidence intervals. Looking at the interval alongside the point estimate provides substantially more information than the P-value alone.

AnalysisPoint estimate95% confidence intervalWhat the interval describes
ANCOVA-10.27-11.97 to -8.57Uncertainty around the treatment difference
MMRM, body weight-12.67-14.34 to -11.00Uncertainty around the MMRM treatment difference
Logistic regressionOR 6.114.04 to 9.26Uncertainty around the odds ratio
MMRM, ≥5% endpointOR 11.677.64 to 17.81Uncertainty around the odds ratio
Generic statistical interpretation
Estimate ± uncertainty → effect size + precision

The point estimate gives the central result from the fitted analysis; the confidence interval communicates how precisely that effect was estimated under the stated assumptions.

11. Understanding the Two Estimand Strategies

One of the most statistically informative features of the registry-reported STEP 3 record is that it reports both treatment-policy and hypothetical estimands.

Treatment-policy question

The registry identifies the ANCOVA and logistic-regression analyses as treatment-policy analyses. This estimand is intended to characterize the treatment effect under the specified treatment-policy strategy rather than restricting the question to an idealized uninterrupted treatment scenario.

Hypothetical question

The registry identifies the MMRM analyses as hypothetical. This changes the target quantity: the analysis addresses the treatment effect under the hypothetical scenario defined by that estimand.

This distinction is important because a numerical difference between two estimates can arise from differences in the statistical question being answered, not necessarily from random fluctuation or a contradiction in the underlying data.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. The available measure is the number affected divided by the number at risk.

Safety measureSemaglutide 2.4 mgPlacebo
Serious adverse events37/4076/204

Semaglutide arm

Serious adverse events affected 37 of 407 participants at risk in the ClinicalTrials.gov record.

Placebo arm

Serious adverse events affected 6 of 204 participants at risk in the ClinicalTrials.gov record.

The serious-adverse-event figures should be kept separate from the primary efficacy analyses. The efficacy estimates address weight-related outcomes, whereas serious adverse events describe an aspect of safety. They are different outcome domains and require different statistical interpretation.

Safety-data limitation: The ClinicalTrials.gov record provides affected/at-risk counts for serious adverse events but do not provide a formal statistical analysis, confidence interval, or P-value for this safety comparison. No additional safety result is inferred here.

13. Randomization and Blinding

STEP 3 is described in the registry as randomized, parallel, and quadruple-masked. These design characteristics are relevant to the interpretation of the treatment comparison.

Randomization
Random allocation creates the basis for comparing the treatment groups under the trial design and reduces systematic allocation differences expected from nonrandom assignment.
Quadruple masking
The registry identifies four levels of masking. The ClinicalTrials.gov record does not specify the individual masked roles, so none are inferred here.

Randomization and masking address different sources of potential bias. Randomization concerns how participants are assigned to treatment groups, while masking concerns knowledge of those assignments during the conduct and assessment of the trial.

14. Analysis Populations and Missing Data

The reported primary analyses use the full analysis set, which the registry defines as all randomized participants. The number analyzed is the number of participants with available data for the particular analysis.

ConceptSTEP 3 registry descriptionStatistical implication
Full analysis setAll randomized participantsMaintains the randomized treatment framework for efficacy analysis
Number analyzedParticipants with available dataCan be smaller than the randomized population for a particular analysis
Intention-to-treat conceptIdentified in the analysis textSupports interpretation according to randomized assignment
EstimandTreatment policy or hypothetical, depending on analysisDefines the treatment effect being targeted

The ClinicalTrials.gov record does not specify a particular missing-data imputation method, a detailed missing-data sensitivity analysis, or a pattern-mixture model. Accordingly, none is attributed to STEP 3 here.

15. Multiplicity and Interim Analysis

The ClinicalTrials.gov record identifies two primary endpoints and four posted primary-endpoint analyses. They also identify superiority as the hypothesis type. However, the ClinicalTrials.gov record does not provide an alpha-spending procedure, interim-analysis boundary, multiplicity-adjustment procedure, or detailed endpoint hierarchy.

Interpretation boundary: Because the ClinicalTrials.gov record does not specify how multiplicity across the two primary endpoints or across the reported analyses was controlled, the four P-values should be reported as posted rather than assigning an unreported familywise-error interpretation to them.

This distinction is especially important when multiple formal analyses are presented. A P-value is interpreted within the testing framework that generated it; without the prespecified multiplicity structure, one should not reconstruct an error-control procedure from the observed results.

16. Statistical Methods: What Each Model Contributes

MethodSTEP 3 roleEffect measureEstimand
ANCOVA Change in Body Weight (%) Treatment difference Treatment policy
MMRM Change in Body Weight (%) Treatment difference Hypothetical
Logistic regression ≥5% body-weight reduction Odds ratio Treatment policy
MMRM ≥5% body-weight reduction Odds ratio Hypothetical

The model family should not be confused with the estimand. ANCOVA and logistic regression describe statistical modeling approaches, while treatment policy and hypothetical describe the target treatment effect. A complete interpretation therefore needs both pieces of information.

17. Interpreting the Odds Ratio

Statistical interpretation

The treatment-policy logistic-regression analysis reported an odds ratio of 6.11 for achieving at least 5% body-weight reduction. This means the estimated odds were 6.11 times as high in the semaglutide group as in the placebo group under that analysis.

The MMRM analysis reported an odds ratio of 11.67 under a hypothetical estimand. Because the two estimates correspond to different analysis specifications, they should not be interpreted as two measurements of a single unchanged odds ratio.

Odds are not probabilities

If an event probability is p, the corresponding odds are p/(1-p). An odds ratio therefore compares odds rather than probabilities. Without the underlying event probabilities, an odds ratio alone cannot provide the absolute difference in the percentage of participants achieving the endpoint.

18. Interpreting the Body-Weight Treatment Differences

The continuous body-weight endpoint provides a different kind of effect measure from the binary ≥5% endpoint.

Continuous endpoint

The ANCOVA and MMRM analyses report treatment differences of -10.27 and -12.67, respectively. These are differences on the scale of the registered percentage-change endpoint.

Threshold endpoint

The logistic-regression and MMRM analyses report odds ratios for crossing the prespecified ≥5% weight-reduction threshold.

A continuous treatment difference and an odds ratio answer different questions. The continuous analysis concerns the estimated difference in the outcome itself, whereas the binary analysis asks about the likelihood of meeting a predefined threshold.

19. Trial Timeline

August 1, 2018

Trial start

The ClinicalTrials.gov record lists 2018-08-01 as the study start date.

March 18, 2020

Primary completion

The ClinicalTrials.gov record lists 2020-03-18 as the primary completion date.

Completed

Registry status

The trial is listed as completed, with results posted and four statistical analyses posted for the two primary endpoints.

20. Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

STEP 3 is a useful teaching example because the ClinicalTrials.gov record contains several important statistical concepts within a relatively compact randomized design: continuous and binary primary endpoints, ANCOVA, MMRM, logistic regression, odds ratios, treatment-policy and hypothetical estimands, intention-to-treat concepts, confidence intervals, and superiority testing.

ConceptHow it appears in STEP 3
RandomizationRandomized allocation in a parallel phase 3 design
BlindingQuadruple-masked design
Intention-to-treatIdentified in the reported primary analyses
Full analysis setAll randomized participants
ANCOVAPrimary analysis of change in body weight (%)
MMRMReported for both primary endpoints
Logistic regressionPrimary analysis of the ≥5% weight-reduction endpoint
Odds ratioEffect measure for the binary ≥5% endpoint
Confidence interval95% two-sided intervals reported for all four primary analyses
P-valueTwo-sided superiority tests reported as <.0001 or <0.0001
EstimandsTreatment-policy and hypothetical estimands reported
Missing dataNumber analyzed defined as participants with available data; detailed imputation strategy not reported

22. Independent Statistical Interpretation

The most direct statistical reading of the ClinicalTrials.gov record is that all four posted primary-endpoint analyses report strong evidence under their respective superiority tests, with two-sided P-values below 0.0001 or .0001 as reported. The magnitude and scale of the effects differ by endpoint and model.

Continuous weight-change endpoint

-10.27 to -12.67

Reported treatment differences across ANCOVA and MMRM analyses

≥5% weight-reduction endpoint

OR 6.11 to 11.67

Reported odds ratios across logistic regression and MMRM analyses

These ranges should not be interpreted as uncertainty intervals for one common effect. They represent estimates from different analytical specifications. The key statistical task is therefore to identify the endpoint, model, estimand, effect measure, confidence interval, and P-value together rather than selecting a single number without its analytical context.

What the reported results do not establish by themselves

The ClinicalTrials.gov record does not provide individual-level treatment effects, the percentage of participants achieving the ≥5% threshold in each group, or a formal comparison of serious adverse-event rates. They also do not establish the results of unreported subgroup analyses, missing-data sensitivity analyses, or multiplicity procedures.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods and calculators connected to randomized trials, longitudinal outcomes, binary endpoints, confidence intervals, and treatment-effect estimation.

26. Record Summary

STEP 3 provides a useful statistical case study because the registry reports two primary endpoints evaluated through multiple modeling approaches. For change in body weight, ANCOVA produced a treatment difference of -10.27 (95% CI -11.97 to -8.57; P < .0001) under a treatment-policy estimand, while MMRM produced a treatment difference of -12.67 (95% CI -14.34 to -11.00; P <0.0001) under a hypothetical estimand. For achievement of at least 5% body-weight reduction after 68 weeks, logistic regression reported an odds ratio of 6.11 (95% CI 4.04 to 9.26; P <0.0001), while the MMRM analysis reported an odds ratio of 11.67 (95% CI 7.64 to 17.81; P <0.0001).

The principal statistical lesson is that a trial result is more than its P-value. The endpoint definition, analysis population, model, effect measure, estimand, confidence interval, and testing framework all contribute to what a reported number means. STEP 3 illustrates this directly by reporting different effect estimates for the same endpoint under different estimand and modeling specifications.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical evidence from the statistical interpretation of that evidence. For STEP 3, the ClinicalTrials.gov record supports a detailed discussion of ANCOVA, MMRM, logistic regression, odds ratios, confidence intervals, intention-to-treat concepts, randomization, masking, and estimands, while leaving unreported baseline, subgroup, imputation, and multiplicity details outside the analysis.