This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
STEP 2 was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating semaglutide in people with type 2 diabetes suffering from overweight or obesity. The registry reports 1210 enrolled participants, three arms, two registered primary endpoints, and formal analyses using ANCOVA, MMRM, and logistic regression.
| Feature | STEP 2 |
|---|---|
| Phase | Phase 3 |
| Therapeutic area | Endocrinology |
| Conditions | Obesity; Overweight |
| Design | Randomized, parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 1210 |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
| Status | Completed |
| ClinicalTrials.gov | NCT03552757 |
2. Clinical Question
The primary statistical comparison reported in the registry is semaglutide 2.4 mg versus placebo. The registered primary endpoints ask whether semaglutide 2.4 mg changes body weight from baseline to week 68 and increases the probability that participants achieve at least a 5% reduction in baseline body weight at week 68.
Population
People with type 2 diabetes suffering from overweight or obesity.
Intervention
Semaglutide 2.4 mg. The broader trial included semaglutide 1.0 mg as an additional study arm.
Comparator
Placebo.
Primary question
Does semaglutide 2.4 mg produce a different change in body weight and a different probability of achieving at least 5% body-weight reduction compared with placebo?
3. Trial Design
Semaglutide 1.0 mg
- Drug intervention
- Included in the three-arm parallel design
- Serious adverse events: 31/402
Semaglutide 2.4 mg
- Drug intervention
- Primary efficacy comparison versus placebo
- Serious adverse events: 40/403
Placebo
- Placebo intervention
- Primary comparator for the reported efficacy analyses
- Serious adverse events: 37/402
The registry lists the interventions as Semaglutide 1.0 mg, Semaglutide 2.4 mg, Placebo I (Semaglutide), and Placebo II (Semaglutide). The posted primary statistical analyses identify the comparison groups as semaglutide 2.4 mg versus placebo.
4. Primary Endpoints
| Endpoint | Time frame | Registry definition / analysis |
|---|---|---|
| Change in Body Weight (%) - Semaglutide 2.4 mg Versus Placebo | Baseline (week 0) to week 68 | Change in body weight (%) from baseline (week 0) to week 68. Results are based on both in-trial and on-treatment observation periods. The posted analyses use ANCOVA for the in-trial period and MMRM for the on-treatment period. |
| Participants Who Achieve (Yes/no): Body Weight Reduction ≥5% - Semaglutide 2.4 mg Versus Placebo | At week 68 | Number of participants who achieved weight reduction ≥5% of baseline body weight (yes/no) at week 68. Results are reported for both in-trial and on-treatment observation periods. |
In-trial versus on-treatment observation
The registry makes an important distinction between two observation periods. The in-trial observation period is described as the uninterrupted interval from the start of randomization at week 0 to the last trial-related subject-site contact, with the registry-reported definition specifying week 75 for the body-weight endpoint. The on-treatment observation period focuses on responses before treatment discontinuation or initiation of other anti-obesity medication or bariatric surgery.
These are not interchangeable estimands. A result based on the in-trial period asks about outcomes during the randomized trial follow-up, whereas an on-treatment analysis restricts attention to observations while the assigned treatment strategy remains relevant under the registry's specified rules.
5. Statistical Methodology
ANCOVA for change in body weight
The in-trial change in body weight was analyzed with an analysis of covariance (ANCOVA). The registry states that the model included randomized treatment, stratification groups, and their interaction as factors, with baseline body weight as a covariate.
The important statistical feature is that the week-68 response is compared after accounting for baseline body weight and the specified stratification structure rather than relying only on an unadjusted difference between observed means.
MMRM for the on-treatment analysis
The on-treatment analysis used a mixed model for repeated measurements (MMRM). The registry states that all responses before first discontinuation of treatment, initiation of another anti-obesity medication, or bariatric surgery were included. The model incorporated randomized treatment, the two stratification groups and their interaction, with baseline body weight as a covariate.
An MMRM uses the longitudinal structure of the data rather than reducing every participant to a single observed value. Its interpretation depends on the model specification and the observations included under the defined on-treatment period.
Logistic regression for ≥5% weight reduction
The binary endpoint was analyzed using logistic regression. The response is whether or not a participant achieved at least a 5% reduction in baseline body weight at week 68.
An odds ratio above 1 indicates higher estimated odds of meeting the binary endpoint in the semaglutide group. The odds ratio is not itself a probability difference or a risk ratio.
Stratified covariate adjustment
The ANCOVA and logistic-regression analyses incorporated the same two named screening stratification factors: oral anti-diabetic (OAD) treatment status and HbA1c category at screening. Their interaction was also included as a factor, and baseline body weight was included as a covariate.
Superiority testing
All four posted primary analyses are identified as superiority analyses. The reported confidence intervals are two-sided 95% confidence intervals, while the ClinicalTrials.gov record does not identify a separate one-sided alpha value.
6. Primary Results: Change in Body Weight
The first registered primary endpoint was change in body weight from baseline (week 0) to week 68, comparing semaglutide 2.4 mg with placebo.
In-trial observation period — ANCOVA
Treatment difference in change in body weight
95% CI: -7.28 to -5.15 · P < 0.0001
Analysis: ANCOVA · 95% CI, two-sided · Superiority
The estimated treatment difference was -6.21 percentage points of body weight for semaglutide 2.4 mg versus placebo under the in-trial observation analysis. Because the estimate is negative, the modeled change in body weight was lower in the semaglutide group relative to placebo under the direction used for this treatment difference.
What the estimate means: the reported treatment difference of -6.21 is the model-based difference in change in body weight, expressed in percentage points of body weight, between semaglutide 2.4 mg and placebo.
What it does not mean: it is not a statement that every participant lost exactly 6.21 percentage points of body weight, nor does it describe the individual treatment response distribution.
Precision: the two-sided 95% confidence interval extends from -7.28 to -5.15. It describes uncertainty around the estimated treatment difference under the specified ANCOVA framework; it does not describe the range of responses among individual participants.
P-value: P < 0.0001 addresses evidence against the null hypothesis under the statistical testing framework. It does not measure the magnitude of the treatment effect. The effect size is described by the treatment difference and its confidence interval.
Model context: the estimate is adjusted for the prespecified stratification groups, their interaction, and baseline body weight. The analysis population was the FAS, comprising all randomized participants, with the number analyzed defined as the number with available data.
On-treatment observation period — MMRM
Treatment difference in change in body weight
95% CI: -8.56 to -6.58 · P < 0.0001
Analysis: MMRM · 95% CI, two-sided · Superiority
The MMRM analysis produced a treatment difference of -7.57 percentage points of body weight for semaglutide 2.4 mg versus placebo. This analysis was based on the on-treatment observation period and included responses before first discontinuation of treatment or initiation of other anti-obesity medication or bariatric surgery.
What the estimate means: the -7.57 estimate represents the reported model-based treatment difference in change in body weight for the on-treatment analysis.
What it does not mean: it should not be read as a guaranteed individual weight change or as an absolute percentage of participants who benefit.
Precision: the 95% confidence interval of -8.56 to -6.58 quantifies uncertainty around this estimated treatment difference under the MMRM framework.
P-value: P < 0.0001 indicates strong statistical evidence against the null hypothesis used for this superiority comparison. It does not tell us that the treatment effect is clinically large, nor does it replace the effect estimate and confidence interval.
Important comparison: the ANCOVA and MMRM estimates should not be treated as two independent replications of exactly the same estimand. They correspond to different observation-period definitions and statistical approaches.
7. Comparing the Two Change-in-Weight Analyses
| Feature | In-trial analysis | On-treatment analysis |
|---|---|---|
| Method | ANCOVA | MMRM |
| Estimate | -6.21 | -7.57 |
| 95% CI | -7.28 to -5.15 | -8.56 to -6.58 |
| P-value | <0.0001 | <0.0001 |
| Observation period | In-trial | On-treatment |
| Covariate / factors | Baseline body weight; OAD treatment status; HbA1c category; interaction between stratification groups | Baseline body weight; OAD treatment status; HbA1c category; interaction between stratification groups |
The two estimates are not interchangeable. The in-trial analysis retains the randomized-trial observation framework, while the on-treatment analysis uses responses before specified treatment-discontinuation or alternative-treatment events. The numerical difference between -6.21 and -7.57 therefore should not be interpreted as a simple measure of statistical disagreement.
8. Primary Results: Participants Achieving ≥5% Body-Weight Reduction
The second registered primary endpoint was the binary outcome of whether a participant achieved a body-weight reduction of at least 5% from baseline at week 68.
In-trial observation period — Logistic Regression
Odds ratio for achieving ≥5% reduction
95% CI: 3.58 to 6.64 · P < 0.0001
Analysis: Logistic regression · 95% CI, two-sided · Superiority
The reported odds ratio of 4.88 means that the estimated odds of achieving at least a 5% reduction in baseline body weight were 4.88 times the odds under placebo, within the specified in-trial logistic-regression model.
What the estimate means: an odds ratio of 4.88 indicates substantially higher estimated odds of meeting the ≥5% endpoint under semaglutide 2.4 mg than under placebo.
What it does not mean: an OR of 4.88 does not mean that 88% of participants achieved the endpoint, nor does it mean that the probability was 4.88 times higher. Odds and probabilities are different quantities.
Precision: the 95% confidence interval from 3.58 to 6.64 quantifies uncertainty around the estimated odds ratio. It does not describe individual treatment responses.
P-value: P < 0.0001 assesses evidence against the null hypothesis for the treatment comparison. It does not measure the size or clinical importance of the effect.
Adjustment: the logistic model included randomized treatment, OAD treatment status, HbA1c category at screening, their interaction, and baseline body weight.
On-treatment observation period — Logistic Regression
Odds ratio for achieving ≥5% reduction
95% CI: 6.31 to 11.97 · P < 0.0001
Analysis: Logistic regression · 95% CI, two-sided · Superiority
For the on-treatment observation period, the reported odds ratio was 8.69. This is the estimated odds ratio from the posted analysis under the on-treatment observation rules.
What the estimate means: an OR of 8.69 indicates that the estimated odds of achieving at least a 5% reduction were 8.69 times those under placebo in the specified on-treatment analysis.
What it does not mean: it is not a risk ratio, probability ratio, percentage-point difference, or statement that 8.69 times as many participants necessarily achieved the endpoint.
Precision: the 95% CI of 6.31 to 11.97 provides the uncertainty interval around the model-based odds ratio.
P-value: P < 0.0001 indicates evidence against the null hypothesis in the specified superiority test. The P-value does not quantify the magnitude of the treatment effect.
Observation-period caution: this estimate reflects the on-treatment analysis rules and therefore should not be directly substituted for the in-trial estimate of 4.88.
9. Summary of Primary Statistical Results
| Primary endpoint | Observation period | Method | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Change in Body Weight (%) | In-trial | ANCOVA | Treatment difference: -6.21 | -7.28 to -5.15 | <0.0001 |
| Change in Body Weight (%) | On-treatment | MMRM | Treatment difference: -7.57 | -8.56 to -6.58 | <0.0001 |
| Body Weight Reduction ≥5% | In-trial | Logistic regression | OR: 4.88 | 3.58 to 6.64 | <0.0001 |
| Body Weight Reduction ≥5% | On-treatment | Logistic regression | OR: 8.69 | 6.31 to 11.97 | <0.0001 |
10. Statistical Methods Explained
Why was ANCOVA used for change in body weight?
ANCOVA is appropriate when the outcome of interest is a continuous measurement and the analysis benefits from adjusting for an important baseline measurement. Here, the posted model uses baseline body weight as a covariate while also accounting for randomized treatment and the specified stratification structure. This can produce a treatment-effect estimate that is adjusted for baseline differences rather than relying solely on an unadjusted comparison of week-68 responses.
Why does baseline body weight appear in the model?
Baseline body weight is measured before the treatment comparison develops, so it can provide information about the expected week-68 response. Including it as a covariate can improve the precision of the treatment comparison and explicitly account for baseline body-weight variation. The registry specifically identifies baseline body weight as a covariate in the ANCOVA and logistic-regression analyses and in the MMRM analysis.
Why use MMRM for the on-treatment analysis?
MMRM is designed for longitudinal data in which participants can contribute repeated observations. Rather than treating repeated measurements from the same participant as unrelated observations, the mixed-model framework accounts for their longitudinal structure. In STEP 2, the registry specifies that responses before particular treatment-discontinuation or alternative-treatment events were included in the on-treatment MMRM.
What does an odds ratio of 4.88 mean?
An odds ratio of 4.88 means the modeled odds of achieving at least a 5% reduction in baseline body weight were 4.88 times the corresponding odds under placebo. If the placebo probability were known, the odds ratio could be translated into a probability comparison. Without those probabilities, however, the OR should not be presented as a probability ratio or percentage-point increase.
Why is the odds ratio of 8.69 different from 4.88?
The two estimates come from different observation periods. The 4.88 estimate is based on the in-trial observation period, while the 8.69 estimate is based on the on-treatment observation period. Changing which observations qualify for the analysis can change the estimated treatment effect even when the underlying randomized comparison is the same.
Why are stratification factors included in the models?
The registry identifies OAD treatment status and HbA1c category at screening as the stratification groups. Incorporating these factors into the analysis aligns the model with the prespecified stratification structure and adjusts the treatment comparison for those variables. Their interaction is also included, meaning the model allows the relationship between treatment and outcome to depend on the combination of the two stratification groups.
What does P < 0.0001 tell us?
It indicates strong evidence against the relevant null hypothesis under the specified statistical test. It does not tell us the probability that the null hypothesis is true, the probability that the treatment is effective in an individual participant, or the magnitude of the treatment effect. For magnitude, the estimate and confidence interval are essential.
11. Confidence Intervals and Effect Size
The four posted primary analyses all include two-sided 95% confidence intervals. These intervals are particularly useful because they put the point estimates into an uncertainty framework.
| Estimate | Interpretive scale | 95% CI |
|---|---|---|
| Treatment difference, in-trial change in body weight | -6.21 percentage points | -7.28 to -5.15 |
| Treatment difference, on-treatment change in body weight | -7.57 percentage points | -8.56 to -6.58 |
| OR, in-trial ≥5% reduction | 4.88 | 3.58 to 6.64 |
| OR, on-treatment ≥5% reduction | 8.69 | 6.31 to 11.97 |
For the treatment-difference endpoints, the entire reported 95% confidence interval is below zero. For the odds-ratio endpoints, the entire reported 95% confidence interval is above one. Those locations relative to the conventional null values are consistent with the reported P-values and the registry's classification of the analyses as superiority tests.
The confidence intervals should still be interpreted as uncertainty around model-based estimates, not as a prediction interval for individual participants or a statement that every possible treatment effect lies within the interval.
12. Analysis Population and Available Data
The posted analyses identify the full analysis set (FAS) as the analysis population. The FAS comprised all randomized participants. The registry further states that the number analyzed was the number of participants with available data.
This distinction matters because "all randomized participants" does not necessarily mean that every randomized participant contributes an observed week-68 value to every analysis. The registry explicitly defines the analyzed number for these results as the number with available data.
13. In-Trial and On-Treatment Estimands
One of the most useful statistical features of STEP 2 is that the registry reports both in-trial and on-treatment analyses. These analyses can be viewed as addressing different questions about the treatment comparison.
In-trial perspective
The analysis uses the in-trial observation period. It is anchored to the randomized trial follow-up and therefore provides an estimate associated with outcomes observed within that trial framework.
On-treatment perspective
The analysis focuses on responses before specified treatment discontinuation or initiation of another anti-obesity medication or bariatric surgery.
Neither perspective should automatically be treated as a replacement for the other. The distinction is substantive: one asks about the randomized trial experience over the defined in-trial period, while the other asks about outcomes while participants remain within the specified on-treatment conditions.
14. Secondary Endpoint Results
The registry reports 41 outcome measures and four posted statistical analyses, all four of which correspond to the two registered primary endpoints. The ClinicalTrials.gov record does not provide additional secondary-endpoint statistical estimates, confidence intervals, or P-values. Accordingly, no secondary efficacy result is added here.
15. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.
| Study arm | Serious adverse events | At risk |
|---|---|---|
| Semaglutide 1.0 mg | 31 | 402 |
| Semaglutide 2.4 mg | 40 | 403 |
| Placebo | 37 | 402 |
These figures describe serious adverse events by arm, but the ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, P-value, event definition details, or exposure-adjusted analysis. Therefore, the counts should be presented descriptively rather than converted into an unsupported inferential safety conclusion.
16. Multiplicity and Multiple Primary Analyses
The registry identifies two registered primary endpoints, and four statistical analyses are posted for those endpoints: two analyses of change in body weight and two analyses of the ≥5% body-weight-reduction endpoint.
| Endpoint | In-trial analysis | On-treatment analysis |
|---|---|---|
| Change in Body Weight (%) | ANCOVA | MMRM |
| Body Weight Reduction ≥5% | Logistic regression | Logistic regression |
The ClinicalTrials.gov record identifies the hypothesis type as superiority, but they do not provide an alpha-allocation or multiplicity-adjustment procedure for the two primary endpoints or the four posted analyses. It would therefore be inappropriate to invent a specific multiplicity strategy.
This is an important distinction between what the statistical record tells us and what a complete statistical analysis plan might contain. The presence of multiple endpoints creates a multiplicity question, but the ClinicalTrials.gov record does not establish how that question was handled.
17. Blinding, Randomization, and Statistical Validity
STEP 2 is identified as randomized and quadruple masked. Randomization is central to the validity of the treatment comparison because it establishes the treatment assignment independently of participants' subsequent outcomes, subject to the usual assumptions of a properly conducted randomized trial.
Randomization
Reduces systematic differences in treatment assignment and provides the foundation for the primary efficacy comparison.
Quadruple masking
Reduces opportunities for knowledge of treatment assignment to influence trial conduct or assessment. The ClinicalTrials.gov record does not identify the four masked roles.
Parallel design
Participants remain associated with their randomized study arm rather than serving as their own randomized comparator in a crossover structure.
Stratification
The efficacy models incorporate OAD treatment status and HbA1c category at screening, plus their interaction.
18. Why the Treatment Difference and Odds Ratio Are Different
The two primary endpoints illustrate two fundamental types of effect measures.
| Measure | Question answered | Null value | STEP 2 estimate |
|---|---|---|---|
| Treatment difference | How different is the modeled change in body weight between treatment groups? | 0 | -6.21 in-trial; -7.57 on-treatment |
| Odds ratio | How do the odds of achieving ≥5% weight reduction compare? | 1 | 4.88 in-trial; 8.69 on-treatment |
A treatment difference is naturally interpreted on the scale of the outcome itself. An odds ratio is a relative measure of odds. Because the underlying endpoints are different, the numerical values should not be placed on a common "effect-size" scale or compared simply by asking which number is larger.
19. Important Limitations and Interpretation Issues
- Different observation periods: in-trial and on-treatment analyses use different observation-period definitions and therefore address different treatment-effect questions.
- Available-data analysis: the registry defines the number analyzed as the number of participants with available data, but the ClinicalTrials.gov record does not provide a missing-data count or a detailed imputation strategy.
- Model dependence: ANCOVA and MMRM estimates depend on their respective model structures, covariates, factors, and analysis populations.
- Odds versus probability: the reported odds ratios cannot be interpreted as probability ratios without knowing the underlying event probabilities.
- Multiplicity: two registered primary endpoints generate a multiplicity issue, but the ClinicalTrials.gov record does not specify the formal error-control strategy.
- Safety inference: serious-adverse-event counts are available by arm, but the ClinicalTrials.gov record does not provide a formal inferential safety analysis.
- Endpoint classification: the registry's inferred binary classification for the primary endpoints should not be confused with the continuous-response statistical models used for change in body weight.
- Generalizability: the ClinicalTrials.gov record identifies the trial population as people with type 2 diabetes suffering from overweight or obesity; they do not provide a broader description of eligibility criteria or participant characteristics from which additional generalizability claims could be derived.
20. Why This Trial Matters Statistically
STEP 2 is a useful teaching example because the same randomized treatment comparison is expressed through several complementary statistical models and two different observation-period perspectives.
| Concept | How it appears in STEP 2 |
|---|---|
| Randomization | The trial is randomized with a parallel-group design. |
| Blinding | The registry identifies the study as quadruple masked. |
| ANCOVA | Used for the in-trial analysis of change in body weight. |
| Covariate adjustment | Baseline body weight is included as a covariate. |
| Stratified analysis | OAD treatment status and HbA1c category at screening are included as stratification groups, with their interaction. |
| MMRM | Used for the on-treatment analysis of change in body weight. |
| Logistic regression | Used for the binary ≥5% body-weight-reduction endpoint. |
| Odds ratio | Quantifies the relative odds of achieving the ≥5% endpoint. |
| Confidence intervals | All four posted primary analyses include two-sided 95% confidence intervals. |
| P-values | All four posted primary analyses report P < 0.0001. |
| Estimand perspective | Both in-trial and on-treatment observation periods are reported. |
| Safety analysis | Serious adverse events are reported descriptively by arm. |
21. Record Timeline
Trial start
The registry lists 2018-06-04 as the study start date.
Primary completion
The registry lists 2020-03-24 as the primary completion date.
Completed
The trial is listed as completed, with results posted in ClinicalTrials.gov.
22. A Statistical Reading of the Four Primary Results
The ANCOVA estimate of -6.21, with a 95% CI of -7.28 to -5.15, describes the modeled difference in change in body weight between semaglutide 2.4 mg and placebo during the in-trial observation period. The P-value was <0.0001.
The MMRM estimate of -7.57, with a 95% CI of -8.56 to -6.58, describes the corresponding model-based treatment difference under the on-treatment observation rules. The P-value was <0.0001.
The logistic-regression OR of 4.88, with a 95% CI of 3.58 to 6.64, indicates higher estimated odds of achieving the binary endpoint under semaglutide 2.4 mg than placebo in the in-trial analysis. The P-value was <0.0001.
The logistic-regression OR of 8.69, with a 95% CI of 6.31 to 11.97, indicates higher estimated odds of achieving the binary endpoint under semaglutide 2.4 mg than placebo in the on-treatment analysis. The P-value was <0.0001.
23. What These Results Do — and Do Not — Establish
Taken together, the four posted primary analyses provide consistent statistical evidence favoring the semaglutide 2.4 mg group under the specified superiority comparisons. The change-in-weight analyses report negative treatment differences, while the ≥5% threshold analyses report odds ratios above one.
That consistency does not eliminate the need to understand what each estimate represents. The treatment difference is a continuous-outcome effect measure; the odds ratio is a binary-outcome effect measure. The in-trial analyses and on-treatment analyses also use different observation-period definitions.
Evidence supported by the ClinicalTrials.gov record
The registry reports four formal primary analyses, each with an estimate, two-sided 95% confidence interval, and P-value of <0.0001.
Information not reported here
The ClinicalTrials.gov record does not provide detailed baseline characteristics, participant-level distributions, secondary-endpoint estimates, missing-data counts, or a formal multiplicity procedure.
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Calculators
26. Sources
- ClinicalTrials.gov: NCT03552757.
- PubMed: PMID 33667417.
- PubMed: PMID 40980163.
- PubMed: PMID 39226070.
- PubMed: PMID 38698650.
- PubMed: PMID 36801984.
Continue with the statistical methods behind STEP 2
Explore the underlying statistical concepts through focused tutorials and practical calculators for clinical-trial analysis.
27. Record Summary
STEP 2 provides a useful example of how a randomized phase 3 trial can express treatment effects through multiple statistical models and observation-period definitions. The registry reports two primary endpoints: change in body weight from baseline to week 68 and achievement of at least a 5% reduction in baseline body weight at week 68. The change endpoint was analyzed using ANCOVA for the in-trial period and MMRM for the on-treatment period, while the binary threshold endpoint was analyzed using logistic regression. All four posted primary analyses report two-sided 95% confidence intervals and P-values of <0.0001.
The most important statistical lesson is that the estimates must be interpreted on their proper scales. Treatment differences describe change in the continuous body-weight outcome, whereas odds ratios describe the relative odds of achieving the binary ≥5% endpoint. The in-trial and on-treatment analyses further illustrate how the definition of the observation period can change the estimand without changing the randomized treatment comparison.