This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the ClinicalTrials.gov record for NCT03611582.
1. Trial at a Glance
STEP 3 was a completed, randomized, quadruple-masked, parallel phase 3 trial in overweight and obesity. The registry reports 611 enrolled participants and two intervention groups: semaglutide and placebo.
| Feature | STEP 3 |
|---|---|
| Trial name | STEP 3 |
| ClinicalTrials.gov identifier | NCT03611582 |
| Phase | Phase 3 |
| Therapeutic area | Endocrinology |
| Conditions | Overweight; Obesity |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 611 |
| Interventions | Semaglutide; placebo |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
| Status | Completed |
2. Clinical Question
The central statistical question was whether semaglutide differed from placebo with respect to weight-related outcomes in participants with overweight or obesity. The registry identifies two primary endpoints: change in body weight from baseline to week 68 and achievement of at least 5% body-weight reduction after 68 weeks.
Population
Participants enrolled in a phase 3 study for overweight or obesity. The ClinicalTrials.gov record reports 611 enrolled participants.
Intervention
Semaglutide. The serious-adverse-event data identify the semaglutide arm as semaglutide 2.4 mg.
Comparator
Placebo, described in the registry data as placebo for semaglutide.
Primary question
How does semaglutide compare with placebo for change in body weight and for achieving at least 5% weight reduction after 68 weeks?
3. Trial Design
Semaglutide 2.4 mg
- Semaglutide
- Serious adverse events reported for 37 of 407 participants at risk
Placebo
- Placebo for semaglutide
- Serious adverse events reported for 6 of 204 participants at risk
4. Endpoints
| Endpoint | Registered time frame | Endpoint type | Reported analysis |
|---|---|---|---|
| Change in Body Weight (%) | Baseline (week 0) to week 68 | Binary in the registry classification | ANCOVA; MMRM |
| Participants Who Achieve (Yes/no): Body Weight Reduction More Than or Equal to 5% | After 68 weeks | Binary | Logistic regression; MMRM |
The registry states that the body-weight endpoint is presented as change in body weight from baseline (week 0) to week 68. For the ≥5% endpoint, “Yes” denotes participants who achieved greater than or equal to 5% weight loss and “No” denotes participants who did not achieve that threshold.
5. Statistical Methodology
Full analysis set and intention-to-treat principle
The reported efficacy analyses use the full analysis set (FAS), which the registry describes as comprising all randomized participants. For each reported analysis, the number analyzed is the number of participants with available data. The registry also identifies intention-to-treat analysis as a concept in the analysis text.
This distinction matters because the randomized population defines the treatment comparison, while the number contributing observed data to a particular analysis can be smaller.
ANCOVA
ANCOVA, or analysis of covariance, is a linear-model approach commonly used when comparing a continuous outcome between randomized treatment groups while accounting for covariates specified in the model. In STEP 3, the registry reports ANCOVA for the baseline-to-week-68 body-weight endpoint and reports a treatment difference as the effect measure.
MMRM
The registry also reports a mixed model for repeated measures (MMRM). MMRM is designed for longitudinal data in which outcomes are measured repeatedly over time. Rather than reducing all observations to a single change score, the model can use the pattern of repeated observations to estimate treatment differences at the relevant time point.
Logistic regression
The ≥5% body-weight-reduction endpoint is binary: participants are classified as achieving the threshold or not achieving it. Logistic regression models the probability of the binary outcome and naturally expresses the treatment comparison through an odds ratio.
Treatment policy and hypothetical estimands
The registry distinguishes the estimand associated with the analyses. The ANCOVA and logistic-regression analyses are identified as using a treatment policy estimand, whereas the corresponding MMRM analyses are identified as using a hypothetical estimand. These are not merely different statistical algorithms; they represent different questions about how the treatment effect is defined in relation to post-randomization events and treatment experience.
Treatment policy
The treatment effect is defined under a strategy in which the outcome is considered regardless of intercurrent events in the way specified by the treatment-policy strategy.
Hypothetical
The treatment effect addresses a specified hypothetical scenario concerning what would have happened under the defined treatment condition.
6. Results: Change in Body Weight (%)
The first registered primary endpoint is change in body weight from baseline (week 0) to week 68. Two formal analyses are posted in the ClinicalTrials.gov record: ANCOVA using a treatment policy estimand and MMRM using a hypothetical estimand.
ANCOVA — treatment policy estimand
Semaglutide 2.4 mg vs placebo
Treatment difference · 95% CI -11.97 to -8.57
Two-sided P < .0001 · Superiority
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Change in Body Weight (%) | ANCOVA | Treatment difference | -10.27 | -11.97 to -8.57 | <.0001 |
The reported treatment difference of -10.27 indicates that the estimated change in the semaglutide group was 10.27 percentage points lower than the corresponding placebo-group value under the reported ANCOVA treatment-policy analysis.
The negative sign identifies the direction of the difference as reported; it does not by itself describe the absolute change within either treatment group.
The two-sided 95% confidence interval, -11.97 to -8.57, describes the statistical uncertainty around the estimated treatment difference under the analysis framework. It is an interval for the treatment-effect estimate, not a range containing individual participants' weight changes.
The P-value of <.0001 addresses the evidence against the null hypothesis specified for the superiority comparison. It does not measure the size of the treatment effect, clinical importance, or the probability that the null hypothesis is true.
The treatment-policy estimand is also important: this estimate answers the question defined by that estimand strategy, rather than automatically answering the hypothetical question represented by the separate MMRM analysis.
MMRM — hypothetical estimand
Semaglutide 2.4 mg vs placebo
Treatment difference · 95% CI -14.34 to -11.00
Two-sided P < 0.0001 · Superiority
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Change in Body Weight (%) | MMRM | Treatment difference | -12.67 | -14.34 to -11.00 | <0.0001 |
The MMRM estimate of -12.67 represents the reported treatment difference under the hypothetical estimand and repeated-measures model. As with the ANCOVA result, the negative value indicates a lower estimated body-weight change in the semaglutide group relative to placebo under the specified analysis.
The 95% CI of -14.34 to -11.00 quantifies uncertainty around this model-based estimate. Its width provides information about precision, while the interval itself does not describe the variability of individual weight changes.
The P-value of <0.0001 provides evidence against the null hypothesis used for the superiority comparison. A P-value should not be read as a measure of effect magnitude.
The estimate differs from the ANCOVA estimate because the analyses are not identical: the MMRM uses repeated-measures modeling and is identified in the registry as addressing a hypothetical estimand, while the ANCOVA result is identified as a treatment-policy analysis. The two estimates therefore should not simply be averaged or treated as independent replications of exactly the same estimand.
7. Results: Participants Achieving at Least 5% Weight Reduction
The second registered primary endpoint is the binary outcome “Participants Who Achieve (Yes/no): Body Weight Reduction More Than or Equal to 5%,” assessed after 68 weeks. The registry defines “Yes” as achieving ≥5% weight loss and “No” as not achieving ≥5% weight loss.
Logistic regression — treatment policy estimand
Odds ratio for achieving ≥5% weight reduction
95% CI: 4.04 to 9.26
Two-sided P < 0.0001 · Superiority
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| ≥5% body-weight reduction after 68 weeks | Logistic regression | Odds ratio | 6.11 | 4.04 to 9.26 | <0.0001 |
An odds ratio of 6.11 means that the reported odds of achieving at least 5% weight reduction were estimated to be 6.11 times as high with semaglutide as with placebo under the treatment-policy logistic-regression analysis.
An odds ratio is not a risk ratio and is not equivalent to saying that six times as many participants achieved the endpoint. The corresponding probabilities depend on the underlying event rates, which are not provided in the ClinicalTrials.gov record.
The 95% CI of 4.04 to 9.26 describes uncertainty around the odds-ratio estimate. It does not describe the range of individual treatment effects.
The two-sided P-value of <0.0001 addresses the hypothesis test; it does not quantify the magnitude or clinical importance of the odds ratio.
Because the analysis is based on a binary endpoint, the interpretation is specifically about achievement of the ≥5% threshold after 68 weeks, not about the continuous magnitude of weight change.
MMRM — hypothetical estimand
Odds ratio for achieving ≥5% weight reduction
95% CI: 7.64 to 17.81
Two-sided P < 0.0001 · Superiority
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| ≥5% body-weight reduction after 68 weeks | MMRM | Odds ratio | 11.67 | 7.64 to 17.81 | <0.0001 |
The reported odds ratio of 11.67 indicates substantially higher estimated odds of achieving the ≥5% threshold under the MMRM analysis and its hypothetical estimand.
The 95% CI of 7.64 to 17.81 provides the uncertainty interval for that estimate under the reported model. The interval is entirely above 1, consistent with the reported superiority test.
Again, this is an odds ratio rather than a probability ratio. Without the underlying event counts or probabilities from the ClinicalTrials.gov record, the odds ratio cannot be converted into an absolute percentage-point difference without additional information.
The difference between the reported OR of 6.11 from logistic regression and 11.67 from the MMRM analysis should be understood in the context of their different modeling and estimand specifications rather than treated as contradictory estimates of exactly the same statistical question.
8. Primary Results Side by Side
| Primary endpoint | Analysis | Estimand | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| Change in Body Weight (%) | ANCOVA | Treatment policy | -10.27 treatment difference | -11.97 to -8.57 | <.0001 |
| Change in Body Weight (%) | MMRM | Hypothetical | -12.67 treatment difference | -14.34 to -11.00 | <0.0001 |
| ≥5% weight reduction | Logistic regression | Treatment policy | OR 6.11 | 4.04 to 9.26 | <0.0001 |
| ≥5% weight reduction | MMRM | Hypothetical | OR 11.67 | 7.64 to 17.81 | <0.0001 |
9. Statistical Methods Explained
Why was ANCOVA used for change in body weight?
ANCOVA is a natural framework for comparing a continuous outcome between randomized groups while incorporating covariate information specified by the analysis model. The registry reports ANCOVA specifically for the baseline-to-week-68 body-weight endpoint and expresses the result as a treatment difference.
What does a treatment difference of -10.27 mean?
It is the reported estimated difference between the semaglutide and placebo groups under the ANCOVA treatment-policy analysis. The negative sign indicates the direction of the contrast as defined by the reported group comparison. It is not an odds ratio and should not be interpreted as a percentage probability.
Why use MMRM?
MMRM is designed for repeated measurements. Longitudinal trials can contain multiple observations per participant, and an MMRM can model those observations jointly rather than relying only on one observed change value. In STEP 3, the registry reports MMRM for both primary endpoints.
What does an odds ratio of 6.11 mean?
An OR of 6.11 means that the estimated odds of achieving the binary ≥5% weight-loss endpoint were 6.11 times as high in the semaglutide group as in the placebo group under the reported logistic-regression analysis. It does not mean that the probability was six times higher.
Why are there two different odds ratios for the same binary endpoint?
The registry reports logistic regression with a treatment-policy estimand and MMRM with a hypothetical estimand. Because these analyses address differently defined statistical questions and use different model structures, their estimates need not be identical.
What does the 95% confidence interval tell us?
The confidence interval describes uncertainty around the estimated treatment effect under the specified statistical model and sampling framework. For example, the ANCOVA estimate is -10.27 with a 95% CI from -11.97 to -8.57. The interval does not represent the range of responses among individual participants.
Why does the P-value not measure effect size?
A P-value quantifies evidence against a specified null hypothesis under the statistical model. It depends on both the estimated effect and the amount of information available. The effect estimate and confidence interval therefore remain necessary for understanding magnitude and precision.
10. Confidence Intervals and Precision
The four reported primary analyses all include two-sided 95% confidence intervals. Looking at the interval alongside the point estimate provides substantially more information than the P-value alone.
| Analysis | Point estimate | 95% confidence interval | What the interval describes |
|---|---|---|---|
| ANCOVA | -10.27 | -11.97 to -8.57 | Uncertainty around the treatment difference |
| MMRM, body weight | -12.67 | -14.34 to -11.00 | Uncertainty around the MMRM treatment difference |
| Logistic regression | OR 6.11 | 4.04 to 9.26 | Uncertainty around the odds ratio |
| MMRM, ≥5% endpoint | OR 11.67 | 7.64 to 17.81 | Uncertainty around the odds ratio |
The point estimate gives the central result from the fitted analysis; the confidence interval communicates how precisely that effect was estimated under the stated assumptions.
11. Understanding the Two Estimand Strategies
One of the most statistically informative features of the registry-reported STEP 3 record is that it reports both treatment-policy and hypothetical estimands.
Treatment-policy question
The registry identifies the ANCOVA and logistic-regression analyses as treatment-policy analyses. This estimand is intended to characterize the treatment effect under the specified treatment-policy strategy rather than restricting the question to an idealized uninterrupted treatment scenario.
Hypothetical question
The registry identifies the MMRM analyses as hypothetical. This changes the target quantity: the analysis addresses the treatment effect under the hypothetical scenario defined by that estimand.
This distinction is important because a numerical difference between two estimates can arise from differences in the statistical question being answered, not necessarily from random fluctuation or a contradiction in the underlying data.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. The available measure is the number affected divided by the number at risk.
| Safety measure | Semaglutide 2.4 mg | Placebo |
|---|---|---|
| Serious adverse events | 37/407 | 6/204 |
Semaglutide arm
Serious adverse events affected 37 of 407 participants at risk in the ClinicalTrials.gov record.
Placebo arm
Serious adverse events affected 6 of 204 participants at risk in the ClinicalTrials.gov record.
The serious-adverse-event figures should be kept separate from the primary efficacy analyses. The efficacy estimates address weight-related outcomes, whereas serious adverse events describe an aspect of safety. They are different outcome domains and require different statistical interpretation.
13. Randomization and Blinding
STEP 3 is described in the registry as randomized, parallel, and quadruple-masked. These design characteristics are relevant to the interpretation of the treatment comparison.
Randomization and masking address different sources of potential bias. Randomization concerns how participants are assigned to treatment groups, while masking concerns knowledge of those assignments during the conduct and assessment of the trial.
14. Analysis Populations and Missing Data
The reported primary analyses use the full analysis set, which the registry defines as all randomized participants. The number analyzed is the number of participants with available data for the particular analysis.
| Concept | STEP 3 registry description | Statistical implication |
|---|---|---|
| Full analysis set | All randomized participants | Maintains the randomized treatment framework for efficacy analysis |
| Number analyzed | Participants with available data | Can be smaller than the randomized population for a particular analysis |
| Intention-to-treat concept | Identified in the analysis text | Supports interpretation according to randomized assignment |
| Estimand | Treatment policy or hypothetical, depending on analysis | Defines the treatment effect being targeted |
The ClinicalTrials.gov record does not specify a particular missing-data imputation method, a detailed missing-data sensitivity analysis, or a pattern-mixture model. Accordingly, none is attributed to STEP 3 here.
15. Multiplicity and Interim Analysis
The ClinicalTrials.gov record identifies two primary endpoints and four posted primary-endpoint analyses. They also identify superiority as the hypothesis type. However, the ClinicalTrials.gov record does not provide an alpha-spending procedure, interim-analysis boundary, multiplicity-adjustment procedure, or detailed endpoint hierarchy.
This distinction is especially important when multiple formal analyses are presented. A P-value is interpreted within the testing framework that generated it; without the prespecified multiplicity structure, one should not reconstruct an error-control procedure from the observed results.
16. Statistical Methods: What Each Model Contributes
| Method | STEP 3 role | Effect measure | Estimand |
|---|---|---|---|
| ANCOVA | Change in Body Weight (%) | Treatment difference | Treatment policy |
| MMRM | Change in Body Weight (%) | Treatment difference | Hypothetical |
| Logistic regression | ≥5% body-weight reduction | Odds ratio | Treatment policy |
| MMRM | ≥5% body-weight reduction | Odds ratio | Hypothetical |
The model family should not be confused with the estimand. ANCOVA and logistic regression describe statistical modeling approaches, while treatment policy and hypothetical describe the target treatment effect. A complete interpretation therefore needs both pieces of information.
17. Interpreting the Odds Ratio
The treatment-policy logistic-regression analysis reported an odds ratio of 6.11 for achieving at least 5% body-weight reduction. This means the estimated odds were 6.11 times as high in the semaglutide group as in the placebo group under that analysis.
The MMRM analysis reported an odds ratio of 11.67 under a hypothetical estimand. Because the two estimates correspond to different analysis specifications, they should not be interpreted as two measurements of a single unchanged odds ratio.
If an event probability is p, the corresponding odds are p/(1-p). An odds ratio therefore compares odds rather than probabilities. Without the underlying event probabilities, an odds ratio alone cannot provide the absolute difference in the percentage of participants achieving the endpoint.
18. Interpreting the Body-Weight Treatment Differences
The continuous body-weight endpoint provides a different kind of effect measure from the binary ≥5% endpoint.
Continuous endpoint
The ANCOVA and MMRM analyses report treatment differences of -10.27 and -12.67, respectively. These are differences on the scale of the registered percentage-change endpoint.
Threshold endpoint
The logistic-regression and MMRM analyses report odds ratios for crossing the prespecified ≥5% weight-reduction threshold.
A continuous treatment difference and an odds ratio answer different questions. The continuous analysis concerns the estimated difference in the outcome itself, whereas the binary analysis asks about the likelihood of meeting a predefined threshold.
19. Trial Timeline
Trial start
The ClinicalTrials.gov record lists 2018-08-01 as the study start date.
Primary completion
The ClinicalTrials.gov record lists 2020-03-18 as the primary completion date.
Registry status
The trial is listed as completed, with results posted and four statistical analyses posted for the two primary endpoints.
20. Limitations and Interpretation Issues
- Different estimands: The treatment-policy and hypothetical analyses target different estimand strategies, so their numerical estimates should not be treated as interchangeable.
- Registry classification: The ClinicalTrials.gov record classifies Change in Body Weight (%) as binary while reporting treatment differences from ANCOVA and MMRM. This page preserves the registry-reported classification rather than silently resolving the discrepancy.
- Available-data analysis: The FAS comprises all randomized participants, but the number analyzed is defined as the number with available data. The ClinicalTrials.gov record does not provide the endpoint-specific numbers analyzed.
- Missing-data methods: The ClinicalTrials.gov record does not specify a particular imputation procedure or sensitivity-analysis strategy. MMRM should not automatically be described as a specific imputation method.
- Multiplicity: Two primary endpoints and four posted primary analyses are reported, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure or alpha-spending framework.
- Safety comparison: Serious adverse-event counts are provided, but no formal inferential comparison is reported.
- Endpoint results: The ClinicalTrials.gov record provides effect estimates, confidence intervals, and P-values but do not provide the underlying group-level body-weight estimates or the numbers achieving the ≥5% threshold.
- Generalizability: The ClinicalTrials.gov record identifies the study population as participants with overweight or obesity, but does not provide the detailed baseline characteristics needed to characterize the enrolled population more fully.
21. Why This Trial Matters Statistically
STEP 3 is a useful teaching example because the ClinicalTrials.gov record contains several important statistical concepts within a relatively compact randomized design: continuous and binary primary endpoints, ANCOVA, MMRM, logistic regression, odds ratios, treatment-policy and hypothetical estimands, intention-to-treat concepts, confidence intervals, and superiority testing.
| Concept | How it appears in STEP 3 |
|---|---|
| Randomization | Randomized allocation in a parallel phase 3 design |
| Blinding | Quadruple-masked design |
| Intention-to-treat | Identified in the reported primary analyses |
| Full analysis set | All randomized participants |
| ANCOVA | Primary analysis of change in body weight (%) |
| MMRM | Reported for both primary endpoints |
| Logistic regression | Primary analysis of the ≥5% weight-reduction endpoint |
| Odds ratio | Effect measure for the binary ≥5% endpoint |
| Confidence interval | 95% two-sided intervals reported for all four primary analyses |
| P-value | Two-sided superiority tests reported as <.0001 or <0.0001 |
| Estimands | Treatment-policy and hypothetical estimands reported |
| Missing data | Number analyzed defined as participants with available data; detailed imputation strategy not reported |
22. Independent Statistical Interpretation
The most direct statistical reading of the ClinicalTrials.gov record is that all four posted primary-endpoint analyses report strong evidence under their respective superiority tests, with two-sided P-values below 0.0001 or .0001 as reported. The magnitude and scale of the effects differ by endpoint and model.
Continuous weight-change endpoint
Reported treatment differences across ANCOVA and MMRM analyses
≥5% weight-reduction endpoint
Reported odds ratios across logistic regression and MMRM analyses
These ranges should not be interpreted as uncertainty intervals for one common effect. They represent estimates from different analytical specifications. The key statistical task is therefore to identify the endpoint, model, estimand, effect measure, confidence interval, and P-value together rather than selecting a single number without its analytical context.
The ClinicalTrials.gov record does not provide individual-level treatment effects, the percentage of participants achieving the ≥5% threshold in each group, or a formal comparison of serious adverse-event rates. They also do not establish the results of unreported subgroup analyses, missing-data sensitivity analyses, or multiplicity procedures.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: NCT03611582 — STEP 3.
- PubMed: PMID 33625476.
- PubMed: PMID 39226070.
- PubMed: PMID 38698650.
- PubMed: PMID 36801984.
- PubMed: PMID 35724304.
Continue through the Clinical Biostats statistical pathway
Explore the statistical methods and calculators connected to randomized trials, longitudinal outcomes, binary endpoints, confidence intervals, and treatment-effect estimation.
26. Record Summary
STEP 3 provides a useful statistical case study because the registry reports two primary endpoints evaluated through multiple modeling approaches. For change in body weight, ANCOVA produced a treatment difference of -10.27 (95% CI -11.97 to -8.57; P < .0001) under a treatment-policy estimand, while MMRM produced a treatment difference of -12.67 (95% CI -14.34 to -11.00; P <0.0001) under a hypothetical estimand. For achievement of at least 5% body-weight reduction after 68 weeks, logistic regression reported an odds ratio of 6.11 (95% CI 4.04 to 9.26; P <0.0001), while the MMRM analysis reported an odds ratio of 11.67 (95% CI 7.64 to 17.81; P <0.0001).
The principal statistical lesson is that a trial result is more than its P-value. The endpoint definition, analysis population, model, effect measure, estimand, confidence interval, and testing framework all contribute to what a reported number means. STEP 3 illustrates this directly by reporting different effect estimates for the same endpoint under different estimand and modeling specifications.