This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
STEP 1 was a randomized, parallel-group, quadruple-masked phase 3 trial in the therapeutic area of endocrinology. The registry describes the study as investigating how well semaglutide works in people suffering from overweight or obesity. The trial enrolled 1,961 participants and compared semaglutide with placebo.
| Feature | STEP 1 |
|---|---|
| Trial name | STEP 1 |
| NCT ID | NCT03548935 |
| Phase | Phase 3 |
| Therapeutic area | Endocrinology |
| Conditions | Metabolism and Nutrition Disorder; Overweight or Obesity |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 1,961 |
| Interventions | Semaglutide (drug); Placebo (semaglutide) (drug) |
| Primary endpoints | 2 |
| Results posted | Yes |
| Outcome measures posted | 42 |
| Statistical analyses posted | 4 |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
2. Clinical Question
The central statistical question is whether semaglutide 2.4 mg produces a different body-weight outcome than placebo in people with overweight or obesity, and whether participants receiving semaglutide are more likely to achieve at least 5% body-weight reduction.
Population
People with overweight or obesity, within the registered conditions of metabolism and nutrition disorder; overweight or obesity.
Intervention
Semaglutide, as registered in the trial's intervention list and statistical analyses as semaglutide 2.4 mg.
Comparator
Placebo, registered as placebo (semaglutide).
Primary questions
What is the treatment difference in change in body weight from baseline to week 68, and how do the odds of achieving 5% or more body-weight reduction compare after week 68?
3. Trial Design
Semaglutide
- Semaglutide
- The primary statistical analyses identify the treatment group as Semaglutide 2.4 mg.
Placebo
- Placebo (semaglutide)
- The primary statistical analyses compare this group with semaglutide 2.4 mg.
Trial timing
Trial start
The registered study start date was 2018-06-04.
Primary completion
The registered primary completion date was 2020-03-30.
Registry status
The trial is listed as completed, with results posted.
4. Endpoints
The registry lists two primary endpoints. Both received formal statistical analyses in the posted results. The endpoint ClinicalTrials.gov record identify the first as a change in body weight percentage and the second as a binary achievement endpoint.
| Primary endpoint | Time frame | Endpoint type | Posted analysis |
|---|---|---|---|
| Change in Body Weight (%) | Baseline (week 0) to week 68 | Binary, as classified in the ClinicalTrials.gov record | ANCOVA |
| Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no) | After week 68 | Binary | Logistic regression |
Endpoint definition and observation periods
For Change in Body Weight (%), the registry states that change from baseline (week 0) to week 68 is presented and that the endpoint was evaluated using both in-trial and on-treatment observation periods. The in-trial observation period is described as the uninterrupted time interval from randomization at week 0 to the last contact with the trial site at week 75. The registry definition for the on-treatment observation period begins with the intervals in which participants were receiving treatment.
For Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no), the registry presents the number of participants achieving weight loss greater than or equal to 5% at week 68. This endpoint was also evaluated using both in-trial and on-treatment observation periods.
5. Analysis Populations
The posted primary analyses use the full analysis set (FAS). The registry description states that the FAS comprised all randomized participants. The number analyzed in an individual analysis is the number of participants with available data.
| Population | Registry description | Role in posted primary analyses |
|---|---|---|
| Full analysis set (FAS) | Comprised all randomized participants. | Primary efficacy analysis population. |
| Number analyzed | Number of participants with available data. | Analysis-specific denominator concept. |
This distinction is important because the phrase "all randomized participants" describes the analysis population, while the analysis-specific number analyzed can depend on data availability. The ClinicalTrials.gov record does not provide separate numerical denominators for each of the four posted statistical analyses.
6. Statistical Methodology
ANCOVA for change in body weight
The first primary endpoint was analyzed using analysis of covariance (ANCOVA). ANCOVA is a linear-model framework that compares treatment groups while accounting for relevant covariate information. In a randomized trial, this can improve precision when baseline measurements are strongly related to the outcome.
Here, Y represents the outcome being analyzed, the treatment indicator distinguishes randomized groups, and baseline information can account for variation in the outcome. The treatment coefficient represents the adjusted treatment difference under the fitted model.
The ClinicalTrials.gov record specifically identify the effect measure as a treatment difference, with the outcome unit given as a percentage point. The posted analyses were based on the FAS and are identified as intention-to-treat analyses.
Logistic regression for the 5% response endpoint
The second primary endpoint is explicitly binary: participants either did or did not achieve 5 or more percent body-weight reduction. The registry reports logistic regression and uses the odds ratio as the effect measure.
Exponentiating the treatment coefficient gives an odds ratio. An odds ratio above 1 indicates higher estimated odds of the specified binary outcome in the treatment group relative to the comparator, under the fitted model.
Superiority testing
All four posted primary-endpoint analyses are identified as superiority analyses. The corresponding p-values therefore provide evidence against the relevant null comparison under the statistical framework used for each endpoint.
Confidence intervals
Each of the four posted primary analyses includes a two-sided 95% confidence interval. A confidence interval provides a range of parameter values compatible with the data and model under the stated statistical framework. It is not a range containing a specified percentage of individual patient outcomes.
Intention-to-treat principle
The statistical analyses are identified as incorporating the intention-to-treat concept. An ITT analysis preserves randomized treatment assignment as the basis of the comparison rather than redefining the comparison according to treatment received. This is particularly important when interpreting a randomized treatment effect.
7. Primary Results: Change in Body Weight (%)
The registry reports two analyses for the primary endpoint Change in Body Weight (%), both comparing semaglutide 2.4 mg with placebo. The difference between the analyses is the estimand: one is identified as a treatment policy estimand, while the other is identified as a hypothetical estimand.
Treatment policy estimand
Treatment difference
95% CI: -13.37 to -11.51 · P < .0001
ANCOVA · Two-sided 95% CI · Superiority
The estimated treatment difference was -12.44 percentage points, comparing semaglutide 2.4 mg with placebo for change in body weight from baseline (week 0) to week 68 under the treatment-policy estimand.
The negative sign indicates that the estimated change was lower in the semaglutide group than in the placebo group according to the direction of the treatment-difference measure reported by the registry. The estimate is a between-group treatment difference; it is not the percentage of participants who responded and it does not mean that every individual experienced a 12.44-percentage-point change attributable to treatment.
The 95% confidence interval extends from -13.37 to -11.51. Its relatively narrow span describes the statistical precision of this estimated treatment difference under the stated model and analysis framework. It does not describe the range of individual treatment responses.
The p-value of < .0001 addresses evidence against the relevant null hypothesis under the superiority analysis. A p-value does not measure the size or clinical importance of an effect; the estimate and its confidence interval are needed to understand magnitude and precision.
Hypothetical estimand
Treatment difference
95% CI: -15.29 to -13.55 · P < 0.0001
ANCOVA · Two-sided 95% CI · Superiority
The estimated treatment difference was -14.42 percentage points under the hypothetical estimand. This is a different statistical question from the treatment-policy analysis even though the same randomized treatment groups and endpoint are involved.
The estimate describes the fitted difference under the hypothetical estimand specified in the registry analysis. It should therefore not be combined with the treatment-policy estimate as though the two were duplicate measurements of one parameter.
The 95% confidence interval of -15.29 to -13.55 quantifies uncertainty around this particular estimated treatment difference. As with the first analysis, it does not describe the distribution of individual patient responses.
The p-value of < 0.0001 indicates strong statistical evidence against the null hypothesis under the reported superiority framework, but it is not an effect-size measure. The treatment difference and confidence interval remain the primary quantities for understanding magnitude and precision.
Why two estimands matter
Treatment policy
The treatment-policy analysis asks a question centered on the treatment strategy assigned at randomization, incorporating the outcome definition specified for that strategy rather than restricting interpretation to an idealized treatment course.
Hypothetical
The hypothetical analysis asks what the treatment comparison would look like under the hypothetical scenario defined by the estimand. Because the question changes, the resulting estimate can differ from the treatment-policy estimate.
8. Primary Results: Achieving 5% or More Body Weight Reduction
The second primary endpoint is binary: whether a participant achieved at least 5% body-weight reduction at week 68. The registry reports two logistic-regression analyses, again corresponding to treatment-policy and hypothetical estimands.
Treatment policy estimand
Odds ratio for achieving 5% or more reduction
95% CI: 8.88 to 14.19 · P < 0.0001
Logistic regression · Two-sided 95% CI · Superiority
An odds ratio of 11.22 means that the estimated odds of achieving the binary endpoint were 11.22 times as high in the semaglutide 2.4 mg group as in the placebo group under the treatment-policy analysis.
An odds ratio is not a probability ratio. An OR of 11.22 does not mean that 11.22 times as many participants achieved the endpoint, nor does it mean that the probability of response increased by 11.22 times. Converting an odds ratio into an absolute probability requires the underlying event probabilities.
The 95% CI of 8.88 to 14.19 describes uncertainty around the estimated odds ratio. It does not describe the range of treatment effects for individual participants.
The p-value of < 0.0001 indicates strong statistical evidence against the null comparison within the reported superiority framework. It does not quantify the magnitude of the treatment effect; the odds ratio and confidence interval do that.
Hypothetical estimand
Odds ratio for achieving 5% or more reduction
95% CI: 28.02 to 48.95 · P < 0.0001
Logistic regression · Two-sided 95% CI · Superiority
An odds ratio of 37.03 means that the estimated odds of achieving the binary endpoint were 37.03 times as high in the semaglutide 2.4 mg group as in the placebo group under the hypothetical estimand.
This should not be interpreted as a 37.03-fold increase in probability. Odds and probabilities are related but are not the same quantity, particularly when the event is not rare.
The 95% confidence interval of 28.02 to 48.95 gives the statistical uncertainty around the estimated odds ratio. The interval is entirely above 1, consistent with the superiority result reported in the registry.
The p-value of < 0.0001 addresses the null hypothesis; it does not tell the reader whether the odds ratio is clinically large, nor does it substitute for the confidence interval.
9. Comparing the Four Primary Analyses
| Primary endpoint | Estimand | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Change in Body Weight (%) | Treatment policy | ANCOVA | Treatment difference | -12.44 | -13.37 to -11.51 | < .0001 |
| Change in Body Weight (%) | Hypothetical | ANCOVA | Treatment difference | -14.42 | -15.29 to -13.55 | < 0.0001 |
| Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no) | Treatment policy | Logistic regression | Odds ratio | 11.22 | 8.88 to 14.19 | <0.0001 |
| Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no) | Hypothetical | Logistic regression | Odds ratio | 37.03 | 28.02 to 48.95 | <0.0001 |
The important statistical feature is that the two endpoints and the two estimands answer different questions. The ANCOVA results quantify a treatment difference in the continuous-style body-weight change measure reported by the registry, whereas the logistic-regression results quantify relative odds for a binary threshold outcome. Within each endpoint, changing the estimand changes the target of inference and therefore can change the estimate.
10. Secondary Endpoint Results
The ClinicalTrials.gov record reports 42 outcome measures, but the statistical analyses provided for this page contain four primary-endpoint analyses and no additional secondary-endpoint statistical analyses. Accordingly, this page does not infer secondary treatment effects from the existence of posted outcome measures.
11. Safety
The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk:
| Safety measure | Semaglutide 2.4 mg | Placebo |
|---|---|---|
| Serious adverse events | 128 / 1306 | 42 / 655 |
The denominators differ between the two arms, so the affected-participant counts alone should not be compared as if they represented equal numbers of participants at risk. The ClinicalTrials.gov record reports the affected/at-risk pairs directly and do not provide a formal statistical comparison or confidence interval for serious adverse events.
12. Statistical Methods Explained
Why was ANCOVA used for change in body weight?
ANCOVA is designed for comparing a quantitative outcome between groups while incorporating baseline information. In a randomized study, baseline adjustment can improve statistical efficiency because participants who start at different baseline levels may otherwise contribute additional unexplained variability. The registry identifies ANCOVA as the method for both posted analyses of Change in Body Weight (%).
What does a treatment difference of -12.44 mean?
A treatment difference is a subtraction-based comparison between the randomized treatment groups under the specified model and estimand. The value -12.44 indicates the direction and magnitude of the estimated between-group difference as defined by the registry's effect measure. It is not an individual patient's change and it is not a response rate.
Why are there two estimates for the same primary endpoint?
The two estimates correspond to different estimands. The first is labeled a treatment-policy estimand and the second a hypothetical estimand. An estimand specifies what treatment effect the analysis is intended to estimate, including how treatment exposure and relevant post-randomization circumstances are handled. Changing that target can change the numerical estimate.
What does an odds ratio of 11.22 mean?
An odds ratio of 11.22 means the estimated odds of achieving the specified binary outcome are 11.22 times as high in the semaglutide 2.4 mg group as in the placebo group under that analysis. Odds are calculated as probability divided by one minus probability. Because odds are not probabilities, an OR of 11.22 should not be described as an 11.22-fold increase in probability.
Why is the odds ratio of 37.03 different from 11.22?
Both values concern the same binary endpoint, but they correspond to different estimands. The 11.22 estimate is from the treatment-policy analysis, whereas 37.03 is from the hypothetical analysis. They therefore target different statistical quantities.
What does a p-value below 0.0001 tell us?
Under the specified statistical model and null hypothesis, a p-value below 0.0001 indicates strong evidence against the null comparison. It does not measure effect size, clinical importance, or the probability that the null hypothesis is true. Those questions require the estimated effect, confidence interval, and clinical context.
What does a 95% confidence interval tell us?
A 95% confidence interval expresses statistical uncertainty around the estimated treatment effect under the relevant model and sampling framework. For example, the treatment-policy body-weight analysis has a 95% CI from -13.37 to -11.51. The interval is about uncertainty in the estimated treatment difference, not variability among individual participants.
13. Interpreting the Treatment Differences
For a treatment difference reported as -12.44 or -14.42, the negative sign is part of the definition of the between-group contrast. It indicates the direction of the semaglutide-versus-placebo difference under the reported effect-measure convention.
The magnitude of an effect is represented by the treatment-difference estimate or odds ratio. Statistical significance is addressed by the hypothesis test and p-value. A very small p-value does not make an effect numerically larger, and a large effect estimate should not be interpreted without its confidence interval.
The width of a confidence interval provides information about precision. The intervals around all four posted primary estimates are reported as two-sided 95% confidence intervals, allowing the reader to see the uncertainty associated with each model-based estimate.
The treatment differences and odds ratios should not be placed on the same numerical scale. A treatment difference describes a difference in the endpoint's reported unit, whereas an odds ratio is a multiplicative comparison of odds for a binary outcome.
14. Estimands: Treatment Policy vs Hypothetical
One of the most important statistical features of the registry-reported STEP 1 record is that both primary endpoints have analyses under two estimand concepts. This provides a useful illustration of why modern clinical-trial interpretation should begin by asking what treatment effect is being estimated, not simply which number is largest or has the smallest p-value.
| Endpoint | Treatment-policy analysis | Hypothetical analysis |
|---|---|---|
| Change in Body Weight (%) | ANCOVA; treatment difference -12.44; 95% CI -13.37 to -11.51; P < .0001 | ANCOVA; treatment difference -14.42; 95% CI -15.29 to -13.55; P < 0.0001 |
| 5% or more body-weight reduction | Logistic regression; OR 11.22; 95% CI 8.88 to 14.19; P <0.0001 | Logistic regression; OR 37.03; 95% CI 28.02 to 48.95; P <0.0001 |
The numerical differences between the two sets of estimates are not evidence of a statistical contradiction. Rather, they illustrate that different estimands can produce different answers because they define different targets of inference.
15. Randomization and Blinding
Randomization is the core design feature that establishes the basis for a causal treatment comparison. By assigning participants randomly, treatment assignment is separated from baseline characteristics in expectation, allowing differences in outcomes between randomized groups to be interpreted within the randomized-trial framework.
STEP 1 is registered as quadruple-masked. Masking is intended to reduce the potential influence of knowledge of treatment assignment on participant behavior, clinical management, outcome assessment, or other trial processes, depending on which parties are masked under the trial's operational definition.
Randomization
Creates the primary comparison between the semaglutide and placebo groups and supports an intention-to-treat analysis.
Quadruple masking
Reduces the potential for treatment knowledge to influence trial conduct or outcome-related processes covered by the masking procedure.
16. Missing Data and Analysis Interpretation
The registry definitions state that the FAS comprised all randomized participants, while the number analyzed represents participants with available data. This distinction means that the statistical interpretation depends not only on randomization but also on how available observations enter the particular analysis.
The ClinicalTrials.gov record identifies the body-weight analyses as treatment-policy and hypothetical estimand analyses, but do not provide a detailed imputation algorithm or a complete missing-data sensitivity-analysis specification. It would therefore be inappropriate to infer a particular imputation method from the existence of the ANCOVA results alone.
17. Multiplicity and Multiple Primary Analyses
STEP 1 has two registered primary endpoints, and the registry-reported statistical-analyses data contain four primary analyses: two estimand-specific analyses for each endpoint. The ClinicalTrials.gov record identifies all four as superiority analyses and provide a two-sided 95% confidence interval and p-value for each.
| Feature | Registry information reported |
|---|---|
| Registered primary endpoints | 2 |
| Posted primary-endpoint analyses | 4 |
| Hypothesis type | Superiority |
| Confidence interval | 95%, two-sided |
| Formal multiplicity adjustment details | Not reported in the ClinicalTrials.gov record |
Multiplicity is important whenever multiple hypotheses are tested because the probability of observing at least one apparently positive result can increase as the number of tests grows. The ClinicalTrials.gov record does not state a multiplicity-control procedure, alpha-allocation scheme, hierarchical testing strategy, or gatekeeping procedure. No such procedure should therefore be attributed to the trial from the ClinicalTrials.gov record.
18. Interim Analysis and Other Design Features
The registry-reported statistical profile does not report an interim-analysis procedure, alpha-spending method, non-inferiority margin, crossover design, factorial structure, or Bayesian analysis.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Interim analysis | No interim-analysis method is reported. |
| Non-inferiority margin | Not applicable to the reported superiority analyses; no margin is reported. |
| Crossover | No crossover feature is reported. |
| Factorial design | The design is registered as parallel, not factorial. |
| Multiplicity procedure | No specific adjustment procedure is reported. |
| Bayesian methods | No Bayesian method is reported. |
| Stratification factors | No stratification factors are reported in the ClinicalTrials.gov record. |
19. Results in Statistical Context
The four posted primary analyses tell a coherent statistical story within the ClinicalTrials.gov record: both primary endpoints have estimates favoring the semaglutide comparison under the reported effect-measure conventions, all four confidence intervals are provided as two-sided 95% intervals, and all four p-values are below 0.0001.
However, the four numbers should not be treated as four independent measurements of one effect. There are two distinct endpoints and two distinct estimands. The treatment-difference estimates quantify a difference in body-weight change, while the odds ratios quantify relative odds of crossing a prespecified binary threshold.
The distinction between effect size, precision, and statistical evidence is especially important here:
Effect size
-12.44 and -14.42 are treatment differences; 11.22 and 37.03 are odds ratios.
Precision
The 95% confidence intervals describe uncertainty around each estimated effect.
Statistical evidence
The four reported p-values are all below 0.0001 under their respective superiority analyses.
Target of inference
The treatment-policy and hypothetical analyses target different estimands and therefore should be interpreted separately.
20. Important Limitations and Interpretation Issues
- Estimand distinction: treatment-policy and hypothetical analyses address different targets of inference and should not be treated as duplicate estimates of one parameter.
- Binary endpoint interpretation: odds ratios are not risk ratios or probability ratios. The reported ORs cannot be converted into absolute response probabilities without additional event-rate information.
- Analysis population: the FAS comprised all randomized participants, while the analysis-specific number analyzed was defined as participants with available data. The ClinicalTrials.gov record does not provide the individual denominators for the four analyses.
- Missing-data details: the ClinicalTrials.gov record does not provide a full imputation or sensitivity-analysis specification.
- Multiplicity: two primary endpoints and four primary analyses are reported, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure.
- Secondary analyses: although 42 outcome measures are posted, the ClinicalTrials.gov record contains only four primary-endpoint analyses. Secondary treatment effects should not be inferred.
- Safety comparisons: serious adverse events are reported as affected/at-risk counts, but no formal statistical comparison is provided in the ClinicalTrials.gov record.
- Statistical-method detail: the registry identifies ANCOVA and logistic regression but does not provide enough information here to reproduce the complete fitted models, including all covariates or estimation details.
- Generalizability: the ClinicalTrials.gov record identifies the trial population only through its registered conditions and brief title; additional eligibility and baseline-characteristic information is not included in the ClinicalTrials.gov record.
21. Why This Trial Matters Statistically
STEP 1 is a useful teaching example because its primary results illustrate several fundamental concepts without relying on a single effect measure. The same randomized comparison is examined using both a treatment-difference framework and a binary-outcome odds-ratio framework, while two estimands are reported for each endpoint.
| Concept | How it appears in STEP 1 |
|---|---|
| Randomization | Registered as randomized allocation in a parallel-group phase 3 trial. |
| Blinding | Registered as quadruple masked. |
| Intention-to-treat analysis | Identified in the posted primary analyses. |
| ANCOVA | Used for Change in Body Weight (%) under both reported estimands. |
| Logistic regression | Used for the binary endpoint of achieving 5 or more percent body-weight reduction. |
| Treatment difference | Reported as -12.44 and -14.42 for the body-weight endpoint. |
| Odds ratio | Reported as 11.22 and 37.03 for the binary endpoint. |
| Confidence intervals | Two-sided 95% intervals are reported for all four primary analyses. |
| P-values | All four primary analyses report p-values below 0.0001. |
| Estimands | Treatment-policy and hypothetical analyses are both reported. |
| Safety analysis | Serious adverse events are reported as affected/at-risk counts by arm. |
22. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: NCT03548935 — STEP 1.
- Linked publication: PubMed PMID 33567185.
- Linked publication: PubMed PMID 40980163.
- Linked publication: PubMed PMID 39316288.
- Linked publication: PubMed PMID 39226070.
- Linked publication: PubMed PMID 38698650.
Continue through the Clinical Biostats statistical pathway
Explore the statistical methods and calculators connected to randomized trials, continuous outcomes, binary endpoints, effect measures, and statistical inference.
25. Record Summary
STEP 1 provides a useful statistical example of a randomized phase 3 trial with two primary endpoints and four posted primary analyses. The body-weight endpoint was analyzed with ANCOVA using treatment differences, while the binary endpoint of achieving 5 or more percent body-weight reduction was analyzed with logistic regression using odds ratios. Each endpoint was analyzed under both treatment-policy and hypothetical estimands.
The reported body-weight treatment differences were -12.44 with a 95% CI of -13.37 to -11.51 for the treatment-policy estimand and -14.42 with a 95% CI of -15.29 to -13.55 for the hypothetical estimand. The corresponding odds ratios for achieving 5 or more percent body-weight reduction were 11.22 with a 95% CI of 8.88 to 14.19 and 37.03 with a 95% CI of 28.02 to 48.95. All four analyses reported superiority hypotheses and p-values below 0.0001.
The main statistical lesson is that an appropriate interpretation requires more than reading the p-value. The reader must identify the endpoint, effect measure, estimand, analysis population, and confidence interval. In this trial, those distinctions explain why the same randomized comparison can yield different numerical estimates without representing conflicting analyses.