← Clinical Trials
Overweight or Obesity Phase 3 Completed NCT03548935

STEP 1: Complete Statistical Analysis of Semaglutide in Overweight or Obesity

An independent statistical review of the randomized phase 3 STEP 1 trial evaluating semaglutide 2.4 mg versus placebo in people with overweight or obesity, with emphasis on ANCOVA, logistic regression, treatment differences, odds ratios, confidence intervals, and estimand-specific interpretation.

Trial start: June 4, 2018  ·  Primary completion: March 30, 2020  ·  Enrollment: 1,961
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

STEP 1 was a randomized, parallel-group, quadruple-masked phase 3 trial in the therapeutic area of endocrinology. The registry describes the study as investigating how well semaglutide works in people suffering from overweight or obesity. The trial enrolled 1,961 participants and compared semaglutide with placebo.

1,961
Enrollment
Randomized trial
2
Arms
Semaglutide vs placebo
3
Phase
Phase 3
4
Primary analyses
Estimate + 95% CI
FeatureSTEP 1
Trial nameSTEP 1
NCT IDNCT03548935
PhasePhase 3
Therapeutic areaEndocrinology
ConditionsMetabolism and Nutrition Disorder; Overweight or Obesity
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment1,961
InterventionsSemaglutide (drug); Placebo (semaglutide) (drug)
Primary endpoints2
Results postedYes
Outcome measures posted42
Statistical analyses posted4
Lead sponsorNovo Nordisk A/S
Sponsor typeIndustry

2. Clinical Question

The central statistical question is whether semaglutide 2.4 mg produces a different body-weight outcome than placebo in people with overweight or obesity, and whether participants receiving semaglutide are more likely to achieve at least 5% body-weight reduction.

Population

People with overweight or obesity, within the registered conditions of metabolism and nutrition disorder; overweight or obesity.

Intervention

Semaglutide, as registered in the trial's intervention list and statistical analyses as semaglutide 2.4 mg.

Comparator

Placebo, registered as placebo (semaglutide).

Primary questions

What is the treatment difference in change in body weight from baseline to week 68, and how do the odds of achieving 5% or more body-weight reduction compare after week 68?

3. Trial Design

01
Randomize1,961 enrolled
02
Parallel groupsSemaglutide vs placebo
03
Quadruple maskingRegistered masking
04
Week 68Body-weight endpoint
05
Primary analysesANCOVA and logistic regression
Allocation
Randomized allocation in a parallel-group design.
Masking
Quadruple masking was registered for the trial.
Primary purpose
Treatment was the registered primary purpose.
Hypothesis type
The posted primary analyses use a superiority hypothesis.
ARM 1

Semaglutide

  • Semaglutide
  • The primary statistical analyses identify the treatment group as Semaglutide 2.4 mg.
ARM 2

Placebo

  • Placebo (semaglutide)
  • The primary statistical analyses compare this group with semaglutide 2.4 mg.

Trial timing

June 4, 2018

Trial start

The registered study start date was 2018-06-04.

March 30, 2020

Primary completion

The registered primary completion date was 2020-03-30.

Completed

Registry status

The trial is listed as completed, with results posted.

4. Endpoints

The registry lists two primary endpoints. Both received formal statistical analyses in the posted results. The endpoint ClinicalTrials.gov record identify the first as a change in body weight percentage and the second as a binary achievement endpoint.

Primary endpointTime frameEndpoint typePosted analysis
Change in Body Weight (%) Baseline (week 0) to week 68 Binary, as classified in the ClinicalTrials.gov record ANCOVA
Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no) After week 68 Binary Logistic regression

Endpoint definition and observation periods

For Change in Body Weight (%), the registry states that change from baseline (week 0) to week 68 is presented and that the endpoint was evaluated using both in-trial and on-treatment observation periods. The in-trial observation period is described as the uninterrupted time interval from randomization at week 0 to the last contact with the trial site at week 75. The registry definition for the on-treatment observation period begins with the intervals in which participants were receiving treatment.

For Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no), the registry presents the number of participants achieving weight loss greater than or equal to 5% at week 68. This endpoint was also evaluated using both in-trial and on-treatment observation periods.

Registry terminology matters. The ClinicalTrials.gov record distinguishes treatment-policy and hypothetical estimands for the posted analyses. These are not interchangeable analyses: they answer different questions about how outcomes are conceptualized when treatment exposure or post-treatment events are relevant.

5. Analysis Populations

The posted primary analyses use the full analysis set (FAS). The registry description states that the FAS comprised all randomized participants. The number analyzed in an individual analysis is the number of participants with available data.

PopulationRegistry descriptionRole in posted primary analyses
Full analysis set (FAS) Comprised all randomized participants. Primary efficacy analysis population.
Number analyzed Number of participants with available data. Analysis-specific denominator concept.

This distinction is important because the phrase "all randomized participants" describes the analysis population, while the analysis-specific number analyzed can depend on data availability. The ClinicalTrials.gov record does not provide separate numerical denominators for each of the four posted statistical analyses.

6. Statistical Methodology

ANCOVA for change in body weight

The first primary endpoint was analyzed using analysis of covariance (ANCOVA). ANCOVA is a linear-model framework that compares treatment groups while accounting for relevant covariate information. In a randomized trial, this can improve precision when baseline measurements are strongly related to the outcome.

Conceptual ANCOVA form
Y = β0 + β1(Treatment) + β2(Baseline) + ε

Here, Y represents the outcome being analyzed, the treatment indicator distinguishes randomized groups, and baseline information can account for variation in the outcome. The treatment coefficient represents the adjusted treatment difference under the fitted model.

The ClinicalTrials.gov record specifically identify the effect measure as a treatment difference, with the outcome unit given as a percentage point. The posted analyses were based on the FAS and are identified as intention-to-treat analyses.

Logistic regression for the 5% response endpoint

The second primary endpoint is explicitly binary: participants either did or did not achieve 5 or more percent body-weight reduction. The registry reports logistic regression and uses the odds ratio as the effect measure.

Conceptual logistic model
logit[P(Y=1)] = β0 + β1(Treatment) + ...

Exponentiating the treatment coefficient gives an odds ratio. An odds ratio above 1 indicates higher estimated odds of the specified binary outcome in the treatment group relative to the comparator, under the fitted model.

Superiority testing

All four posted primary-endpoint analyses are identified as superiority analyses. The corresponding p-values therefore provide evidence against the relevant null comparison under the statistical framework used for each endpoint.

Confidence intervals

Each of the four posted primary analyses includes a two-sided 95% confidence interval. A confidence interval provides a range of parameter values compatible with the data and model under the stated statistical framework. It is not a range containing a specified percentage of individual patient outcomes.

Intention-to-treat principle

The statistical analyses are identified as incorporating the intention-to-treat concept. An ITT analysis preserves randomized treatment assignment as the basis of the comparison rather than redefining the comparison according to treatment received. This is particularly important when interpreting a randomized treatment effect.

7. Primary Results: Change in Body Weight (%)

The registry reports two analyses for the primary endpoint Change in Body Weight (%), both comparing semaglutide 2.4 mg with placebo. The difference between the analyses is the estimand: one is identified as a treatment policy estimand, while the other is identified as a hypothetical estimand.

Treatment policy estimand

Treatment difference

-12.44

95% CI: -13.37 to -11.51   ·   P < .0001

ANCOVA  ·  Two-sided 95% CI  ·  Superiority

Clinical Biostats interpretation

The estimated treatment difference was -12.44 percentage points, comparing semaglutide 2.4 mg with placebo for change in body weight from baseline (week 0) to week 68 under the treatment-policy estimand.

The negative sign indicates that the estimated change was lower in the semaglutide group than in the placebo group according to the direction of the treatment-difference measure reported by the registry. The estimate is a between-group treatment difference; it is not the percentage of participants who responded and it does not mean that every individual experienced a 12.44-percentage-point change attributable to treatment.

The 95% confidence interval extends from -13.37 to -11.51. Its relatively narrow span describes the statistical precision of this estimated treatment difference under the stated model and analysis framework. It does not describe the range of individual treatment responses.

The p-value of < .0001 addresses evidence against the relevant null hypothesis under the superiority analysis. A p-value does not measure the size or clinical importance of an effect; the estimate and its confidence interval are needed to understand magnitude and precision.

Hypothetical estimand

Treatment difference

-14.42

95% CI: -15.29 to -13.55   ·   P < 0.0001

ANCOVA  ·  Two-sided 95% CI  ·  Superiority

Clinical Biostats interpretation

The estimated treatment difference was -14.42 percentage points under the hypothetical estimand. This is a different statistical question from the treatment-policy analysis even though the same randomized treatment groups and endpoint are involved.

The estimate describes the fitted difference under the hypothetical estimand specified in the registry analysis. It should therefore not be combined with the treatment-policy estimate as though the two were duplicate measurements of one parameter.

The 95% confidence interval of -15.29 to -13.55 quantifies uncertainty around this particular estimated treatment difference. As with the first analysis, it does not describe the distribution of individual patient responses.

The p-value of < 0.0001 indicates strong statistical evidence against the null hypothesis under the reported superiority framework, but it is not an effect-size measure. The treatment difference and confidence interval remain the primary quantities for understanding magnitude and precision.

Why two estimands matter

Treatment policy

The treatment-policy analysis asks a question centered on the treatment strategy assigned at randomization, incorporating the outcome definition specified for that strategy rather than restricting interpretation to an idealized treatment course.

Hypothetical

The hypothetical analysis asks what the treatment comparison would look like under the hypothetical scenario defined by the estimand. Because the question changes, the resulting estimate can differ from the treatment-policy estimate.

8. Primary Results: Achieving 5% or More Body Weight Reduction

The second primary endpoint is binary: whether a participant achieved at least 5% body-weight reduction at week 68. The registry reports two logistic-regression analyses, again corresponding to treatment-policy and hypothetical estimands.

Treatment policy estimand

Odds ratio for achieving 5% or more reduction

11.22

95% CI: 8.88 to 14.19   ·   P < 0.0001

Logistic regression  ·  Two-sided 95% CI  ·  Superiority

Clinical Biostats interpretation

An odds ratio of 11.22 means that the estimated odds of achieving the binary endpoint were 11.22 times as high in the semaglutide 2.4 mg group as in the placebo group under the treatment-policy analysis.

An odds ratio is not a probability ratio. An OR of 11.22 does not mean that 11.22 times as many participants achieved the endpoint, nor does it mean that the probability of response increased by 11.22 times. Converting an odds ratio into an absolute probability requires the underlying event probabilities.

The 95% CI of 8.88 to 14.19 describes uncertainty around the estimated odds ratio. It does not describe the range of treatment effects for individual participants.

The p-value of < 0.0001 indicates strong statistical evidence against the null comparison within the reported superiority framework. It does not quantify the magnitude of the treatment effect; the odds ratio and confidence interval do that.

Hypothetical estimand

Odds ratio for achieving 5% or more reduction

37.03

95% CI: 28.02 to 48.95   ·   P < 0.0001

Logistic regression  ·  Two-sided 95% CI  ·  Superiority

Clinical Biostats interpretation

An odds ratio of 37.03 means that the estimated odds of achieving the binary endpoint were 37.03 times as high in the semaglutide 2.4 mg group as in the placebo group under the hypothetical estimand.

This should not be interpreted as a 37.03-fold increase in probability. Odds and probabilities are related but are not the same quantity, particularly when the event is not rare.

The 95% confidence interval of 28.02 to 48.95 gives the statistical uncertainty around the estimated odds ratio. The interval is entirely above 1, consistent with the superiority result reported in the registry.

The p-value of < 0.0001 addresses the null hypothesis; it does not tell the reader whether the odds ratio is clinically large, nor does it substitute for the confidence interval.

9. Comparing the Four Primary Analyses

Primary endpointEstimandMethodEffect measureEstimate95% CIP-value
Change in Body Weight (%) Treatment policy ANCOVA Treatment difference -12.44 -13.37 to -11.51 < .0001
Change in Body Weight (%) Hypothetical ANCOVA Treatment difference -14.42 -15.29 to -13.55 < 0.0001
Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no) Treatment policy Logistic regression Odds ratio 11.22 8.88 to 14.19 <0.0001
Participants Who Achieve 5 or More Percent Body Weight Reduction (Yes/no) Hypothetical Logistic regression Odds ratio 37.03 28.02 to 48.95 <0.0001

The important statistical feature is that the two endpoints and the two estimands answer different questions. The ANCOVA results quantify a treatment difference in the continuous-style body-weight change measure reported by the registry, whereas the logistic-regression results quantify relative odds for a binary threshold outcome. Within each endpoint, changing the estimand changes the target of inference and therefore can change the estimate.

10. Secondary Endpoint Results

The ClinicalTrials.gov record reports 42 outcome measures, but the statistical analyses provided for this page contain four primary-endpoint analyses and no additional secondary-endpoint statistical analyses. Accordingly, this page does not infer secondary treatment effects from the existence of posted outcome measures.

Why this distinction matters: an outcome measure being posted in a registry is not the same as a formal statistical comparison being posted for that outcome. A reported numerical outcome and a treatment-effect estimate answer different questions.

11. Safety

The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk:

Safety measureSemaglutide 2.4 mgPlacebo
Serious adverse events128 / 130642 / 655
Serious adverse events: affected participants
Semaglutide 2.4 mg
128
Placebo
42

The denominators differ between the two arms, so the affected-participant counts alone should not be compared as if they represented equal numbers of participants at risk. The ClinicalTrials.gov record reports the affected/at-risk pairs directly and do not provide a formal statistical comparison or confidence interval for serious adverse events.

Safety interpretation: the serious-adverse-event figures are descriptive counts with their corresponding at-risk denominators in the ClinicalTrials.gov record. They should not be converted into a treatment-effect claim without an appropriate definition of the safety estimand, analysis population, event window, and statistical method.

12. Statistical Methods Explained

Why was ANCOVA used for change in body weight?

ANCOVA is designed for comparing a quantitative outcome between groups while incorporating baseline information. In a randomized study, baseline adjustment can improve statistical efficiency because participants who start at different baseline levels may otherwise contribute additional unexplained variability. The registry identifies ANCOVA as the method for both posted analyses of Change in Body Weight (%).

What does a treatment difference of -12.44 mean?

A treatment difference is a subtraction-based comparison between the randomized treatment groups under the specified model and estimand. The value -12.44 indicates the direction and magnitude of the estimated between-group difference as defined by the registry's effect measure. It is not an individual patient's change and it is not a response rate.

Why are there two estimates for the same primary endpoint?

The two estimates correspond to different estimands. The first is labeled a treatment-policy estimand and the second a hypothetical estimand. An estimand specifies what treatment effect the analysis is intended to estimate, including how treatment exposure and relevant post-randomization circumstances are handled. Changing that target can change the numerical estimate.

What does an odds ratio of 11.22 mean?

An odds ratio of 11.22 means the estimated odds of achieving the specified binary outcome are 11.22 times as high in the semaglutide 2.4 mg group as in the placebo group under that analysis. Odds are calculated as probability divided by one minus probability. Because odds are not probabilities, an OR of 11.22 should not be described as an 11.22-fold increase in probability.

Why is the odds ratio of 37.03 different from 11.22?

Both values concern the same binary endpoint, but they correspond to different estimands. The 11.22 estimate is from the treatment-policy analysis, whereas 37.03 is from the hypothetical analysis. They therefore target different statistical quantities.

What does a p-value below 0.0001 tell us?

Under the specified statistical model and null hypothesis, a p-value below 0.0001 indicates strong evidence against the null comparison. It does not measure effect size, clinical importance, or the probability that the null hypothesis is true. Those questions require the estimated effect, confidence interval, and clinical context.

What does a 95% confidence interval tell us?

A 95% confidence interval expresses statistical uncertainty around the estimated treatment effect under the relevant model and sampling framework. For example, the treatment-policy body-weight analysis has a 95% CI from -13.37 to -11.51. The interval is about uncertainty in the estimated treatment difference, not variability among individual participants.

13. Interpreting the Treatment Differences

Direction of effect

For a treatment difference reported as -12.44 or -14.42, the negative sign is part of the definition of the between-group contrast. It indicates the direction of the semaglutide-versus-placebo difference under the reported effect-measure convention.

Magnitude versus significance

The magnitude of an effect is represented by the treatment-difference estimate or odds ratio. Statistical significance is addressed by the hypothesis test and p-value. A very small p-value does not make an effect numerically larger, and a large effect estimate should not be interpreted without its confidence interval.

Precision

The width of a confidence interval provides information about precision. The intervals around all four posted primary estimates are reported as two-sided 95% confidence intervals, allowing the reader to see the uncertainty associated with each model-based estimate.

Different endpoint scales

The treatment differences and odds ratios should not be placed on the same numerical scale. A treatment difference describes a difference in the endpoint's reported unit, whereas an odds ratio is a multiplicative comparison of odds for a binary outcome.

14. Estimands: Treatment Policy vs Hypothetical

One of the most important statistical features of the registry-reported STEP 1 record is that both primary endpoints have analyses under two estimand concepts. This provides a useful illustration of why modern clinical-trial interpretation should begin by asking what treatment effect is being estimated, not simply which number is largest or has the smallest p-value.

EndpointTreatment-policy analysisHypothetical analysis
Change in Body Weight (%) ANCOVA; treatment difference -12.44; 95% CI -13.37 to -11.51; P < .0001 ANCOVA; treatment difference -14.42; 95% CI -15.29 to -13.55; P < 0.0001
5% or more body-weight reduction Logistic regression; OR 11.22; 95% CI 8.88 to 14.19; P <0.0001 Logistic regression; OR 37.03; 95% CI 28.02 to 48.95; P <0.0001

The numerical differences between the two sets of estimates are not evidence of a statistical contradiction. Rather, they illustrate that different estimands can produce different answers because they define different targets of inference.

15. Randomization and Blinding

Randomization is the core design feature that establishes the basis for a causal treatment comparison. By assigning participants randomly, treatment assignment is separated from baseline characteristics in expectation, allowing differences in outcomes between randomized groups to be interpreted within the randomized-trial framework.

STEP 1 is registered as quadruple-masked. Masking is intended to reduce the potential influence of knowledge of treatment assignment on participant behavior, clinical management, outcome assessment, or other trial processes, depending on which parties are masked under the trial's operational definition.

Randomization

Creates the primary comparison between the semaglutide and placebo groups and supports an intention-to-treat analysis.

Quadruple masking

Reduces the potential for treatment knowledge to influence trial conduct or outcome-related processes covered by the masking procedure.

16. Missing Data and Analysis Interpretation

The registry definitions state that the FAS comprised all randomized participants, while the number analyzed represents participants with available data. This distinction means that the statistical interpretation depends not only on randomization but also on how available observations enter the particular analysis.

The ClinicalTrials.gov record identifies the body-weight analyses as treatment-policy and hypothetical estimand analyses, but do not provide a detailed imputation algorithm or a complete missing-data sensitivity-analysis specification. It would therefore be inappropriate to infer a particular imputation method from the existence of the ANCOVA results alone.

Important limitation: the ClinicalTrials.gov record does not provide enough detail to reconstruct the full missing-data procedure, including any specific imputation model, assumptions about missingness, or sensitivity-analysis framework. The posted treatment estimates should therefore be interpreted according to the estimands and analysis descriptions that are actually reported.

17. Multiplicity and Multiple Primary Analyses

STEP 1 has two registered primary endpoints, and the registry-reported statistical-analyses data contain four primary analyses: two estimand-specific analyses for each endpoint. The ClinicalTrials.gov record identifies all four as superiority analyses and provide a two-sided 95% confidence interval and p-value for each.

FeatureRegistry information reported
Registered primary endpoints2
Posted primary-endpoint analyses4
Hypothesis typeSuperiority
Confidence interval95%, two-sided
Formal multiplicity adjustment detailsNot reported in the ClinicalTrials.gov record

Multiplicity is important whenever multiple hypotheses are tested because the probability of observing at least one apparently positive result can increase as the number of tests grows. The ClinicalTrials.gov record does not state a multiplicity-control procedure, alpha-allocation scheme, hierarchical testing strategy, or gatekeeping procedure. No such procedure should therefore be attributed to the trial from the ClinicalTrials.gov record.

18. Interim Analysis and Other Design Features

The registry-reported statistical profile does not report an interim-analysis procedure, alpha-spending method, non-inferiority margin, crossover design, factorial structure, or Bayesian analysis.

Design topicWhat the ClinicalTrials.gov record supports
Interim analysisNo interim-analysis method is reported.
Non-inferiority marginNot applicable to the reported superiority analyses; no margin is reported.
CrossoverNo crossover feature is reported.
Factorial designThe design is registered as parallel, not factorial.
Multiplicity procedureNo specific adjustment procedure is reported.
Bayesian methodsNo Bayesian method is reported.
Stratification factorsNo stratification factors are reported in the ClinicalTrials.gov record.

19. Results in Statistical Context

The four posted primary analyses tell a coherent statistical story within the ClinicalTrials.gov record: both primary endpoints have estimates favoring the semaglutide comparison under the reported effect-measure conventions, all four confidence intervals are provided as two-sided 95% intervals, and all four p-values are below 0.0001.

However, the four numbers should not be treated as four independent measurements of one effect. There are two distinct endpoints and two distinct estimands. The treatment-difference estimates quantify a difference in body-weight change, while the odds ratios quantify relative odds of crossing a prespecified binary threshold.

The distinction between effect size, precision, and statistical evidence is especially important here:

Effect size

-12.44 and -14.42 are treatment differences; 11.22 and 37.03 are odds ratios.

Precision

The 95% confidence intervals describe uncertainty around each estimated effect.

Statistical evidence

The four reported p-values are all below 0.0001 under their respective superiority analyses.

Target of inference

The treatment-policy and hypothetical analyses target different estimands and therefore should be interpreted separately.

20. Important Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

STEP 1 is a useful teaching example because its primary results illustrate several fundamental concepts without relying on a single effect measure. The same randomized comparison is examined using both a treatment-difference framework and a binary-outcome odds-ratio framework, while two estimands are reported for each endpoint.

ConceptHow it appears in STEP 1
RandomizationRegistered as randomized allocation in a parallel-group phase 3 trial.
BlindingRegistered as quadruple masked.
Intention-to-treat analysisIdentified in the posted primary analyses.
ANCOVAUsed for Change in Body Weight (%) under both reported estimands.
Logistic regressionUsed for the binary endpoint of achieving 5 or more percent body-weight reduction.
Treatment differenceReported as -12.44 and -14.42 for the body-weight endpoint.
Odds ratioReported as 11.22 and 37.03 for the binary endpoint.
Confidence intervalsTwo-sided 95% intervals are reported for all four primary analyses.
P-valuesAll four primary analyses report p-values below 0.0001.
EstimandsTreatment-policy and hypothetical analyses are both reported.
Safety analysisSerious adverse events are reported as affected/at-risk counts by arm.

22. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

23. Related Statistical Calculators

24. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods and calculators connected to randomized trials, continuous outcomes, binary endpoints, effect measures, and statistical inference.

25. Record Summary

STEP 1 provides a useful statistical example of a randomized phase 3 trial with two primary endpoints and four posted primary analyses. The body-weight endpoint was analyzed with ANCOVA using treatment differences, while the binary endpoint of achieving 5 or more percent body-weight reduction was analyzed with logistic regression using odds ratios. Each endpoint was analyzed under both treatment-policy and hypothetical estimands.

The reported body-weight treatment differences were -12.44 with a 95% CI of -13.37 to -11.51 for the treatment-policy estimand and -14.42 with a 95% CI of -15.29 to -13.55 for the hypothetical estimand. The corresponding odds ratios for achieving 5 or more percent body-weight reduction were 11.22 with a 95% CI of 8.88 to 14.19 and 37.03 with a 95% CI of 28.02 to 48.95. All four analyses reported superiority hypotheses and p-values below 0.0001.

The main statistical lesson is that an appropriate interpretation requires more than reading the p-value. The reader must identify the endpoint, effect measure, estimand, analysis population, and confidence interval. In this trial, those distinctions explain why the same randomized comparison can yield different numerical estimates without representing conflicting analyses.

Clinical Biostats methodology: A trial-results page should distinguish the numerical results actually posted in the registry from statistical interpretation. For STEP 1, the most important interpretive distinctions are between treatment differences and odds ratios, and between treatment-policy and hypothetical estimands.