← Clinical Trials
Obesity / Overweight Phase 3 Completed NCT04657016

SURMOUNT-3: Complete Statistical Analysis of Tirzepatide in Obesity or Overweight

An independent statistical review of the randomized, double-masked phase 3 SURMOUNT-3 trial evaluating tirzepatide versus placebo in participants with obesity or overweight after a lifestyle weight loss program.

Trial start: 2021-03-29  ·  Primary completion: 2023-04-20  ·  Enrollment: 579
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SURMOUNT-3 was a randomized, double-masked, parallel-group phase 3 trial evaluating tirzepatide versus placebo in participants with obesity or overweight after a lifestyle weight loss program. The registry reports 579 participants, two treatment arms, two primary endpoints, 22 posted outcome measures, and 22 posted statistical analyses.

579
Enrolled
2 treatment arms
3
Phase
Randomized parallel design
-24.5
Primary weight effect
95% CI -26.1 to -22.8
130.36
Primary odds ratio
95% CI 69.98 to 242.84
FeatureSURMOUNT-3
Trial nameSURMOUNT-3
NCT identifierNCT04657016
PhasePhase 3
Therapeutic areaEndocrinology
ConditionsObesity; Overweight
DesignRandomized, parallel-group, double-masked
AllocationRandomized
Primary purposeTreatment
Enrollment579
Arms2
InterventionsTirzepatide; placebo
Primary endpoints2 binary endpoints as categorized in the registry data
Hypothesis typeSuperiority
Lead sponsorEli Lilly and Company
Sponsor typeIndustry
StatusCompleted

2. Clinical Question

The central statistical question is whether tirzepatide produces different outcomes from placebo in participants with obesity or overweight after a lifestyle weight loss program. The registry classifies the primary hypotheses as superiority.

Population

Participants with obesity or overweight who entered the trial after a lifestyle weight loss program.

Intervention

Tirzepatide.

Comparator

Placebo.

Primary question

Does tirzepatide produce superior body-weight outcomes compared with placebo at the prespecified assessment times?

3. Trial Design

01
Enroll 579 participants
02
Randomize 2 parallel arms
03
Mask Double-masked
04
Assess Baseline and follow-up
05
Analyze Mixed models and logistic regression
ARM A · TIRZEPATIDE

Tirzepatide

  • Randomized intervention arm.
  • Included in the superiority comparison.
  • 287 participants are identified as the at-risk denominator for the reported serious-adverse-event summary.
ARM B · PLACEBO

Placebo

  • Randomized comparator arm.
  • Included in the superiority comparison.
  • 292 participants are identified as the at-risk denominator for the reported serious-adverse-event summary.
What the registry establishes: the study is randomized, parallel, double-masked, and designed for treatment comparison. The ClinicalTrials.gov record does not provide a randomization ratio, dosing schedule, lifestyle-program protocol details, crossover information, or an interim-analysis description, so those design features are not inferred here.

4. Endpoints

The registry reports two primary endpoints. Both have formal statistical analyses posted. The definitions below preserve the registry's endpoint wording and time frames.

EndpointTime frameRegistered definition / analysis
Percent Change From Baseline in Body Weight Baseline, 72 Weeks Percent change from baseline in body weight. Least Squares (LS) mean was determined by mixed-model repeated measures (MMRM) model for post-baseline measures: Variable = Baseline + Analysis Country + Sex + Treatment + Time + Treatment*Time (Type III sum of squares).
Percentage of Participants With Greater Than or Equal to (≥) 5% Body Weight Reduction Week 72 Percentage of participants with ≥5% body weight reduction was analysed by Logistic regression model using imputed data with baseline body weight, Analysis Country, Sex, Treatment as factors.

Primary endpoint structure

The first endpoint is a continuous change-from-baseline measure analyzed with a repeated-measures mixed model. The second converts the weight response into a binary threshold: whether a participant achieved at least 5% body weight reduction at Week 72. That distinction matters because the two endpoints answer related but different questions.

Continuous endpoint

The body-weight percent-change endpoint preserves the magnitude of change and uses information from post-baseline repeated measurements through the MMRM framework.

Binary endpoint

The ≥5% endpoint asks whether each participant crosses a clinically defined response threshold and is summarized through logistic regression and an odds ratio.

5. Statistical Methodology

Mixed-model repeated measures

The primary percent-change endpoint was analyzed using a mixed-model repeated measures approach. The registered model includes baseline body weight, Analysis Country, Sex, Treatment, Time, and the Treatment × Time interaction, with Type III sums of squares.

Registered MMRM structure
Outcome = Baseline + Analysis Country + Sex + Treatment + Time + Treatment × Time

The model therefore adjusts for baseline and specified categorical factors while estimating treatment effects across post-baseline time points. The treatment-by-time interaction permits the treatment difference to vary with time rather than forcing a single identical effect at every measurement.

Least-squares means and mean difference

The registry reports a Mean Difference (Net) for the primary continuous endpoint. In a covariate-adjusted longitudinal model, the estimated treatment contrast represents the model-based difference between treatment groups after accounting for the variables included in the model.

Logistic regression

The ≥5% weight-reduction endpoint was analyzed with logistic regression using imputed data. The registry identifies baseline body weight, Analysis Country, Sex, and Treatment as model factors.

Conceptual logistic model
logit[P(Y = 1)] = β0 + β1Treatment + covariate terms

The odds ratio compares the odds of achieving the specified binary response between treatment groups, conditional on the variables included in the model.

Odds ratio

An odds ratio above 1 for the tirzepatide-versus-placebo comparison indicates higher estimated odds of achieving the specified response in the tirzepatide group. An odds ratio is not the same as a risk ratio or a difference in percentages.

ANCOVA

Two secondary quality-of-life outcomes were analyzed using ANCOVA. The registry reports ANCOVA for change from baseline in the Short Form 36 Version 2 Health Survey Version 2 Acute Form Physical Functioning Domain Score and for change from baseline in the Impact of Weight on Quality of Life Lite Clinical Trials Version Physical Function Composite Score.

Analysis population

For the posted analyses, the registry-defined population was: all randomly assigned participants who took at least 1 dose of study drug, had a baseline and at least 1 post-baseline value for the outcome, excluding data after discontinuation of study drug.

Important population distinction: this is not simply the full randomized enrollment of 579 participants. The formal analyses use a defined subset based on treatment exposure and availability of baseline and post-baseline outcome information, with data after study-drug discontinuation excluded.

6. Primary Results

Percent Change From Baseline in Body Weight

The primary continuous endpoint was assessed from baseline through 72 weeks using the registry's mixed-model repeated measures approach. The reported comparison is the net mean difference between placebo and tirzepatide.

Net mean difference in percent body-weight change

-24.5

95% CI: -26.1 to -22.8   ·   P < 0.001

Method: mixed-effects model / mixed-model repeated measures

Clinical Biostats interpretation

The estimate of -24.5 means that the model-based difference in percent change from baseline in body weight between the compared groups was 24.5 percentage points in the direction of greater reduction for the treatment comparison as reported by the registry.

The estimate is a between-group mean difference; it is not the percentage of participants who lost 24.5% of body weight, and it does not mean that every participant experienced that amount of change.

The 95% confidence interval of -26.1 to -22.8 describes statistical uncertainty around the estimated mean treatment difference under the specified model and analysis framework. It is not an interval containing the individual treatment effect for 95% of participants.

The P < 0.001 result addresses evidence against the null hypothesis used for the superiority comparison. It does not measure the magnitude of the treatment effect; the estimate and confidence interval provide that information.

Because this is a repeated-measures model, interpretation depends on the specified model structure, covariate adjustment, treatment-by-time interaction, available observations, and handling of observations after study-drug discontinuation.

Percentage of Participants With ≥5% Body Weight Reduction

The second primary endpoint was analyzed at Week 72 with logistic regression using imputed data and the registered baseline body weight, Analysis Country, Sex, and Treatment factors.

Odds ratio for ≥5% body weight reduction

130.36

95% CI: 69.98 to 242.84   ·   P < 0.001

Method: logistic regression

Clinical Biostats interpretation

An odds ratio of 130.36 means that the estimated odds of achieving at least 5% body weight reduction were substantially higher in the tirzepatide group than in the placebo group under the specified logistic regression model.

This is an odds ratio, not a probability ratio. It does not mean that 130.36 times as many participants achieved the endpoint, nor can it be converted directly into a percentage-point difference without the underlying response probabilities.

The 95% confidence interval of 69.98 to 242.84 is wide in absolute terms, even though the entire interval is above 1. This indicates considerable uncertainty in the exact magnitude of the odds ratio while maintaining the same directional interpretation under the model.

The P < 0.001 value addresses the evidence against the null treatment comparison; it is not an effect-size measure and should not be interpreted as saying that the treatment effect is "130 times statistically significant."

The use of imputed data is particularly relevant. The result depends not only on the observed outcomes but also on the prespecified or implemented assumptions and procedure used to handle missing outcome information.

7. Secondary Endpoint Results

The registry contains 20 secondary statistical analyses in addition to the two primary endpoint analyses. They cover maintenance of weight loss, additional weight-loss thresholds, anthropometric measures, blood pressure, lipid measures, glycemic measures, insulin, and patient-reported physical-function outcomes.

Weight-loss threshold outcomes

Secondary endpointTime frameEffect estimate95% CIP-value
Percentage of Participants Who Maintain ≥80% of the Body Weight Lost During Intensive Lifestyle Program 72 Weeks OR 101.60 39.17 to 263.55 <0.001
Percentage of Participants Who Achieve ≥10% Body Weight Reduction 72 Weeks OR 153.95 78.90 to 300.37 <0.001
Percentage of Participants Who Achieve ≥15% Body Weight Reduction 72 Weeks OR 144.48 62.65 to 333.21 <0.001
Percentage of Participants Who Achieve ≥20% Body Weight Reduction 72 Weeks OR 118.06 40.08 to 347.74 <0.001

These four endpoints all use logistic regression and odds ratios. Their estimates are not interchangeable with the primary continuous mean difference because each asks a threshold-based binary question.

Anthropometric and cardiovascular measures

Secondary endpointEffect measureEstimate95% CIP-value
Change From Baseline in Waist Circumference Mean difference -17.9 -19.5 to -16.3 <0.001
Change From Baseline in Body Weight Mean difference -25.0 -26.9 to -23.2 <0.001
Change From Baseline in Body Mass Index (BMI) Mean difference -8.9 -9.6 to -8.3 <0.001
Change From Baseline in Systolic Blood Pressure (SBP) Mean difference -10.2 -12.2 to -8.1 <0.001
Change From Baseline in Diastolic Blood Pressure (DBP) Mean difference -5.7 -7.2 to -4.3 <0.001

All five measures were analyzed using mixed-effects models. The effect estimates are on their respective original scales: centimeters for waist circumference, kilograms for body weight, kilograms per meter squared for BMI, and mmHg for blood pressure.

Lipid and metabolic outcomes

Secondary endpointEffect measureEstimate95% CIP-value
Percent Change From Baseline in Total Cholesterol Mean difference -7.79 -10.40 to -5.10 <0.001
Percent Change From Baseline in High Density Lipoprotein (HDL) Cholesterol Mean difference 11.4 8.2 to 14.7 <0.001
Percent Change From Baseline in Low Density Lipoprotein (LDL) Cholesterol Mean difference -11.50 -15.30 to -7.53 <0.001
Percent Change From Baseline in Very Low-Density Lipoprotein (VLDL) Cholesterol Mean difference -27.8 -32.1 to -23.2 <0.001
Percent Change From Baseline in Triglycerides Mean difference -28.0 -32.3 to -23.4 <0.001
Percent Change From Baseline in Free Fatty Acids Mean difference -21.3 -28.4 to -13.6 <0.001
Change From Baseline in Fasting Glucose Mean difference -11.2 -13.5 to -8.8 <0.001
Change From Baseline in Hemoglobin A1c (HbA1c) Mean difference -0.47 -0.53 to -0.42 <0.001
Percent Change From Baseline in Fasting Insulin Mean difference -48.1 -53.7 to -41.7 <0.001

The registry reports mixed-effects models for all of these laboratory and metabolic endpoints. The direction of the estimate must be read in conjunction with the endpoint definition: for example, a negative mean difference in a percent-change endpoint and a positive mean difference for HDL are not interpreted in the same way merely because their numerical signs differ.

Patient-reported physical-function outcomes

Secondary endpointMethodEffect measureEstimate95% CIP-value
Change From Baseline in Short Form 36 Version 2 Health Survey Version 2 (SF 36v2) Acute Form Physical Functioning Domain Score ANCOVA Mean difference 3.8 2.8 to 4.9 <0.001
Change From Baseline in Impact of Weight on Quality of Life Lite Clinical Trials Version (IWQOL-Lite-CT) Physical Function Composite Score ANCOVA Mean difference 12.8 9.7 to 16.0 <0.001

8. Results by Statistical Method

One useful way to understand SURMOUNT-3 is to group the posted analyses by model rather than by clinical domain. The registry reports three principal statistical method families: mixed-effects models, logistic regression, and ANCOVA.

MethodRole in SURMOUNT-3Examples
Mixed-effects model / MMRM Repeated continuous outcomes measured from baseline through 72 weeks Body weight, waist circumference, BMI, blood pressure, lipids, glucose, HbA1c, fasting insulin
Logistic regression Binary weight-response endpoints ≥5%, ≥10%, ≥15%, and ≥20% body-weight reduction; maintenance of ≥80% of weight lost during the intensive lifestyle program
ANCOVA Continuous change-from-baseline patient-reported physical-function outcomes SF 36v2 physical functioning; IWQOL-Lite-CT physical function composite

9. Secondary Results: What the Estimates Actually Measure

Mean differences

A mean difference compares the modeled average outcome between randomized treatment groups. For a change-from-baseline endpoint, it is a difference in changes; for a percent-change endpoint, it is a difference in percentage-point changes. It is not automatically a relative percent difference between treatment groups.

Odds ratios

The binary endpoints are summarized by odds ratios. An odds ratio of 1 represents equal odds under the model. Values above 1 indicate higher estimated odds in the treatment group, while values below 1 indicate lower estimated odds. The odds ratio can become much larger than the corresponding risk ratio when the event is common, so the two measures should not be substituted for one another.

Odds and probability are different
Odds = p / (1 − p)

An odds ratio compares odds, not probabilities directly. To translate an odds ratio into absolute probabilities, a baseline probability or other relevant reference probability is required.

Why the confidence intervals matter

The confidence intervals provide a range of values compatible with the statistical model and data under the stated confidence framework. A narrow interval indicates greater precision around the estimate; a wider interval indicates greater uncertainty about its exact magnitude.

Why the P-values do not measure effect size

All posted analyses in the ClinicalTrials.gov record has P-values reported as <0.001. That tells the reader that the calculated P-value is below that threshold; it does not distinguish an estimate of modest magnitude from one of large magnitude. Effect estimates and confidence intervals are therefore essential to interpretation.

10. Missing Data and Imputation

The registry explicitly states that the ≥5% body-weight-reduction endpoint was analyzed using imputed data. The ClinicalTrials.gov record does not specify the imputation algorithm, number of imputations, imputation model, or sensitivity-analysis strategy, so those details are not inferred here.

Why this matters: a binary endpoint can change materially when missing observations are handled differently. The logistic-regression result therefore needs to be understood as the result of the stated imputed-data analysis, rather than as a simple calculation from only the observed Week 72 responses.

The continuous primary endpoint has a different structure: its registry definition specifies an MMRM analysis using post-baseline measures and excluding data after discontinuation of study drug. The ClinicalTrials.gov record does not provide the covariance structure or other implementation details of that model.

11. Analysis Population and Treatment Discontinuation

The posted statistical analyses use participants who were randomly assigned, took at least one dose of study drug, had baseline and at least one post-baseline value for the relevant outcome, and whose data after discontinuation of study drug were excluded.

Why randomization still matters

Randomization establishes the original treatment comparison. It provides the design foundation for comparing groups even though the formal posted analysis population has additional treatment-exposure and outcome-availability requirements.

Why exclusion after discontinuation matters

Removing post-discontinuation observations defines the estimand represented by the reported analysis and can affect how the treatment contrast should be interpreted.

The ClinicalTrials.gov record does not describe a separate intention-to-treat analysis, a treatment-policy estimand, a hypothetical estimand, or another explicit intercurrent-event strategy. Those approaches should therefore not be attributed to this trial from the available information.

12. Statistical Methods Explained

Why was an MMRM used for the primary percent-change endpoint?

The primary body-weight endpoint is measured repeatedly after baseline. An MMRM allows the analysis to use the longitudinal structure of those measurements rather than reducing the entire follow-up period to one isolated observation. The registered model includes Time and Treatment × Time, so the estimated treatment contrast can account for how outcomes evolve across time.

What does a mean difference of -24.5 mean?

It is the reported net mean difference for percent change from baseline in body weight between the compared groups. The negative sign indicates the direction of the reported treatment contrast. It does not mean that every participant lost 24.5% of body weight, and it is not a ratio.

What does an odds ratio of 130.36 mean?

It means that the estimated odds of achieving at least 5% body weight reduction were 130.36 times the comparator odds under the specified logistic regression model. It is not a 130.36-fold increase in probability and cannot by itself provide the absolute percentage of responders.

Why use logistic regression for the ≥5% endpoint?

The endpoint converts weight change into a binary response: a participant either reaches the threshold or does not. Logistic regression is designed for this type of binary outcome and can incorporate baseline and other specified factors.

Why does the confidence interval matter when P < 0.001?

The P-value and confidence interval answer different questions. The P-value addresses evidence against a null hypothesis, whereas the confidence interval shows the range of treatment-effect values compatible with the model and data under the stated confidence framework. For the ≥5% endpoint, the interval from 69.98 to 242.84 shows that the exact magnitude of the odds ratio is substantially less certain than the direction of the comparison.

Why was ANCOVA used for some secondary outcomes?

ANCOVA is a linear-model framework commonly used for continuous outcomes when the analysis compares groups while accounting for baseline information. In SURMOUNT-3, the registry specifically identifies ANCOVA for the two patient-reported physical-function outcomes. The ClinicalTrials.gov record does not give the complete covariate specification for those ANCOVA models.

Why are the many P-values important to interpret carefully?

The registry reports many secondary outcomes, and every registry-reported secondary analysis has a P-value of <0.001. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure or endpoint hierarchy for these 22 posted analyses. Therefore, the individual nominal P-values should not automatically be treated as establishing a separately controlled familywise error rate across all posted endpoints.

13. Multiplicity and Multiple Endpoints

SURMOUNT-3 has two registered primary endpoints and 20 secondary statistical analyses in the ClinicalTrials.gov record. This creates an important distinction between identifying a statistical association for an individual endpoint and establishing formal error control across a family of hypotheses.

Endpoint groupNumber in the ClinicalTrials.gov recordStatistical issue
Primary endpoints 2 Both have formal statistical analyses, estimates, 95% confidence intervals, and P-values.
Secondary analyses 20 Multiple endpoints are evaluated; the ClinicalTrials.gov record does not specify an adjustment procedure.
Total outcome measures posted 22 The registry contains a broader set of outcomes than the two primary endpoints.

The appropriate interpretation of a P-value depends on the hypothesis family and prespecified testing strategy. Because the ClinicalTrials.gov record does not describe a multiplicity procedure, this page does not assign one retrospectively.

14. Superiority Testing

The registry identifies the hypothesis type for all registry-reported statistical analyses as superiority. That means the formal question is whether the treatment groups differ in the specified direction under the corresponding statistical test, rather than whether a treatment is merely no worse than a comparator by a prespecified non-inferiority margin.

Superiority versus non-inferiority
Superiority: H0 = no treatment difference   vs   HA = treatment difference

The ClinicalTrials.gov record does not report a non-inferiority margin, so no non-inferiority interpretation should be applied to these results.

15. Safety Results

The ClinicalTrials.gov record provides serious adverse-event counts by treatment arm. These are reported as affected participants divided by participants at risk.

Safety measurePlaceboTirzepatide
Serious adverse events 14 / 292 17 / 287

These figures describe the number of participants affected and the corresponding at-risk denominator. The ClinicalTrials.gov record does not provide a statistical comparison, confidence interval, P-value, event-specific breakdown, severity distribution, or exposure-adjusted rate.

Clinical Biostats interpretation

The serious-adverse-event summary should be interpreted descriptively from the information reported. The difference between 14/292 and 17/287 cannot by itself establish whether there is a statistically meaningful treatment-group difference because the ClinicalTrials.gov record does not report a formal comparison.

Safety interpretation also differs conceptually from efficacy interpretation. The efficacy analyses have prespecified model structures and reported effect estimates with confidence intervals and P-values; the serious-adverse-event ClinicalTrials.gov record are counts and denominators only.

16. Randomization and Blinding

The trial is classified as randomized with double masking and a parallel design. These design features are important statistical safeguards because randomization is intended to balance prognostic factors across treatment groups in expectation, while masking can reduce differential behavior, assessment, or reporting related to treatment assignment.

Randomization

Random allocation creates the primary basis for the between-group treatment comparison. The registry does not provide the randomization ratio in the ClinicalTrials.gov record.

Double masking

Double masking means treatment assignment was not openly known to the relevant masked parties. The ClinicalTrials.gov record does not identify which specific parties were masked.

Parallel groups

Participants were assigned to treatment groups that were compared in parallel rather than through a crossover design.

No crossover information reported

The ClinicalTrials.gov record does not report a crossover treatment strategy or crossover rate, so none is assumed in the interpretation.

17. Stratification and Covariate Adjustment

The registered primary MMRM explicitly includes Analysis Country, Sex, Treatment, Time, and the Treatment × Time interaction, in addition to baseline body weight. The registered logistic-regression analysis also includes baseline body weight, Analysis Country, Sex, and Treatment.

ModelVariables explicitly reported in registry-reported registry definition
Primary MMRM Baseline; Analysis Country; Sex; Treatment; Time; Treatment × Time
≥5% logistic regression Baseline body weight; Analysis Country; Sex; Treatment

These variables are not all equivalent to randomization stratification factors. In particular, the ClinicalTrials.gov record describes them as components of the analysis models rather than as a list of stratification variables used during randomization.

18. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The two primary analyses report estimates in the prespecified superiority framework, with 95% confidence intervals and P-values <0.001. One is a model-based mean difference in percent body-weight change; the other is a logistic-regression odds ratio for achieving at least 5% body-weight reduction.

Clinical interpretation

The trial evaluates both the magnitude of weight change and the probability of crossing clinically defined weight-loss thresholds. These outcomes capture complementary aspects of response rather than a single universal measure of benefit.

The secondary results extend the analysis to maintenance of prior weight loss, larger weight-loss thresholds, waist circumference, BMI, blood pressure, metabolic biomarkers, and physical-function measures. Each should be interpreted on its own measurement scale and with attention to its statistical model.

19. Limitations and Interpretation Issues

20. Why This Trial Matters Statistically

SURMOUNT-3 is a useful teaching case because it combines longitudinal continuous outcomes with threshold-based binary outcomes within the same randomized trial. The statistical methods therefore illustrate why the form of the endpoint determines the appropriate effect measure and model.

ConceptHow it appears in SURMOUNT-3
RandomizationRandomized parallel-group treatment comparison
BlindingDouble-masked design
Mixed-effects modelingPrimary body-weight percent-change analysis and many secondary longitudinal outcomes
Repeated measuresPost-baseline body-weight measurements analyzed through MMRM
Logistic regressionBinary weight-loss threshold outcomes
Odds ratioPrimary ≥5% response endpoint and four additional binary secondary endpoints
ANCOVATwo patient-reported physical-function outcomes
Confidence intervalsReported for all registry-reported statistical analyses
P-valuesAll registry-reported statistical analyses report P <0.001
Missing-data handlingImputed data explicitly reported for the ≥5% endpoint
MultiplicityTwo primary endpoints plus 20 secondary statistical analyses
Superiority testingHypothesis type reported as superiority for the posted analyses

21. A Practical Reading of the Primary Results

The two primary endpoints should be read together because they measure different aspects of the same clinical domain.

Magnitude of change

The body-weight percent-change endpoint preserves the size of the change and produces a model-based mean difference of -24.5 with a 95% CI of -26.1 to -22.8.

Threshold response

The ≥5% endpoint asks whether participants crossed a predefined threshold and produces an odds ratio of 130.36 with a 95% CI of 69.98 to 242.84.

Different scales

The mean difference and odds ratio cannot be compared numerically. One is expressed in percentage-point change; the other compares odds.

Different assumptions

The MMRM and logistic regression models make different assumptions and address different outcome structures, so each result should be interpreted within its own analysis framework.

22. Statistical Interpretation of the Confidence Intervals

The primary continuous endpoint has a 95% confidence interval from -26.1 to -22.8. The entire interval lies on the same side of zero, which is consistent with the reported superiority P-value.

The primary binary endpoint has a 95% confidence interval from 69.98 to 242.84. The entire interval lies above 1, which is the null value for an odds ratio. At the same time, the interval is broad, emphasizing that statistical evidence for a directional difference does not imply precise knowledge of the exact odds-ratio magnitude.

What confidence intervals do not mean

A 95% confidence interval does not mean that there is a 95% probability that the fixed treatment effect is inside the displayed interval. It is a frequentist interval produced by a procedure designed to have 95% coverage under repeated sampling when the model and assumptions are appropriate.

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Limitations of the Available Registry Record

The ClinicalTrials.gov record is sufficient to reconstruct the major statistical architecture and all 22 posted statistical analyses, but they do not contain every detail that would normally appear in a full statistical analysis plan.

AvailableNot specified in the ClinicalTrials.gov record
Primary endpoint definitions and time frames Complete SAP-level derivation algorithms
Analysis populations Complete covariance-structure specification for MMRM
Methods and effect measures Detailed imputation algorithm and sensitivity analyses
Estimates, 95% CIs, and P-values Multiplicity-adjustment procedure
Serious adverse-event counts by arm Formal statistical safety comparison

Accordingly, this page interprets the statistical information that is actually registry-reported rather than reconstructing undocumented protocol or SAP details.

26. Sources

Continue through Clinical Biostats

Use the related statistical concepts to move from the trial results to deeper explanations of the models and effect measures used in randomized clinical research.

27. Record Summary

SURMOUNT-3 provides a useful statistical example of how a randomized, double-masked phase 3 trial can combine several complementary outcome frameworks. The primary body-weight endpoint uses an MMRM with baseline adjustment, country, sex, treatment, time, and treatment-by-time interaction, producing a reported net mean difference of -24.5 with a 95% CI of -26.1 to -22.8 and P <0.001. The second primary endpoint converts weight change into a binary ≥5% response and uses logistic regression, producing an odds ratio of 130.36 with a 95% CI of 69.98 to 242.84 and P <0.001.

The secondary analyses extend the same statistical framework across larger weight-loss thresholds, maintenance of prior weight loss, anthropometric measures, cardiovascular measures, metabolic biomarkers, and patient-reported physical function. The most important interpretive lesson is that these estimates cannot be treated as interchangeable: mean differences, odds ratios, and confidence intervals answer different statistical questions and must be read on their respective scales.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical evidence from statistical interpretation. For SURMOUNT-3, that means preserving the registry's endpoint definitions, analysis populations, models, estimates, confidence intervals, and P-values while making clear where the ClinicalTrials.gov record does not specify additional SAP-level assumptions or procedures.