This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
SURMOUNT-3 was a randomized, double-masked, parallel-group phase 3 trial evaluating tirzepatide versus placebo in participants with obesity or overweight after a lifestyle weight loss program. The registry reports 579 participants, two treatment arms, two primary endpoints, 22 posted outcome measures, and 22 posted statistical analyses.
| Feature | SURMOUNT-3 |
|---|---|
| Trial name | SURMOUNT-3 |
| NCT identifier | NCT04657016 |
| Phase | Phase 3 |
| Therapeutic area | Endocrinology |
| Conditions | Obesity; Overweight |
| Design | Randomized, parallel-group, double-masked |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 579 |
| Arms | 2 |
| Interventions | Tirzepatide; placebo |
| Primary endpoints | 2 binary endpoints as categorized in the registry data |
| Hypothesis type | Superiority |
| Lead sponsor | Eli Lilly and Company |
| Sponsor type | Industry |
| Status | Completed |
2. Clinical Question
The central statistical question is whether tirzepatide produces different outcomes from placebo in participants with obesity or overweight after a lifestyle weight loss program. The registry classifies the primary hypotheses as superiority.
Population
Participants with obesity or overweight who entered the trial after a lifestyle weight loss program.
Intervention
Tirzepatide.
Comparator
Placebo.
Primary question
Does tirzepatide produce superior body-weight outcomes compared with placebo at the prespecified assessment times?
3. Trial Design
Tirzepatide
- Randomized intervention arm.
- Included in the superiority comparison.
- 287 participants are identified as the at-risk denominator for the reported serious-adverse-event summary.
Placebo
- Randomized comparator arm.
- Included in the superiority comparison.
- 292 participants are identified as the at-risk denominator for the reported serious-adverse-event summary.
4. Endpoints
The registry reports two primary endpoints. Both have formal statistical analyses posted. The definitions below preserve the registry's endpoint wording and time frames.
| Endpoint | Time frame | Registered definition / analysis |
|---|---|---|
| Percent Change From Baseline in Body Weight | Baseline, 72 Weeks | Percent change from baseline in body weight. Least Squares (LS) mean was determined by mixed-model repeated measures (MMRM) model for post-baseline measures: Variable = Baseline + Analysis Country + Sex + Treatment + Time + Treatment*Time (Type III sum of squares). |
| Percentage of Participants With Greater Than or Equal to (≥) 5% Body Weight Reduction | Week 72 | Percentage of participants with ≥5% body weight reduction was analysed by Logistic regression model using imputed data with baseline body weight, Analysis Country, Sex, Treatment as factors. |
Primary endpoint structure
The first endpoint is a continuous change-from-baseline measure analyzed with a repeated-measures mixed model. The second converts the weight response into a binary threshold: whether a participant achieved at least 5% body weight reduction at Week 72. That distinction matters because the two endpoints answer related but different questions.
Continuous endpoint
The body-weight percent-change endpoint preserves the magnitude of change and uses information from post-baseline repeated measurements through the MMRM framework.
Binary endpoint
The ≥5% endpoint asks whether each participant crosses a clinically defined response threshold and is summarized through logistic regression and an odds ratio.
5. Statistical Methodology
Mixed-model repeated measures
The primary percent-change endpoint was analyzed using a mixed-model repeated measures approach. The registered model includes baseline body weight, Analysis Country, Sex, Treatment, Time, and the Treatment × Time interaction, with Type III sums of squares.
The model therefore adjusts for baseline and specified categorical factors while estimating treatment effects across post-baseline time points. The treatment-by-time interaction permits the treatment difference to vary with time rather than forcing a single identical effect at every measurement.
Least-squares means and mean difference
The registry reports a Mean Difference (Net) for the primary continuous endpoint. In a covariate-adjusted longitudinal model, the estimated treatment contrast represents the model-based difference between treatment groups after accounting for the variables included in the model.
Logistic regression
The ≥5% weight-reduction endpoint was analyzed with logistic regression using imputed data. The registry identifies baseline body weight, Analysis Country, Sex, and Treatment as model factors.
The odds ratio compares the odds of achieving the specified binary response between treatment groups, conditional on the variables included in the model.
Odds ratio
An odds ratio above 1 for the tirzepatide-versus-placebo comparison indicates higher estimated odds of achieving the specified response in the tirzepatide group. An odds ratio is not the same as a risk ratio or a difference in percentages.
ANCOVA
Two secondary quality-of-life outcomes were analyzed using ANCOVA. The registry reports ANCOVA for change from baseline in the Short Form 36 Version 2 Health Survey Version 2 Acute Form Physical Functioning Domain Score and for change from baseline in the Impact of Weight on Quality of Life Lite Clinical Trials Version Physical Function Composite Score.
Analysis population
For the posted analyses, the registry-defined population was: all randomly assigned participants who took at least 1 dose of study drug, had a baseline and at least 1 post-baseline value for the outcome, excluding data after discontinuation of study drug.
6. Primary Results
Percent Change From Baseline in Body Weight
The primary continuous endpoint was assessed from baseline through 72 weeks using the registry's mixed-model repeated measures approach. The reported comparison is the net mean difference between placebo and tirzepatide.
Net mean difference in percent body-weight change
95% CI: -26.1 to -22.8 · P < 0.001
Method: mixed-effects model / mixed-model repeated measures
The estimate of -24.5 means that the model-based difference in percent change from baseline in body weight between the compared groups was 24.5 percentage points in the direction of greater reduction for the treatment comparison as reported by the registry.
The estimate is a between-group mean difference; it is not the percentage of participants who lost 24.5% of body weight, and it does not mean that every participant experienced that amount of change.
The 95% confidence interval of -26.1 to -22.8 describes statistical uncertainty around the estimated mean treatment difference under the specified model and analysis framework. It is not an interval containing the individual treatment effect for 95% of participants.
The P < 0.001 result addresses evidence against the null hypothesis used for the superiority comparison. It does not measure the magnitude of the treatment effect; the estimate and confidence interval provide that information.
Because this is a repeated-measures model, interpretation depends on the specified model structure, covariate adjustment, treatment-by-time interaction, available observations, and handling of observations after study-drug discontinuation.
Percentage of Participants With ≥5% Body Weight Reduction
The second primary endpoint was analyzed at Week 72 with logistic regression using imputed data and the registered baseline body weight, Analysis Country, Sex, and Treatment factors.
Odds ratio for ≥5% body weight reduction
95% CI: 69.98 to 242.84 · P < 0.001
Method: logistic regression
An odds ratio of 130.36 means that the estimated odds of achieving at least 5% body weight reduction were substantially higher in the tirzepatide group than in the placebo group under the specified logistic regression model.
This is an odds ratio, not a probability ratio. It does not mean that 130.36 times as many participants achieved the endpoint, nor can it be converted directly into a percentage-point difference without the underlying response probabilities.
The 95% confidence interval of 69.98 to 242.84 is wide in absolute terms, even though the entire interval is above 1. This indicates considerable uncertainty in the exact magnitude of the odds ratio while maintaining the same directional interpretation under the model.
The P < 0.001 value addresses the evidence against the null treatment comparison; it is not an effect-size measure and should not be interpreted as saying that the treatment effect is "130 times statistically significant."
The use of imputed data is particularly relevant. The result depends not only on the observed outcomes but also on the prespecified or implemented assumptions and procedure used to handle missing outcome information.
7. Secondary Endpoint Results
The registry contains 20 secondary statistical analyses in addition to the two primary endpoint analyses. They cover maintenance of weight loss, additional weight-loss thresholds, anthropometric measures, blood pressure, lipid measures, glycemic measures, insulin, and patient-reported physical-function outcomes.
Weight-loss threshold outcomes
| Secondary endpoint | Time frame | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|
| Percentage of Participants Who Maintain ≥80% of the Body Weight Lost During Intensive Lifestyle Program | 72 Weeks | OR 101.60 | 39.17 to 263.55 | <0.001 |
| Percentage of Participants Who Achieve ≥10% Body Weight Reduction | 72 Weeks | OR 153.95 | 78.90 to 300.37 | <0.001 |
| Percentage of Participants Who Achieve ≥15% Body Weight Reduction | 72 Weeks | OR 144.48 | 62.65 to 333.21 | <0.001 |
| Percentage of Participants Who Achieve ≥20% Body Weight Reduction | 72 Weeks | OR 118.06 | 40.08 to 347.74 | <0.001 |
These four endpoints all use logistic regression and odds ratios. Their estimates are not interchangeable with the primary continuous mean difference because each asks a threshold-based binary question.
Anthropometric and cardiovascular measures
| Secondary endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change From Baseline in Waist Circumference | Mean difference | -17.9 | -19.5 to -16.3 | <0.001 |
| Change From Baseline in Body Weight | Mean difference | -25.0 | -26.9 to -23.2 | <0.001 |
| Change From Baseline in Body Mass Index (BMI) | Mean difference | -8.9 | -9.6 to -8.3 | <0.001 |
| Change From Baseline in Systolic Blood Pressure (SBP) | Mean difference | -10.2 | -12.2 to -8.1 | <0.001 |
| Change From Baseline in Diastolic Blood Pressure (DBP) | Mean difference | -5.7 | -7.2 to -4.3 | <0.001 |
All five measures were analyzed using mixed-effects models. The effect estimates are on their respective original scales: centimeters for waist circumference, kilograms for body weight, kilograms per meter squared for BMI, and mmHg for blood pressure.
Lipid and metabolic outcomes
| Secondary endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Percent Change From Baseline in Total Cholesterol | Mean difference | -7.79 | -10.40 to -5.10 | <0.001 |
| Percent Change From Baseline in High Density Lipoprotein (HDL) Cholesterol | Mean difference | 11.4 | 8.2 to 14.7 | <0.001 |
| Percent Change From Baseline in Low Density Lipoprotein (LDL) Cholesterol | Mean difference | -11.50 | -15.30 to -7.53 | <0.001 |
| Percent Change From Baseline in Very Low-Density Lipoprotein (VLDL) Cholesterol | Mean difference | -27.8 | -32.1 to -23.2 | <0.001 |
| Percent Change From Baseline in Triglycerides | Mean difference | -28.0 | -32.3 to -23.4 | <0.001 |
| Percent Change From Baseline in Free Fatty Acids | Mean difference | -21.3 | -28.4 to -13.6 | <0.001 |
| Change From Baseline in Fasting Glucose | Mean difference | -11.2 | -13.5 to -8.8 | <0.001 |
| Change From Baseline in Hemoglobin A1c (HbA1c) | Mean difference | -0.47 | -0.53 to -0.42 | <0.001 |
| Percent Change From Baseline in Fasting Insulin | Mean difference | -48.1 | -53.7 to -41.7 | <0.001 |
The registry reports mixed-effects models for all of these laboratory and metabolic endpoints. The direction of the estimate must be read in conjunction with the endpoint definition: for example, a negative mean difference in a percent-change endpoint and a positive mean difference for HDL are not interpreted in the same way merely because their numerical signs differ.
Patient-reported physical-function outcomes
| Secondary endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Change From Baseline in Short Form 36 Version 2 Health Survey Version 2 (SF 36v2) Acute Form Physical Functioning Domain Score | ANCOVA | Mean difference | 3.8 | 2.8 to 4.9 | <0.001 |
| Change From Baseline in Impact of Weight on Quality of Life Lite Clinical Trials Version (IWQOL-Lite-CT) Physical Function Composite Score | ANCOVA | Mean difference | 12.8 | 9.7 to 16.0 | <0.001 |
8. Results by Statistical Method
One useful way to understand SURMOUNT-3 is to group the posted analyses by model rather than by clinical domain. The registry reports three principal statistical method families: mixed-effects models, logistic regression, and ANCOVA.
| Method | Role in SURMOUNT-3 | Examples |
|---|---|---|
| Mixed-effects model / MMRM | Repeated continuous outcomes measured from baseline through 72 weeks | Body weight, waist circumference, BMI, blood pressure, lipids, glucose, HbA1c, fasting insulin |
| Logistic regression | Binary weight-response endpoints | ≥5%, ≥10%, ≥15%, and ≥20% body-weight reduction; maintenance of ≥80% of weight lost during the intensive lifestyle program |
| ANCOVA | Continuous change-from-baseline patient-reported physical-function outcomes | SF 36v2 physical functioning; IWQOL-Lite-CT physical function composite |
9. Secondary Results: What the Estimates Actually Measure
Mean differences
A mean difference compares the modeled average outcome between randomized treatment groups. For a change-from-baseline endpoint, it is a difference in changes; for a percent-change endpoint, it is a difference in percentage-point changes. It is not automatically a relative percent difference between treatment groups.
Odds ratios
The binary endpoints are summarized by odds ratios. An odds ratio of 1 represents equal odds under the model. Values above 1 indicate higher estimated odds in the treatment group, while values below 1 indicate lower estimated odds. The odds ratio can become much larger than the corresponding risk ratio when the event is common, so the two measures should not be substituted for one another.
An odds ratio compares odds, not probabilities directly. To translate an odds ratio into absolute probabilities, a baseline probability or other relevant reference probability is required.
Why the confidence intervals matter
The confidence intervals provide a range of values compatible with the statistical model and data under the stated confidence framework. A narrow interval indicates greater precision around the estimate; a wider interval indicates greater uncertainty about its exact magnitude.
Why the P-values do not measure effect size
All posted analyses in the ClinicalTrials.gov record has P-values reported as <0.001. That tells the reader that the calculated P-value is below that threshold; it does not distinguish an estimate of modest magnitude from one of large magnitude. Effect estimates and confidence intervals are therefore essential to interpretation.
10. Missing Data and Imputation
The registry explicitly states that the ≥5% body-weight-reduction endpoint was analyzed using imputed data. The ClinicalTrials.gov record does not specify the imputation algorithm, number of imputations, imputation model, or sensitivity-analysis strategy, so those details are not inferred here.
The continuous primary endpoint has a different structure: its registry definition specifies an MMRM analysis using post-baseline measures and excluding data after discontinuation of study drug. The ClinicalTrials.gov record does not provide the covariance structure or other implementation details of that model.
11. Analysis Population and Treatment Discontinuation
The posted statistical analyses use participants who were randomly assigned, took at least one dose of study drug, had baseline and at least one post-baseline value for the relevant outcome, and whose data after discontinuation of study drug were excluded.
Why randomization still matters
Randomization establishes the original treatment comparison. It provides the design foundation for comparing groups even though the formal posted analysis population has additional treatment-exposure and outcome-availability requirements.
Why exclusion after discontinuation matters
Removing post-discontinuation observations defines the estimand represented by the reported analysis and can affect how the treatment contrast should be interpreted.
The ClinicalTrials.gov record does not describe a separate intention-to-treat analysis, a treatment-policy estimand, a hypothetical estimand, or another explicit intercurrent-event strategy. Those approaches should therefore not be attributed to this trial from the available information.
12. Statistical Methods Explained
Why was an MMRM used for the primary percent-change endpoint?
The primary body-weight endpoint is measured repeatedly after baseline. An MMRM allows the analysis to use the longitudinal structure of those measurements rather than reducing the entire follow-up period to one isolated observation. The registered model includes Time and Treatment × Time, so the estimated treatment contrast can account for how outcomes evolve across time.
What does a mean difference of -24.5 mean?
It is the reported net mean difference for percent change from baseline in body weight between the compared groups. The negative sign indicates the direction of the reported treatment contrast. It does not mean that every participant lost 24.5% of body weight, and it is not a ratio.
What does an odds ratio of 130.36 mean?
It means that the estimated odds of achieving at least 5% body weight reduction were 130.36 times the comparator odds under the specified logistic regression model. It is not a 130.36-fold increase in probability and cannot by itself provide the absolute percentage of responders.
Why use logistic regression for the ≥5% endpoint?
The endpoint converts weight change into a binary response: a participant either reaches the threshold or does not. Logistic regression is designed for this type of binary outcome and can incorporate baseline and other specified factors.
Why does the confidence interval matter when P < 0.001?
The P-value and confidence interval answer different questions. The P-value addresses evidence against a null hypothesis, whereas the confidence interval shows the range of treatment-effect values compatible with the model and data under the stated confidence framework. For the ≥5% endpoint, the interval from 69.98 to 242.84 shows that the exact magnitude of the odds ratio is substantially less certain than the direction of the comparison.
Why was ANCOVA used for some secondary outcomes?
ANCOVA is a linear-model framework commonly used for continuous outcomes when the analysis compares groups while accounting for baseline information. In SURMOUNT-3, the registry specifically identifies ANCOVA for the two patient-reported physical-function outcomes. The ClinicalTrials.gov record does not give the complete covariate specification for those ANCOVA models.
Why are the many P-values important to interpret carefully?
The registry reports many secondary outcomes, and every registry-reported secondary analysis has a P-value of <0.001. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure or endpoint hierarchy for these 22 posted analyses. Therefore, the individual nominal P-values should not automatically be treated as establishing a separately controlled familywise error rate across all posted endpoints.
13. Multiplicity and Multiple Endpoints
SURMOUNT-3 has two registered primary endpoints and 20 secondary statistical analyses in the ClinicalTrials.gov record. This creates an important distinction between identifying a statistical association for an individual endpoint and establishing formal error control across a family of hypotheses.
| Endpoint group | Number in the ClinicalTrials.gov record | Statistical issue |
|---|---|---|
| Primary endpoints | 2 | Both have formal statistical analyses, estimates, 95% confidence intervals, and P-values. |
| Secondary analyses | 20 | Multiple endpoints are evaluated; the ClinicalTrials.gov record does not specify an adjustment procedure. |
| Total outcome measures posted | 22 | The registry contains a broader set of outcomes than the two primary endpoints. |
The appropriate interpretation of a P-value depends on the hypothesis family and prespecified testing strategy. Because the ClinicalTrials.gov record does not describe a multiplicity procedure, this page does not assign one retrospectively.
14. Superiority Testing
The registry identifies the hypothesis type for all registry-reported statistical analyses as superiority. That means the formal question is whether the treatment groups differ in the specified direction under the corresponding statistical test, rather than whether a treatment is merely no worse than a comparator by a prespecified non-inferiority margin.
The ClinicalTrials.gov record does not report a non-inferiority margin, so no non-inferiority interpretation should be applied to these results.
15. Safety Results
The ClinicalTrials.gov record provides serious adverse-event counts by treatment arm. These are reported as affected participants divided by participants at risk.
| Safety measure | Placebo | Tirzepatide |
|---|---|---|
| Serious adverse events | 14 / 292 | 17 / 287 |
These figures describe the number of participants affected and the corresponding at-risk denominator. The ClinicalTrials.gov record does not provide a statistical comparison, confidence interval, P-value, event-specific breakdown, severity distribution, or exposure-adjusted rate.
The serious-adverse-event summary should be interpreted descriptively from the information reported. The difference between 14/292 and 17/287 cannot by itself establish whether there is a statistically meaningful treatment-group difference because the ClinicalTrials.gov record does not report a formal comparison.
Safety interpretation also differs conceptually from efficacy interpretation. The efficacy analyses have prespecified model structures and reported effect estimates with confidence intervals and P-values; the serious-adverse-event ClinicalTrials.gov record are counts and denominators only.
16. Randomization and Blinding
The trial is classified as randomized with double masking and a parallel design. These design features are important statistical safeguards because randomization is intended to balance prognostic factors across treatment groups in expectation, while masking can reduce differential behavior, assessment, or reporting related to treatment assignment.
Randomization
Random allocation creates the primary basis for the between-group treatment comparison. The registry does not provide the randomization ratio in the ClinicalTrials.gov record.
Double masking
Double masking means treatment assignment was not openly known to the relevant masked parties. The ClinicalTrials.gov record does not identify which specific parties were masked.
Parallel groups
Participants were assigned to treatment groups that were compared in parallel rather than through a crossover design.
No crossover information reported
The ClinicalTrials.gov record does not report a crossover treatment strategy or crossover rate, so none is assumed in the interpretation.
17. Stratification and Covariate Adjustment
The registered primary MMRM explicitly includes Analysis Country, Sex, Treatment, Time, and the Treatment × Time interaction, in addition to baseline body weight. The registered logistic-regression analysis also includes baseline body weight, Analysis Country, Sex, and Treatment.
| Model | Variables explicitly reported in registry-reported registry definition |
|---|---|
| Primary MMRM | Baseline; Analysis Country; Sex; Treatment; Time; Treatment × Time |
| ≥5% logistic regression | Baseline body weight; Analysis Country; Sex; Treatment |
These variables are not all equivalent to randomization stratification factors. In particular, the ClinicalTrials.gov record describes them as components of the analysis models rather than as a list of stratification variables used during randomization.
18. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The two primary analyses report estimates in the prespecified superiority framework, with 95% confidence intervals and P-values <0.001. One is a model-based mean difference in percent body-weight change; the other is a logistic-regression odds ratio for achieving at least 5% body-weight reduction.
Clinical interpretation
The trial evaluates both the magnitude of weight change and the probability of crossing clinically defined weight-loss thresholds. These outcomes capture complementary aspects of response rather than a single universal measure of benefit.
The secondary results extend the analysis to maintenance of prior weight loss, larger weight-loss thresholds, waist circumference, BMI, blood pressure, metabolic biomarkers, and physical-function measures. Each should be interpreted on its own measurement scale and with attention to its statistical model.
19. Limitations and Interpretation Issues
- Analysis population: the posted analyses are restricted to randomly assigned participants who took at least one dose, had baseline and post-baseline outcome information, and exclude data after study-drug discontinuation.
- Missing-data assumptions: the ≥5% endpoint uses imputed data, but the ClinicalTrials.gov record does not specify the imputation method or sensitivity analyses.
- Multiplicity: the ClinicalTrials.gov record includes two primary endpoints and 20 secondary statistical analyses, but do not specify a multiplicity-adjustment procedure.
- Odds-ratio interpretation: odds ratios are not risk ratios or percentage-point differences. The large OR estimates should therefore not be translated directly into percentages without the underlying probabilities.
- Model dependence: the primary continuous result depends on the MMRM specification, including baseline adjustment, time, treatment-by-time interaction, and the treatment-discontinuation rule.
- Confidence-interval precision: the binary endpoint confidence intervals are broad on the odds-ratio scale, so the exact magnitude of the estimated association is less certain than its direction.
- Registry metadata: some outcome measures in the ClinicalTrials.gov record is categorized as "Binary" even when their units are continuous measures such as percent change. The analysis method and effect measure are therefore more informative for understanding the statistical procedure than the inferred endpoint-type label alone.
- Safety information: the registry-reported serious-adverse-event data are descriptive counts only and do not include a formal between-group analysis.
- Incomplete methodological detail: the ClinicalTrials.gov record does not report all implementation details of the covariance structure, imputation procedure, multiplicity strategy, or complete ANCOVA specifications.
20. Why This Trial Matters Statistically
SURMOUNT-3 is a useful teaching case because it combines longitudinal continuous outcomes with threshold-based binary outcomes within the same randomized trial. The statistical methods therefore illustrate why the form of the endpoint determines the appropriate effect measure and model.
| Concept | How it appears in SURMOUNT-3 |
|---|---|
| Randomization | Randomized parallel-group treatment comparison |
| Blinding | Double-masked design |
| Mixed-effects modeling | Primary body-weight percent-change analysis and many secondary longitudinal outcomes |
| Repeated measures | Post-baseline body-weight measurements analyzed through MMRM |
| Logistic regression | Binary weight-loss threshold outcomes |
| Odds ratio | Primary ≥5% response endpoint and four additional binary secondary endpoints |
| ANCOVA | Two patient-reported physical-function outcomes |
| Confidence intervals | Reported for all registry-reported statistical analyses |
| P-values | All registry-reported statistical analyses report P <0.001 |
| Missing-data handling | Imputed data explicitly reported for the ≥5% endpoint |
| Multiplicity | Two primary endpoints plus 20 secondary statistical analyses |
| Superiority testing | Hypothesis type reported as superiority for the posted analyses |
21. A Practical Reading of the Primary Results
The two primary endpoints should be read together because they measure different aspects of the same clinical domain.
Magnitude of change
The body-weight percent-change endpoint preserves the size of the change and produces a model-based mean difference of -24.5 with a 95% CI of -26.1 to -22.8.
Threshold response
The ≥5% endpoint asks whether participants crossed a predefined threshold and produces an odds ratio of 130.36 with a 95% CI of 69.98 to 242.84.
Different scales
The mean difference and odds ratio cannot be compared numerically. One is expressed in percentage-point change; the other compares odds.
Different assumptions
The MMRM and logistic regression models make different assumptions and address different outcome structures, so each result should be interpreted within its own analysis framework.
22. Statistical Interpretation of the Confidence Intervals
The primary continuous endpoint has a 95% confidence interval from -26.1 to -22.8. The entire interval lies on the same side of zero, which is consistent with the reported superiority P-value.
The primary binary endpoint has a 95% confidence interval from 69.98 to 242.84. The entire interval lies above 1, which is the null value for an odds ratio. At the same time, the interval is broad, emphasizing that statistical evidence for a directional difference does not imply precise knowledge of the exact odds-ratio magnitude.
A 95% confidence interval does not mean that there is a 95% probability that the fixed treatment effect is inside the displayed interval. It is a frequentist interval produced by a procedure designed to have 95% coverage under repeated sampling when the model and assumptions are appropriate.
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Limitations of the Available Registry Record
The ClinicalTrials.gov record is sufficient to reconstruct the major statistical architecture and all 22 posted statistical analyses, but they do not contain every detail that would normally appear in a full statistical analysis plan.
| Available | Not specified in the ClinicalTrials.gov record |
|---|---|
| Primary endpoint definitions and time frames | Complete SAP-level derivation algorithms |
| Analysis populations | Complete covariance-structure specification for MMRM |
| Methods and effect measures | Detailed imputation algorithm and sensitivity analyses |
| Estimates, 95% CIs, and P-values | Multiplicity-adjustment procedure |
| Serious adverse-event counts by arm | Formal statistical safety comparison |
Accordingly, this page interprets the statistical information that is actually registry-reported rather than reconstructing undocumented protocol or SAP details.
26. Sources
- ClinicalTrials.gov: SURMOUNT-3, NCT04657016.
- PubMed record: PMID 41885866.
- PubMed record: PMID 41640675.
- PubMed record: PMID 41537305.
- PubMed record: PMID 41187013.
- PubMed record: PMID 40717199.
Continue through Clinical Biostats
Use the related statistical concepts to move from the trial results to deeper explanations of the models and effect measures used in randomized clinical research.
27. Record Summary
SURMOUNT-3 provides a useful statistical example of how a randomized, double-masked phase 3 trial can combine several complementary outcome frameworks. The primary body-weight endpoint uses an MMRM with baseline adjustment, country, sex, treatment, time, and treatment-by-time interaction, producing a reported net mean difference of -24.5 with a 95% CI of -26.1 to -22.8 and P <0.001. The second primary endpoint converts weight change into a binary ≥5% response and uses logistic regression, producing an odds ratio of 130.36 with a 95% CI of 69.98 to 242.84 and P <0.001.
The secondary analyses extend the same statistical framework across larger weight-loss thresholds, maintenance of prior weight loss, anthropometric measures, cardiovascular measures, metabolic biomarkers, and patient-reported physical function. The most important interpretive lesson is that these estimates cannot be treated as interchangeable: mean differences, odds ratios, and confidence intervals answer different statistical questions and must be read on their respective scales.