← Clinical Trials
Type 2 Diabetes Phase 3 Double-Blind NCT04039503

SURPASS-5: Complete Statistical Analysis of Tirzepatide in Type 2 Diabetes

An independent statistical analysis of the randomized phase 3 SURPASS-5 trial evaluating tirzepatide versus placebo in participants with type 2 diabetes inadequately controlled on insulin glargine with or without metformin.

Trial status: COMPLETED  ·  Enrollment: 475  ·  Primary completion: December 22, 2020
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SURPASS-5 was a randomized, double-blind, parallel phase 3 trial evaluating three tirzepatide doses against placebo in participants with type 2 diabetes inadequately controlled on insulin glargine with or without metformin. The registry reports 475 enrolled participants, four treatment arms, 11 posted outcome measures, and 21 posted statistical analyses.

475
Enrolled
4 treatment arms
3
Tirzepatide doses
5, 10, and 15 mg
40
Primary time point
Week 40
<0.001
Primary P-values
Both comparisons
FeatureSURPASS-5
PhasePhase 3
ConditionType 2 Diabetes
Brief trial descriptionTirzepatide versus placebo in participants with type 2 diabetes inadequately controlled on insulin glargine with or without metformin
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment475
Arms4
Lead sponsorEli Lilly and Company
Sponsor typeIndustry
Trial startAugust 30, 2019
Primary completionDecember 22, 2020

2. Clinical Question

The central statistical question was whether adding tirzepatide at 5 mg, 10 mg, or 15 mg produced a greater improvement in HbA1c at Week 40 than placebo in participants with type 2 diabetes inadequately controlled on insulin glargine with or without metformin.

Population

Participants with type 2 diabetes inadequately controlled on insulin glargine with or without metformin.

Intervention

Tirzepatide, evaluated at 5 mg, 10 mg, and 15 mg treatment levels.

Comparator

Placebo.

Primary question

Does tirzepatide produce a greater change from baseline in HbA1c at Week 40 than placebo?

3. Trial Design

01
Randomize 475 participants
02
4 arms Three tirzepatide doses + placebo
03
Double-blind Parallel-group design
04
Follow-up Baseline through Week 40
05
Analyze MMRM and logistic regression
ARM 1

5 mg Tirzepatide

Tirzepatide treatment at 5 mg.

ARM 2

10 mg Tirzepatide

Tirzepatide treatment at 10 mg.

ARM 3

15 mg Tirzepatide

Tirzepatide treatment at 15 mg.

ARM 4

Placebo

Placebo comparator.

The design is statistically informative because the three active-dose comparisons create several parallel treatment-versus-placebo questions within the same randomized trial. The registry classifies the hypothesis type for the posted analyses as superiority.

Design interpretation: Randomization establishes the intended framework for comparing treatment groups, while double masking reduces the opportunity for knowledge of assignment to influence participant or investigator behavior and outcome assessment. The registry does not provide additional design details such as a crossover scheme, factorial structure, interim-analysis plan, or non-inferiority margin, so those features are not inferred here.

4. Endpoints

EndpointTime frameRegistry definition / analysis description
Change From Baseline in Hemoglobin A1c (HbA1c) (10 mg and 15 mg) Baseline, Week 40 HbA1c is the glycosylated fraction of hemoglobin A. HbA1c is measured primarily to identify average plasma glucose concentration over prolonged periods of time. LS mean was determined by mixed-model repeated measures (MMRM) with Baseline + Baseline Metformin Use (Yes, No) + Pooled Country + Treatment + Time + Treatment*Time (Type III sum of squares).
Change From Baseline in HbA1c (5 mg) Baseline, Week 40 Posted as a secondary endpoint and analyzed with a mixed models analysis.
Change From Baseline in Body Weight Baseline, Week 40 Continuous outcome measured in kilograms and analyzed with a mixed models analysis.
Percentage of Participants Achieving an HbA1c Target Value of <7% Week 40 Binary outcome analyzed using logistic regression.
Change From Baseline in Fasting Serum Glucose Baseline, Week 40 Continuous outcome measured in mg/dL and analyzed with a mixed models analysis.
Percentage of Participants Who Achieved Weight Loss ≥5% Week 40 Binary outcome analyzed using logistic regression.
Percentage Change From Baseline in Daily Mean Insulin Glargine Dose Baseline, Week 40 Percentage change in daily mean insulin glargine dose. The posted method is not reported.
Percentage of Participants Achieving an HbA1c Target Value of <5.7% Week 40 Binary outcome analyzed using logistic regression.

The registry identifies one registered primary endpoint, with two posted primary statistical comparisons: 10 mg versus placebo and 15 mg versus placebo. The 5 mg HbA1c comparison is posted as a secondary analysis.

5. Analysis Populations

The registry-defined populations for the posted analyses are not simply the full enrollment total. The primary HbA1c analyses use participants who were randomized, received at least one dose of study drug, and had a baseline and at least one post-baseline value, with an exclusion for patients discontinuing study drug due to inadvertent enrollment.

AnalysisPopulation described in the registry
Primary HbA1c analyses All randomized participants who received at least 1 dose of study drug and had a baseline and at least 1 post-baseline value, excluding patients discontinuing study drug due to inadvertent enrollment.
Secondary HbA1c 5 mg analysis All randomly assigned participants who received at least one dose of 5 mg tirzepatide or placebo and had a baseline and at least 1 post-baseline value, with the registry exclusion for inadvertent enrollment.
Body weight analyses All randomly assigned participants who took at least 1 dose of study drug and had a baseline and at least 1 post-baseline value, excluding patients discontinuing study drug due to inadvertent enrollment.
Binary endpoint analyses Generally all randomly assigned participants who took at least 1 dose of study drug and had a baseline and at least 1 post-baseline value, with the registry exclusion for inadvertent enrollment.

This distinction matters because the enrolled population of 475 is not automatically identical to every analysis population. Missing baseline or post-baseline measurements, treatment exposure, and the registry's specified exclusions can change the number of observations contributing to an individual analysis.

6. Statistical Methodology

Mixed-model repeated measures

The primary endpoint uses a mixed-model repeated measures framework. The registry states that the least-squares mean was determined with baseline HbA1c, baseline metformin use, pooled country, treatment, time, and treatment-by-time interaction included in the model, with Type III sums of squares.

Conceptual model
Outcome = Baseline + Metformin Use + Pooled Country + Treatment + Time + Treatment × Time + repeated-measures structure

The treatment-by-time interaction allows the estimated treatment difference to vary according to the scheduled assessment time. The reported Week 40 result is expressed as a least-squares mean difference rather than a raw arithmetic difference between observed group means.

Least-squares mean difference

The reported primary effect measure is the LS Mean Difference. This is a model-adjusted contrast between treatment groups. It should not automatically be interpreted as the simple difference between two unadjusted sample means because the model incorporates the covariates and repeated-measure structure specified by the registry.

Logistic regression

Binary outcomes such as achieving an HbA1c target of <7%, achieving an HbA1c target of <5.7%, and achieving at least 5% weight loss were analyzed using logistic regression. The effect measure reported by the registry is an odds ratio.

Odds-ratio interpretation
OR = odds of outcome in treatment group ÷ odds of outcome in comparator group

An odds ratio above 1 indicates higher odds of the specified binary outcome in the treatment group relative to the comparator. It is not itself a probability, risk ratio, or percentage-point difference.

Confidence intervals

The registry reports two-sided 95% confidence intervals for the posted treatment effects. For a mean difference, the interval describes uncertainty around the estimated difference. For an odds ratio, the interval describes uncertainty around the estimated ratio of odds.

P-values

The primary comparisons have P-values of <0.001. A P-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model and testing framework. It does not quantify the size, clinical importance, or probability of the treatment effect.

7. Primary Results: HbA1c at Week 40

The registered primary endpoint was change from baseline in HbA1c at Week 40 for the 10 mg and 15 mg comparisons. Both posted analyses used mixed models and reported least-squares mean differences with two-sided 95% confidence intervals.

10 mg Tirzepatide vs Placebo

LS mean difference in HbA1c change

-1.66

95% CI: -1.88 to -1.43   ·   P < 0.001

Outcome unit: Percentage of HbA1c  ·  Baseline to Week 40

Clinical Biostats interpretation

The reported LS mean difference of -1.66 means that the model-estimated change from baseline in HbA1c was 1.66 percentage points lower for 10 mg tirzepatide than for placebo at the specified Week 40 analysis.

The negative sign is important: the treatment-group estimate is lower than the placebo-group estimate for this change-from-baseline outcome. It does not mean that every participant experienced a 1.66-point reduction, nor does it establish the individual treatment response for any particular participant.

The 95% confidence interval of -1.88 to -1.43 describes statistical uncertainty around the model-based treatment contrast. It does not mean that 95% of individual participants had treatment effects inside that interval.

The P-value of <0.001 indicates strong incompatibility with the null hypothesis of no treatment difference under the stated analysis. It does not say that there is a <0.1% probability that the null hypothesis is true, and it does not measure the magnitude of the HbA1c effect.

Because the analysis is based on a repeated-measures model, interpretation also depends on the model specification, the handling of repeated observations, the analysis population, and the assumptions underlying the model. This is not a time-to-event analysis, so proportional-hazards assumptions are not relevant to this particular estimate.

15 mg Tirzepatide vs Placebo

LS mean difference in HbA1c change

-1.65

95% CI: -1.88 to -1.43   ·   P < 0.001

Outcome unit: Percentage of HbA1c  ·  Baseline to Week 40

Clinical Biostats interpretation

The reported LS mean difference of -1.65 indicates a model-estimated HbA1c change from baseline that was 1.65 percentage points lower with 15 mg tirzepatide than with placebo at Week 40.

Again, this is a population-level model contrast rather than a prediction of an individual's response. The estimate summarizes the treatment comparison under the analysis population and model specified in the registry.

The 95% confidence interval of -1.88 to -1.43 gives the statistical precision of the estimated treatment difference. The relatively narrow interval indicates that the registry estimate is substantially more precise than it would be if the interval were very wide, but it still does not describe individual-patient variability.

The P-value of <0.001 addresses the null-hypothesis testing question, not the size or clinical relevance of the treatment effect. Effect size and precision should therefore be considered from the estimate and confidence interval rather than from the P-value alone.

As with the 10 mg comparison, the analysis relies on the mixed-model framework and its assumptions. The posted registry information does not provide enough detail to reconstruct every covariance or missing-data specification, so those details should not be inferred.

Primary comparisonLS mean difference95% CIP-valueMethod
10 mg Tirzepatide vs Placebo-1.66-1.88 to -1.43<0.001Mixed Models Analysis
15 mg Tirzepatide vs Placebo-1.65-1.88 to -1.43<0.001Mixed Models Analysis
Important statistical point: the two primary comparisons are separate treatment-versus-placebo contrasts. The posted P-values demonstrate evidence against their respective null hypotheses under the registry's analysis framework; they should not be interpreted as evidence that the 15 mg dose is statistically superior to the 10 mg dose because no direct 15 mg-versus-10 mg primary comparison is reported here.

8. Secondary HbA1c Result: 5 mg Tirzepatide

The 5 mg tirzepatide comparison was posted as a secondary endpoint using the same general mixed-model approach.

LS mean difference in HbA1c change

-1.30

95% CI: -1.52 to -1.07   ·   P < 0.001

5 mg Tirzepatide vs Placebo  ·  Baseline to Week 40

The estimate of -1.30 represents the model-based difference in HbA1c change between 5 mg tirzepatide and placebo. Its two-sided 95% confidence interval extends from -1.52 to -1.07, and the reported P-value is <0.001.

Because this is a secondary endpoint rather than the registered primary endpoint, its inferential role is different from the two primary comparisons. The statistical estimate can be interpreted directly, but the broader evidentiary hierarchy depends on the trial's prespecified multiplicity strategy, which is not reported in the ClinicalTrials.gov record.

9. Secondary Results: Body Weight

Body weight was analyzed from baseline to Week 40 using mixed models. The registry reports LS mean differences for all three tirzepatide doses versus placebo.

ComparisonLS mean difference95% CIP-value
5 mg Tirzepatide vs Placebo-7.8 kg-9.4 to -6.3<0.001
10 mg Tirzepatide vs Placebo-9.9 kg-11.5 to -8.3<0.001
15 mg Tirzepatide vs Placebo-12.6 kg-14.2 to -11.0<0.001

5 mg

The model-estimated difference in change from baseline was -7.8 kg, with a 95% CI of -9.4 to -6.3 and P < 0.001.

10 mg

The model-estimated difference was -9.9 kg, with a 95% CI of -11.5 to -8.3 and P < 0.001.

15 mg

The model-estimated difference was -12.6 kg, with a 95% CI of -14.2 to -11.0 and P < 0.001.

Interpretation

These are model-based treatment contrasts. They are not direct evidence that the dose levels were formally compared with one another.

10. Secondary Results: HbA1c Target <7%

The registry reports logistic-regression analyses for the percentage of participants achieving an HbA1c target value of <7% at Week 40.

ComparisonOdds ratio95% CIP-value
5 mg Tirzepatide vs Placebo37.7715.23 to 93.70<0.001
10 mg Tirzepatide vs Placebo100.0730.02 to 333.62<0.001
15 mg Tirzepatide vs Placebo43.3116.92 to 110.83<0.001
How to interpret these odds ratios

An odds ratio of 37.77 for the 5 mg comparison means the estimated odds of achieving the specified HbA1c target were 37.77 times the corresponding odds in the placebo group under the posted logistic-regression analysis.

An odds ratio of 100.07 does not mean that 100.07% of participants achieved the target, nor does it mean that the probability was 100.07 times higher. Odds and probabilities are mathematically related but are not interchangeable.

The same principle applies to the 15 mg estimate of 43.31. The wide confidence intervals, particularly for the larger odds-ratio estimates, also show that ratio estimates can be substantially less precise than their point estimates may initially suggest.

11. Secondary Results: Fasting Serum Glucose

Fasting serum glucose was analyzed as a change-from-baseline outcome at Week 40 using mixed models.

ComparisonLS mean difference95% CIP-value
5 mg Tirzepatide vs Placebo-22.5 mg/dL-29.5 to -15.4<0.001
10 mg Tirzepatide vs Placebo-29.0 mg/dL-36.0 to -22.0<0.001
15 mg Tirzepatide vs Placebo-28.8 mg/dL-35.9 to -21.6<0.001

The negative estimates indicate lower model-estimated change in fasting serum glucose relative to placebo. The reported confidence intervals for all three comparisons remain below zero, and all three P-values are <0.001.

12. Secondary Results: Weight Loss ≥5%

The binary endpoint for achieving weight loss of at least 5% was analyzed using logistic regression.

ComparisonOdds ratio95% CIP-value
5 mg Tirzepatide vs Placebo17.157.55 to 38.93<0.001
10 mg Tirzepatide vs Placebo27.2411.87 to 62.55<0.001
15 mg Tirzepatide vs Placebo79.6132.76 to 193.44<0.001

The odds ratios are large relative measures, but they should not be converted directly into probability differences without the underlying event probabilities. The confidence intervals also demonstrate substantial uncertainty around the precise magnitude of the odds ratios, despite the direction of every interval being above 1.

13. Secondary Results: Daily Mean Insulin Glargine Dose

The registry reports percentage change from baseline in daily mean insulin glargine dose at Week 40. For these comparisons, the record does not name a statistical method, so none is inferred here.

ComparisonEstimate difference95% CIMethod
5 mg Tirzepatide vs 10 mg Tirzepatide-35.4-46.0 to -22.8Not reported
10 mg Tirzepatide vs Placebo-38.2-48.3 to -26.1Not reported
15 mg Tirzepatide vs Placebo-49.3-57.7 to -39.4Not reported
Methodological caution: these estimates should not be described as mixed-model, logistic-regression, or another specific model because the ClinicalTrials.gov record does not report the method. The correct statistical interpretation is therefore limited to the reported estimate and confidence interval.

14. Secondary Results: HbA1c Target <5.7%

The registry also reports logistic-regression analyses for the percentage of participants achieving an HbA1c target value of <5.7% at Week 40.

ComparisonOdds ratio95% CIP-value
5 mg Tirzepatide vs Placebo12.223.93 to 38.00<0.001
10 mg Tirzepatide vs Placebo32.3610.52 to 99.49<0.001
15 mg Tirzepatide vs Placebo56.2618.27 to 173.26<0.001

These results use the same odds-ratio framework as the <7% endpoint. The odds-ratio scale can produce very large numbers when the outcome is substantially more common in one group than another, so interpretation should remain on the odds scale unless the underlying event probabilities are also reported.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These counts should be interpreted as affected participants divided by the number at risk in each arm.

ArmSerious adverse eventsAffected / at risk
5 mg TirzepatideSerious adverse events9 / 116
10 mg TirzepatideSerious adverse events13 / 119
15 mg TirzepatideSerious adverse events9 / 120
PlaceboSerious adverse events10 / 120

5 mg Tirzepatide

Serious adverse events were reported in 9 of 116 participants at risk.

10 mg Tirzepatide

Serious adverse events were reported in 13 of 119 participants at risk.

15 mg Tirzepatide

Serious adverse events were reported in 9 of 120 participants at risk.

Placebo

Serious adverse events were reported in 10 of 120 participants at risk.

The ClinicalTrials.gov record does not report a formal statistical comparison of serious adverse-event rates. Accordingly, the counts are presented descriptively rather than converted into an inferred treatment effect or P-value.

16. Multiplicity and Multiple Comparisons

SURPASS-5 contains several treatment comparisons and several outcome measures. The registry data identifies the primary endpoint as the change from baseline in HbA1c for the 10 mg and 15 mg comparisons, while additional HbA1c, body-weight, glucose, target-achievement, and insulin-dose analyses are posted as secondary outcomes.

Analysis familyPosted roleStatistical method reported
HbA1c change, 10 mg vs placeboPrimaryMixed Models Analysis
HbA1c change, 15 mg vs placeboPrimaryMixed Models Analysis
HbA1c change, 5 mg vs placeboSecondaryMixed Models Analysis
Body weightSecondaryMixed Models Analysis
HbA1c <7%SecondaryLogistic regression
Fasting serum glucoseSecondaryMixed Models Analysis
Weight loss ≥5%SecondaryLogistic regression
HbA1c <5.7%SecondaryLogistic regression
Daily mean insulin glargine doseSecondaryNot reported

Multiplicity is important because multiple hypothesis tests create more opportunities for statistically significant findings to occur by chance. The ClinicalTrials.gov record does not provide a multiplicity-adjustment procedure, alpha-allocation scheme, hierarchical testing sequence, or interim-analysis plan. Therefore, no such procedure is inferred.

Interpretation boundary: the fact that many posted analyses have P-values of <0.001 does not establish that every secondary endpoint was independently confirmatory or that a particular familywise error rate was maintained. The role of each result should be understood from its registered endpoint role and the statistical design information actually reported.

17. Stratification and Covariate Adjustment

The primary endpoint definition provides unusually useful detail about the covariates incorporated into the MMRM. The model included:

This is different from simply comparing the arithmetic mean change in HbA1c between arms. The least-squares means are model-adjusted estimates that account for the covariates specified in the analysis.

Why baseline adjustment can help
Adjusted treatment contrast = model-estimated outcome difference conditional on specified covariates

Including baseline HbA1c can improve precision because the analysis accounts for an important predictor of the outcome. Including baseline metformin use and pooled country allows the treatment comparison to account for those prespecified factors within the model.

18. Missing Data and Repeated Measures

The primary analysis population requires a baseline and at least one post-baseline value. This is important because a mixed-model repeated-measures analysis uses available longitudinal measurements rather than requiring every participant to contribute an observation at every scheduled time point.

However, the ClinicalTrials.gov record does not specify the missing-data assumption, covariance structure, imputation sensitivity analyses, or estimand framework used for the MMRM beyond the model terms explicitly reported in the endpoint definition.

Why this matters: MMRM is not synonymous with "no missing-data assumptions." The validity of inference depends on the assumptions connecting observed longitudinal data with unobserved outcomes. Because the ClinicalTrials.gov record does not specify those details, they should not be reconstructed from convention alone.

19. Statistical Methods Explained

Why was a mixed-effects model used for HbA1c?

HbA1c is measured longitudinally, with the registered primary endpoint defined from baseline to Week 40 and the model including time and treatment-by-time interaction. A mixed-model repeated-measures framework is suited to repeated observations because it models the longitudinal outcome while estimating treatment contrasts at the relevant assessment time.

What does an LS mean difference of -1.66 mean?

It means the model-estimated change from baseline in HbA1c was 1.66 percentage points lower for 10 mg tirzepatide than for placebo at the specified Week 40 comparison. It is an adjusted model contrast, not an individual prediction and not necessarily the same as the difference between two unadjusted observed means.

What does an odds ratio of 100.07 mean?

It means the estimated odds of achieving the HbA1c target of <7% were 100.07 times the odds in the placebo group for the 10 mg comparison under the posted logistic-regression analysis. It does not mean that 100.07% of participants achieved the target and should not be interpreted as a probability ratio.

Why are confidence intervals important when the P-value is <0.001?

The P-value addresses evidence against the null hypothesis, whereas the confidence interval describes the statistical precision of the estimated treatment effect. For example, the 10 mg primary HbA1c estimate is -1.66 with a 95% CI of -1.88 to -1.43. The interval tells the reader considerably more about the range of model-compatible effect estimates than the P-value alone.

Why doesn't a P-value measure effect size?

A P-value depends on the estimated effect, its uncertainty, and the amount of information contributing to the analysis. A small P-value can accompany a modest effect in a large, precise study, while a larger effect can fail to achieve a small P-value when uncertainty is substantial. Effect estimates and confidence intervals should therefore be read alongside the P-value.

Why distinguish primary from secondary endpoints?

Primary endpoints define the main confirmatory questions of a trial. Secondary endpoints provide additional evidence but can involve many additional statistical comparisons. Without knowing the prespecified multiplicity strategy, one should not automatically treat every statistically significant secondary result as having the same confirmatory status as a primary endpoint.

Why does randomization matter for interpretation?

Randomization creates the trial's intended basis for comparing treatment groups. It helps balance measured and unmeasured factors in expectation and supports causal interpretation of the randomized treatment contrast, subject to the trial's design, conduct, analysis population, and missing-data assumptions.

20. Understanding the Primary HbA1c Effect

Effect size

The primary estimates were -1.66 for 10 mg tirzepatide versus placebo and -1.65 for 15 mg tirzepatide versus placebo. On the change-from-baseline scale, both estimates indicate a lower model-estimated HbA1c change in the tirzepatide groups relative to placebo.

Precision

The 10 mg 95% confidence interval was -1.88 to -1.43, while the 15 mg 95% confidence interval was also -1.88 to -1.43. A confidence interval is an interval for the statistical treatment contrast, not an interval containing 95% of individual treatment responses.

Inference

Both primary comparisons have P-values of <0.001. That result addresses the null-hypothesis testing question under the posted mixed-model analysis. It does not establish the probability that the treatment works, the probability that the estimate is exactly correct, or the clinical importance of the observed effect.

21. Comparing Dose Groups Without Overinterpreting Them

The posted results include three active-dose comparisons with placebo, but the presence of three estimates does not by itself create a formal dose-response test.

Outcome5 mg vs placebo10 mg vs placebo15 mg vs placebo
HbA1c change-1.30-1.66-1.65
Body weight change-7.8 kg-9.9 kg-12.6 kg
Fasting serum glucose change-22.5 mg/dL-29.0 mg/dL-28.8 mg/dL
Weight loss ≥5% OR17.1527.2479.61

The pattern across the posted point estimates can be described, but a formal claim that one dose is statistically superior to another would require a direct dose-to-dose comparison and an appropriate inferential framework. The ClinicalTrials.gov record does not provide such a formal comparison for these outcomes.

22. Why This Trial Matters Statistically

SURPASS-5 is a useful teaching example because it combines randomized treatment allocation, multiple active doses, repeated continuous outcomes, binary responder endpoints, model-adjusted mean contrasts, odds ratios, confidence intervals, and multiple secondary analyses within one phase 3 trial.

ConceptHow it appears in SURPASS-5
RandomizationParticipants were randomly allocated in a parallel-group design.
BlindingThe trial was double-masked.
Multiple treatment armsThree tirzepatide dose groups were compared with placebo.
Mixed-effects modelingUsed for the primary HbA1c endpoint and several continuous secondary outcomes.
Repeated measuresThe primary endpoint uses a mixed-model repeated-measures framework with time and treatment-by-time interaction.
Covariate adjustmentBaseline HbA1c, baseline metformin use, and pooled country were included in the primary model.
Least-squares meansPrimary and several secondary continuous outcomes are reported as LS mean differences.
Logistic regressionUsed for HbA1c targets and weight-loss responder outcomes.
Odds ratiosBinary endpoints are reported using odds ratios with 95% confidence intervals.
MultiplicityOne primary endpoint generates two primary dose-versus-placebo comparisons, alongside numerous secondary analyses.
Analysis populationsPosted analyses require treatment exposure and baseline/post-baseline observations rather than automatically using all 475 enrolled participants.
Safety analysisSerious adverse events are reported descriptively by treatment arm.

23. Important Limitations and Interpretation Issues

24. What the Odds Ratio Does — and Does Not — Mean

Statistical interpretation

For the 10 mg comparison on the HbA1c <7% endpoint, the reported odds ratio was 100.07. This means the estimated odds of achieving the specified endpoint were 100.07 times the odds under placebo in the posted logistic-regression analysis.

It does not mean that the probability of achieving the endpoint was 100.07 times higher, that 100.07% achieved the endpoint, or that every participant experienced that relative effect.

Why the confidence interval matters

The corresponding 95% confidence interval was 30.02 to 333.62. The interval is broad relative to the point estimate, which means that the precise magnitude of the odds ratio is substantially less certain than the direction of the association indicated by the estimate and interval.

Why probability and odds should not be mixed

Odds are defined as probability divided by one minus probability. Because that transformation is nonlinear, a large odds ratio cannot be interpreted as the same numerical increase in probability. Absolute responder rates are required for a direct probability-based interpretation.

25. Planned Analysis Features Not Reported in the Supplied Data

Several design features often discussed on clinical-trial statistical pages are not provided in the trial data for SURPASS-5. They are therefore not reconstructed here.

FeatureInformation reportedInterpretation
Non-inferiority marginNot reportedNot applicable to the registry-reported superiority analyses; no margin is inferred.
CrossoverNot reportedNo crossover structure is inferred.
Factorial designNot reportedThe registry explicitly describes a parallel design; no factorial structure is inferred.
Interim analysisNot reportedNo interim-testing or alpha-spending procedure is inferred.
Multiplicity adjustmentNot reportedNo correction procedure or hierarchical strategy is inferred.
Missing-data imputationNot reportedNo imputation method is assigned beyond the reported MMRM framework.
Bayesian methodsNot reportedThe posted analyses are described using mixed models and logistic regression, not Bayesian methods.

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Calculators

28. Sources

Continue with the statistical methods

Explore the broader Clinical Biostats tutorials, statistical calculators, and clinical-trial analyses connected to the methods used in SURPASS-5.

29. Record Summary

SURPASS-5 provides a useful statistical example of how a randomized phase 3 trial can combine repeated-measures modeling with categorical responder analyses. The primary HbA1c endpoint was evaluated using a mixed-model repeated-measures framework incorporating baseline HbA1c, baseline metformin use, pooled country, treatment, time, and treatment-by-time interaction. The two primary treatment comparisons produced LS mean differences of -1.66 for 10 mg tirzepatide versus placebo and -1.65 for 15 mg tirzepatide versus placebo, each with a two-sided 95% confidence interval of -1.88 to -1.43 and a P-value of <0.001.

The secondary analyses extend the statistical story across body weight, fasting serum glucose, HbA1c targets, weight-loss response, and insulin glargine dose. Continuous outcomes were generally reported as LS mean differences, while binary outcomes were analyzed with logistic regression and expressed as odds ratios. The registry also reports serious adverse events descriptively by arm.

Clinical Biostats methodology: A trial-results page should separate the reported numerical evidence from statistical interpretation. For SURPASS-5, that means preserving the registry's endpoint definitions, analysis populations, effect measures, confidence intervals, and P-values while avoiding unsupported assumptions about multiplicity, missing-data handling, interim analyses, or unreported statistical methods.