← Clinical Trials
Type 2 Diabetes Phase 3 Randomized NCT03730662

SURPASS-4: Complete Statistical Analysis of Tirzepatide in Type 2 Diabetes

An independent statistical analysis of the randomized phase 3 SURPASS-4 trial comparing once-weekly tirzepatide with once-daily insulin glargine in participants with type 2 diabetes and increased cardiovascular risk.

Trial status: COMPLETED  ·  Enrollment: 2002  ·  Primary completion: January 22, 2021
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses posted in the ClinicalTrials.gov record. The registry provides the official trial record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SURPASS-4 was a randomized, parallel, unmasked phase 3 treatment trial comparing three doses of once-weekly tirzepatide with once-daily insulin glargine in participants with type 2 diabetes mellitus and increased cardiovascular risk. The ClinicalTrials.gov record contains two formal primary analyses for the Week 52 HbA1c endpoint and additional secondary analyses for body weight, HbA1c target attainment, and fasting serum glucose.

2002
Enrolled
Phase 3 trial
4
Arms
3 tirzepatide doses + insulin glargine
-0.99
10 mg HbA1c difference
97.5% CI -1.13 to -0.86
-1.14
15 mg HbA1c difference
97.5% CI -1.28 to -1.00
FeatureSURPASS-4
Trial nameSURPASS-4
ClinicalTrials.gov identifierNCT03730662
PhasePhase 3
StatusCOMPLETED
Therapeutic areaEndocrinology
ConditionType 2 Diabetes Mellitus
AllocationRANDOMIZED
Design modelPARALLEL
MaskingNONE
Primary purposeTREATMENT
Enrollment2002
InterventionsTirzepatide; Insulin Glargine
Lead sponsorEli Lilly and Company
Sponsor typeINDUSTRY
StartNovember 20, 2018
Primary completionJanuary 22, 2021

2. Clinical Question

The clinical question was whether once-weekly tirzepatide, at the evaluated 10 mg and 15 mg doses for the primary endpoint and at 5 mg for a secondary endpoint, could produce a favorable change in HbA1c compared with once-daily insulin glargine at Week 52 in participants with type 2 diabetes mellitus.

Population

Participants with type 2 diabetes mellitus and increased cardiovascular risk, as described by the trial's brief title.

Intervention

Once-weekly tirzepatide at 5 mg, 10 mg, or 15 mg.

Comparator

Once-daily insulin glargine.

Primary question

For the registered Week 52 HbA1c endpoint, do the 10 mg and 15 mg tirzepatide comparisons establish non-inferiority relative to insulin glargine, with the design also guided by a superiority objective?

3. Trial Design

01
Randomize2002 enrolled
02
4 armsThree tirzepatide doses + insulin glargine
03
TreatmentOnce-weekly tirzepatide or once-daily insulin glargine
04
Week 52HbA1c and other outcome measures
05
AnalysisMixed-effects and logistic regression models
Allocation
Randomized allocation in a parallel-group design.
Masking
NONE according to the registry record.
Primary purpose
TREATMENT.
Statistical methods posted
Mixed-effects model and logistic regression.
ARM · 5 MG TIRZEPATIDE

Tirzepatide 5 mg

  • Once-weekly tirzepatide.
  • Secondary analyses included HbA1c, body weight, percentage with HbA1c <7.0%, and fasting serum glucose.
ARM · 10 MG TIRZEPATIDE

Tirzepatide 10 mg

  • Once-weekly tirzepatide.
  • Week 52 change in HbA1c was a primary analysis.
ARM · 15 MG TIRZEPATIDE

Tirzepatide 15 mg

  • Once-weekly tirzepatide.
  • Week 52 change in HbA1c was a primary analysis.
ARM · INSULIN GLARGINE

Insulin glargine

  • Once-daily insulin glargine.
  • Comparator for the three tirzepatide dose groups.

4. Trial Timing and Registry Results

2018-11-20

Trial start

The registry profile lists November 20, 2018 as the study start date.

2021-01-22

Primary completion

The registry profile lists January 22, 2021 as the primary completion date.

COMPLETED

Results posted

The trial is listed as completed, with 8 outcome measures and 12 statistical analyses posted in the ClinicalTrials.gov record.

5. Endpoints

The registered primary endpoint was change from baseline in hemoglobin A1c (HbA1c) at Week 52 for the 10 mg and 15 mg comparisons. The registry definition posted on ClinicalTrials.gov for the endpoint states that HbA1c is the glycosylated fraction of hemoglobin A and is measured primarily to identify average plasma glucose concentration over prolonged periods of time. It further states that the least-squares mean was determined by a mixed-model repeated-measures model for post-baseline measures.

EndpointTime frameAnalysisEffect measure
Change From Baseline in Hemoglobin A1c (HbA1c) (10 mg and 15 mg) Baseline, Week 52 Mixed Models Analysis Mean Difference (Net)
Change From Baseline in HbA1c (5 mg) Baseline, Week 52 Mixed Models Analysis Mean Difference (Net)
Change From Baseline in Body Weight Baseline, Week 52 Mixed Models Analysis Mean Difference (Net)
Percentage of Participants With HbA1c of <7.0% Week 52 Regression, Logistic Odds Ratio (OR)
Change From Baseline in Fasting Serum Glucose Baseline, Week 52 Mixed Models Analysis Mean Difference (Net)
Endpoint terminology: The registry's statistical-analysis records label the HbA1c analyses with an endpoint type of "Binary," even though the reported effect measure is a mean difference in change from baseline and the methodology is a mixed-effects model. This page preserves the reported analysis method and effect measure rather than attempting to reinterpret the registry's endpoint-type field.

6. Analysis Population

The statistical analyses posted on ClinicalTrials.gov use the same general analysis-population definition: all randomized participants who received at least one dose of study drug and had a baseline and at least 1 post-baseline value, excluding participants discontinuing study drug due to inadvertent enrolment.

FeatureRegistry-specified analysis population
Randomization All randomized participants meeting the subsequent analysis criteria.
Treatment exposure Participants must have received at least one dose of study drug.
Baseline measurement A baseline value was required.
Post-baseline information At least 1 post-baseline value was required.
Exclusion Participants discontinuing study drug due to inadvertent enrolment were excluded.

This population definition is important when interpreting the estimates. The reported analyses are not simply a comparison of every enrolled participant regardless of whether they received treatment or registry-reported post-baseline measurements. The analysis population is explicitly conditioned on treatment exposure and availability of baseline and post-baseline data.

7. Primary Results: HbA1c at Week 52

The registry contains two formal primary statistical analyses for the registered endpoint "Change From Baseline in Hemoglobin A1c (HbA1c) (10 mg and 15 mg)" at Baseline and Week 52. Both use a mixed-effects model and report a net mean difference relative to insulin glargine. The hypothesis type recorded for both comparisons is non-inferiority.

10 mg Tirzepatide vs Insulin Glargine

Net mean difference in change from baseline HbA1c

-0.99

97.5% two-sided CI: -1.13 to -0.86   ·   P < 0.001

Time frame: Baseline, Week 52   ·   Hypothesis type: Non-inferiority

MeasureReported value
Comparison10 mg Tirzepatide vs Insulin Glargine
Effect measureMean Difference (Net)
Estimate-0.99
Confidence interval97.5% two-sided CI -1.13 to -0.86
P-value<0.001
MethodMixed Models Analysis / mixed-effects model
Time frameBaseline, Week 52
Clinical Biostats interpretation

The estimated net mean difference of -0.99 means that the model-estimated change from baseline in HbA1c was 0.99 percentage points lower for the 10 mg tirzepatide comparison than for insulin glargine, based on the registry's reported mean-difference convention.

The negative value describes the direction of the between-group difference; it does not mean that every participant experienced a 0.99-percentage-point reduction, nor does it describe the individual treatment response of any particular participant.

The 97.5% two-sided confidence interval of -1.13 to -0.86 describes statistical uncertainty around the estimated between-group difference under the specified analysis. It is not a range containing 97.5% of individual patient responses.

The P < 0.001 result addresses evidence against the statistical null hypothesis under the analysis framework. A p-value does not measure the size, clinical importance, or probability of the observed effect, and it should not be interpreted as the probability that the treatment works.

The registry identifies this comparison as a non-inferiority analysis and states that sample-size selection was guided by the objective of establishing superiority. The ClinicalTrials.gov record does not provide a numerical non-inferiority margin, so a margin-specific conclusion cannot be reconstructed here. In a non-inferiority analysis, the key question is whether the confidence interval is compatible with an effect that remains within the prespecified acceptable margin, rather than whether a p-value alone is small.

15 mg Tirzepatide vs Insulin Glargine

Net mean difference in change from baseline HbA1c

-1.14

97.5% two-sided CI: -1.28 to -1.00   ·   P < 0.001

Time frame: Baseline, Week 52   ·   Hypothesis type: Non-inferiority

MeasureReported value
Comparison15 mg Tirzepatide vs Insulin Glargine
Effect measureMean Difference (Net)
Estimate-1.14
Confidence interval97.5% two-sided CI -1.28 to -1.00
P-value<0.001
MethodMixed Models Analysis / mixed-effects model
Time frameBaseline, Week 52
Clinical Biostats interpretation

The estimated net mean difference of -1.14 indicates a model-estimated change from baseline in HbA1c that was 1.14 percentage points lower for the 15 mg tirzepatide comparison than for insulin glargine, using the registry's reported effect-measure convention.

The estimate is a between-group mean difference. It does not imply that each individual participant experienced a 1.14-percentage-point difference, and it does not provide an individual-level probability of benefit.

The 97.5% two-sided CI of -1.28 to -1.00 quantifies uncertainty around the model-based estimate. The relatively narrow interval provides information about the precision of this particular estimated comparison, but it does not describe individual variability in HbA1c response.

The P < 0.001 indicates strong statistical evidence under the reported hypothesis-testing framework. It is not an effect-size measure. The magnitude of the estimate and its confidence interval are required to understand the size and precision of the difference.

For the non-inferiority objective, interpretation should be tied to the prespecified non-inferiority margin. Because the ClinicalTrials.gov record does not state that numerical margin, this page does not manufacture one or perform a reconstructed margin calculation.

8. Primary HbA1c Results Side by Side

ComparisonNet mean difference97.5% two-sided CIP-valueHypothesis
10 mg Tirzepatide vs Insulin Glargine-0.99-1.13 to -0.86<0.001Non-inferiority
15 mg Tirzepatide vs Insulin Glargine-1.14-1.28 to -1.00<0.001Non-inferiority

The two primary estimates are both negative, with the 15 mg comparison showing a larger reported difference than the 10 mg comparison. That comparison is descriptive: the two estimates come from separate dose-versus-control analyses and should not automatically be interpreted as a formal statistical comparison of 10 mg versus 15 mg.

Non-inferiority caution: The ClinicalTrials.gov record states that the primary objective was to establish noninferiority and that sample-size selection was guided by superiority objectives, but they do not provide the numerical non-inferiority margin. A complete non-inferiority interpretation requires the prespecified margin and its direction. The p-values reported here should therefore not be substituted for the margin-based decision rule.

9. Secondary HbA1c Result: 5 mg Tirzepatide

The 5 mg tirzepatide comparison was recorded as a secondary endpoint rather than a primary endpoint. It used the same general mixed-model approach and the same analysis-population definition.

Net mean difference in change from baseline HbA1c

-0.80

97.5% two-sided CI: -0.93 to -0.66   ·   P < 0.001

5 mg Tirzepatide vs Insulin Glargine   ·   Hypothesis type: Non-inferiority

Clinical Biostats interpretation

The reported estimate of -0.80 is the net mean difference in change from baseline HbA1c for 5 mg tirzepatide versus insulin glargine. The negative direction indicates a lower model-estimated change for the tirzepatide group under the registry's effect-measure convention.

The 97.5% CI of -0.93 to -0.66 describes uncertainty around that estimate. Because the confidence interval is based on the statistical model, it should not be read as an interval containing individual participants' HbA1c changes.

The P < 0.001 is evidence from the specified statistical test and is not itself a measure of treatment magnitude. The estimate and confidence interval provide the corresponding information about the size and precision of the difference.

10. Secondary Body-Weight Results

Change from baseline in body weight was evaluated at Baseline and Week 52 using a mixed-effects model. The registry reports superiority as the hypothesis type for all three dose comparisons.

ComparisonNet mean difference95% two-sided CIP-value
5 mg Tirzepatide vs Insulin Glargine-9.0 kg-9.8 to -8.3 kg<0.001
10 mg Tirzepatide vs Insulin Glargine-11.4 kg-12.1 to -10.6 kg<0.001
15 mg Tirzepatide vs Insulin Glargine-13.5 kg-14.3 to -12.8 kg<0.001

5 mg comparison

The reported net mean difference was -9.0 kg, with a 95% two-sided CI of -9.8 to -8.3 kg and P < 0.001.

10 mg comparison

The reported net mean difference was -11.4 kg, with a 95% two-sided CI of -12.1 to -10.6 kg and P < 0.001.

15 mg comparison

The reported net mean difference was -13.5 kg, with a 95% two-sided CI of -14.3 to -12.8 kg and P < 0.001.

Statistical role

These are secondary superiority analyses. The registry does not supply additional information here about a separate dose-response hypothesis test.

The pattern of estimates is descriptively more negative at the higher tirzepatide doses. However, comparing the numerical estimates across doses is not the same as conducting a formal dose-response test. The statistical analyses posted on ClinicalTrials.gov contain dose-versus-insulin-glargine comparisons, not a separate formal test of one tirzepatide dose against another.

11. Secondary HbA1c Target-Attainment Results

The registry also reports the percentage of participants with HbA1c of <7.0% at Week 52. These binary outcomes were analyzed using logistic regression, with odds ratio as the effect measure and superiority as the hypothesis type.

ComparisonOdds ratio95% two-sided CIP-value
5 mg Tirzepatide vs Insulin Glargine4.783.47 to 6.58<0.001
10 mg Tirzepatide vs Insulin Glargine9.236.31 to 13.49<0.001
15 mg Tirzepatide vs Insulin Glargine11.877.88 to 17.89<0.001
Odds-ratio interpretation
OR = odds of HbA1c < 7.0% in tirzepatide group ÷ odds of HbA1c < 7.0% in insulin-glargine group

An odds ratio above 1 indicates higher estimated odds of the specified binary outcome in the tirzepatide comparison group. The odds ratio is not itself a risk ratio or a percentage-point difference in the probability of achieving HbA1c <7.0%.

Clinical Biostats interpretation

The reported odds ratio of 4.78 for 5 mg means that the estimated odds of HbA1c <7.0% were 4.78 times the corresponding odds under the logistic-regression comparison with insulin glargine. The 95% CI of 3.47 to 6.58 describes uncertainty around that odds-ratio estimate.

The 9.23 estimate for 10 mg and 11.87 estimate for 15 mg are interpreted using the same odds framework. They do not mean that 92.3% or 1187% of participants achieved the target, nor do they directly state the absolute probability of achieving HbA1c <7.0%.

Because the registry supplies odds ratios rather than the underlying event proportions, absolute risk differences and numbers needed to treat cannot be calculated from the ClinicalTrials.gov record without introducing additional information.

12. Secondary Fasting Serum Glucose Results

Change from baseline in fasting serum glucose was assessed at Baseline and Week 52 using a mixed-effects model. The registry reports superiority as the hypothesis type.

ComparisonNet mean difference95% two-sided CIP-value
5 mg Tirzepatide vs Insulin Glargine1.0 mg/dL-3.7 to 5.7 mg/dL0.672
10 mg Tirzepatide vs Insulin Glargine-3.6 mg/dL-8.2 to 1.1 mg/dL0.134
15 mg Tirzepatide vs Insulin Glargine-8.0 mg/dL-12.6 to -3.4 mg/dL<0.001

5 mg

The estimate was 1.0 mg/dL, with a 95% CI of -3.7 to 5.7 mg/dL and P = 0.672.

10 mg

The estimate was -3.6 mg/dL, with a 95% CI of -8.2 to 1.1 mg/dL and P = 0.134.

15 mg

The estimate was -8.0 mg/dL, with a 95% CI of -12.6 to -3.4 mg/dL and P < 0.001.

Why the CI matters

The confidence interval communicates the range of values compatible with the statistical estimate under the model; it provides information that a p-value alone cannot provide.

The 5 mg and 10 mg confidence intervals include zero, while the 15 mg confidence interval does not. That distinction corresponds to the reported p-values, but the confidence intervals additionally communicate the direction and precision of the estimated differences.

13. Statistical Methodology

Mixed-effects model

The registry reports a mixed-effects model for the continuous longitudinal endpoints, including change from baseline in HbA1c, body weight, and fasting serum glucose. For the primary HbA1c endpoint, the registry definition specifically states that the least-squares mean was determined by a mixed-model repeated-measures (MMRM) model for post-baseline measures.

Conceptual longitudinal model
Outcome at follow-up = baseline information + treatment information + other model terms + within-participant correlation

A repeated-measures framework can use the information contributed by multiple post-baseline observations while accounting for the fact that measurements from the same participant are correlated.

The registry-reported primary-endpoint definition identifies baseline, pooled country, baseline sodium-glucose co-transporter-2 inhibitor (SGLT-2i) use flag, and treatment among the model terms before the text becomes truncated in the registry extract. This page therefore does not reconstruct any additional model terms beyond those explicitly present in the ClinicalTrials.gov record.

Why use a mixed-effects approach?

HbA1c, body weight, and fasting serum glucose are measured repeatedly over time. Treating repeated observations from the same participant as if they were independent would ignore an important feature of the data. A mixed-effects or repeated-measures model is designed to account for within-participant dependence while estimating adjusted between-group differences.

Least-squares means

The primary registry definition states that the least-squares mean was determined by the MMRM. A least-squares mean is a model-adjusted mean rather than simply the arithmetic average of observed values. It represents the model-based expected outcome after accounting for the covariates and treatment terms included in the specified model.

Logistic regression

The Week 52 endpoint "Percentage of Participants With HbA1c of <7.0%" is binary: each participant either meets the specified threshold or does not. The registry reports logistic regression and an odds ratio for these comparisons.

Logistic model
logit[P(Y=1)] = β₀ + β₁(Treatment) + …

Exponentiating a treatment coefficient gives an odds ratio. An OR of 1 represents equal odds between the comparison groups; an OR above 1 indicates higher estimated odds in the numerator group.

Mean difference

The reported mean difference is a contrast between model-estimated treatment-group means or changes. A negative value indicates that the treatment-group estimate is lower than the comparator estimate under the reported subtraction convention.

Confidence intervals

The primary HbA1c analyses use 97.5% two-sided confidence intervals, while the reported secondary analyses use 95% two-sided confidence intervals. These different confidence levels should not be silently treated as interchangeable. The width of an interval reflects both statistical precision and the chosen confidence level.

14. Non-Inferiority and Superiority Logic

The ClinicalTrials.gov record identifies the primary 10 mg and 15 mg HbA1c comparisons as non-inferiority analyses. The registry comment states: "Although the primary objective is to establish noninferiority, sample size selection is guided by the objective of establishing superiority." It also states that the chosen sample size and randomization ratio provides greater than 90% power to establish superiority of the 10 mg and 15 mg doses to insulin glargine.

Non-inferiority question

Is the treatment effect sufficiently favorable that it does not fall beyond the prespecified non-inferiority margin relative to the comparator?

Superiority question

Is the treatment effect statistically distinguishable from the comparator in the favorable direction under the prespecified superiority test?

Margin not reported: The ClinicalTrials.gov record does not contain the numerical non-inferiority margin. Because that margin determines the formal non-inferiority decision boundary, it would be inappropriate to infer or manufacture a value from the observed confidence interval.

15. Multiplicity

The ClinicalTrials.gov record contains multiple comparisons: the two primary HbA1c dose comparisons, a secondary HbA1c comparison, three body-weight comparisons, three HbA1c-target comparisons, and three fasting-glucose comparisons. The profile identifies the overall hypothesis types as non-inferiority and superiority, but it does not provide a complete multiplicity-adjustment strategy for all posted analyses.

Analysis familyComparisons reportedHypothesis type
Primary HbA1c10 mg vs insulin glargine; 15 mg vs insulin glargineNon-inferiority
Secondary HbA1c5 mg vs insulin glargineNon-inferiority
Body weight5 mg, 10 mg, 15 mg vs insulin glargineSuperiority
HbA1c <7.0%5 mg, 10 mg, 15 mg vs insulin glargineSuperiority
Fasting serum glucose5 mg, 10 mg, 15 mg vs insulin glargineSuperiority

When multiple hypotheses are tested, the interpretation of individual p-values depends on the prespecified testing hierarchy and type-I-error control. The ClinicalTrials.gov record does not provide enough detail to reconstruct a complete alpha-allocation or gatekeeping strategy, so this page does not assign one retrospectively.

16. Interim Analysis, Crossover, and Bayesian Methods

Interim analysis

The ClinicalTrials.gov record does not provide an interim-analysis strategy or alpha-spending description. No such methodology is inferred here.

Crossover

The ClinicalTrials.gov record does not report a crossover design or crossover results. No crossover effect is assumed.

Bayesian methods

No Bayesian method is listed among the normalized statistical methods. The reported analyses are frequentist mixed-effects and logistic-regression analyses.

Factorial design

The registry identifies the design model as PARALLEL rather than a factorial design.

These omissions are analytically important. A trial-results page should distinguish between a method that is actually documented and a method that might be common in another trial. The registry-reported SURPASS-4 data support discussion of mixed-effects modeling, logistic regression, confidence intervals, odds ratios, non-inferiority, superiority, and randomization; they do not support adding an interim-monitoring, crossover, factorial, or Bayesian analysis.

17. Missing Data and Longitudinal Interpretation

The analysis population requires a baseline value and at least one post-baseline value. That criterion means participants without the required measurement history are not represented in the reported analysis population.

What the ClinicalTrials.gov record does not establish: The registry extract does not state a specific missing-data imputation method, such as multiple imputation, last observation carried forward, or a particular missing-at-random sensitivity analysis. The use of an MMRM does not justify assigning a specific imputation method to the trial. Accordingly, no additional imputation strategy is claimed here.

This distinction matters because a longitudinal mixed model and an imputation procedure are not synonymous. An MMRM can estimate treatment effects using available longitudinal observations under its statistical assumptions without requiring the analyst to fill every missing value with a single deterministic replacement. The actual assumptions and sensitivity analyses must be taken from the trial's complete statistical analysis plan if they are to be described.

18. Secondary Results: Statistical Summary

EndpointComparisonEstimateCIP-valueMethod
HbA1c change5 mg vs insulin glargine-0.8097.5% CI -0.93 to -0.66<0.001Mixed-effects model
Body weight change5 mg vs insulin glargine-9.0 kg95% CI -9.8 to -8.3<0.001Mixed-effects model
Body weight change10 mg vs insulin glargine-11.4 kg95% CI -12.1 to -10.6<0.001Mixed-effects model
Body weight change15 mg vs insulin glargine-13.5 kg95% CI -14.3 to -12.8<0.001Mixed-effects model
HbA1c <7.0%5 mg vs insulin glargineOR 4.7895% CI 3.47 to 6.58<0.001Logistic regression
HbA1c <7.0%10 mg vs insulin glargineOR 9.2395% CI 6.31 to 13.49<0.001Logistic regression
HbA1c <7.0%15 mg vs insulin glargineOR 11.8795% CI 7.88 to 17.89<0.001Logistic regression
Fasting serum glucose change5 mg vs insulin glargine1.0 mg/dL95% CI -3.7 to 5.70.672Mixed-effects model
Fasting serum glucose change10 mg vs insulin glargine-3.6 mg/dL95% CI -8.2 to 1.10.134Mixed-effects model
Fasting serum glucose change15 mg vs insulin glargine-8.0 mg/dL95% CI -12.6 to -3.4<0.001Mixed-effects model

19. Safety Results

The ClinicalTrials.gov record provides serious adverse-event counts by treatment arm. These are reported as affected participants divided by participants at risk.

Treatment armSerious adverse eventsAffected / at risk
5 mg Tirzepatide4848 / 329
10 mg Tirzepatide5454 / 328
15 mg Tirzepatide4141 / 338
Insulin Glargine193193 / 1000

The affected/at-risk figures should be read as the registry reports them rather than converted into newly calculated percentages. The denominators differ across arms, so the raw event counts alone are not directly comparable measures of event frequency.

Safety interpretation: These serious-adverse-event figures describe the number affected and number at risk reported in the ClinicalTrials.gov record. They do not establish causality for individual events, and the ClinicalTrials.gov record does not provide the detailed event categories, exposure duration, severity distribution, or treatment-relatedness needed for a fuller safety analysis.

20. Statistical Methods Explained

Why was an MMRM or mixed-effects model used for HbA1c?

HbA1c is a longitudinal measurement. Participants can contribute information at multiple post-baseline visits, and observations from the same participant are correlated. A mixed-model repeated-measures framework is designed to estimate treatment differences while accounting for that within-participant structure. The registry specifically identifies MMRM as the method used to determine the least-squares mean for post-baseline measures.

What does a mean difference of -1.14 mean?

It means that the model-estimated change from baseline HbA1c for the 15 mg tirzepatide comparison was 1.14 percentage points lower than the corresponding insulin-glargine estimate under the reported subtraction convention. It does not mean every participant experienced exactly that difference.

Why is a confidence interval more informative than a p-value alone?

The p-value describes evidence against a null hypothesis under the specified statistical model. A confidence interval adds information about the estimated effect's precision and direction. For example, the 15 mg HbA1c analysis reports a net mean difference of -1.14 with a 97.5% two-sided CI of -1.28 to -1.00. The interval communicates substantially more about the estimated magnitude than P < 0.001 alone.

What does an odds ratio of 11.87 mean?

An odds ratio of 11.87 means the estimated odds of having HbA1c <7.0% were 11.87 times the corresponding odds for the insulin-glargine comparison group. Odds are not probabilities, so the OR cannot be read as an 11.87-fold increase in the percentage of participants achieving the endpoint.

Why does non-inferiority use a margin rather than only a p-value?

Non-inferiority asks whether the treatment effect remains within a prespecified clinically acceptable loss relative to the comparator. That requires a margin. A small p-value against equality does not by itself establish non-inferiority, because the non-inferiority question is directional and margin-based. The registry-reported SURPASS-4 data identify the non-inferiority hypothesis but do not provide the numerical margin.

Why shouldn't the three tirzepatide doses be treated as a formal dose-response test?

The reported analyses compare each tirzepatide dose with insulin glargine. The fact that estimates differ numerically across 5 mg, 10 mg, and 15 mg does not itself constitute a formal test of dose-response. Such a conclusion would require an explicitly specified dose-response analysis or a direct statistical comparison among doses.

What does randomization contribute statistically?

Randomization creates the framework for comparing outcomes between treatment assignments while reducing systematic allocation differences in expectation. It does not guarantee that every baseline characteristic or post-baseline measurement will be identical between groups. The validity of the treatment comparison also depends on adherence to the prespecified analysis and appropriate handling of follow-up information.

21. Confidence Intervals, Estimates, and P-values

Estimate

The estimate is the point summary of the treatment contrast. For continuous endpoints it is a net mean difference; for the HbA1c threshold endpoint it is an odds ratio.

Confidence interval

The interval communicates uncertainty and precision around the model-based estimate. Primary HbA1c analyses use 97.5% two-sided intervals; the registry-reported secondary analyses use 95% two-sided intervals.

P-value

The p-value quantifies statistical evidence against a specified null hypothesis under the model. It is not the probability that the null hypothesis is true.

Clinical magnitude

Statistical significance and clinical importance are different concepts. The size of the estimate, its confidence interval, the endpoint definition, and the context of the comparison all matter.

22. What These Analyses Do — and Do Not — Establish

Continuous outcomes

The reported HbA1c, body-weight, and fasting-serum-glucose analyses estimate between-group differences in changes from baseline. These are population-level model estimates rather than individual treatment responses.

Binary outcome

The HbA1c <7.0% analyses use logistic regression and odds ratios. The OR describes relative odds, not absolute percentages. Without the underlying event counts or probabilities in the ClinicalTrials.gov record, absolute risk differences cannot be derived without adding information.

Non-inferiority

The primary HbA1c comparisons are labeled non-inferiority analyses. Formal non-inferiority interpretation requires the prespecified margin. The ClinicalTrials.gov record does not include that numerical margin, so this page reports the observed estimate, confidence interval, and p-value without reconstructing a margin-based decision.

Multiple comparisons

The registry contains several dose-versus-comparator analyses across multiple endpoints. Individual p-values should be interpreted in the context of the trial's prespecified hypothesis hierarchy and type-I-error strategy. The registry-reported extract does not contain enough information to reconstruct that complete strategy.

23. Important Limitations and Interpretation Issues

24. Why This Trial Matters Statistically

SURPASS-4 is a useful statistical teaching case because it combines randomized parallel-group design with two distinct classes of outcome analysis. Continuous longitudinal outcomes are handled with mixed-effects modeling, while a clinically defined binary HbA1c threshold is analyzed with logistic regression. The primary HbA1c endpoint additionally introduces the distinction between non-inferiority and superiority hypotheses.

ConceptHow it appears in SURPASS-4
RandomizationRandomized parallel-group phase 3 design.
Multiple treatment doses5 mg, 10 mg, and 15 mg tirzepatide compared with insulin glargine.
Longitudinal modelingMixed-effects model / MMRM for change from baseline endpoints.
Least-squares meansRegistry states that LS means were determined by MMRM for post-baseline measures.
Mean differencePrimary HbA1c, body weight, and fasting serum glucose treatment contrasts.
Logistic regressionWeek 52 HbA1c <7.0% endpoint.
Odds ratioEffect measure for the binary HbA1c threshold endpoint.
Confidence intervals97.5% two-sided intervals for primary HbA1c analyses and 95% two-sided intervals for registry-reported secondary analyses.
Non-inferiorityPrimary 10 mg and 15 mg HbA1c comparisons.
SuperiorityReported secondary body-weight, HbA1c-target, and fasting-glucose analyses.
MultiplicityMultiple dose-versus-control comparisons across several endpoints.
Safety populationsSerious adverse events are reported as affected/at-risk counts by treatment arm.

25. A Statistical Reading of the Complete Results

The strongest statistical feature of the registry-reported SURPASS-4 results is not any single p-value. It is the combination of an explicitly randomized comparison, prespecified Week 52 endpoints, model-based estimates, confidence intervals, and different methods matched to different outcome types.

For the primary HbA1c endpoint, both the 10 mg and 15 mg comparisons have negative net mean differences with 97.5% two-sided confidence intervals entirely below zero and P < 0.001. The registry labels these analyses as non-inferiority analyses and states that superiority was also an objective guiding sample-size selection. Because the numerical non-inferiority margin is absent from the ClinicalTrials.gov record, the most defensible interpretation is to report the estimates and uncertainty without retroactively supplying a decision boundary.

The secondary results illustrate why the choice of effect measure matters. Body weight is summarized as a mean difference, while achieving HbA1c <7.0% is summarized as an odds ratio. These quantities answer different statistical questions. A mean difference compares average changes, whereas an odds ratio compares the odds of a binary outcome.

The fasting-serum-glucose results also demonstrate the importance of looking beyond statistical significance. The 5 mg estimate is 1.0 mg/dL with a 95% CI from -3.7 to 5.7 mg/dL and P = 0.672. The 10 mg estimate is -3.6 mg/dL with a 95% CI from -8.2 to 1.1 mg/dL and P = 0.134. The 15 mg estimate is -8.0 mg/dL with a 95% CI from -12.6 to -3.4 mg/dL and P < 0.001. Presenting all three estimates and intervals gives substantially more information than simply classifying the p-values as significant or nonsignificant.

Finally, the serious-adverse-event data show why safety should be analyzed separately from efficacy. The registry-reported counts are 48/329, 54/328, 41/338, and 193/1000 across the four treatment arms. Without additional information on event types, exposure time, severity, and relatedness, those counts should remain descriptive rather than being converted into a broader safety conclusion.

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Calculators

28. Limitations of the Statistical Record

The ClinicalTrials.gov record is sufficiently detailed to reconstruct the principal statistical story of SURPASS-4, but they do not constitute the complete statistical analysis plan. Several details that would normally be important for a full statistical review are not present in the registry-reported extract, including the numerical non-inferiority margin, the complete multiplicity hierarchy, detailed missing-data assumptions and sensitivity analyses, and any interim-analysis framework.

Similarly, the ClinicalTrials.gov record is limited to the 12 statistical analyses included in the trial data. No additional baseline characteristics, subgroup estimates, long-term follow-up results, or unreported outcomes have been imported from external publications. This is deliberate: a reproducible trial-results page should distinguish what is contained in the specified registry record from what might be available elsewhere.

Reproducibility principle: An educational statistical analysis should not silently fill gaps in a registry extract with remembered results from a publication. When a design feature or numerical parameter is not reported, the appropriate response is to identify the limitation rather than manufacture precision.

29. Sources

The numerical results and trial-design facts presented on this page are restricted to the registry-reported SURPASS-4 trial data. The linked publications are provided as publication references identified in that data; no additional numerical results from those publications are incorporated into this page.

Continue through the Clinical Biostats statistical library

Explore the statistical concepts behind randomized trials, longitudinal models, binary outcomes, confidence intervals, and non-inferiority analysis.

30. Record Summary

SURPASS-4 is a randomized phase 3 parallel-group trial with 2002 enrolled participants evaluating once-weekly tirzepatide versus once-daily insulin glargine in participants with type 2 diabetes mellitus and increased cardiovascular risk. The ClinicalTrials.gov record reports mixed-effects analyses for longitudinal continuous outcomes and logistic regression for the binary endpoint of HbA1c <7.0% at Week 52.

The two primary HbA1c analyses report net mean differences of -0.99 for 10 mg tirzepatide versus insulin glargine and -1.14 for 15 mg versus insulin glargine, each with a 97.5% two-sided confidence interval and P < 0.001. The registry identifies these analyses as non-inferiority hypotheses and states that sample-size selection was guided by a superiority objective, while the ClinicalTrials.gov record does not provide the numerical non-inferiority margin.

The secondary analyses extend the statistical picture: body-weight changes are reported as mean differences, achievement of HbA1c <7.0% is reported using odds ratios, and fasting serum glucose is reported using mean differences. Together, these results demonstrate why trial interpretation requires attention to the endpoint definition, effect measure, confidence interval, hypothesis type, analysis population, and multiplicity rather than relying on p-values alone.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical story from the documented evidence while clearly separating reported results from educational interpretation. For SURPASS-4, the central statistical lessons are mixed-effects modeling for longitudinal outcomes, logistic regression for binary outcomes, interpretation of mean differences and odds ratios, confidence-interval reasoning, and the margin-based logic of non-inferiority trials.