This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses posted in the ClinicalTrials.gov record. The registry provides the official trial record.
1. Trial at a Glance
SURPASS-4 was a randomized, parallel, unmasked phase 3 treatment trial comparing three doses of once-weekly tirzepatide with once-daily insulin glargine in participants with type 2 diabetes mellitus and increased cardiovascular risk. The ClinicalTrials.gov record contains two formal primary analyses for the Week 52 HbA1c endpoint and additional secondary analyses for body weight, HbA1c target attainment, and fasting serum glucose.
| Feature | SURPASS-4 |
|---|---|
| Trial name | SURPASS-4 |
| ClinicalTrials.gov identifier | NCT03730662 |
| Phase | Phase 3 |
| Status | COMPLETED |
| Therapeutic area | Endocrinology |
| Condition | Type 2 Diabetes Mellitus |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | NONE |
| Primary purpose | TREATMENT |
| Enrollment | 2002 |
| Interventions | Tirzepatide; Insulin Glargine |
| Lead sponsor | Eli Lilly and Company |
| Sponsor type | INDUSTRY |
| Start | November 20, 2018 |
| Primary completion | January 22, 2021 |
2. Clinical Question
The clinical question was whether once-weekly tirzepatide, at the evaluated 10 mg and 15 mg doses for the primary endpoint and at 5 mg for a secondary endpoint, could produce a favorable change in HbA1c compared with once-daily insulin glargine at Week 52 in participants with type 2 diabetes mellitus.
Population
Participants with type 2 diabetes mellitus and increased cardiovascular risk, as described by the trial's brief title.
Intervention
Once-weekly tirzepatide at 5 mg, 10 mg, or 15 mg.
Comparator
Once-daily insulin glargine.
Primary question
For the registered Week 52 HbA1c endpoint, do the 10 mg and 15 mg tirzepatide comparisons establish non-inferiority relative to insulin glargine, with the design also guided by a superiority objective?
3. Trial Design
Tirzepatide 5 mg
- Once-weekly tirzepatide.
- Secondary analyses included HbA1c, body weight, percentage with HbA1c <7.0%, and fasting serum glucose.
Tirzepatide 10 mg
- Once-weekly tirzepatide.
- Week 52 change in HbA1c was a primary analysis.
Tirzepatide 15 mg
- Once-weekly tirzepatide.
- Week 52 change in HbA1c was a primary analysis.
Insulin glargine
- Once-daily insulin glargine.
- Comparator for the three tirzepatide dose groups.
4. Trial Timing and Registry Results
Trial start
The registry profile lists November 20, 2018 as the study start date.
Primary completion
The registry profile lists January 22, 2021 as the primary completion date.
Results posted
The trial is listed as completed, with 8 outcome measures and 12 statistical analyses posted in the ClinicalTrials.gov record.
5. Endpoints
The registered primary endpoint was change from baseline in hemoglobin A1c (HbA1c) at Week 52 for the 10 mg and 15 mg comparisons. The registry definition posted on ClinicalTrials.gov for the endpoint states that HbA1c is the glycosylated fraction of hemoglobin A and is measured primarily to identify average plasma glucose concentration over prolonged periods of time. It further states that the least-squares mean was determined by a mixed-model repeated-measures model for post-baseline measures.
| Endpoint | Time frame | Analysis | Effect measure |
|---|---|---|---|
| Change From Baseline in Hemoglobin A1c (HbA1c) (10 mg and 15 mg) | Baseline, Week 52 | Mixed Models Analysis | Mean Difference (Net) |
| Change From Baseline in HbA1c (5 mg) | Baseline, Week 52 | Mixed Models Analysis | Mean Difference (Net) |
| Change From Baseline in Body Weight | Baseline, Week 52 | Mixed Models Analysis | Mean Difference (Net) |
| Percentage of Participants With HbA1c of <7.0% | Week 52 | Regression, Logistic | Odds Ratio (OR) |
| Change From Baseline in Fasting Serum Glucose | Baseline, Week 52 | Mixed Models Analysis | Mean Difference (Net) |
6. Analysis Population
The statistical analyses posted on ClinicalTrials.gov use the same general analysis-population definition: all randomized participants who received at least one dose of study drug and had a baseline and at least 1 post-baseline value, excluding participants discontinuing study drug due to inadvertent enrolment.
| Feature | Registry-specified analysis population |
|---|---|
| Randomization | All randomized participants meeting the subsequent analysis criteria. |
| Treatment exposure | Participants must have received at least one dose of study drug. |
| Baseline measurement | A baseline value was required. |
| Post-baseline information | At least 1 post-baseline value was required. |
| Exclusion | Participants discontinuing study drug due to inadvertent enrolment were excluded. |
This population definition is important when interpreting the estimates. The reported analyses are not simply a comparison of every enrolled participant regardless of whether they received treatment or registry-reported post-baseline measurements. The analysis population is explicitly conditioned on treatment exposure and availability of baseline and post-baseline data.
7. Primary Results: HbA1c at Week 52
The registry contains two formal primary statistical analyses for the registered endpoint "Change From Baseline in Hemoglobin A1c (HbA1c) (10 mg and 15 mg)" at Baseline and Week 52. Both use a mixed-effects model and report a net mean difference relative to insulin glargine. The hypothesis type recorded for both comparisons is non-inferiority.
10 mg Tirzepatide vs Insulin Glargine
Net mean difference in change from baseline HbA1c
97.5% two-sided CI: -1.13 to -0.86 · P < 0.001
Time frame: Baseline, Week 52 · Hypothesis type: Non-inferiority
| Measure | Reported value |
|---|---|
| Comparison | 10 mg Tirzepatide vs Insulin Glargine |
| Effect measure | Mean Difference (Net) |
| Estimate | -0.99 |
| Confidence interval | 97.5% two-sided CI -1.13 to -0.86 |
| P-value | <0.001 |
| Method | Mixed Models Analysis / mixed-effects model |
| Time frame | Baseline, Week 52 |
The estimated net mean difference of -0.99 means that the model-estimated change from baseline in HbA1c was 0.99 percentage points lower for the 10 mg tirzepatide comparison than for insulin glargine, based on the registry's reported mean-difference convention.
The negative value describes the direction of the between-group difference; it does not mean that every participant experienced a 0.99-percentage-point reduction, nor does it describe the individual treatment response of any particular participant.
The 97.5% two-sided confidence interval of -1.13 to -0.86 describes statistical uncertainty around the estimated between-group difference under the specified analysis. It is not a range containing 97.5% of individual patient responses.
The P < 0.001 result addresses evidence against the statistical null hypothesis under the analysis framework. A p-value does not measure the size, clinical importance, or probability of the observed effect, and it should not be interpreted as the probability that the treatment works.
The registry identifies this comparison as a non-inferiority analysis and states that sample-size selection was guided by the objective of establishing superiority. The ClinicalTrials.gov record does not provide a numerical non-inferiority margin, so a margin-specific conclusion cannot be reconstructed here. In a non-inferiority analysis, the key question is whether the confidence interval is compatible with an effect that remains within the prespecified acceptable margin, rather than whether a p-value alone is small.
15 mg Tirzepatide vs Insulin Glargine
Net mean difference in change from baseline HbA1c
97.5% two-sided CI: -1.28 to -1.00 · P < 0.001
Time frame: Baseline, Week 52 · Hypothesis type: Non-inferiority
| Measure | Reported value |
|---|---|
| Comparison | 15 mg Tirzepatide vs Insulin Glargine |
| Effect measure | Mean Difference (Net) |
| Estimate | -1.14 |
| Confidence interval | 97.5% two-sided CI -1.28 to -1.00 |
| P-value | <0.001 |
| Method | Mixed Models Analysis / mixed-effects model |
| Time frame | Baseline, Week 52 |
The estimated net mean difference of -1.14 indicates a model-estimated change from baseline in HbA1c that was 1.14 percentage points lower for the 15 mg tirzepatide comparison than for insulin glargine, using the registry's reported effect-measure convention.
The estimate is a between-group mean difference. It does not imply that each individual participant experienced a 1.14-percentage-point difference, and it does not provide an individual-level probability of benefit.
The 97.5% two-sided CI of -1.28 to -1.00 quantifies uncertainty around the model-based estimate. The relatively narrow interval provides information about the precision of this particular estimated comparison, but it does not describe individual variability in HbA1c response.
The P < 0.001 indicates strong statistical evidence under the reported hypothesis-testing framework. It is not an effect-size measure. The magnitude of the estimate and its confidence interval are required to understand the size and precision of the difference.
For the non-inferiority objective, interpretation should be tied to the prespecified non-inferiority margin. Because the ClinicalTrials.gov record does not state that numerical margin, this page does not manufacture one or perform a reconstructed margin calculation.
8. Primary HbA1c Results Side by Side
| Comparison | Net mean difference | 97.5% two-sided CI | P-value | Hypothesis |
|---|---|---|---|---|
| 10 mg Tirzepatide vs Insulin Glargine | -0.99 | -1.13 to -0.86 | <0.001 | Non-inferiority |
| 15 mg Tirzepatide vs Insulin Glargine | -1.14 | -1.28 to -1.00 | <0.001 | Non-inferiority |
The two primary estimates are both negative, with the 15 mg comparison showing a larger reported difference than the 10 mg comparison. That comparison is descriptive: the two estimates come from separate dose-versus-control analyses and should not automatically be interpreted as a formal statistical comparison of 10 mg versus 15 mg.
9. Secondary HbA1c Result: 5 mg Tirzepatide
The 5 mg tirzepatide comparison was recorded as a secondary endpoint rather than a primary endpoint. It used the same general mixed-model approach and the same analysis-population definition.
Net mean difference in change from baseline HbA1c
97.5% two-sided CI: -0.93 to -0.66 · P < 0.001
5 mg Tirzepatide vs Insulin Glargine · Hypothesis type: Non-inferiority
The reported estimate of -0.80 is the net mean difference in change from baseline HbA1c for 5 mg tirzepatide versus insulin glargine. The negative direction indicates a lower model-estimated change for the tirzepatide group under the registry's effect-measure convention.
The 97.5% CI of -0.93 to -0.66 describes uncertainty around that estimate. Because the confidence interval is based on the statistical model, it should not be read as an interval containing individual participants' HbA1c changes.
The P < 0.001 is evidence from the specified statistical test and is not itself a measure of treatment magnitude. The estimate and confidence interval provide the corresponding information about the size and precision of the difference.
10. Secondary Body-Weight Results
Change from baseline in body weight was evaluated at Baseline and Week 52 using a mixed-effects model. The registry reports superiority as the hypothesis type for all three dose comparisons.
| Comparison | Net mean difference | 95% two-sided CI | P-value |
|---|---|---|---|
| 5 mg Tirzepatide vs Insulin Glargine | -9.0 kg | -9.8 to -8.3 kg | <0.001 |
| 10 mg Tirzepatide vs Insulin Glargine | -11.4 kg | -12.1 to -10.6 kg | <0.001 |
| 15 mg Tirzepatide vs Insulin Glargine | -13.5 kg | -14.3 to -12.8 kg | <0.001 |
5 mg comparison
The reported net mean difference was -9.0 kg, with a 95% two-sided CI of -9.8 to -8.3 kg and P < 0.001.
10 mg comparison
The reported net mean difference was -11.4 kg, with a 95% two-sided CI of -12.1 to -10.6 kg and P < 0.001.
15 mg comparison
The reported net mean difference was -13.5 kg, with a 95% two-sided CI of -14.3 to -12.8 kg and P < 0.001.
Statistical role
These are secondary superiority analyses. The registry does not supply additional information here about a separate dose-response hypothesis test.
The pattern of estimates is descriptively more negative at the higher tirzepatide doses. However, comparing the numerical estimates across doses is not the same as conducting a formal dose-response test. The statistical analyses posted on ClinicalTrials.gov contain dose-versus-insulin-glargine comparisons, not a separate formal test of one tirzepatide dose against another.
11. Secondary HbA1c Target-Attainment Results
The registry also reports the percentage of participants with HbA1c of <7.0% at Week 52. These binary outcomes were analyzed using logistic regression, with odds ratio as the effect measure and superiority as the hypothesis type.
| Comparison | Odds ratio | 95% two-sided CI | P-value |
|---|---|---|---|
| 5 mg Tirzepatide vs Insulin Glargine | 4.78 | 3.47 to 6.58 | <0.001 |
| 10 mg Tirzepatide vs Insulin Glargine | 9.23 | 6.31 to 13.49 | <0.001 |
| 15 mg Tirzepatide vs Insulin Glargine | 11.87 | 7.88 to 17.89 | <0.001 |
An odds ratio above 1 indicates higher estimated odds of the specified binary outcome in the tirzepatide comparison group. The odds ratio is not itself a risk ratio or a percentage-point difference in the probability of achieving HbA1c <7.0%.
The reported odds ratio of 4.78 for 5 mg means that the estimated odds of HbA1c <7.0% were 4.78 times the corresponding odds under the logistic-regression comparison with insulin glargine. The 95% CI of 3.47 to 6.58 describes uncertainty around that odds-ratio estimate.
The 9.23 estimate for 10 mg and 11.87 estimate for 15 mg are interpreted using the same odds framework. They do not mean that 92.3% or 1187% of participants achieved the target, nor do they directly state the absolute probability of achieving HbA1c <7.0%.
Because the registry supplies odds ratios rather than the underlying event proportions, absolute risk differences and numbers needed to treat cannot be calculated from the ClinicalTrials.gov record without introducing additional information.
12. Secondary Fasting Serum Glucose Results
Change from baseline in fasting serum glucose was assessed at Baseline and Week 52 using a mixed-effects model. The registry reports superiority as the hypothesis type.
| Comparison | Net mean difference | 95% two-sided CI | P-value |
|---|---|---|---|
| 5 mg Tirzepatide vs Insulin Glargine | 1.0 mg/dL | -3.7 to 5.7 mg/dL | 0.672 |
| 10 mg Tirzepatide vs Insulin Glargine | -3.6 mg/dL | -8.2 to 1.1 mg/dL | 0.134 |
| 15 mg Tirzepatide vs Insulin Glargine | -8.0 mg/dL | -12.6 to -3.4 mg/dL | <0.001 |
5 mg
The estimate was 1.0 mg/dL, with a 95% CI of -3.7 to 5.7 mg/dL and P = 0.672.
10 mg
The estimate was -3.6 mg/dL, with a 95% CI of -8.2 to 1.1 mg/dL and P = 0.134.
15 mg
The estimate was -8.0 mg/dL, with a 95% CI of -12.6 to -3.4 mg/dL and P < 0.001.
Why the CI matters
The confidence interval communicates the range of values compatible with the statistical estimate under the model; it provides information that a p-value alone cannot provide.
The 5 mg and 10 mg confidence intervals include zero, while the 15 mg confidence interval does not. That distinction corresponds to the reported p-values, but the confidence intervals additionally communicate the direction and precision of the estimated differences.
13. Statistical Methodology
Mixed-effects model
The registry reports a mixed-effects model for the continuous longitudinal endpoints, including change from baseline in HbA1c, body weight, and fasting serum glucose. For the primary HbA1c endpoint, the registry definition specifically states that the least-squares mean was determined by a mixed-model repeated-measures (MMRM) model for post-baseline measures.
A repeated-measures framework can use the information contributed by multiple post-baseline observations while accounting for the fact that measurements from the same participant are correlated.
The registry-reported primary-endpoint definition identifies baseline, pooled country, baseline sodium-glucose co-transporter-2 inhibitor (SGLT-2i) use flag, and treatment among the model terms before the text becomes truncated in the registry extract. This page therefore does not reconstruct any additional model terms beyond those explicitly present in the ClinicalTrials.gov record.
Why use a mixed-effects approach?
HbA1c, body weight, and fasting serum glucose are measured repeatedly over time. Treating repeated observations from the same participant as if they were independent would ignore an important feature of the data. A mixed-effects or repeated-measures model is designed to account for within-participant dependence while estimating adjusted between-group differences.
Least-squares means
The primary registry definition states that the least-squares mean was determined by the MMRM. A least-squares mean is a model-adjusted mean rather than simply the arithmetic average of observed values. It represents the model-based expected outcome after accounting for the covariates and treatment terms included in the specified model.
Logistic regression
The Week 52 endpoint "Percentage of Participants With HbA1c of <7.0%" is binary: each participant either meets the specified threshold or does not. The registry reports logistic regression and an odds ratio for these comparisons.
Exponentiating a treatment coefficient gives an odds ratio. An OR of 1 represents equal odds between the comparison groups; an OR above 1 indicates higher estimated odds in the numerator group.
Mean difference
The reported mean difference is a contrast between model-estimated treatment-group means or changes. A negative value indicates that the treatment-group estimate is lower than the comparator estimate under the reported subtraction convention.
Confidence intervals
The primary HbA1c analyses use 97.5% two-sided confidence intervals, while the reported secondary analyses use 95% two-sided confidence intervals. These different confidence levels should not be silently treated as interchangeable. The width of an interval reflects both statistical precision and the chosen confidence level.
14. Non-Inferiority and Superiority Logic
The ClinicalTrials.gov record identifies the primary 10 mg and 15 mg HbA1c comparisons as non-inferiority analyses. The registry comment states: "Although the primary objective is to establish noninferiority, sample size selection is guided by the objective of establishing superiority." It also states that the chosen sample size and randomization ratio provides greater than 90% power to establish superiority of the 10 mg and 15 mg doses to insulin glargine.
Non-inferiority question
Is the treatment effect sufficiently favorable that it does not fall beyond the prespecified non-inferiority margin relative to the comparator?
Superiority question
Is the treatment effect statistically distinguishable from the comparator in the favorable direction under the prespecified superiority test?
15. Multiplicity
The ClinicalTrials.gov record contains multiple comparisons: the two primary HbA1c dose comparisons, a secondary HbA1c comparison, three body-weight comparisons, three HbA1c-target comparisons, and three fasting-glucose comparisons. The profile identifies the overall hypothesis types as non-inferiority and superiority, but it does not provide a complete multiplicity-adjustment strategy for all posted analyses.
| Analysis family | Comparisons reported | Hypothesis type |
|---|---|---|
| Primary HbA1c | 10 mg vs insulin glargine; 15 mg vs insulin glargine | Non-inferiority |
| Secondary HbA1c | 5 mg vs insulin glargine | Non-inferiority |
| Body weight | 5 mg, 10 mg, 15 mg vs insulin glargine | Superiority |
| HbA1c <7.0% | 5 mg, 10 mg, 15 mg vs insulin glargine | Superiority |
| Fasting serum glucose | 5 mg, 10 mg, 15 mg vs insulin glargine | Superiority |
When multiple hypotheses are tested, the interpretation of individual p-values depends on the prespecified testing hierarchy and type-I-error control. The ClinicalTrials.gov record does not provide enough detail to reconstruct a complete alpha-allocation or gatekeeping strategy, so this page does not assign one retrospectively.
16. Interim Analysis, Crossover, and Bayesian Methods
Interim analysis
The ClinicalTrials.gov record does not provide an interim-analysis strategy or alpha-spending description. No such methodology is inferred here.
Crossover
The ClinicalTrials.gov record does not report a crossover design or crossover results. No crossover effect is assumed.
Bayesian methods
No Bayesian method is listed among the normalized statistical methods. The reported analyses are frequentist mixed-effects and logistic-regression analyses.
Factorial design
The registry identifies the design model as PARALLEL rather than a factorial design.
These omissions are analytically important. A trial-results page should distinguish between a method that is actually documented and a method that might be common in another trial. The registry-reported SURPASS-4 data support discussion of mixed-effects modeling, logistic regression, confidence intervals, odds ratios, non-inferiority, superiority, and randomization; they do not support adding an interim-monitoring, crossover, factorial, or Bayesian analysis.
17. Missing Data and Longitudinal Interpretation
The analysis population requires a baseline value and at least one post-baseline value. That criterion means participants without the required measurement history are not represented in the reported analysis population.
This distinction matters because a longitudinal mixed model and an imputation procedure are not synonymous. An MMRM can estimate treatment effects using available longitudinal observations under its statistical assumptions without requiring the analyst to fill every missing value with a single deterministic replacement. The actual assumptions and sensitivity analyses must be taken from the trial's complete statistical analysis plan if they are to be described.
18. Secondary Results: Statistical Summary
| Endpoint | Comparison | Estimate | CI | P-value | Method |
|---|---|---|---|---|---|
| HbA1c change | 5 mg vs insulin glargine | -0.80 | 97.5% CI -0.93 to -0.66 | <0.001 | Mixed-effects model |
| Body weight change | 5 mg vs insulin glargine | -9.0 kg | 95% CI -9.8 to -8.3 | <0.001 | Mixed-effects model |
| Body weight change | 10 mg vs insulin glargine | -11.4 kg | 95% CI -12.1 to -10.6 | <0.001 | Mixed-effects model |
| Body weight change | 15 mg vs insulin glargine | -13.5 kg | 95% CI -14.3 to -12.8 | <0.001 | Mixed-effects model |
| HbA1c <7.0% | 5 mg vs insulin glargine | OR 4.78 | 95% CI 3.47 to 6.58 | <0.001 | Logistic regression |
| HbA1c <7.0% | 10 mg vs insulin glargine | OR 9.23 | 95% CI 6.31 to 13.49 | <0.001 | Logistic regression |
| HbA1c <7.0% | 15 mg vs insulin glargine | OR 11.87 | 95% CI 7.88 to 17.89 | <0.001 | Logistic regression |
| Fasting serum glucose change | 5 mg vs insulin glargine | 1.0 mg/dL | 95% CI -3.7 to 5.7 | 0.672 | Mixed-effects model |
| Fasting serum glucose change | 10 mg vs insulin glargine | -3.6 mg/dL | 95% CI -8.2 to 1.1 | 0.134 | Mixed-effects model |
| Fasting serum glucose change | 15 mg vs insulin glargine | -8.0 mg/dL | 95% CI -12.6 to -3.4 | <0.001 | Mixed-effects model |
19. Safety Results
The ClinicalTrials.gov record provides serious adverse-event counts by treatment arm. These are reported as affected participants divided by participants at risk.
| Treatment arm | Serious adverse events | Affected / at risk |
|---|---|---|
| 5 mg Tirzepatide | 48 | 48 / 329 |
| 10 mg Tirzepatide | 54 | 54 / 328 |
| 15 mg Tirzepatide | 41 | 41 / 338 |
| Insulin Glargine | 193 | 193 / 1000 |
The affected/at-risk figures should be read as the registry reports them rather than converted into newly calculated percentages. The denominators differ across arms, so the raw event counts alone are not directly comparable measures of event frequency.
20. Statistical Methods Explained
Why was an MMRM or mixed-effects model used for HbA1c?
HbA1c is a longitudinal measurement. Participants can contribute information at multiple post-baseline visits, and observations from the same participant are correlated. A mixed-model repeated-measures framework is designed to estimate treatment differences while accounting for that within-participant structure. The registry specifically identifies MMRM as the method used to determine the least-squares mean for post-baseline measures.
What does a mean difference of -1.14 mean?
It means that the model-estimated change from baseline HbA1c for the 15 mg tirzepatide comparison was 1.14 percentage points lower than the corresponding insulin-glargine estimate under the reported subtraction convention. It does not mean every participant experienced exactly that difference.
Why is a confidence interval more informative than a p-value alone?
The p-value describes evidence against a null hypothesis under the specified statistical model. A confidence interval adds information about the estimated effect's precision and direction. For example, the 15 mg HbA1c analysis reports a net mean difference of -1.14 with a 97.5% two-sided CI of -1.28 to -1.00. The interval communicates substantially more about the estimated magnitude than P < 0.001 alone.
What does an odds ratio of 11.87 mean?
An odds ratio of 11.87 means the estimated odds of having HbA1c <7.0% were 11.87 times the corresponding odds for the insulin-glargine comparison group. Odds are not probabilities, so the OR cannot be read as an 11.87-fold increase in the percentage of participants achieving the endpoint.
Why does non-inferiority use a margin rather than only a p-value?
Non-inferiority asks whether the treatment effect remains within a prespecified clinically acceptable loss relative to the comparator. That requires a margin. A small p-value against equality does not by itself establish non-inferiority, because the non-inferiority question is directional and margin-based. The registry-reported SURPASS-4 data identify the non-inferiority hypothesis but do not provide the numerical margin.
Why shouldn't the three tirzepatide doses be treated as a formal dose-response test?
The reported analyses compare each tirzepatide dose with insulin glargine. The fact that estimates differ numerically across 5 mg, 10 mg, and 15 mg does not itself constitute a formal test of dose-response. Such a conclusion would require an explicitly specified dose-response analysis or a direct statistical comparison among doses.
What does randomization contribute statistically?
Randomization creates the framework for comparing outcomes between treatment assignments while reducing systematic allocation differences in expectation. It does not guarantee that every baseline characteristic or post-baseline measurement will be identical between groups. The validity of the treatment comparison also depends on adherence to the prespecified analysis and appropriate handling of follow-up information.
21. Confidence Intervals, Estimates, and P-values
Estimate
The estimate is the point summary of the treatment contrast. For continuous endpoints it is a net mean difference; for the HbA1c threshold endpoint it is an odds ratio.
Confidence interval
The interval communicates uncertainty and precision around the model-based estimate. Primary HbA1c analyses use 97.5% two-sided intervals; the registry-reported secondary analyses use 95% two-sided intervals.
P-value
The p-value quantifies statistical evidence against a specified null hypothesis under the model. It is not the probability that the null hypothesis is true.
Clinical magnitude
Statistical significance and clinical importance are different concepts. The size of the estimate, its confidence interval, the endpoint definition, and the context of the comparison all matter.
22. What These Analyses Do — and Do Not — Establish
The reported HbA1c, body-weight, and fasting-serum-glucose analyses estimate between-group differences in changes from baseline. These are population-level model estimates rather than individual treatment responses.
The HbA1c <7.0% analyses use logistic regression and odds ratios. The OR describes relative odds, not absolute percentages. Without the underlying event counts or probabilities in the ClinicalTrials.gov record, absolute risk differences cannot be derived without adding information.
The primary HbA1c comparisons are labeled non-inferiority analyses. Formal non-inferiority interpretation requires the prespecified margin. The ClinicalTrials.gov record does not include that numerical margin, so this page reports the observed estimate, confidence interval, and p-value without reconstructing a margin-based decision.
The registry contains several dose-versus-comparator analyses across multiple endpoints. Individual p-values should be interpreted in the context of the trial's prespecified hypothesis hierarchy and type-I-error strategy. The registry-reported extract does not contain enough information to reconstruct that complete strategy.
23. Important Limitations and Interpretation Issues
- Analysis population: the reported analyses require at least one dose, a baseline value, and at least one post-baseline value, with exclusion of participants discontinuing study drug due to inadvertent enrolment. This differs from simply analyzing every enrolled participant.
- Missing-data strategy: the registry extract does not identify a specific imputation method or detailed missing-data sensitivity analysis.
- Non-inferiority margin: the numerical margin is not present in the ClinicalTrials.gov record, so a complete margin-based non-inferiority reconstruction is not possible.
- Multiplicity: many treatment-versus-control comparisons are reported across doses and endpoints, while the ClinicalTrials.gov record does not provide the full multiplicity-adjustment hierarchy.
- Dose comparisons: different estimates among 5 mg, 10 mg, and 15 mg do not by themselves constitute a formal dose-response test.
- Model dependence: mixed-effects estimates depend on the specified longitudinal model and its assumptions, including assumptions concerning the covariance structure and missing observations.
- Odds-ratio interpretation: odds ratios should not be substituted for risk ratios or absolute risk differences, particularly when the outcome is not rare.
- Registry detail: the registry endpoint definition is partially truncated after the listed model terms. This page does not infer omitted covariates or model specifications.
- Safety detail: the ClinicalTrials.gov record contains serious-adverse-event counts by arm but not the detailed event-level information required for a comprehensive safety analysis.
24. Why This Trial Matters Statistically
SURPASS-4 is a useful statistical teaching case because it combines randomized parallel-group design with two distinct classes of outcome analysis. Continuous longitudinal outcomes are handled with mixed-effects modeling, while a clinically defined binary HbA1c threshold is analyzed with logistic regression. The primary HbA1c endpoint additionally introduces the distinction between non-inferiority and superiority hypotheses.
| Concept | How it appears in SURPASS-4 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design. |
| Multiple treatment doses | 5 mg, 10 mg, and 15 mg tirzepatide compared with insulin glargine. |
| Longitudinal modeling | Mixed-effects model / MMRM for change from baseline endpoints. |
| Least-squares means | Registry states that LS means were determined by MMRM for post-baseline measures. |
| Mean difference | Primary HbA1c, body weight, and fasting serum glucose treatment contrasts. |
| Logistic regression | Week 52 HbA1c <7.0% endpoint. |
| Odds ratio | Effect measure for the binary HbA1c threshold endpoint. |
| Confidence intervals | 97.5% two-sided intervals for primary HbA1c analyses and 95% two-sided intervals for registry-reported secondary analyses. |
| Non-inferiority | Primary 10 mg and 15 mg HbA1c comparisons. |
| Superiority | Reported secondary body-weight, HbA1c-target, and fasting-glucose analyses. |
| Multiplicity | Multiple dose-versus-control comparisons across several endpoints. |
| Safety populations | Serious adverse events are reported as affected/at-risk counts by treatment arm. |
25. A Statistical Reading of the Complete Results
The strongest statistical feature of the registry-reported SURPASS-4 results is not any single p-value. It is the combination of an explicitly randomized comparison, prespecified Week 52 endpoints, model-based estimates, confidence intervals, and different methods matched to different outcome types.
For the primary HbA1c endpoint, both the 10 mg and 15 mg comparisons have negative net mean differences with 97.5% two-sided confidence intervals entirely below zero and P < 0.001. The registry labels these analyses as non-inferiority analyses and states that superiority was also an objective guiding sample-size selection. Because the numerical non-inferiority margin is absent from the ClinicalTrials.gov record, the most defensible interpretation is to report the estimates and uncertainty without retroactively supplying a decision boundary.
The secondary results illustrate why the choice of effect measure matters. Body weight is summarized as a mean difference, while achieving HbA1c <7.0% is summarized as an odds ratio. These quantities answer different statistical questions. A mean difference compares average changes, whereas an odds ratio compares the odds of a binary outcome.
The fasting-serum-glucose results also demonstrate the importance of looking beyond statistical significance. The 5 mg estimate is 1.0 mg/dL with a 95% CI from -3.7 to 5.7 mg/dL and P = 0.672. The 10 mg estimate is -3.6 mg/dL with a 95% CI from -8.2 to 1.1 mg/dL and P = 0.134. The 15 mg estimate is -8.0 mg/dL with a 95% CI from -12.6 to -3.4 mg/dL and P < 0.001. Presenting all three estimates and intervals gives substantially more information than simply classifying the p-values as significant or nonsignificant.
Finally, the serious-adverse-event data show why safety should be analyzed separately from efficacy. The registry-reported counts are 48/329, 54/328, 41/338, and 193/1000 across the four treatment arms. Without additional information on event types, exposure time, severity, and relatedness, those counts should remain descriptive rather than being converted into a broader safety conclusion.
26. Related Tutorials
Learn more about the methods used in this trial:
27. Related Calculators
28. Limitations of the Statistical Record
The ClinicalTrials.gov record is sufficiently detailed to reconstruct the principal statistical story of SURPASS-4, but they do not constitute the complete statistical analysis plan. Several details that would normally be important for a full statistical review are not present in the registry-reported extract, including the numerical non-inferiority margin, the complete multiplicity hierarchy, detailed missing-data assumptions and sensitivity analyses, and any interim-analysis framework.
Similarly, the ClinicalTrials.gov record is limited to the 12 statistical analyses included in the trial data. No additional baseline characteristics, subgroup estimates, long-term follow-up results, or unreported outcomes have been imported from external publications. This is deliberate: a reproducible trial-results page should distinguish what is contained in the specified registry record from what might be available elsewhere.
29. Sources
- ClinicalTrials.gov: NCT03730662 — SURPASS-4.
- Linked publication: PubMed PMID 39531161.
- Linked publication: PubMed PMID 37668888.
- Linked publication: PubMed PMID 37526908.
- Linked publication: PubMed PMID 36152639.
- Linked publication: PubMed PMID 35210595.
The numerical results and trial-design facts presented on this page are restricted to the registry-reported SURPASS-4 trial data. The linked publications are provided as publication references identified in that data; no additional numerical results from those publications are incorporated into this page.
Continue through the Clinical Biostats statistical library
Explore the statistical concepts behind randomized trials, longitudinal models, binary outcomes, confidence intervals, and non-inferiority analysis.
30. Record Summary
SURPASS-4 is a randomized phase 3 parallel-group trial with 2002 enrolled participants evaluating once-weekly tirzepatide versus once-daily insulin glargine in participants with type 2 diabetes mellitus and increased cardiovascular risk. The ClinicalTrials.gov record reports mixed-effects analyses for longitudinal continuous outcomes and logistic regression for the binary endpoint of HbA1c <7.0% at Week 52.
The two primary HbA1c analyses report net mean differences of -0.99 for 10 mg tirzepatide versus insulin glargine and -1.14 for 15 mg versus insulin glargine, each with a 97.5% two-sided confidence interval and P < 0.001. The registry identifies these analyses as non-inferiority hypotheses and states that sample-size selection was guided by a superiority objective, while the ClinicalTrials.gov record does not provide the numerical non-inferiority margin.
The secondary analyses extend the statistical picture: body-weight changes are reported as mean differences, achievement of HbA1c <7.0% is reported using odds ratios, and fasting serum glucose is reported using mean differences. Together, these results demonstrate why trial interpretation requires attention to the endpoint definition, effect measure, confidence interval, hypothesis type, analysis population, and multiplicity rather than relying on p-values alone.