← Clinical Trials
Type 2 Diabetes Phase 3 Randomized NCT01885208

SUSTAIN-3: Complete Statistical Analysis of Semaglutide in Type 2 Diabetes

An independent statistical review of the randomized phase 3 SUSTAIN-3 trial comparing once-weekly semaglutide 1.0 mg with once-weekly exenatide ER 2.0 mg as add-on treatment to 1–2 oral antidiabetic drugs in subjects with type 2 diabetes.

Trial period: 2013-12-02 to 2015-07-13  ·  Enrollment: 813  ·  Sponsor: Novo Nordisk A/S
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SUSTAIN-3 was a randomized, parallel-group, phase 3 trial evaluating the efficacy and safety of once-weekly semaglutide 1.0 mg versus once-weekly exenatide ER 2.0 mg in subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs.

813
Enrolled
2 randomized arms
2
Arms
Parallel design
−0.62
Treatment difference
95% CI −0.8 to −0.44
< 0.0001
P-value
Two-sided 95% CI
FeatureSUSTAIN-3
PhasePhase 3
Therapeutic areaDiabetes
ConditionDiabetes; Diabetes Mellitus, Type 2
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
Enrollment813
InterventionsSemaglutide; exenatide
Primary endpointChange From Baseline in HbA1c (Glycosylated Haemoglobin)
Results statusCompleted; results posted
Lead sponsorNovo Nordisk A/S
Sponsor typeIndustry
ClinicalTrials.govNCT01885208

2. Clinical Question

The primary question was whether once-weekly semaglutide 1.0 mg differed from once-weekly exenatide ER 2.0 mg in the change in HbA1c from baseline to week 56 among subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs.

Population

Subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs.

Intervention

Semaglutide 1.0 mg administered once weekly.

Comparator

Exenatide ER 2.0 mg administered once weekly.

Primary question

What is the treatment difference in mean change in HbA1c from baseline to week 56, and does it meet the prespecified non-inferiority and superiority criteria?

3. Trial Design

01
Randomize 813 subjects
02
2 arms Semaglutide or exenatide ER
03
Weekly treatment 1.0 mg vs 2.0 mg
04
Week 56 HbA1c assessment
05
Compare Treatment difference
ARM 1

Semaglutide

  • Semaglutide 1.0 mg
  • Once weekly
  • Add-on to 1–2 oral antidiabetic drugs
ARM 2

Exenatide ER

  • Exenatide ER 2.0 mg
  • Once weekly
  • Add-on to 1–2 oral antidiabetic drugs
Allocation
Randomized
Design model
Parallel
Masking
None
Primary purpose
Treatment

The trial began on 2013-12-02 and reached primary completion on 2015-07-13. The registry reports 813 enrolled subjects and two treatment arms.

4. Endpoints

EndpointTime frameRegistry definition
Change From Baseline in HbA1c (Glycosylated Haemoglobin) Week 0, week 56 Mean change in HbA1c from baseline to week 56.

The registered primary endpoint is a change-from-baseline continuous outcome expressed as the percentage of glycosylated haemoglobin. The registry analysis uses the treatment difference between semaglutide 1.0 mg and exenatide ER 2.0 mg as the effect measure.

Endpoint interpretation: because the reported effect is a treatment difference in change from baseline, a negative value means that the estimated change was lower in the semaglutide group than in the exenatide ER group. The treatment difference should not be interpreted as a ratio or a percentage reduction.

5. Analysis Population

The primary statistical analyses used the full analysis set (FAS). The registry defines this population as all randomized subjects who had received at least one dose of randomized semaglutide 1.0 mg or exenatide ER 2.0 mg.

PopulationRegistry definition / role
Full analysis set All randomized subjects who had received at least one dose of randomized semaglutide 1.0 mg or exenatide ER 2.0 mg.

This is important because the analysis remains anchored to randomized treatment assignment while requiring exposure to at least one dose. The registry describes the analysis as incorporating intention-to-treat analysis among the concepts used in the statistical analysis.

6. Statistical Methodology

Mixed model for repeated measurements

The registry reports a mixed model for repeated measurements for the primary endpoint. Post-baseline responses were analyzed using a mixed model with treatment and country as fixed factors and baseline value as a covariate, with these factors nested within visit.

Model structure reported by the registry
Post-baseline HbA1c response ~ treatment + country + baseline HbA1c, nested within visit

The important statistical feature is that the analysis uses repeated post-baseline measurements rather than treating the week 56 observation as an isolated outcome. Baseline HbA1c is incorporated as a covariate, while treatment and country are included as fixed factors.

Covariate adjustment

Baseline HbA1c was included as a covariate. Covariate adjustment can improve precision when the baseline measurement is strongly related to the outcome and can account for baseline differences in the statistical model without changing the randomized treatment comparison.

Country adjustment

Country was included as a fixed factor. In a multicountry clinical trial, this allows the model to account for systematic differences associated with country while estimating the treatment contrast.

Visit structure

The registry states that treatment, country, and baseline value were all nested within visit. This reflects the longitudinal nature of the post-baseline responses and allows the treatment comparison to be modeled within the repeated-visit framework rather than relying only on one isolated measurement.

Effect measure

The effect measure reported by the registry is a treatment difference. Unlike a hazard ratio or risk ratio, a treatment difference retains the original measurement scale of the outcome. Here, that scale is percentage of glycosylated haemoglobin.

7. Primary Result: Non-Inferiority Analysis

The registry reports a formal non-inferiority analysis for the primary endpoint, comparing semaglutide 1.0 mg with exenatide ER 2.0 mg.

Treatment difference in change from baseline in HbA1c

−0.62

95% CI: −0.8 to −0.44   ·   P < 0.0001

Two-sided 95% confidence interval · Full analysis set · Mixed model for repeated measurements

ComponentReported result
EndpointChange From Baseline in HbA1c (Glycosylated Haemoglobin)
Time frameWeek 0, week 56
ComparisonSemaglutide 1.0 mg vs Exenatide ER 2.0 mg
Estimate−0.62
95% CI−0.8 to −0.44
P-value< 0.0001
HypothesisNon-inferiority
Non-inferiority margin0.3 %

The prespecified non-inferiority rule was that non-inferiority would be concluded if the upper limit of the two-sided 95% confidence interval for the estimated treatment difference was below the non-inferiority margin of 0.3 %.

Clinical Biostats interpretation

The estimated treatment difference was −0.62 percentage points on the HbA1c scale. In the direction used for the reported comparison, this indicates a lower estimated HbA1c change for semaglutide relative to exenatide ER by 0.62 percentage points.

The estimate does not mean that every individual subject experienced a 0.62-percentage-point difference. It is a population-level treatment contrast estimated from the specified mixed model.

The two-sided 95% confidence interval extends from −0.8 to −0.44. For the prespecified non-inferiority rule, the important feature is the upper confidence-limit value of −0.44: it is below the specified 0.3 % margin.

The reported P < 0.0001 indicates strong statistical evidence against the null hypothesis associated with the reported analysis. It does not measure the magnitude of the treatment difference, and it does not replace the confidence-interval-based non-inferiority criterion.

The result is model-based and depends on the full analysis set and repeated-measures framework reported by the registry. The confidence interval describes statistical uncertainty around the estimated treatment difference; it is not a prediction interval for individual subjects.

8. Primary Result: Superiority Analysis

The same primary endpoint was also evaluated under a superiority hypothesis.

Superiority treatment difference

−0.62

95% CI: −0.8 to −0.44   ·   P < 0.0001

Two-sided 95% confidence interval · Full analysis set · Mixed model for repeated measurements

ComponentReported result
EndpointChange From Baseline in HbA1c (Glycosylated Haemoglobin)
Time frameWeek 0, week 56
ComparisonSemaglutide 1.0 mg vs Exenatide ER 2.0 mg
Estimate−0.62
95% CI−0.8 to −0.44
P-value< 0.0001
HypothesisSuperiority
Superiority margin0 %

The registry states that superiority was concluded if the upper limit of the two-sided 95% confidence interval for the estimated treatment difference was below the prespecified superiority margin of 0 %.

Clinical Biostats interpretation

The treatment estimate is again −0.62, because the superiority analysis uses the same estimated treatment contrast and confidence interval. The difference is the hypothesis being evaluated, not a different observed effect.

The entire reported 95% confidence interval, −0.8 to −0.44, lies below the superiority margin of 0 %. Under the registry's stated decision rule, that provides evidence for superiority in the direction of the semaglutide treatment difference.

The P < 0.0001 result is evidence against the relevant null hypothesis; it should not be interpreted as the probability that the treatment difference is real, nor as the probability that the treatment is superior for an individual subject.

The precision of the estimate is conveyed by the confidence interval. The interval is narrower than the distance between the estimate and zero, but it remains an interval estimate rather than a guarantee of the exact population effect.

Because this is a repeated-measures analysis, the interpretation depends on the statistical model specified by the registry, including baseline adjustment, country adjustment, and the visit structure. The result should therefore be understood as a model-based estimated treatment difference rather than a simple arithmetic comparison of two raw means.

9. Understanding the Non-Inferiority and Superiority Logic

SUSTAIN-3 is statistically useful because the same treatment contrast was evaluated against two different prespecified reference margins.

QuestionDecision rule reported in the registryObserved CI
Non-inferiority Upper limit of the two-sided 95% CI below 0.3 % −0.44, which is below 0.3 %
Superiority Upper limit of the two-sided 95% CI below 0 % −0.44, which is below 0 %

The distinction is conceptual. A non-inferiority trial asks whether the new treatment is not unacceptably worse than the comparator according to a prespecified margin. A superiority test asks whether the estimated treatment difference crosses the null value in the prespecified direction.

Here, the reported confidence interval is entirely below both reference values. Thus the same estimated treatment effect satisfies both registry decision rules. The statistical conclusion comes from the prespecified confidence-interval criteria rather than from the numerical size of the P-value alone.

10. Why a Mixed-Effects Model Was Used

The primary endpoint is measured repeatedly over time, with post-baseline responses incorporated into a model that accounts for the visit structure. A mixed-effects or mixed-model framework is therefore well suited to the longitudinal data structure described in the registry.

Repeated observations

Each subject can contribute post-baseline measurements across visits. Modeling the longitudinal observations can use more of the available information than analyzing only one isolated measurement.

Baseline adjustment

Baseline HbA1c is included as a covariate, allowing the treatment comparison to account statistically for the starting HbA1c measurement.

Country adjustment

Country is included as a fixed factor, allowing systematic country-related differences to be incorporated into the model.

Visit structure

The registry specifies that treatment and country factors and baseline value were nested within visit, reflecting the repeated longitudinal assessment framework.

Conceptual treatment contrast
Treatment difference = estimated mean change under semaglutide − estimated mean change under exenatide ER

The reported value of −0.62 is therefore a model-based contrast on the HbA1c percentage scale, not a hazard ratio, odds ratio, or relative risk.

11. Statistical Methods Explained

Why was a mixed model used?

A mixed model was used because the registry describes post-baseline responses in a repeated-measures framework. The approach allows the analysis to represent measurements collected across visits while estimating the treatment difference after accounting for baseline HbA1c and country.

What does the treatment difference of −0.62 mean?

It means that the model-estimated treatment contrast between semaglutide 1.0 mg and exenatide ER 2.0 mg was −0.62 percentage points of glycosylated haemoglobin. The negative sign indicates the direction of the reported semaglutide-versus-exenatide comparison.

Why was baseline HbA1c included as a covariate?

Baseline adjustment can improve the precision of a treatment comparison by incorporating information about where subjects started. The model does not simply compare unadjusted week 56 values; it includes baseline HbA1c in the specified statistical model.

Why is the non-inferiority margin important?

Non-inferiority is not established merely because the P-value is small. The registry specifies a 0.3 % non-inferiority margin and states that the upper limit of the two-sided 95% confidence interval must be below that margin. The margin defines the amount of potential disadvantage considered acceptable under the trial's prespecified rule.

Why can the same estimate be used for both non-inferiority and superiority?

The underlying estimated treatment difference does not change simply because the hypothesis changes. What changes is the reference value against which the confidence interval is evaluated: 0.3 % for the reported non-inferiority criterion and 0 % for the reported superiority criterion.

What does the P-value add if the confidence interval is already available?

The P-value summarizes evidence against the relevant null hypothesis under the specified statistical test. The confidence interval adds information about the magnitude and precision of the estimated treatment difference. Neither quantity alone describes the complete clinical importance of the result.

Why is the full analysis set important?

The registry defines the FAS as all randomized subjects who received at least one dose of randomized treatment. Anchoring the efficacy analysis to randomized treatment assignment helps preserve the randomized comparison, while the dose requirement defines the analysis population used in the reported model.

12. Confidence Intervals and Precision

The primary treatment difference was −0.62, with a two-sided 95% confidence interval from −0.8 to −0.44.

Reported treatment difference and 95% CI
Lower limit
−0.8
Estimate
−0.62
Upper limit
−0.44

The confidence interval provides a range of values compatible with the model and sampling framework at the stated confidence level. It should not be described as saying that there is a 95% probability that the true treatment effect lies inside this particular interval.

Precision versus statistical significance

The P-value of < 0.0001 addresses evidence against a null hypothesis. The interval from −0.8 to −0.44 addresses uncertainty around the treatment estimate. These are related but different pieces of statistical information.

13. Safety

The registry provides serious adverse event counts by randomized treatment arm. These are reported as affected subjects over subjects at risk.

Safety measureSemaglutide 1.0 mgExenatide ER
Serious adverse events 38/404 24/405
38/404
Semaglutide
Serious adverse events
24/405
Exenatide ER
Serious adverse events
404
Semaglutide at risk
Registry denominator
405
Exenatide at risk
Registry denominator

These counts should be interpreted separately from the HbA1c efficacy analysis. The serious-adverse-event comparison describes a safety outcome, whereas the primary endpoint estimates a treatment difference in a longitudinal glycemic measure.

Safety interpretation: the ClinicalTrials.gov record reports serious adverse events by arm, but do not provide a formal statistical comparison for this safety measure. The counts should therefore not be converted into a comparative P-value or relative effect without additional analysis data.

14. Non-Inferiority Design: A Closer Statistical Look

Non-inferiority analysis is particularly sensitive to how the estimand, analysis population, endpoint, and margin are defined. SUSTAIN-3 provides a clear example because the registry explicitly states the decision rule and margin.

Prespecified non-inferiority criterion
Upper limit of two-sided 95% CI < 0.3 %

The reported upper confidence limit was −0.44, which is below the stated non-inferiority margin.

A common misconception is that a non-inferiority trial is simply a conventional superiority trial with a different P-value. It is not. The key question is whether the uncertainty interval excludes treatment differences that would exceed the prespecified clinically acceptable disadvantage.

Another important distinction is that the margin is not estimated from the observed data. It is specified in advance as part of the trial's design. The registry supplies the margin of 0.3 % for this analysis.

15. Superiority After Non-Inferiority

The trial's statistical analysis also evaluated superiority using a superiority margin of 0 %. Because the reported two-sided 95% confidence interval was −0.8 to −0.44, its upper limit was below zero.

Non-inferiority question

Is the upper confidence limit below the prespecified non-inferiority margin of 0.3 %?

Superiority question

Is the upper confidence limit below the superiority margin of 0 %?

Observed upper limit

The reported upper limit was −0.44 for the two-sided 95% confidence interval.

Statistical implication

The same reported interval is below both prespecified reference values.

This illustrates why a confidence interval is especially informative in non-inferiority and superiority settings: it simultaneously communicates the direction, magnitude, precision, and relationship to the prespecified decision threshold.

16. What the P-Value Does — and Does Not — Mean

Statistical interpretation

The reported P < 0.0001 indicates strong evidence against the null hypothesis associated with the specified analysis. It does not mean that there is a probability of less than 0.0001 that the null hypothesis is true.

Effect size is separate

The magnitude of the observed treatment difference is −0.62. The P-value does not tell us whether an effect is small, moderate, or large on the HbA1c scale. That question requires examination of the treatment estimate and its clinical context.

Precision is separate

The 95% confidence interval of −0.8 to −0.44 communicates uncertainty around the estimate. A P-value alone cannot provide that range.

17. Intention-to-Treat and the Full Analysis Set

The registry identifies intention-to-treat analysis as a concept in the primary statistical analysis and defines the full analysis set as all randomized subjects who received at least one dose of randomized treatment.

The randomized treatment assignment is central to causal inference in a randomized trial. If patients deviate from treatment, discontinue treatment, or otherwise have incomplete observations, analyzing subjects according to the randomized comparison can help preserve the balance created by randomization.

At the same time, the exact FAS definition matters. In SUSTAIN-3, the registry's FAS is not described simply as every randomized subject; it specifically requires receipt of at least one dose of randomized semaglutide 1.0 mg or exenatide ER 2.0 mg.

Why this matters: analysis populations are part of the estimand. Two analyses can use the same endpoint and treatment groups but answer somewhat different questions if their inclusion rules differ. The registry-defined FAS should therefore be preserved when describing the reported primary result.

18. Missing Data and Repeated Measurements

The registry analysis states that post-baseline responses were analyzed with a mixed model for repeated measurements. However, the ClinicalTrials.gov record does not specify a particular missing-data imputation method.

For that reason, this page does not attribute a specific imputation procedure to SUSTAIN-3. A mixed-model analysis of repeated measurements is not synonymous with a simple last-observation-carried-forward procedure, and the statistical handling of incomplete longitudinal data should not be inferred when the registry data do not specify it.

Data boundary: the ClinicalTrials.gov record does not report a specific missing-data or imputation strategy. No particular imputation method is therefore attributed to the trial.

19. Design Features the Registry Does Not Report in the Supplied Data

Several design topics can be important in clinical-trial interpretation, but the registry-reported SUSTAIN-3 data do not provide enough information to describe them as trial-specific features.

Design topicWhat can be stated from the ClinicalTrials.gov record
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designThe design model is reported as parallel; no factorial structure is reported.
Interim analysisNot reported in the ClinicalTrials.gov record.
Multiplicity strategyNot reported in the ClinicalTrials.gov record beyond the separate non-inferiority and superiority analyses.
Stratification factorsNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported; the registry-reported analysis is a frequentist mixed-model analysis with confidence intervals and P-values.
Missing-data imputationNo specific imputation method is reported in the ClinicalTrials.gov record.

These omissions are important because a statistical analysis page should distinguish between what the registry actually reports and what might be common practice in other trials. The absence of a registry-reported detail is not a basis for reconstructing it from memory.

20. Trial Timeline

2013-12-02

Trial start

SUSTAIN-3 began as a randomized phase 3 study of semaglutide versus exenatide ER.

Phase 3 · Randomized

Parallel treatment comparison

The study used a randomized, parallel design with two treatment arms and no masking.

Week 0 → Week 56

Primary endpoint window

The registered primary endpoint was mean change in HbA1c from baseline to week 56.

2015-07-13

Primary completion

The registry reports primary completion on 2015-07-13.

21. Interpreting the Treatment Difference

What the estimate means

A treatment difference of −0.62 means that the model-estimated change in HbA1c for semaglutide 1.0 mg was 0.62 percentage points lower than the corresponding estimated change for exenatide ER 2.0 mg, using the treatment contrast reported by the registry.

What the estimate does not mean

It does not mean that every subject had a 0.62-percentage-point difference. It is not a hazard ratio, relative risk, odds ratio, or percentage reduction in risk. It is a treatment difference on the HbA1c percentage scale.

What the confidence interval adds

The two-sided 95% confidence interval of −0.8 to −0.44 describes statistical uncertainty around the model-estimated treatment difference. It also provides the basis for the registry's reported non-inferiority and superiority decision rules.

What the P-value adds

The P < 0.0001 result summarizes evidence against the relevant null hypothesis under the specified analysis. It does not quantify the size or clinical importance of the effect.

22. Limitations

23. Why This Trial Matters Statistically

SUSTAIN-3 is a useful teaching example because its primary result illustrates several core principles of modern clinical-trial statistics without requiring a time-to-event endpoint.

ConceptHow it appears in SUSTAIN-3
RandomizationThe trial uses randomized allocation in a two-arm parallel design.
Intention-to-treat principleThe statistical analysis identifies intention-to-treat analysis as an analysis concept and uses a registry-defined full analysis set.
Covariate adjustmentBaseline HbA1c is included as a covariate.
Repeated measurementsPost-baseline responses are analyzed with a mixed model for repeated measurements.
Fixed effectsTreatment and country are specified as fixed factors.
Confidence intervalsA two-sided 95% CI is reported for the treatment difference.
Non-inferiorityThe upper confidence limit is compared with a prespecified 0.3 % margin.
SuperiorityThe upper confidence limit is compared with a 0 % superiority margin.
P-valuesThe primary analyses report P < 0.0001.
Safety analysisSerious adverse events are reported separately by treatment arm.

The trial therefore provides a particularly clear example of why statistical interpretation should go beyond asking whether a P-value is "significant." The treatment estimate, confidence interval, analysis population, model structure, and prespecified non-inferiority or superiority margin all contribute to the interpretation.

24. Clinical Biostats Perspective: What to Look at First

Start with the estimand

The primary question concerns the difference in change in HbA1c from baseline to week 56 between two randomized treatments.

Then inspect the model

The registry specifies a mixed model for repeated measurements with treatment, country, baseline adjustment, and visit structure.

Then inspect uncertainty

The treatment estimate is −0.62 with a two-sided 95% CI of −0.8 to −0.44.

Finally inspect the decision rule

The same confidence interval is evaluated against the prespecified non-inferiority margin of 0.3 % and superiority margin of 0 %.

This sequence prevents a common analytical mistake: starting with the P-value and working backward. In a well-specified clinical trial, the endpoint, estimand, analysis population, model, confidence interval, and decision threshold should be considered together.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Calculators

27. Sources

Continue through Clinical Biostats

Connect this trial's statistical methods with deeper tutorials and practical statistical calculators.

28. Record Summary

SUSTAIN-3 is a randomized phase 3 comparison of once-weekly semaglutide 1.0 mg and once-weekly exenatide ER 2.0 mg in subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs. Its registered primary endpoint was mean change in HbA1c from baseline to week 56. The reported analysis used a mixed model for repeated measurements with treatment and country as fixed factors and baseline value as a covariate, with the factors nested within visit.

The reported treatment difference was −0.62, with a two-sided 95% confidence interval of −0.8 to −0.44 and P < 0.0001. The registry reports both non-inferiority and superiority analyses. Non-inferiority was assessed against a prespecified margin of 0.3 %, while superiority was assessed against a margin of 0 %. In both cases, the upper limit of the reported confidence interval was below the relevant margin.

The most important statistical lesson is that the result should be read as a complete inferential statement rather than as a P-value alone: randomized comparison + defined analysis population + longitudinal mixed model + covariate adjustment + treatment difference + confidence interval + prespecified decision margin. The registry also reports serious adverse events of 38/404 for semaglutide 1.0 mg and 24/405 for exenatide ER, which should be interpreted separately from the primary efficacy analysis.

Clinical Biostats methodology: This page separates reported trial evidence from statistical interpretation. Where the ClinicalTrials.gov record does not report a design feature, endpoint result, subgroup analysis, or missing-data procedure, the page does not infer one from outside knowledge.