This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
SUSTAIN-3 was a randomized, parallel-group, phase 3 trial evaluating the efficacy and safety of once-weekly semaglutide 1.0 mg versus once-weekly exenatide ER 2.0 mg in subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs.
| Feature | SUSTAIN-3 |
|---|---|
| Phase | Phase 3 |
| Therapeutic area | Diabetes |
| Condition | Diabetes; Diabetes Mellitus, Type 2 |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 813 |
| Interventions | Semaglutide; exenatide |
| Primary endpoint | Change From Baseline in HbA1c (Glycosylated Haemoglobin) |
| Results status | Completed; results posted |
| Lead sponsor | Novo Nordisk A/S |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT01885208 |
2. Clinical Question
The primary question was whether once-weekly semaglutide 1.0 mg differed from once-weekly exenatide ER 2.0 mg in the change in HbA1c from baseline to week 56 among subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs.
Population
Subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs.
Intervention
Semaglutide 1.0 mg administered once weekly.
Comparator
Exenatide ER 2.0 mg administered once weekly.
Primary question
What is the treatment difference in mean change in HbA1c from baseline to week 56, and does it meet the prespecified non-inferiority and superiority criteria?
3. Trial Design
Semaglutide
- Semaglutide 1.0 mg
- Once weekly
- Add-on to 1–2 oral antidiabetic drugs
Exenatide ER
- Exenatide ER 2.0 mg
- Once weekly
- Add-on to 1–2 oral antidiabetic drugs
The trial began on 2013-12-02 and reached primary completion on 2015-07-13. The registry reports 813 enrolled subjects and two treatment arms.
4. Endpoints
| Endpoint | Time frame | Registry definition |
|---|---|---|
| Change From Baseline in HbA1c (Glycosylated Haemoglobin) | Week 0, week 56 | Mean change in HbA1c from baseline to week 56. |
The registered primary endpoint is a change-from-baseline continuous outcome expressed as the percentage of glycosylated haemoglobin. The registry analysis uses the treatment difference between semaglutide 1.0 mg and exenatide ER 2.0 mg as the effect measure.
5. Analysis Population
The primary statistical analyses used the full analysis set (FAS). The registry defines this population as all randomized subjects who had received at least one dose of randomized semaglutide 1.0 mg or exenatide ER 2.0 mg.
| Population | Registry definition / role |
|---|---|
| Full analysis set | All randomized subjects who had received at least one dose of randomized semaglutide 1.0 mg or exenatide ER 2.0 mg. |
This is important because the analysis remains anchored to randomized treatment assignment while requiring exposure to at least one dose. The registry describes the analysis as incorporating intention-to-treat analysis among the concepts used in the statistical analysis.
6. Statistical Methodology
Mixed model for repeated measurements
The registry reports a mixed model for repeated measurements for the primary endpoint. Post-baseline responses were analyzed using a mixed model with treatment and country as fixed factors and baseline value as a covariate, with these factors nested within visit.
The important statistical feature is that the analysis uses repeated post-baseline measurements rather than treating the week 56 observation as an isolated outcome. Baseline HbA1c is incorporated as a covariate, while treatment and country are included as fixed factors.
Covariate adjustment
Baseline HbA1c was included as a covariate. Covariate adjustment can improve precision when the baseline measurement is strongly related to the outcome and can account for baseline differences in the statistical model without changing the randomized treatment comparison.
Country adjustment
Country was included as a fixed factor. In a multicountry clinical trial, this allows the model to account for systematic differences associated with country while estimating the treatment contrast.
Visit structure
The registry states that treatment, country, and baseline value were all nested within visit. This reflects the longitudinal nature of the post-baseline responses and allows the treatment comparison to be modeled within the repeated-visit framework rather than relying only on one isolated measurement.
Effect measure
The effect measure reported by the registry is a treatment difference. Unlike a hazard ratio or risk ratio, a treatment difference retains the original measurement scale of the outcome. Here, that scale is percentage of glycosylated haemoglobin.
7. Primary Result: Non-Inferiority Analysis
The registry reports a formal non-inferiority analysis for the primary endpoint, comparing semaglutide 1.0 mg with exenatide ER 2.0 mg.
Treatment difference in change from baseline in HbA1c
95% CI: −0.8 to −0.44 · P < 0.0001
Two-sided 95% confidence interval · Full analysis set · Mixed model for repeated measurements
| Component | Reported result |
|---|---|
| Endpoint | Change From Baseline in HbA1c (Glycosylated Haemoglobin) |
| Time frame | Week 0, week 56 |
| Comparison | Semaglutide 1.0 mg vs Exenatide ER 2.0 mg |
| Estimate | −0.62 |
| 95% CI | −0.8 to −0.44 |
| P-value | < 0.0001 |
| Hypothesis | Non-inferiority |
| Non-inferiority margin | 0.3 % |
The prespecified non-inferiority rule was that non-inferiority would be concluded if the upper limit of the two-sided 95% confidence interval for the estimated treatment difference was below the non-inferiority margin of 0.3 %.
The estimated treatment difference was −0.62 percentage points on the HbA1c scale. In the direction used for the reported comparison, this indicates a lower estimated HbA1c change for semaglutide relative to exenatide ER by 0.62 percentage points.
The estimate does not mean that every individual subject experienced a 0.62-percentage-point difference. It is a population-level treatment contrast estimated from the specified mixed model.
The two-sided 95% confidence interval extends from −0.8 to −0.44. For the prespecified non-inferiority rule, the important feature is the upper confidence-limit value of −0.44: it is below the specified 0.3 % margin.
The reported P < 0.0001 indicates strong statistical evidence against the null hypothesis associated with the reported analysis. It does not measure the magnitude of the treatment difference, and it does not replace the confidence-interval-based non-inferiority criterion.
The result is model-based and depends on the full analysis set and repeated-measures framework reported by the registry. The confidence interval describes statistical uncertainty around the estimated treatment difference; it is not a prediction interval for individual subjects.
8. Primary Result: Superiority Analysis
The same primary endpoint was also evaluated under a superiority hypothesis.
Superiority treatment difference
95% CI: −0.8 to −0.44 · P < 0.0001
Two-sided 95% confidence interval · Full analysis set · Mixed model for repeated measurements
| Component | Reported result |
|---|---|
| Endpoint | Change From Baseline in HbA1c (Glycosylated Haemoglobin) |
| Time frame | Week 0, week 56 |
| Comparison | Semaglutide 1.0 mg vs Exenatide ER 2.0 mg |
| Estimate | −0.62 |
| 95% CI | −0.8 to −0.44 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
| Superiority margin | 0 % |
The registry states that superiority was concluded if the upper limit of the two-sided 95% confidence interval for the estimated treatment difference was below the prespecified superiority margin of 0 %.
The treatment estimate is again −0.62, because the superiority analysis uses the same estimated treatment contrast and confidence interval. The difference is the hypothesis being evaluated, not a different observed effect.
The entire reported 95% confidence interval, −0.8 to −0.44, lies below the superiority margin of 0 %. Under the registry's stated decision rule, that provides evidence for superiority in the direction of the semaglutide treatment difference.
The P < 0.0001 result is evidence against the relevant null hypothesis; it should not be interpreted as the probability that the treatment difference is real, nor as the probability that the treatment is superior for an individual subject.
The precision of the estimate is conveyed by the confidence interval. The interval is narrower than the distance between the estimate and zero, but it remains an interval estimate rather than a guarantee of the exact population effect.
Because this is a repeated-measures analysis, the interpretation depends on the statistical model specified by the registry, including baseline adjustment, country adjustment, and the visit structure. The result should therefore be understood as a model-based estimated treatment difference rather than a simple arithmetic comparison of two raw means.
9. Understanding the Non-Inferiority and Superiority Logic
SUSTAIN-3 is statistically useful because the same treatment contrast was evaluated against two different prespecified reference margins.
| Question | Decision rule reported in the registry | Observed CI |
|---|---|---|
| Non-inferiority | Upper limit of the two-sided 95% CI below 0.3 % | −0.44, which is below 0.3 % |
| Superiority | Upper limit of the two-sided 95% CI below 0 % | −0.44, which is below 0 % |
The distinction is conceptual. A non-inferiority trial asks whether the new treatment is not unacceptably worse than the comparator according to a prespecified margin. A superiority test asks whether the estimated treatment difference crosses the null value in the prespecified direction.
Here, the reported confidence interval is entirely below both reference values. Thus the same estimated treatment effect satisfies both registry decision rules. The statistical conclusion comes from the prespecified confidence-interval criteria rather than from the numerical size of the P-value alone.
10. Why a Mixed-Effects Model Was Used
The primary endpoint is measured repeatedly over time, with post-baseline responses incorporated into a model that accounts for the visit structure. A mixed-effects or mixed-model framework is therefore well suited to the longitudinal data structure described in the registry.
Repeated observations
Each subject can contribute post-baseline measurements across visits. Modeling the longitudinal observations can use more of the available information than analyzing only one isolated measurement.
Baseline adjustment
Baseline HbA1c is included as a covariate, allowing the treatment comparison to account statistically for the starting HbA1c measurement.
Country adjustment
Country is included as a fixed factor, allowing systematic country-related differences to be incorporated into the model.
Visit structure
The registry specifies that treatment and country factors and baseline value were nested within visit, reflecting the repeated longitudinal assessment framework.
The reported value of −0.62 is therefore a model-based contrast on the HbA1c percentage scale, not a hazard ratio, odds ratio, or relative risk.
11. Statistical Methods Explained
Why was a mixed model used?
A mixed model was used because the registry describes post-baseline responses in a repeated-measures framework. The approach allows the analysis to represent measurements collected across visits while estimating the treatment difference after accounting for baseline HbA1c and country.
What does the treatment difference of −0.62 mean?
It means that the model-estimated treatment contrast between semaglutide 1.0 mg and exenatide ER 2.0 mg was −0.62 percentage points of glycosylated haemoglobin. The negative sign indicates the direction of the reported semaglutide-versus-exenatide comparison.
Why was baseline HbA1c included as a covariate?
Baseline adjustment can improve the precision of a treatment comparison by incorporating information about where subjects started. The model does not simply compare unadjusted week 56 values; it includes baseline HbA1c in the specified statistical model.
Why is the non-inferiority margin important?
Non-inferiority is not established merely because the P-value is small. The registry specifies a 0.3 % non-inferiority margin and states that the upper limit of the two-sided 95% confidence interval must be below that margin. The margin defines the amount of potential disadvantage considered acceptable under the trial's prespecified rule.
Why can the same estimate be used for both non-inferiority and superiority?
The underlying estimated treatment difference does not change simply because the hypothesis changes. What changes is the reference value against which the confidence interval is evaluated: 0.3 % for the reported non-inferiority criterion and 0 % for the reported superiority criterion.
What does the P-value add if the confidence interval is already available?
The P-value summarizes evidence against the relevant null hypothesis under the specified statistical test. The confidence interval adds information about the magnitude and precision of the estimated treatment difference. Neither quantity alone describes the complete clinical importance of the result.
Why is the full analysis set important?
The registry defines the FAS as all randomized subjects who received at least one dose of randomized treatment. Anchoring the efficacy analysis to randomized treatment assignment helps preserve the randomized comparison, while the dose requirement defines the analysis population used in the reported model.
12. Confidence Intervals and Precision
The primary treatment difference was −0.62, with a two-sided 95% confidence interval from −0.8 to −0.44.
The confidence interval provides a range of values compatible with the model and sampling framework at the stated confidence level. It should not be described as saying that there is a 95% probability that the true treatment effect lies inside this particular interval.
The P-value of < 0.0001 addresses evidence against a null hypothesis. The interval from −0.8 to −0.44 addresses uncertainty around the treatment estimate. These are related but different pieces of statistical information.
13. Safety
The registry provides serious adverse event counts by randomized treatment arm. These are reported as affected subjects over subjects at risk.
| Safety measure | Semaglutide 1.0 mg | Exenatide ER |
|---|---|---|
| Serious adverse events | 38/404 | 24/405 |
These counts should be interpreted separately from the HbA1c efficacy analysis. The serious-adverse-event comparison describes a safety outcome, whereas the primary endpoint estimates a treatment difference in a longitudinal glycemic measure.
14. Non-Inferiority Design: A Closer Statistical Look
Non-inferiority analysis is particularly sensitive to how the estimand, analysis population, endpoint, and margin are defined. SUSTAIN-3 provides a clear example because the registry explicitly states the decision rule and margin.
The reported upper confidence limit was −0.44, which is below the stated non-inferiority margin.
A common misconception is that a non-inferiority trial is simply a conventional superiority trial with a different P-value. It is not. The key question is whether the uncertainty interval excludes treatment differences that would exceed the prespecified clinically acceptable disadvantage.
Another important distinction is that the margin is not estimated from the observed data. It is specified in advance as part of the trial's design. The registry supplies the margin of 0.3 % for this analysis.
15. Superiority After Non-Inferiority
The trial's statistical analysis also evaluated superiority using a superiority margin of 0 %. Because the reported two-sided 95% confidence interval was −0.8 to −0.44, its upper limit was below zero.
Non-inferiority question
Is the upper confidence limit below the prespecified non-inferiority margin of 0.3 %?
Superiority question
Is the upper confidence limit below the superiority margin of 0 %?
Observed upper limit
The reported upper limit was −0.44 for the two-sided 95% confidence interval.
Statistical implication
The same reported interval is below both prespecified reference values.
This illustrates why a confidence interval is especially informative in non-inferiority and superiority settings: it simultaneously communicates the direction, magnitude, precision, and relationship to the prespecified decision threshold.
16. What the P-Value Does — and Does Not — Mean
The reported P < 0.0001 indicates strong evidence against the null hypothesis associated with the specified analysis. It does not mean that there is a probability of less than 0.0001 that the null hypothesis is true.
The magnitude of the observed treatment difference is −0.62. The P-value does not tell us whether an effect is small, moderate, or large on the HbA1c scale. That question requires examination of the treatment estimate and its clinical context.
The 95% confidence interval of −0.8 to −0.44 communicates uncertainty around the estimate. A P-value alone cannot provide that range.
17. Intention-to-Treat and the Full Analysis Set
The registry identifies intention-to-treat analysis as a concept in the primary statistical analysis and defines the full analysis set as all randomized subjects who received at least one dose of randomized treatment.
The randomized treatment assignment is central to causal inference in a randomized trial. If patients deviate from treatment, discontinue treatment, or otherwise have incomplete observations, analyzing subjects according to the randomized comparison can help preserve the balance created by randomization.
At the same time, the exact FAS definition matters. In SUSTAIN-3, the registry's FAS is not described simply as every randomized subject; it specifically requires receipt of at least one dose of randomized semaglutide 1.0 mg or exenatide ER 2.0 mg.
18. Missing Data and Repeated Measurements
The registry analysis states that post-baseline responses were analyzed with a mixed model for repeated measurements. However, the ClinicalTrials.gov record does not specify a particular missing-data imputation method.
For that reason, this page does not attribute a specific imputation procedure to SUSTAIN-3. A mixed-model analysis of repeated measurements is not synonymous with a simple last-observation-carried-forward procedure, and the statistical handling of incomplete longitudinal data should not be inferred when the registry data do not specify it.
19. Design Features the Registry Does Not Report in the Supplied Data
Several design topics can be important in clinical-trial interpretation, but the registry-reported SUSTAIN-3 data do not provide enough information to describe them as trial-specific features.
| Design topic | What can be stated from the ClinicalTrials.gov record |
|---|---|
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Factorial design | The design model is reported as parallel; no factorial structure is reported. |
| Interim analysis | Not reported in the ClinicalTrials.gov record. |
| Multiplicity strategy | Not reported in the ClinicalTrials.gov record beyond the separate non-inferiority and superiority analyses. |
| Stratification factors | Not reported in the ClinicalTrials.gov record. |
| Bayesian methods | Not reported; the registry-reported analysis is a frequentist mixed-model analysis with confidence intervals and P-values. |
| Missing-data imputation | No specific imputation method is reported in the ClinicalTrials.gov record. |
These omissions are important because a statistical analysis page should distinguish between what the registry actually reports and what might be common practice in other trials. The absence of a registry-reported detail is not a basis for reconstructing it from memory.
20. Trial Timeline
Trial start
SUSTAIN-3 began as a randomized phase 3 study of semaglutide versus exenatide ER.
Parallel treatment comparison
The study used a randomized, parallel design with two treatment arms and no masking.
Primary endpoint window
The registered primary endpoint was mean change in HbA1c from baseline to week 56.
Primary completion
The registry reports primary completion on 2015-07-13.
21. Interpreting the Treatment Difference
A treatment difference of −0.62 means that the model-estimated change in HbA1c for semaglutide 1.0 mg was 0.62 percentage points lower than the corresponding estimated change for exenatide ER 2.0 mg, using the treatment contrast reported by the registry.
It does not mean that every subject had a 0.62-percentage-point difference. It is not a hazard ratio, relative risk, odds ratio, or percentage reduction in risk. It is a treatment difference on the HbA1c percentage scale.
The two-sided 95% confidence interval of −0.8 to −0.44 describes statistical uncertainty around the model-estimated treatment difference. It also provides the basis for the registry's reported non-inferiority and superiority decision rules.
The P < 0.0001 result summarizes evidence against the relevant null hypothesis under the specified analysis. It does not quantify the size or clinical importance of the effect.
22. Limitations
- Single primary endpoint: the ClinicalTrials.gov record provides one registered primary endpoint, change in HbA1c from baseline to week 56. They do not provide additional efficacy endpoint results for this page.
- Model dependence: the primary estimate is generated by a mixed model for repeated measurements with specified treatment, country, baseline, and visit structure. It is therefore a model-based estimate rather than an unadjusted arithmetic difference.
- Analysis-population definition: the reported result uses the registry-defined full analysis set, which requires randomized treatment assignment and at least one dose.
- Non-inferiority interpretation: the conclusion depends on the prespecified 0.3 % margin and the confidence-interval rule reported by the registry.
- Superiority interpretation: the superiority conclusion depends on the reported 0 % margin and the same two-sided 95% confidence interval.
- Missing-data details: the ClinicalTrials.gov record does not identify a specific imputation procedure, so no particular missing-data method is attributed here.
- Safety detail: the ClinicalTrials.gov record reports serious adverse events by arm but do not provide a formal statistical comparison for that safety endpoint.
- Limited secondary-endpoint information: the registry reports six outcome measures in total, but the ClinicalTrials.gov record does not provide their definitions or statistical results. This page therefore does not infer or reproduce additional endpoint results.
- External validity: the ClinicalTrials.gov record identifies the population as subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs. Applicability outside that defined population cannot be established from the ClinicalTrials.gov record alone.
23. Why This Trial Matters Statistically
SUSTAIN-3 is a useful teaching example because its primary result illustrates several core principles of modern clinical-trial statistics without requiring a time-to-event endpoint.
| Concept | How it appears in SUSTAIN-3 |
|---|---|
| Randomization | The trial uses randomized allocation in a two-arm parallel design. |
| Intention-to-treat principle | The statistical analysis identifies intention-to-treat analysis as an analysis concept and uses a registry-defined full analysis set. |
| Covariate adjustment | Baseline HbA1c is included as a covariate. |
| Repeated measurements | Post-baseline responses are analyzed with a mixed model for repeated measurements. |
| Fixed effects | Treatment and country are specified as fixed factors. |
| Confidence intervals | A two-sided 95% CI is reported for the treatment difference. |
| Non-inferiority | The upper confidence limit is compared with a prespecified 0.3 % margin. |
| Superiority | The upper confidence limit is compared with a 0 % superiority margin. |
| P-values | The primary analyses report P < 0.0001. |
| Safety analysis | Serious adverse events are reported separately by treatment arm. |
The trial therefore provides a particularly clear example of why statistical interpretation should go beyond asking whether a P-value is "significant." The treatment estimate, confidence interval, analysis population, model structure, and prespecified non-inferiority or superiority margin all contribute to the interpretation.
24. Clinical Biostats Perspective: What to Look at First
Start with the estimand
The primary question concerns the difference in change in HbA1c from baseline to week 56 between two randomized treatments.
Then inspect the model
The registry specifies a mixed model for repeated measurements with treatment, country, baseline adjustment, and visit structure.
Then inspect uncertainty
The treatment estimate is −0.62 with a two-sided 95% CI of −0.8 to −0.44.
Finally inspect the decision rule
The same confidence interval is evaluated against the prespecified non-inferiority margin of 0.3 % and superiority margin of 0 %.
This sequence prevents a common analytical mistake: starting with the P-value and working backward. In a well-specified clinical trial, the endpoint, estimand, analysis population, model, confidence interval, and decision threshold should be considered together.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Calculators
27. Sources
- ClinicalTrials.gov: SUSTAIN-3, NCT01885208. Trial profile, registered endpoint, statistical analyses, analysis population, and reported serious adverse events.
- PubMed record: PMID 29246950.
- PubMed record: PMID 29687620.
- PubMed record: PMID 29748996.
- PubMed record: PMID 29766634.
- PubMed record: PMID 29862621.
Continue through Clinical Biostats
Connect this trial's statistical methods with deeper tutorials and practical statistical calculators.
28. Record Summary
SUSTAIN-3 is a randomized phase 3 comparison of once-weekly semaglutide 1.0 mg and once-weekly exenatide ER 2.0 mg in subjects with type 2 diabetes receiving 1–2 oral antidiabetic drugs. Its registered primary endpoint was mean change in HbA1c from baseline to week 56. The reported analysis used a mixed model for repeated measurements with treatment and country as fixed factors and baseline value as a covariate, with the factors nested within visit.
The reported treatment difference was −0.62, with a two-sided 95% confidence interval of −0.8 to −0.44 and P < 0.0001. The registry reports both non-inferiority and superiority analyses. Non-inferiority was assessed against a prespecified margin of 0.3 %, while superiority was assessed against a margin of 0 %. In both cases, the upper limit of the reported confidence interval was below the relevant margin.
The most important statistical lesson is that the result should be read as a complete inferential statement rather than as a P-value alone: randomized comparison + defined analysis population + longitudinal mixed model + covariate adjustment + treatment difference + confidence interval + prespecified decision margin. The registry also reports serious adverse events of 38/404 for semaglutide 1.0 mg and 24/405 for exenatide ER, which should be interpreted separately from the primary efficacy analysis.