← Clinical Trials
Type 2 Diabetes Phase 3 Non-Inferiority NCT01624259

AWARD-6: Complete Statistical Analysis of Dulaglutide in Type 2 Diabetes

An independent statistical analysis of the randomized phase 3 AWARD-6 trial comparing LY2189265 (dulaglutide) with liraglutide in participants with type 2 diabetes, centered on the prespecified 26-week HbA1c non-inferiority analysis and its supporting secondary outcomes.

Trial status: COMPLETED  ·  Enrollment: 599  ·  Primary completion: November 2013
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the statistical analyses posted in the ClinicalTrials.gov record. The registry record is the official source for the trial information.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

AWARD-6 was a randomized, parallel, open-label phase 3 trial comparing LY2189265 (dulaglutide) with liraglutide in type 2 diabetes. The registered primary endpoint was change from baseline to 26 weeks in glycosylated hemoglobin (HbA1c). The posted statistical analyses evaluated non-inferiority and superiority for that endpoint, together with several secondary efficacy measures.

599
Enrolled
Phase 3
2
Arms
Parallel design
26
Weeks
Primary endpoint
-0.06
HbA1c LS mean difference
95% CI -0.19 to 0.07
FeatureAWARD-6
Trial nameAWARD-6
NCT identifierNCT01624259
PhasePhase 3
ConditionType 2 Diabetes
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment599
InterventionsLY2189265 (drug); Liraglutide (drug); Metformin (drug)
Lead sponsorEli Lilly and Company
Sponsor typeIndustry
StartJune 2012
Primary completionNovember 2013
Results postedYes
Outcome measures posted23
Statistical analyses posted9

2. Clinical Question

The primary statistical question was whether 1.5 mg LY2189265 was non-inferior to 1.8 mg liraglutide for change from baseline in HbA1c at 26 weeks. A superiority analysis was also posted for the same endpoint.

Population

Participants with type 2 diabetes enrolled in the randomized phase 3 AWARD-6 trial.

Intervention

LY2189265, identified in the trial data as dulaglutide, at 1.5 mg in the primary statistical comparison.

Comparator

Liraglutide at 1.8 mg in the primary statistical comparison.

Primary question

Was the 26-week HbA1c change with 1.5 mg LY2189265 sufficiently close to that with 1.8 mg liraglutide to satisfy the prespecified non-inferiority criterion?

3. Trial Design

01
Randomize599 participants enrolled
02
Parallel armsLY2189265 vs liraglutide
03
26 weeksPrimary HbA1c endpoint
04
ModelREML-based MMRM
05
CompareNon-inferiority and superiority
ARM 1

1.5 mg LY2189265

  • LY2189265 was the drug intervention evaluated in the primary comparison.
  • The primary statistical analysis compared 1.5 mg LY2189265 with 1.8 mg liraglutide.
  • The serious adverse event count posted on ClinicalTrials.gov for this arm was 5 affected participants among 299 at risk.
ARM 2

1.8 mg Liraglutide

  • Liraglutide was the comparator drug intervention.
  • The primary statistical analysis compared 1.8 mg liraglutide with 1.5 mg LY2189265.
  • The serious adverse event count posted on ClinicalTrials.gov for this arm was 11 affected participants among 300 at risk.
Open-label design: the registry describes the masking as none. That is important when interpreting outcomes that could be affected by participant or investigator knowledge of treatment assignment. The primary HbA1c analysis nevertheless used a prespecified mixed-effects model rather than an unadjusted comparison of observed changes.

4. Registered Primary Endpoint

EndpointTime frameRegistered definition
Change From Baseline to 26 Weeks Endpoint in Glycosylated Hemoglobin (HbA1c) Baseline, 26 Weeks Least Squares means of HbA1c change from baseline to Week 26 were adjusted by fixed effects of treatment, country, visit, treatment-by-visit interaction, participant as random effect, and baseline HbA1c as covariates, using a mixed-effects model for repeated measures (MMRM) with restricted maximum likelihood (REML).

The ClinicalTrials.gov record identifies one registered primary endpoint, but two statistical analyses were posted for it: one testing non-inferiority and one testing superiority. This distinction matters because the same estimated treatment difference can support different inferential questions depending on the prespecified hypothesis and decision rule.

5. Statistical Methodology

Mixed-effects model for repeated measures

The primary endpoint was analyzed with a mixed-effects model for repeated measures. The registry definition specifies treatment, country, visit, and treatment-by-visit interaction as fixed effects, participant as a random effect, and baseline HbA1c as a covariate. The model used restricted maximum likelihood (REML).

Primary model structure
HbA1c change ~ treatment + country + visit + treatment×visit + baseline HbA1c + participant random effect

This structure allows the analysis to account for repeated measurements within participants while adjusting for baseline HbA1c and incorporating the prespecified treatment-by-visit interaction.

Least-squares mean difference

The treatment effect was reported as an LS Mean Difference. This is an adjusted difference between treatment groups derived from the fitted model rather than simply subtracting two raw arithmetic means.

For the primary analysis, the estimated difference was -0.06 percentage points of glycosylated hemoglobin, with the comparison defined as 1.5 mg LY2189265 minus 1.8 mg liraglutide.

Restricted maximum likelihood

REML is an estimation approach commonly used for variance components in mixed-effects models. In this trial, the registry specifically identifies REML as the estimation method for the MMRM analysis.

Analysis population

The primary analyses used participants who were randomized and received at least 1 dose of LY2189265 or liraglutide with evaluable HbA1c data. This is narrower than simply saying "all randomized participants," because treatment exposure and evaluable HbA1c data were part of the posted analysis-population definition.

Secondary modeling

The posted secondary analyses used three main statistical families: ANCOVA for continuous outcomes such as body weight, BMI, fasting plasma glucose, and HOMA2-%B; mixed-effects models for repeated measurements such as 7-point self-monitored plasma glucose; and logistic regression for binary HbA1c achievement thresholds.

6. Primary Results: HbA1c at 26 Weeks

Non-inferiority analysis

LS mean difference in HbA1c change

-0.06

95% CI: -0.19 to 0.07   ·   P < 0.001

Comparison: 1.5 mg LY2189265 vs 1.8 mg liraglutide

The registry analysis states that non-inferiority was assessed using a 0.4% margin. The posted design description states that non-inferiority would be demonstrated if the upper bound of the two-sided 95% confidence interval for the treatment difference was below the non-inferiority margin.

Clinical Biostats interpretation

The estimated LS mean difference of -0.06 means that the model-estimated change in HbA1c for the 1.5 mg LY2189265 group was 0.06 percentage points lower than the corresponding estimate for the 1.8 mg liraglutide group, using the stated model and comparison direction.

The estimate does not mean that every participant experienced a 0.06-point difference, nor does it establish that the two treatments produce identical HbA1c responses. It is an adjusted population-level treatment contrast.

The two-sided 95% CI of -0.19 to 0.07 describes the statistical uncertainty around that estimated difference under the analysis framework. Because the upper confidence bound is 0.07, it lies below the reported 0.4% non-inferiority margin.

The P-value of <0.001 should not be interpreted as a measure of how large the treatment effect is. A P-value addresses the compatibility of the data with a specified hypothesis; the effect estimate and confidence interval provide the more direct description of magnitude and precision.

The non-inferiority conclusion depends on the prespecified margin and analysis framework rather than simply on whether a conventional superiority P-value is below 0.05. The registry also states that family-wise Type I error was controlled using a serial gatekeeping strategy.

Superiority analysis

LS mean difference in HbA1c change

-0.06

95% CI: -0.19 to 0.07   ·   P = 0.186

Comparison: 1.5 mg LY2189265 vs 1.8 mg liraglutide

Clinical Biostats interpretation

The superiority analysis uses the same estimated LS mean difference, -0.06, because it is based on the same primary endpoint comparison and fitted model. The inferential question is different: superiority asks whether the observed difference provides sufficient evidence of a treatment difference in the specified direction.

The 95% CI of -0.19 to 0.07 includes zero. Thus, the interval is compatible with both a modestly lower and a modestly higher adjusted HbA1c change for LY2189265 relative to liraglutide within the range represented by the interval.

The superiority P-value of 0.186 does not provide evidence for a statistically significant superiority effect under the posted superiority analysis. It also does not establish equivalence or prove that the treatments have exactly the same effect.

This distinction is central to non-inferiority trials: failure to demonstrate superiority is not the same statistical statement as demonstration of non-inferiority.

7. The Non-Inferiority Margin

The registry-reported statistical analysis specifies a 0.4% non-inferiority margin for the HbA1c comparison. The design description also states that the sample-size calculation assumed a 0 difference between treatments and a common standard deviation of 1.3% for change from baseline in HbA1c. The planned design required 222 completers, or 444 total participants, at 26 weeks to provide 90% power under those assumptions.

Non-inferiority decision logic
Upper bound of 95% CI < 0.4%  →  non-inferiority criterion satisfied

The posted estimate was -0.06 with a two-sided 95% CI of -0.19 to 0.07. The relevant comparison is therefore between the confidence-limit boundary and the prespecified margin, not between the P-value and a generic significance threshold alone.

The statistical direction also matters. With the treatment difference defined as LY2189265 minus liraglutide, a positive upper confidence limit indicates the largest treatment disadvantage for LY2189265 represented by the interval. The non-inferiority margin specifies how large that disadvantage could be while still satisfying the prespecified criterion.

Non-inferiority is not equivalence: a non-inferiority conclusion does not establish that the two treatments have identical effects. It establishes that the estimated disadvantage of the investigational treatment is sufficiently limited according to the prespecified margin and analysis framework.

8. Secondary Results

The registry contains seven posted secondary statistical analyses in addition to the two primary analyses. These cover body weight, BMI, fasting plasma glucose, 7-point self-monitored plasma glucose, two HbA1c achievement thresholds, and HOMA2-%B.

Secondary endpointMethodEffect95% CIP-value
Change From Baseline in Body Weight at 26 Weeks ANCOVA LS Mean Difference 0.71 kg 0.17 to 1.26 0.010
Change From Baseline in Body Mass Index (BMI) at 26 Weeks ANCOVA LS Mean Difference 0.25 kg/m2 0.05 to 0.45 0.013
Change From Baseline in Fasting Plasma Glucose (FPG) at 26 Weeks ANCOVA LS Mean Difference -0.57 mg/dL -5.69 to 4.56 0.828
Change From Baseline in 7-Point Self Monitored Plasma Glucose (SMPG) at 26 Weeks Mixed Models Analysis LS Mean Difference -2.25 mg/dL -5.91 to 1.41 0.228
Percentage of Participants Achieving HbA1c ≤6.5% at 26 Weeks Logistic regression Odds Ratio 1.23 0.81 to 1.86 0.322
Percentage of Participants Achieving HbA1c <7% at 26 Weeks Logistic regression Odds Ratio 1.02 0.64 to 1.63 0.925
Change From Baseline in HOMA2-%B at 26 Weeks ANCOVA LS Mean Difference 1.43% -4.06 to 6.92 0.608

Body weight

The body-weight analysis estimated an LS mean difference of 0.71 kg, with a 95% CI of 0.17 to 1.26 and P = 0.010. The comparison is defined as 1.5 mg LY2189265 minus 1.8 mg liraglutide, so the positive estimate indicates a higher adjusted change in body weight for LY2189265 under the posted ANCOVA analysis.

Body mass index

The BMI analysis produced an LS mean difference of 0.25 kg/m2 with a 95% CI of 0.05 to 0.45 and P = 0.013. As with body weight, the estimate is an adjusted between-group contrast rather than a statement about the individual change experienced by every participant.

Fasting plasma glucose

The estimated LS mean difference was -0.57 mg/dL, with a 95% CI of -5.69 to 4.56 and P = 0.828. The confidence interval includes zero and spans both negative and positive values, indicating substantial uncertainty about the direction of the adjusted between-group difference.

Seven-point self-monitored plasma glucose

The mixed-model analysis estimated an LS mean difference of -2.25 mg/dL, with a 95% CI of -5.91 to 1.41 and P = 0.228. The registry specifies that the P-value came from a pairwise comparison of LS means at 26 weeks from a REML-based MMRM.

HbA1c achievement thresholds

For the HbA1c ≤6.5% threshold, logistic regression produced an odds ratio of 1.23 with a 95% CI of 0.81 to 1.86 and P = 0.322. For the HbA1c <7% threshold, the odds ratio was 1.02 with a 95% CI of 0.64 to 1.63 and P = 0.925.

An odds ratio of 1.23 does not mean that the probability of achieving the threshold was 23 percentage points higher. It means the estimated odds were 1.23 times as large in the LY2189265 group relative to the liraglutide group under the specified logistic model. Odds and probabilities are related but are not interchangeable.

HOMA2-%B

The HOMA2-%B analysis estimated an LS mean difference of 1.43%, with a 95% CI of -4.06 to 6.92 and P = 0.608. The interval crosses zero and is relatively broad compared with the point estimate, so the estimate should be interpreted primarily as a model-based treatment contrast with considerable uncertainty.

9. Missing Data, Rescue Measurements, and Analysis Definitions

Several secondary endpoint analysis populations explicitly state that only pre-rescue measurements were used. The registry also reports the use of last observation carried forward (LOCF) for missing postbaseline measurements in the body-weight, BMI, fasting-plasma-glucose, HbA1c-threshold, and HOMA2-%B analyses.

Why pre-rescue measurements matter

Measurements obtained before rescue treatment are intended to preserve the comparison under the trial's specified treatment pathway. After rescue, the observed outcome can reflect additional treatment rather than only the randomized intervention.

What LOCF does

LOCF replaces a missing later measurement with the participant's last available observation under the specified analysis rule. It is a deterministic imputation convention, not a direct observation of the missing value.

The ClinicalTrials.gov record also identify multiple-imputation or missing-data concepts in several analysis records, but the detailed statistical-analysis fields provided here do not give enough information to reconstruct a separate multiple-imputation procedure or compare it quantitatively with LOCF. Accordingly, no additional imputation results are inferred on this page.

Interpretive caution: imputation can affect both estimates and uncertainty. The appropriate interpretation depends on assumptions about why measurements are missing and on the specific estimand being targeted. The ClinicalTrials.gov record establishes that LOCF was used for several secondary analyses; it does not provide enough detail here to reconstruct the assumptions behind that procedure.

10. Multiplicity and Serial Gatekeeping

The primary non-inferiority analysis explicitly states that the family-wise Type I error rate was controlled using a serial gatekeeping strategy. This is an important design feature because the trial included more than one inferential question for the primary endpoint.

Why gatekeeping matters
Multiple planned hypotheses → prespecified testing sequence → controlled family-wise Type I error

A gatekeeping procedure defines how later hypotheses can be tested while maintaining the desired overall error rate. The registry text identifies the strategy but does not provide the complete testing sequence or all of its numerical allocation details.

By contrast, several secondary analyses explicitly state no adjustment for multiplicity. This means their nominal P-values should not automatically be interpreted as independent confirmatory evidence across the entire collection of secondary endpoints.

AnalysisMultiplicity information reportedInterpretive implication
Primary HbA1c non-inferiorityFamily-wise Type I error controlled by serial gatekeepingInterpret within the prespecified inferential framework.
Primary HbA1c superiorityPosted as a separate superiority analysisDifferent hypothesis from non-inferiority.
HbA1c achievement ≤6.5%No adjustment for multiplicityNominal P-value should not be treated as if it were independently multiplicity-adjusted.
HOMA2-%BNo adjustment for multiplicityInterpret as a secondary analysis rather than an independently protected confirmatory test.

11. Statistical Methods Explained

Why was an MMRM used for the primary endpoint?

The primary endpoint involved change in HbA1c measured over visits, and the registry specifies a mixed-effects model for repeated measures. Such a model can represent correlations among repeated observations from the same participant while estimating treatment differences through the fixed effects in the model. The treatment-by-visit interaction also allows the treatment contrast to vary across visits rather than forcing one common effect at every measurement occasion.

What does an LS mean difference of -0.06 mean?

The LS mean difference is the model-adjusted difference between the two treatment groups, defined here as 1.5 mg LY2189265 minus 1.8 mg liraglutide. A value of -0.06 indicates a model-estimated difference of -0.06 percentage points in HbA1c change. It is not an individual-level treatment effect and is not equivalent to saying that every participant's HbA1c changed by exactly that amount.

Why is non-inferiority judged against a margin?

A non-inferiority trial asks whether the new treatment is not unacceptably worse than the comparator according to a clinically specified boundary. In AWARD-6, the registry-reported margin is 0.4%. The decision therefore depends on whether the confidence interval excludes treatment disadvantages larger than that margin. This is fundamentally different from asking only whether the confidence interval excludes zero.

Why can the non-inferiority and superiority conclusions differ?

They test different null hypotheses. Superiority asks whether the treatment difference is distinguishable from zero in the specified direction. Non-inferiority asks whether the treatment disadvantage is smaller than the prespecified margin. A confidence interval can include zero while remaining entirely below a non-inferiority margin.

What does an odds ratio of 1.23 mean?

An odds ratio of 1.23 indicates that the estimated odds of achieving the specified HbA1c threshold were 1.23 times the odds in the comparator group under the logistic model. It does not mean a 23-percentage-point increase in probability. The distinction becomes especially important when the outcome is not rare.

Why is ANCOVA useful for continuous secondary endpoints?

ANCOVA can compare treatment groups while adjusting for prespecified covariates, including baseline values when included in the model. In this trial, the posted body-weight, BMI, fasting-plasma-glucose, and HOMA2-%B analyses used ANCOVA. The resulting LS mean difference therefore represents an adjusted model-based treatment comparison rather than a simple unadjusted difference between observed means.

Why does multiplicity matter?

If many hypotheses are tested, the probability of obtaining at least one small P-value by chance increases unless the testing strategy accounts for the number and relationship of the hypotheses. AWARD-6 provides both sides of this issue: the primary non-inferiority analysis specifies family-wise error control through serial gatekeeping, while some secondary analyses explicitly report no multiplicity adjustment.

12. Primary Analysis: Precision and Statistical Evidence

Estimate

The primary LS mean difference was -0.06. This is the central point estimate from the posted MMRM analysis.

Confidence interval

The two-sided 95% CI was -0.19 to 0.07. The interval describes uncertainty around the estimated treatment contrast under the model. It does not describe the range of individual participant responses.

Non-inferiority

The upper confidence bound of 0.07 is below the registry-reported non-inferiority margin of 0.4%. This is the key numerical relationship underlying the posted non-inferiority analysis.

Superiority

The superiority P-value was 0.186, and the confidence interval includes zero. The posted data therefore do not provide a statistically significant superiority result for the primary HbA1c comparison under that analysis.

13. Safety Results

The ClinicalTrials.gov record provides serious adverse events by randomized arm as affected participants divided by participants at risk.

Safety measure1.5 mg LY21892651.8 mg Liraglutide
Serious adverse events5 / 29911 / 300
Participants at risk299300
Serious adverse events: affected participants
1.5 mg LY2189265
5
1.8 mg Liraglutide
11

The serious-adverse-event counts should be interpreted as descriptive safety information. The ClinicalTrials.gov record does not provide a formal between-group hypothesis test, confidence interval, exposure-adjusted rate, or adjudication details for these events, so none is inferred here.

14. Trial Timeline

June 2012

Trial start

The registry identifies June 2012 as the study start.

November 2013

Primary completion

The registry identifies November 2013 as the primary completion date.

Completed

Trial status

The ClinicalTrials.gov record lists the trial status as COMPLETED, with results posted.

15. What the Primary Confidence Interval Does — and Does Not — Mean

Magnitude

The estimate of -0.06 describes the adjusted average treatment contrast in HbA1c change under the specified MMRM. It does not describe the magnitude of response for an individual participant.

Precision

The 95% CI of -0.19 to 0.07 indicates the range of treatment contrasts compatible with the statistical uncertainty represented by the model and sampling framework. It is not a range containing 95% of individual treatment effects.

Non-inferiority logic

The clinically important comparison is between the upper CI bound, 0.07, and the non-inferiority margin, 0.4%. Because the upper bound is below the margin, the posted non-inferiority criterion is satisfied.

P-value

The non-inferiority P-value of <0.001 and the superiority P-value of 0.186 answer different inferential questions. Neither P-value is a measure of clinical importance or effect magnitude.

16. Design Topics Relevant to Interpretation

Randomization

Allocation was randomized, which provides the design basis for comparing treatment groups without relying solely on post hoc adjustment to create comparability.

Open-label treatment

The trial was not masked. Knowledge of treatment assignment can matter particularly for subjective outcomes, treatment behavior, and other measurements susceptible to participant or investigator expectations.

Non-inferiority margin

The registry-reported primary analysis uses a 0.4% margin. That margin is part of the statistical definition of the non-inferiority question.

Repeated measurements

The primary HbA1c analysis used an MMRM with visit and treatment-by-visit interaction, recognizing the longitudinal structure of the data.

The ClinicalTrials.gov record does not provide enough information to characterize a factorial design, crossover design, Bayesian analysis, or a formal interim-analysis procedure. Those topics are therefore not attributed to AWARD-6 on this page.

17. Limitations

18. Why This Trial Matters Statistically

AWARD-6 is a useful teaching case because the primary endpoint demonstrates a central distinction in clinical-trial inference: non-inferiority and superiority are not interchangeable hypotheses. The same treatment contrast can have a confidence interval that includes zero and still satisfy a non-inferiority margin.

ConceptHow it appears in AWARD-6
RandomizationRandomized parallel-group phase 3 design.
Non-inferiority designPrimary HbA1c analysis used a 0.4% non-inferiority margin.
Mixed-effects modelPrimary HbA1c endpoint analyzed with an MMRM using REML.
Covariate adjustmentBaseline HbA1c was included as a covariate in the primary model.
Repeated measurementsVisit and treatment-by-visit interaction were included in the primary model.
Least-squares meansTreatment effect reported as an LS mean difference.
Confidence intervalsPrimary estimate accompanied by a two-sided 95% CI.
Superiority testingA separate superiority analysis was posted for the same primary endpoint.
ANCOVAUsed for several continuous secondary outcomes.
Logistic regressionUsed for binary HbA1c achievement thresholds.
Missing-data methodsLOCF was used for several secondary analyses; multiple-imputation/missing-data concepts also appear in the posted analysis descriptions.
MultiplicitySerial gatekeeping controlled family-wise Type I error for the primary non-inferiority framework; some secondary analyses had no multiplicity adjustment.

19. Related Tutorials

Learn more about the methods used in this trial:

20. Related Statistical Calculators

Use these calculator pathways to explore the statistical concepts underlying the trial:

21. Sources

Continue through Clinical Biostats

Explore the statistical concepts behind randomized trials, longitudinal models, non-inferiority designs, categorical outcomes, and clinical-trial inference.

22. Record Summary

AWARD-6 provides a clear example of how a clinical trial can produce a non-inferiority conclusion without demonstrating superiority. The primary HbA1c analysis estimated an LS mean difference of -0.06 with a two-sided 95% CI of -0.19 to 0.07. The registry-reported statistical design used a 0.4% non-inferiority margin, and the registry reports non-inferiority with a P-value of <0.001. The separate superiority analysis produced the same estimate and confidence interval but a P-value of 0.186.

The secondary analyses extend the statistical picture through ANCOVA, mixed-effects models, and logistic regression. Their interpretation requires attention to confidence intervals, missing-data procedures, pre-rescue measurements, and multiplicity. The result is a useful illustration of why a clinical-trial statistical analysis should examine not only whether a P-value is small, but also which hypothesis was tested, which estimand was estimated, what margin was prespecified, how uncertainty was quantified, and how the analysis handled repeated and missing observations.

Clinical Biostats methodology: A trial-results page should not merely repeat a reported result. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and avoiding conclusions that are not supported by the ClinicalTrials.gov record.