This page separates reported trial results from statistical interpretation. Numerical results are restricted to the statistical analyses posted in the ClinicalTrials.gov record. The registry record is the official source for the trial information.
1. Trial at a Glance
AWARD-6 was a randomized, parallel, open-label phase 3 trial comparing LY2189265 (dulaglutide) with liraglutide in type 2 diabetes. The registered primary endpoint was change from baseline to 26 weeks in glycosylated hemoglobin (HbA1c). The posted statistical analyses evaluated non-inferiority and superiority for that endpoint, together with several secondary efficacy measures.
| Feature | AWARD-6 |
|---|---|
| Trial name | AWARD-6 |
| NCT identifier | NCT01624259 |
| Phase | Phase 3 |
| Condition | Type 2 Diabetes |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 599 |
| Interventions | LY2189265 (drug); Liraglutide (drug); Metformin (drug) |
| Lead sponsor | Eli Lilly and Company |
| Sponsor type | Industry |
| Start | June 2012 |
| Primary completion | November 2013 |
| Results posted | Yes |
| Outcome measures posted | 23 |
| Statistical analyses posted | 9 |
2. Clinical Question
The primary statistical question was whether 1.5 mg LY2189265 was non-inferior to 1.8 mg liraglutide for change from baseline in HbA1c at 26 weeks. A superiority analysis was also posted for the same endpoint.
Population
Participants with type 2 diabetes enrolled in the randomized phase 3 AWARD-6 trial.
Intervention
LY2189265, identified in the trial data as dulaglutide, at 1.5 mg in the primary statistical comparison.
Comparator
Liraglutide at 1.8 mg in the primary statistical comparison.
Primary question
Was the 26-week HbA1c change with 1.5 mg LY2189265 sufficiently close to that with 1.8 mg liraglutide to satisfy the prespecified non-inferiority criterion?
3. Trial Design
1.5 mg LY2189265
- LY2189265 was the drug intervention evaluated in the primary comparison.
- The primary statistical analysis compared 1.5 mg LY2189265 with 1.8 mg liraglutide.
- The serious adverse event count posted on ClinicalTrials.gov for this arm was 5 affected participants among 299 at risk.
1.8 mg Liraglutide
- Liraglutide was the comparator drug intervention.
- The primary statistical analysis compared 1.8 mg liraglutide with 1.5 mg LY2189265.
- The serious adverse event count posted on ClinicalTrials.gov for this arm was 11 affected participants among 300 at risk.
4. Registered Primary Endpoint
| Endpoint | Time frame | Registered definition |
|---|---|---|
| Change From Baseline to 26 Weeks Endpoint in Glycosylated Hemoglobin (HbA1c) | Baseline, 26 Weeks | Least Squares means of HbA1c change from baseline to Week 26 were adjusted by fixed effects of treatment, country, visit, treatment-by-visit interaction, participant as random effect, and baseline HbA1c as covariates, using a mixed-effects model for repeated measures (MMRM) with restricted maximum likelihood (REML). |
The ClinicalTrials.gov record identifies one registered primary endpoint, but two statistical analyses were posted for it: one testing non-inferiority and one testing superiority. This distinction matters because the same estimated treatment difference can support different inferential questions depending on the prespecified hypothesis and decision rule.
5. Statistical Methodology
Mixed-effects model for repeated measures
The primary endpoint was analyzed with a mixed-effects model for repeated measures. The registry definition specifies treatment, country, visit, and treatment-by-visit interaction as fixed effects, participant as a random effect, and baseline HbA1c as a covariate. The model used restricted maximum likelihood (REML).
This structure allows the analysis to account for repeated measurements within participants while adjusting for baseline HbA1c and incorporating the prespecified treatment-by-visit interaction.
Least-squares mean difference
The treatment effect was reported as an LS Mean Difference. This is an adjusted difference between treatment groups derived from the fitted model rather than simply subtracting two raw arithmetic means.
For the primary analysis, the estimated difference was -0.06 percentage points of glycosylated hemoglobin, with the comparison defined as 1.5 mg LY2189265 minus 1.8 mg liraglutide.
Restricted maximum likelihood
REML is an estimation approach commonly used for variance components in mixed-effects models. In this trial, the registry specifically identifies REML as the estimation method for the MMRM analysis.
Analysis population
The primary analyses used participants who were randomized and received at least 1 dose of LY2189265 or liraglutide with evaluable HbA1c data. This is narrower than simply saying "all randomized participants," because treatment exposure and evaluable HbA1c data were part of the posted analysis-population definition.
Secondary modeling
The posted secondary analyses used three main statistical families: ANCOVA for continuous outcomes such as body weight, BMI, fasting plasma glucose, and HOMA2-%B; mixed-effects models for repeated measurements such as 7-point self-monitored plasma glucose; and logistic regression for binary HbA1c achievement thresholds.
6. Primary Results: HbA1c at 26 Weeks
Non-inferiority analysis
LS mean difference in HbA1c change
95% CI: -0.19 to 0.07 · P < 0.001
Comparison: 1.5 mg LY2189265 vs 1.8 mg liraglutide
The registry analysis states that non-inferiority was assessed using a 0.4% margin. The posted design description states that non-inferiority would be demonstrated if the upper bound of the two-sided 95% confidence interval for the treatment difference was below the non-inferiority margin.
The estimated LS mean difference of -0.06 means that the model-estimated change in HbA1c for the 1.5 mg LY2189265 group was 0.06 percentage points lower than the corresponding estimate for the 1.8 mg liraglutide group, using the stated model and comparison direction.
The estimate does not mean that every participant experienced a 0.06-point difference, nor does it establish that the two treatments produce identical HbA1c responses. It is an adjusted population-level treatment contrast.
The two-sided 95% CI of -0.19 to 0.07 describes the statistical uncertainty around that estimated difference under the analysis framework. Because the upper confidence bound is 0.07, it lies below the reported 0.4% non-inferiority margin.
The P-value of <0.001 should not be interpreted as a measure of how large the treatment effect is. A P-value addresses the compatibility of the data with a specified hypothesis; the effect estimate and confidence interval provide the more direct description of magnitude and precision.
The non-inferiority conclusion depends on the prespecified margin and analysis framework rather than simply on whether a conventional superiority P-value is below 0.05. The registry also states that family-wise Type I error was controlled using a serial gatekeeping strategy.
Superiority analysis
LS mean difference in HbA1c change
95% CI: -0.19 to 0.07 · P = 0.186
Comparison: 1.5 mg LY2189265 vs 1.8 mg liraglutide
The superiority analysis uses the same estimated LS mean difference, -0.06, because it is based on the same primary endpoint comparison and fitted model. The inferential question is different: superiority asks whether the observed difference provides sufficient evidence of a treatment difference in the specified direction.
The 95% CI of -0.19 to 0.07 includes zero. Thus, the interval is compatible with both a modestly lower and a modestly higher adjusted HbA1c change for LY2189265 relative to liraglutide within the range represented by the interval.
The superiority P-value of 0.186 does not provide evidence for a statistically significant superiority effect under the posted superiority analysis. It also does not establish equivalence or prove that the treatments have exactly the same effect.
This distinction is central to non-inferiority trials: failure to demonstrate superiority is not the same statistical statement as demonstration of non-inferiority.
7. The Non-Inferiority Margin
The registry-reported statistical analysis specifies a 0.4% non-inferiority margin for the HbA1c comparison. The design description also states that the sample-size calculation assumed a 0 difference between treatments and a common standard deviation of 1.3% for change from baseline in HbA1c. The planned design required 222 completers, or 444 total participants, at 26 weeks to provide 90% power under those assumptions.
The posted estimate was -0.06 with a two-sided 95% CI of -0.19 to 0.07. The relevant comparison is therefore between the confidence-limit boundary and the prespecified margin, not between the P-value and a generic significance threshold alone.
The statistical direction also matters. With the treatment difference defined as LY2189265 minus liraglutide, a positive upper confidence limit indicates the largest treatment disadvantage for LY2189265 represented by the interval. The non-inferiority margin specifies how large that disadvantage could be while still satisfying the prespecified criterion.
8. Secondary Results
The registry contains seven posted secondary statistical analyses in addition to the two primary analyses. These cover body weight, BMI, fasting plasma glucose, 7-point self-monitored plasma glucose, two HbA1c achievement thresholds, and HOMA2-%B.
| Secondary endpoint | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|
| Change From Baseline in Body Weight at 26 Weeks | ANCOVA | LS Mean Difference 0.71 kg | 0.17 to 1.26 | 0.010 |
| Change From Baseline in Body Mass Index (BMI) at 26 Weeks | ANCOVA | LS Mean Difference 0.25 kg/m2 | 0.05 to 0.45 | 0.013 |
| Change From Baseline in Fasting Plasma Glucose (FPG) at 26 Weeks | ANCOVA | LS Mean Difference -0.57 mg/dL | -5.69 to 4.56 | 0.828 |
| Change From Baseline in 7-Point Self Monitored Plasma Glucose (SMPG) at 26 Weeks | Mixed Models Analysis | LS Mean Difference -2.25 mg/dL | -5.91 to 1.41 | 0.228 |
| Percentage of Participants Achieving HbA1c ≤6.5% at 26 Weeks | Logistic regression | Odds Ratio 1.23 | 0.81 to 1.86 | 0.322 |
| Percentage of Participants Achieving HbA1c <7% at 26 Weeks | Logistic regression | Odds Ratio 1.02 | 0.64 to 1.63 | 0.925 |
| Change From Baseline in HOMA2-%B at 26 Weeks | ANCOVA | LS Mean Difference 1.43% | -4.06 to 6.92 | 0.608 |
Body weight
The body-weight analysis estimated an LS mean difference of 0.71 kg, with a 95% CI of 0.17 to 1.26 and P = 0.010. The comparison is defined as 1.5 mg LY2189265 minus 1.8 mg liraglutide, so the positive estimate indicates a higher adjusted change in body weight for LY2189265 under the posted ANCOVA analysis.
Body mass index
The BMI analysis produced an LS mean difference of 0.25 kg/m2 with a 95% CI of 0.05 to 0.45 and P = 0.013. As with body weight, the estimate is an adjusted between-group contrast rather than a statement about the individual change experienced by every participant.
Fasting plasma glucose
The estimated LS mean difference was -0.57 mg/dL, with a 95% CI of -5.69 to 4.56 and P = 0.828. The confidence interval includes zero and spans both negative and positive values, indicating substantial uncertainty about the direction of the adjusted between-group difference.
Seven-point self-monitored plasma glucose
The mixed-model analysis estimated an LS mean difference of -2.25 mg/dL, with a 95% CI of -5.91 to 1.41 and P = 0.228. The registry specifies that the P-value came from a pairwise comparison of LS means at 26 weeks from a REML-based MMRM.
HbA1c achievement thresholds
For the HbA1c ≤6.5% threshold, logistic regression produced an odds ratio of 1.23 with a 95% CI of 0.81 to 1.86 and P = 0.322. For the HbA1c <7% threshold, the odds ratio was 1.02 with a 95% CI of 0.64 to 1.63 and P = 0.925.
An odds ratio of 1.23 does not mean that the probability of achieving the threshold was 23 percentage points higher. It means the estimated odds were 1.23 times as large in the LY2189265 group relative to the liraglutide group under the specified logistic model. Odds and probabilities are related but are not interchangeable.
HOMA2-%B
The HOMA2-%B analysis estimated an LS mean difference of 1.43%, with a 95% CI of -4.06 to 6.92 and P = 0.608. The interval crosses zero and is relatively broad compared with the point estimate, so the estimate should be interpreted primarily as a model-based treatment contrast with considerable uncertainty.
9. Missing Data, Rescue Measurements, and Analysis Definitions
Several secondary endpoint analysis populations explicitly state that only pre-rescue measurements were used. The registry also reports the use of last observation carried forward (LOCF) for missing postbaseline measurements in the body-weight, BMI, fasting-plasma-glucose, HbA1c-threshold, and HOMA2-%B analyses.
Why pre-rescue measurements matter
Measurements obtained before rescue treatment are intended to preserve the comparison under the trial's specified treatment pathway. After rescue, the observed outcome can reflect additional treatment rather than only the randomized intervention.
What LOCF does
LOCF replaces a missing later measurement with the participant's last available observation under the specified analysis rule. It is a deterministic imputation convention, not a direct observation of the missing value.
The ClinicalTrials.gov record also identify multiple-imputation or missing-data concepts in several analysis records, but the detailed statistical-analysis fields provided here do not give enough information to reconstruct a separate multiple-imputation procedure or compare it quantitatively with LOCF. Accordingly, no additional imputation results are inferred on this page.
10. Multiplicity and Serial Gatekeeping
The primary non-inferiority analysis explicitly states that the family-wise Type I error rate was controlled using a serial gatekeeping strategy. This is an important design feature because the trial included more than one inferential question for the primary endpoint.
A gatekeeping procedure defines how later hypotheses can be tested while maintaining the desired overall error rate. The registry text identifies the strategy but does not provide the complete testing sequence or all of its numerical allocation details.
By contrast, several secondary analyses explicitly state no adjustment for multiplicity. This means their nominal P-values should not automatically be interpreted as independent confirmatory evidence across the entire collection of secondary endpoints.
| Analysis | Multiplicity information reported | Interpretive implication |
|---|---|---|
| Primary HbA1c non-inferiority | Family-wise Type I error controlled by serial gatekeeping | Interpret within the prespecified inferential framework. |
| Primary HbA1c superiority | Posted as a separate superiority analysis | Different hypothesis from non-inferiority. |
| HbA1c achievement ≤6.5% | No adjustment for multiplicity | Nominal P-value should not be treated as if it were independently multiplicity-adjusted. |
| HOMA2-%B | No adjustment for multiplicity | Interpret as a secondary analysis rather than an independently protected confirmatory test. |
11. Statistical Methods Explained
Why was an MMRM used for the primary endpoint?
The primary endpoint involved change in HbA1c measured over visits, and the registry specifies a mixed-effects model for repeated measures. Such a model can represent correlations among repeated observations from the same participant while estimating treatment differences through the fixed effects in the model. The treatment-by-visit interaction also allows the treatment contrast to vary across visits rather than forcing one common effect at every measurement occasion.
What does an LS mean difference of -0.06 mean?
The LS mean difference is the model-adjusted difference between the two treatment groups, defined here as 1.5 mg LY2189265 minus 1.8 mg liraglutide. A value of -0.06 indicates a model-estimated difference of -0.06 percentage points in HbA1c change. It is not an individual-level treatment effect and is not equivalent to saying that every participant's HbA1c changed by exactly that amount.
Why is non-inferiority judged against a margin?
A non-inferiority trial asks whether the new treatment is not unacceptably worse than the comparator according to a clinically specified boundary. In AWARD-6, the registry-reported margin is 0.4%. The decision therefore depends on whether the confidence interval excludes treatment disadvantages larger than that margin. This is fundamentally different from asking only whether the confidence interval excludes zero.
Why can the non-inferiority and superiority conclusions differ?
They test different null hypotheses. Superiority asks whether the treatment difference is distinguishable from zero in the specified direction. Non-inferiority asks whether the treatment disadvantage is smaller than the prespecified margin. A confidence interval can include zero while remaining entirely below a non-inferiority margin.
What does an odds ratio of 1.23 mean?
An odds ratio of 1.23 indicates that the estimated odds of achieving the specified HbA1c threshold were 1.23 times the odds in the comparator group under the logistic model. It does not mean a 23-percentage-point increase in probability. The distinction becomes especially important when the outcome is not rare.
Why is ANCOVA useful for continuous secondary endpoints?
ANCOVA can compare treatment groups while adjusting for prespecified covariates, including baseline values when included in the model. In this trial, the posted body-weight, BMI, fasting-plasma-glucose, and HOMA2-%B analyses used ANCOVA. The resulting LS mean difference therefore represents an adjusted model-based treatment comparison rather than a simple unadjusted difference between observed means.
Why does multiplicity matter?
If many hypotheses are tested, the probability of obtaining at least one small P-value by chance increases unless the testing strategy accounts for the number and relationship of the hypotheses. AWARD-6 provides both sides of this issue: the primary non-inferiority analysis specifies family-wise error control through serial gatekeeping, while some secondary analyses explicitly report no multiplicity adjustment.
12. Primary Analysis: Precision and Statistical Evidence
The primary LS mean difference was -0.06. This is the central point estimate from the posted MMRM analysis.
The two-sided 95% CI was -0.19 to 0.07. The interval describes uncertainty around the estimated treatment contrast under the model. It does not describe the range of individual participant responses.
The upper confidence bound of 0.07 is below the registry-reported non-inferiority margin of 0.4%. This is the key numerical relationship underlying the posted non-inferiority analysis.
The superiority P-value was 0.186, and the confidence interval includes zero. The posted data therefore do not provide a statistically significant superiority result for the primary HbA1c comparison under that analysis.
13. Safety Results
The ClinicalTrials.gov record provides serious adverse events by randomized arm as affected participants divided by participants at risk.
| Safety measure | 1.5 mg LY2189265 | 1.8 mg Liraglutide |
|---|---|---|
| Serious adverse events | 5 / 299 | 11 / 300 |
| Participants at risk | 299 | 300 |
The serious-adverse-event counts should be interpreted as descriptive safety information. The ClinicalTrials.gov record does not provide a formal between-group hypothesis test, confidence interval, exposure-adjusted rate, or adjudication details for these events, so none is inferred here.
14. Trial Timeline
Trial start
The registry identifies June 2012 as the study start.
Primary completion
The registry identifies November 2013 as the primary completion date.
Trial status
The ClinicalTrials.gov record lists the trial status as COMPLETED, with results posted.
15. What the Primary Confidence Interval Does — and Does Not — Mean
The estimate of -0.06 describes the adjusted average treatment contrast in HbA1c change under the specified MMRM. It does not describe the magnitude of response for an individual participant.
The 95% CI of -0.19 to 0.07 indicates the range of treatment contrasts compatible with the statistical uncertainty represented by the model and sampling framework. It is not a range containing 95% of individual treatment effects.
The clinically important comparison is between the upper CI bound, 0.07, and the non-inferiority margin, 0.4%. Because the upper bound is below the margin, the posted non-inferiority criterion is satisfied.
The non-inferiority P-value of <0.001 and the superiority P-value of 0.186 answer different inferential questions. Neither P-value is a measure of clinical importance or effect magnitude.
16. Design Topics Relevant to Interpretation
Randomization
Allocation was randomized, which provides the design basis for comparing treatment groups without relying solely on post hoc adjustment to create comparability.
Open-label treatment
The trial was not masked. Knowledge of treatment assignment can matter particularly for subjective outcomes, treatment behavior, and other measurements susceptible to participant or investigator expectations.
Non-inferiority margin
The registry-reported primary analysis uses a 0.4% margin. That margin is part of the statistical definition of the non-inferiority question.
Repeated measurements
The primary HbA1c analysis used an MMRM with visit and treatment-by-visit interaction, recognizing the longitudinal structure of the data.
The ClinicalTrials.gov record does not provide enough information to characterize a factorial design, crossover design, Bayesian analysis, or a formal interim-analysis procedure. Those topics are therefore not attributed to AWARD-6 on this page.
17. Limitations
- Open-label design: the registry specifies no masking. This can introduce behavioral or assessment-related influences for outcomes susceptible to treatment awareness.
- Analysis-population restriction: the primary analysis population required randomization, at least one dose of LY2189265 or liraglutide, and evaluable HbA1c data. It therefore is not identical in wording to an unrestricted all-randomized analysis population.
- Non-inferiority depends on the margin: the conclusion is meaningful in relation to the prespecified 0.4% margin. A different margin would define a different non-inferiority question.
- Secondary multiplicity: some secondary analyses explicitly report no multiplicity adjustment, so their nominal P-values should not be interpreted as a family-wise protected collection of confirmatory tests.
- LOCF: several secondary analyses used last observation carried forward for missing postbaseline values. Imputation assumptions can influence estimated treatment effects.
- Secondary endpoint precision: some confidence intervals span both negative and positive values, indicating uncertainty about the direction of the treatment contrast.
- Limited registry detail: the ClinicalTrials.gov record does not provide the complete statistical analysis plan, all model coefficients, detailed covariance structures, or full missing-data assumptions. Those details should not be reconstructed from the summary fields.
- Safety inference: the registry-reported serious-adverse-event data are descriptive counts by arm and do not include a formal between-group statistical comparison.
18. Why This Trial Matters Statistically
AWARD-6 is a useful teaching case because the primary endpoint demonstrates a central distinction in clinical-trial inference: non-inferiority and superiority are not interchangeable hypotheses. The same treatment contrast can have a confidence interval that includes zero and still satisfy a non-inferiority margin.
| Concept | How it appears in AWARD-6 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design. |
| Non-inferiority design | Primary HbA1c analysis used a 0.4% non-inferiority margin. |
| Mixed-effects model | Primary HbA1c endpoint analyzed with an MMRM using REML. |
| Covariate adjustment | Baseline HbA1c was included as a covariate in the primary model. |
| Repeated measurements | Visit and treatment-by-visit interaction were included in the primary model. |
| Least-squares means | Treatment effect reported as an LS mean difference. |
| Confidence intervals | Primary estimate accompanied by a two-sided 95% CI. |
| Superiority testing | A separate superiority analysis was posted for the same primary endpoint. |
| ANCOVA | Used for several continuous secondary outcomes. |
| Logistic regression | Used for binary HbA1c achievement thresholds. |
| Missing-data methods | LOCF was used for several secondary analyses; multiple-imputation/missing-data concepts also appear in the posted analysis descriptions. |
| Multiplicity | Serial gatekeeping controlled family-wise Type I error for the primary non-inferiority framework; some secondary analyses had no multiplicity adjustment. |
19. Related Tutorials
Learn more about the methods used in this trial:
20. Related Statistical Calculators
Use these calculator pathways to explore the statistical concepts underlying the trial:
21. Sources
- ClinicalTrials.gov: AWARD-6, NCT01624259.
- Linked PubMed record: PMID 27161178.
- Linked PubMed record: PMID 26691396.
- Linked PubMed record: PMID 25018121.
Continue through Clinical Biostats
Explore the statistical concepts behind randomized trials, longitudinal models, non-inferiority designs, categorical outcomes, and clinical-trial inference.
22. Record Summary
AWARD-6 provides a clear example of how a clinical trial can produce a non-inferiority conclusion without demonstrating superiority. The primary HbA1c analysis estimated an LS mean difference of -0.06 with a two-sided 95% CI of -0.19 to 0.07. The registry-reported statistical design used a 0.4% non-inferiority margin, and the registry reports non-inferiority with a P-value of <0.001. The separate superiority analysis produced the same estimate and confidence interval but a P-value of 0.186.
The secondary analyses extend the statistical picture through ANCOVA, mixed-effects models, and logistic regression. Their interpretation requires attention to confidence intervals, missing-data procedures, pre-rescue measurements, and multiplicity. The result is a useful illustration of why a clinical-trial statistical analysis should examine not only whether a P-value is small, but also which hypothesis was tested, which estimand was estimated, what margin was prespecified, how uncertainty was quantified, and how the analysis handled repeated and missing observations.