This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data posted on ClinicalTrials.gov for CAROLINA.
1. Trial at a Glance
CAROLINA was a randomized, double-masked, parallel-group phase 3 trial evaluating linagliptin versus glimepiride in patients with type 2 diabetes. The primary registered endpoint was the first 3-point major adverse cardiovascular event (3P-MACE), analyzed as a time-to-event outcome using a Cox proportional-hazards model.
| Feature | CAROLINA |
|---|---|
| Trial name | CAROLINA: Cardiovascular Outcome Study of Linagliptin Versus Glimepiride in Patients With Type 2 Diabetes |
| Phase | Phase 3 |
| Status | COMPLETED |
| Therapeutic area | Endocrinology |
| Condition | Diabetes Mellitus, Type 2 |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | DOUBLE |
| Primary purpose | TREATMENT |
| Enrollment | 6103.0 |
| ClinicalTrials.gov | NCT01243424 |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | INDUSTRY |
2. Clinical Question
The primary clinical question was whether linagliptin could be shown to be non-inferior to glimepiride with respect to the time to the first 3-point major adverse cardiovascular event, defined using Clinical Event Committee (CEC)-confirmed adjudicated events.
Population
Patients with Diabetes Mellitus, Type 2 enrolled in the CAROLINA phase 3 randomized trial.
Intervention
Linagliptin, with linagliptin placebo and glimepiride placebo included among the registered drug interventions supporting the masked comparison.
Comparator
Glimepiride, with placebo components used as part of the double-masked trial design.
Primary question
Is the hazard of the first 3-point MACE with linagliptin sufficiently similar to glimepiride to satisfy the prespecified non-inferiority criterion?
3. Trial Design
Linagliptin
- Linagliptin was the active treatment comparison.
- Linagliptin placebo was included among the registered interventions.
Glimepiride
- Glimepiride was the active comparator.
- Glimepiride placebo was included among the registered interventions.
4. Timeline and Follow-up
Trial start
CAROLINA began on November 11, 2010.
Primary completion
The registered primary completion date was August 21, 2018.
Primary cardiovascular follow-up
The first 3-point MACE endpoint was assessed from randomization until the individual day of trial completion, up to 432 weeks.
Selected secondary follow-up
Some secondary outcomes used follow-up extending to 433 weeks, including the time-to-first occurrence composite endpoint and the accelerated cognitive decline endpoint.
5. Primary Endpoint
| Endpoint | Definition / time frame | Primary analysis |
|---|---|---|
| The First 3-point Major Adverse Cardiovascular Events (3P-MACE) | From randomization until individual day of trial completion, up to 432 weeks. | Cox proportional-hazards model comparing linagliptin with glimepiride. |
The registered endpoint is a composite time-to-event outcome. The first occurrence of any CEC-confirmed adjudicated component counts as the event: CV death, including fatal stroke and fatal myocardial infarction (MI); non-fatal MI, excluding silent MI; or nonfatal stroke.
6. Statistical Methodology
| Method | Role in the posted analyses | Effect measure |
|---|---|---|
| Cox proportional-hazards model | Primary 3P-MACE and other time-to-event endpoints | Hazard ratio |
| Logistic regression | Binary composite and cognitive outcomes | Odds ratio |
| ANCOVA | Continuous outcomes including HbA1c, FPG, lipids, creatinine, eGFR, UACR, and insulin secretion rate | Mean difference or ratio of geometric means |
The statistical profile of CAROLINA is notable because the registry contains a sequence of analyses rather than a single isolated comparison. The primary endpoint was evaluated first for non-inferiority and then for superiority. Subsequent analyses are identified in the registry as later steps in a pre-defined hierarchical testing approach.
The primary Cox model
The registry states that a Cox proportional-hazard model with treatment as a factor was applied to compare linagliptin with glimepiride.
An HR of 1 would indicate equal estimated hazards. An HR below 1 indicates a lower estimated hazard in the linagliptin group, whereas an HR above 1 indicates a higher estimated hazard.
7. Primary Results: 3-point MACE
Non-inferiority analysis
Hazard ratio for first 3-point MACE
95.47% CI: 0.84–1.14 · P < 0.0001
Two-sided confidence interval. Treated set analysis. Hypothesis type: non-inferiority.
The estimated hazard ratio of 0.98 means that, in this analysis, the estimated hazard of the first 3-point MACE with linagliptin was approximately 2% lower than the estimated hazard with glimepiride. This is a relative hazard comparison, not a statement that 2% fewer participants experienced an event.
The 95.47% confidence interval of 0.84 to 1.14 describes the precision of the estimated hazard ratio under the analysis framework. It spans values below and above 1, so it includes both a lower and a higher estimated hazard for linagliptin relative to glimepiride.
The P < 0.0001 value belongs to the non-inferiority hypothesis test. It does not measure the size of the treatment effect and should not be interpreted as the probability that the treatments are equivalent or as the probability that the null hypothesis is true.
Most importantly, the registry states that this was the first step in a pre-defined hierarchical testing approach. The upper confidence bound of the HR was compared with the prespecified non-inferiority margin. The registry extract does not provide the numerical value of that margin, so the margin itself is not reproduced here.
Because this is a time-to-event analysis, the interpretation also depends on the Cox model framework and its assumptions. The hazard ratio is not a direct ratio of cumulative event probabilities, and censoring and the proportional-hazards assumption are relevant to interpretation.
Superiority analysis
Hazard ratio for first 3-point MACE
95.47% CI: 0.84–1.14 · P = 0.3813
Two-sided confidence interval. Treated set analysis. Hypothesis type: superiority.
The superiority analysis uses the same estimated hazard ratio, 0.98, and the same 95.47% confidence interval of 0.84 to 1.14. The registry reports a superiority p-value of 0.3813.
The p-value is a measure of compatibility with the superiority hypothesis under the specified testing framework; it is not an effect-size measure. The HR and confidence interval provide the information about the magnitude and precision of the estimated relative hazard.
The registry explicitly identifies this as the second step in the pre-defined hierarchical testing approach. Therefore, the result should be understood within that sequential testing framework rather than treated as an isolated hypothesis test.
The confidence interval remains important: it extends from 0.84 to 1.14, encompassing a range of plausible relative hazard values on both sides of 1. As with the non-inferiority analysis, the HR describes relative instantaneous event rates under the Cox model rather than a simple difference in cumulative risks.
8. Secondary Results: Hierarchical Testing Pathway
The registry identifies several subsequent analyses as steps in the pre-defined hierarchical testing approach. These analyses cover cardiovascular outcomes, composite glycaemic/weight/hypoglycaemia outcomes, glycaemic measures, laboratory outcomes, insulin secretion, and cognition.
4-point MACE
Hazard ratio for first 4-point MACE
95.47% CI: 0.86–1.14 · P = 0.4334
Two-sided confidence interval. Treated set. Hypothesis type: superiority.
The first 4-point MACE was analyzed as a time-to-event endpoint using a Cox proportional-hazards model with treatment as a factor. The registry identifies this as the third step in the pre-defined hierarchical testing approach.
The HR of 0.99 is close to 1, indicating that the estimated hazards were very similar under this analysis. The 95.47% CI of 0.86 to 1.14 quantifies the uncertainty around that estimate. The p-value of 0.4334 is a hypothesis-test result and should not be confused with a measure of effect magnitude.
Glycaemic control without rescue, excessive weight gain, and moderate/severe hypoglycaemia
Odds ratio
95.47% CI: 1.43–1.96 · P < 0.0001
Logistic regression; two-sided confidence interval; TS without duplicates with non-completers considered failure.
This binary endpoint measured the percentage of participants taking trial medication at trial end who maintained HbA1c ≤7.0% without rescue medication, without >2% weight gain, and without moderate/severe hypoglycaemic episodes during the maintenance phase from Visit 6 (Week 16) to the final visit (Week 432).
An odds ratio of 1.68 means the estimated odds of meeting the composite criterion were 1.68 times as high for linagliptin relative to glimepiride under the logistic model. Odds are not the same quantity as probabilities or risk ratios.
The 95.47% CI of 1.43 to 1.96 lies above 1. The p-value of <0.0001 addresses the statistical hypothesis associated with the comparison; it does not say that there is a 0.01% or smaller probability that the observed effect is due to chance, nor does it quantify clinical importance.
The analysis also has an important population definition: patients who were off-drug or died before regular study stop were handled as non-completers and considered failures. That rule directly affects the binary outcome being analyzed.
Glycaemic control without rescue medication and without >2% weight gain
Odds ratio
95.47% CI: 1.11–1.48 · P = 0.0004
Logistic regression; two-sided confidence interval; TS without duplicates with non-completers considered failure.
This endpoint used the same maintenance-phase framework but did not include the requirement concerning moderate/severe hypoglycaemic episodes. The registry identifies it as the fifth step in the pre-defined hierarchical testing approach and notes multiplicity adjustment in the analysis text.
The OR of 1.29 indicates higher estimated odds of meeting this composite criterion with linagliptin under the fitted logistic model. The 95.47% CI of 1.11 to 1.48 provides the stated precision around the odds-ratio estimate.
The p-value of 0.0004 is evidence from the specified hypothesis test; it is not a measure of how large or clinically important the difference is. The hierarchical sequence and multiplicity adjustment matter because multiple related outcomes were examined.
9. Additional Cardiovascular and Clinical Outcomes
| Outcome | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Time to first occurrence of any components of the composite endpoint of all adjudication-confirmed events | HR 0.96 | 0.85–1.09 | 0.5249 | Cox proportional-hazards model |
| Accelerated cognitive decline at end of follow-up | OR 1.01 | 0.86–1.18 | 0.9112 | Logistic regression |
The adjudication-confirmed time-to-event endpoint was measured from the start of treatment until 7 days after the end of treatment, up to 433 weeks. The cognitive endpoint was assessed at 433 weeks. The registry identifies multiplicity adjustment in both analysis pathways and intention-to-treat analysis among the concepts associated with the cognitive analysis.
For the adjudication-confirmed time-to-event endpoint, the HR of 0.96 is close to 1, with a 95% CI of 0.85 to 1.09. For accelerated cognitive decline, the OR of 1.01 is also close to 1, with a 95% CI of 0.86 to 1.18. These are estimates of different statistical quantities and should not be directly compared as though they represented the same endpoint scale.
10. Glycaemic Outcomes
HbA1c
Mean difference in final HbA1c
95% CI: -0.15 to -0.03 · P = 0.0023
Mean difference = Linagliptin mean − Glimepiride mean.
The analysis evaluated change from baseline to the final visit in HbA1c at baseline and Week 432. The ANCOVA model included treatment as a fixed categorical effect and baseline HbA1c as a continuous covariate. The analysis used the TS without duplicates considering all available data.
The negative mean difference means that the final HbA1c mean was estimated to be 0.09 percentage points lower with linagliptin than with glimepiride under the stated ANCOVA model. The 95% CI extends from -0.15 to -0.03.
The important statistical feature is the use of baseline covariate adjustment. ANCOVA compares treatment groups while accounting for the continuous baseline HbA1c measurement in the model, rather than simply comparing raw final means.
Fasting Plasma Glucose
Mean difference in final FPG
95% CI: -9.7 to -4.8 mg/dL · P < 0.0001
Mean difference = Linagliptin mean − Glimepiride mean.
Fasting plasma glucose was assessed at baseline and Week 432. The ANCOVA model included treatment as a fixed categorical effect and baseline FPG as a continuous covariate.
The estimate of -7.3 mg/dL indicates a lower estimated final FPG value for linagliptin relative to glimepiride under the stated model. The confidence interval, -9.7 to -4.8 mg/dL, gives the registry-reported precision of that estimate.
As with HbA1c, the p-value is not an effect-size measure. The numerical effect estimate and confidence interval are the appropriate quantities for describing magnitude and uncertainty.
11. Lipid and Renal Outcomes
| Outcome | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| LDL cholesterol | Mean difference 0.4 mg/dL | -1.3 to 2.1 | 0.6400 | ANCOVA |
| HDL cholesterol | Mean difference 0.5 mg/dL | 0.0 to 1.0 | 0.0497 | ANCOVA |
| Total cholesterol | Mean difference -0.4 mg/dL | -2.4 to 1.6 | 0.6823 | ANCOVA |
| Triglycerides | Mean difference -3.5 mg/dL | -9.6 to 2.7 | 0.2678 | ANCOVA |
| Creatinine | Mean difference -0.01 mg/dL | -0.03 to 0.01 | 0.5165 | ANCOVA |
| eGFR | Mean difference 1.0 mL/minute/1.73 meter2 | 0.2 to 1.8 | 0.5165 | ANCOVA |
| UACR | Ratio of geometric means 0.97 | 0.91 to 1.03 | 0.2921 | ANCOVA |
These outcomes were generally analyzed using ANCOVA with treatment as a fixed categorical effect and the relevant baseline measurement as a continuous covariate. The registry identifies multiplicity adjustment in these analysis records.
Why a geometric-mean ratio for UACR?
The UACR analysis used a ratio of geometric means rather than an arithmetic mean difference. The registry defines the ratio as:
An estimate of 0.97 therefore represents a multiplicative comparison: the estimated geometric mean for linagliptin was 0.97 times that for glimepiride under the stated analysis. Its 95% CI was 0.91 to 1.03.
12. Insulin Secretion and Equivalence Testing
Mean difference in insulin secretion rate
95% CI: -36.46 to 44.71 · P = 0.8402
ANCOVA; equivalence hypothesis; mean difference = Linagliptin mean − Glimepiride mean.
The insulin secretion rate (ISR) outcome was measured at a fixed glucose concentration at 208 weeks. The analysis used the Meal Tolerance Test last-observation-carried-forward set, comprising randomized and treated patients with one dose of study drug and signed sub-study informed consent with the required assessments described in the registry.
The estimated net mean difference was 4.13 pmol/min/m², with a 95% CI of -36.46 to 44.71. The p-value was 0.8402.
This analysis is explicitly labelled equivalence in the registry. Equivalence is not established merely because a conventional p-value is large. Equivalence requires a prespecified equivalence interval or margins and demonstration that the confidence interval lies within those bounds. The ClinicalTrials.gov record identifies the hypothesis type but does not provide the numerical equivalence margins, so no additional equivalence conclusion is added here.
13. Statistical Methods Explained
Why was a Cox proportional-hazards model used for 3P-MACE?
3P-MACE is a time-to-event endpoint: participants can experience the first event at different times, and some observations may be censored. A Cox model uses the timing information rather than reducing follow-up to a simple yes/no event indicator. Its principal treatment effect measure is the hazard ratio.
What does an HR of 0.98 mean?
For the primary endpoint, an HR of 0.98 means the estimated hazard under linagliptin was 0.98 times the estimated hazard under glimepiride. Equivalently, the estimated relative hazard was about 2% lower. This does not mean that the probability of having a MACE was exactly 2% lower.
Why is non-inferiority judged against a margin?
Non-inferiority asks whether any disadvantage of the new treatment is sufficiently small according to a prespecified clinical/statistical margin. The key comparison is therefore between the confidence bound and that margin, rather than simply asking whether a conventional superiority p-value is below 0.05. CAROLINA's registry explicitly states that the upper confidence bound was compared with the non-inferiority margin.
Why was ANCOVA used for HbA1c and other continuous outcomes?
ANCOVA provides a model-based comparison of treatment groups while adjusting for a prespecified baseline continuous covariate. In CAROLINA, the registry states that baseline HbA1c, FPG, LDL cholesterol, HDL cholesterol, total cholesterol, creatinine, eGFR, UACR, or ISR was included as the corresponding continuous baseline covariate in the relevant analyses.
What does an odds ratio of 1.68 mean?
An OR of 1.68 means that the estimated odds of meeting the specified binary composite criterion were 1.68 times as high with linagliptin as with glimepiride under the logistic regression model. Odds are different from probabilities, so an OR should not automatically be described as a 68% increase in probability.
Why does the hierarchy matter?
The registry identifies a pre-defined hierarchical testing approach. The primary 3P-MACE analysis was the first step for non-inferiority and the second step for superiority. Later endpoints are also identified as subsequent steps. A hierarchy can control how confirmatory claims are sequenced across multiple hypotheses rather than treating every p-value as an independent primary test.
14. Multiplicity, Missing Data, and Analysis Populations
The posted analysis records explicitly identify multiplicity adjustment for several secondary analyses. This is statistically important because CAROLINA contains numerous secondary endpoints and multiple hypothesis-testing steps. Without accounting for multiplicity, interpreting every nominal p-value as though it came from a single prespecified hypothesis could overstate the strength of evidence.
The analysis populations were not identical across endpoints. The primary 3P-MACE analysis used the treated set, defined as all patients treated with at least one dose of trial drug. Several continuous and binary analyses used the treated set without duplicates. The insulin secretion analysis used a specialized Meal Tolerance Test LOCF set, while the cognitive analysis used a Full Analysis Set cognition population.
The registry also explicitly identifies multiple imputation / missing data as a concept for the insulin secretion analysis. For the binary maintenance-phase outcomes, patients who were off-drug or died before regular study stop were handled as non-completers and considered failures. These rules are part of the estimand being analyzed: changing the missing-data or intercurrent-event strategy can change the meaning of the resulting treatment comparison.
15. Safety
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.
| Treatment | Serious adverse events | At risk |
|---|---|---|
| Linagliptin | 1403 | 3023 |
| Glimepiride | 1448 | 3010 |
The ClinicalTrials.gov record provides affected and at-risk counts but do not provide a formal statistical comparison for these serious adverse events. Accordingly, this page does not add an unreported p-value, confidence interval, risk ratio, or odds ratio.
16. Limitations
Registry-derived scope
This analysis is restricted to the ClinicalTrials.gov record. It does not import numerical results from publications or other sources.
Non-inferiority margin
The registry states that the primary HR confidence bound was compared with a prespecified margin, but the ClinicalTrials.gov record does not provide the numerical margin.
Different analysis populations
Different endpoints use different analysis sets, including TS, TS without duplicates, a cognitive full analysis set, and a specialized ISR LOCF set.
Multiple endpoints
The registry reports 20 outcome measures and 17 statistical analyses. Several analyses are explicitly identified as later steps in a hierarchical testing approach.
A second limitation is that the registry extract reported here does not provide every design detail that might be needed for a complete independent reconstruction of the statistical analysis plan. In particular, the numerical non-inferiority and equivalence margins are not available in the ClinicalTrials.gov record. They are therefore not inferred.
The primary endpoint is a composite cardiovascular endpoint. Composite outcomes can be statistically efficient because several event types contribute information, but the component events may differ in clinical frequency and importance. The ClinicalTrials.gov record defines the components but do not provide component-specific estimates, so no component-level interpretation is added.
17. Why This Trial Matters Statistically
CAROLINA is statistically instructive because it combines several core ideas in clinical-trial methodology within one randomized comparison.
First, the primary cardiovascular outcome is a classic time-to-event endpoint. The Cox model allows the analysis to incorporate both whether and when the first qualifying event occurred, while the hazard ratio provides a relative measure of the instantaneous event rate under the model.
Second, the primary objective illustrates the distinction between non-inferiority and superiority. The registry records non-inferiority as the first testing step and superiority as the second. That sequence is fundamentally different from simply asking whether one treatment has a statistically significant advantage over another.
Third, the trial demonstrates how hierarchical testing and multiplicity adjustment can structure a large family of endpoints. Several later analyses are explicitly described as steps in the pre-defined hierarchy. This matters because a collection of individually small p-values cannot automatically be interpreted as though each were the sole prespecified hypothesis.
Fourth, CAROLINA illustrates why analysis population definitions matter. The treated set is used for the primary cardiovascular analysis, while other endpoints use different populations and rules for missing observations or intercurrent events. The numerical result is inseparable from the population and analysis rule that generated it.
Finally, the trial provides examples of three major regression frameworks: Cox regression for time-to-event outcomes, logistic regression for binary outcomes, and ANCOVA for continuous outcomes. Understanding which estimand and outcome structure each model addresses is more informative than treating all reported p-values as interchangeable measures of evidence.
18. Key Statistical Takeaways
Primary endpoint
First 3-point MACE was analyzed with a Cox proportional-hazards model in the treated set.
Non-inferiority
The primary HR was 0.98 with a 95.47% CI of 0.84–1.14; the registry states that the upper bound was compared with a prespecified non-inferiority margin.
Superiority
The same HR of 0.98 had a superiority p-value of 0.3813 in the second step of the hierarchy.
Model diversity
Cox regression, logistic regression, and ANCOVA were all used, matched to different endpoint structures.
Covariate adjustment
ANCOVA analyses included the relevant baseline measurement as a continuous covariate.
Multiplicity
Several secondary analyses explicitly identify multiplicity adjustment and hierarchical testing.
Explore the statistical methods behind CAROLINA
Use the related tutorials and calculators to study the trial's core concepts, including Cox models, non-inferiority, ANCOVA, odds ratios, confidence intervals, and multiplicity.