← Clinical Trials
Type 2 Diabetes Phase 3 Cardiovascular Outcomes NCT01243424

CAROLINA: Complete Statistical Analysis of Linagliptin in Type 2 Diabetes

An independent statistical analysis of the randomized, double-masked phase 3 CAROLINA trial comparing linagliptin with glimepiride in patients with type 2 diabetes, with emphasis on the non-inferiority analysis of first 3-point major adverse cardiovascular events and the subsequent hierarchical statistical analyses.

CAROLINA  ·  Phase 3  ·  Completed  ·  Enrollment 6103
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data posted on ClinicalTrials.gov for CAROLINA.

1. Trial at a Glance

CAROLINA was a randomized, double-masked, parallel-group phase 3 trial evaluating linagliptin versus glimepiride in patients with type 2 diabetes. The primary registered endpoint was the first 3-point major adverse cardiovascular event (3P-MACE), analyzed as a time-to-event outcome using a Cox proportional-hazards model.

6103
Enrollment
Randomized trial
2
Treatment arms
Linagliptin vs glimepiride
0.98
Primary HR
95.47% CI 0.84–1.14
<0.0001
NI p-value
Primary analysis
FeatureCAROLINA
Trial nameCAROLINA: Cardiovascular Outcome Study of Linagliptin Versus Glimepiride in Patients With Type 2 Diabetes
PhasePhase 3
StatusCOMPLETED
Therapeutic areaEndocrinology
ConditionDiabetes Mellitus, Type 2
AllocationRANDOMIZED
Design modelPARALLEL
MaskingDOUBLE
Primary purposeTREATMENT
Enrollment6103.0
ClinicalTrials.govNCT01243424
Lead sponsorBoehringer Ingelheim
Sponsor typeINDUSTRY

2. Clinical Question

The primary clinical question was whether linagliptin could be shown to be non-inferior to glimepiride with respect to the time to the first 3-point major adverse cardiovascular event, defined using Clinical Event Committee (CEC)-confirmed adjudicated events.

Population

Patients with Diabetes Mellitus, Type 2 enrolled in the CAROLINA phase 3 randomized trial.

Intervention

Linagliptin, with linagliptin placebo and glimepiride placebo included among the registered drug interventions supporting the masked comparison.

Comparator

Glimepiride, with placebo components used as part of the double-masked trial design.

Primary question

Is the hazard of the first 3-point MACE with linagliptin sufficiently similar to glimepiride to satisfy the prespecified non-inferiority criterion?

3. Trial Design

01
Randomize6103 participants
02
TreatLinagliptin or glimepiride
03
FollowCardiovascular and clinical outcomes
04
AdjudicateCEC-confirmed events
05
AnalyzeHierarchical testing
Allocation
Randomized, parallel-group design.
Masking
Double masking was specified in the registry.
Primary endpoint
First 3-point Major Adverse Cardiovascular Events (3P-MACE).
Statistical framework
Cox proportional-hazards, logistic regression, and ANCOVA, with non-inferiority, superiority, equivalence, and other hypothesis types represented in the posted analyses.
ARM A

Linagliptin

  • Linagliptin was the active treatment comparison.
  • Linagliptin placebo was included among the registered interventions.
ARM B

Glimepiride

  • Glimepiride was the active comparator.
  • Glimepiride placebo was included among the registered interventions.
Why masking matters statistically. A double-masked design reduces the opportunity for knowledge of treatment assignment to influence participant behavior, treatment administration, outcome assessment, and other aspects of trial conduct. For a cardiovascular time-to-event endpoint involving adjudicated events, this is particularly relevant because systematic knowledge of treatment assignment could otherwise introduce avoidable sources of bias.

4. Timeline and Follow-up

2010-11-11

Trial start

CAROLINA began on November 11, 2010.

2018-08-21

Primary completion

The registered primary completion date was August 21, 2018.

Up to 432 weeks

Primary cardiovascular follow-up

The first 3-point MACE endpoint was assessed from randomization until the individual day of trial completion, up to 432 weeks.

Up to 433 weeks

Selected secondary follow-up

Some secondary outcomes used follow-up extending to 433 weeks, including the time-to-first occurrence composite endpoint and the accelerated cognitive decline endpoint.

5. Primary Endpoint

EndpointDefinition / time framePrimary analysis
The First 3-point Major Adverse Cardiovascular Events (3P-MACE) From randomization until individual day of trial completion, up to 432 weeks. Cox proportional-hazards model comparing linagliptin with glimepiride.

The registered endpoint is a composite time-to-event outcome. The first occurrence of any CEC-confirmed adjudicated component counts as the event: CV death, including fatal stroke and fatal myocardial infarction (MI); non-fatal MI, excluding silent MI; or nonfatal stroke.

Analysis population. The primary analysis used the treated set (TS), defined as all patients treated with at least one dose of trial drug.

6. Statistical Methodology

MethodRole in the posted analysesEffect measure
Cox proportional-hazards modelPrimary 3P-MACE and other time-to-event endpointsHazard ratio
Logistic regressionBinary composite and cognitive outcomesOdds ratio
ANCOVAContinuous outcomes including HbA1c, FPG, lipids, creatinine, eGFR, UACR, and insulin secretion rateMean difference or ratio of geometric means

The statistical profile of CAROLINA is notable because the registry contains a sequence of analyses rather than a single isolated comparison. The primary endpoint was evaluated first for non-inferiority and then for superiority. Subsequent analyses are identified in the registry as later steps in a pre-defined hierarchical testing approach.

The primary Cox model

The registry states that a Cox proportional-hazard model with treatment as a factor was applied to compare linagliptin with glimepiride.

HR = hazard with linagliptin ÷ hazard with glimepiride

An HR of 1 would indicate equal estimated hazards. An HR below 1 indicates a lower estimated hazard in the linagliptin group, whereas an HR above 1 indicates a higher estimated hazard.

7. Primary Results: 3-point MACE

Non-inferiority analysis

Hazard ratio for first 3-point MACE

0.98

95.47% CI: 0.84–1.14   ·   P < 0.0001

Two-sided confidence interval. Treated set analysis. Hypothesis type: non-inferiority.

Clinical Biostats interpretation

The estimated hazard ratio of 0.98 means that, in this analysis, the estimated hazard of the first 3-point MACE with linagliptin was approximately 2% lower than the estimated hazard with glimepiride. This is a relative hazard comparison, not a statement that 2% fewer participants experienced an event.

The 95.47% confidence interval of 0.84 to 1.14 describes the precision of the estimated hazard ratio under the analysis framework. It spans values below and above 1, so it includes both a lower and a higher estimated hazard for linagliptin relative to glimepiride.

The P < 0.0001 value belongs to the non-inferiority hypothesis test. It does not measure the size of the treatment effect and should not be interpreted as the probability that the treatments are equivalent or as the probability that the null hypothesis is true.

Most importantly, the registry states that this was the first step in a pre-defined hierarchical testing approach. The upper confidence bound of the HR was compared with the prespecified non-inferiority margin. The registry extract does not provide the numerical value of that margin, so the margin itself is not reproduced here.

Because this is a time-to-event analysis, the interpretation also depends on the Cox model framework and its assumptions. The hazard ratio is not a direct ratio of cumulative event probabilities, and censoring and the proportional-hazards assumption are relevant to interpretation.

Superiority analysis

Hazard ratio for first 3-point MACE

0.98

95.47% CI: 0.84–1.14   ·   P = 0.3813

Two-sided confidence interval. Treated set analysis. Hypothesis type: superiority.

Clinical Biostats interpretation

The superiority analysis uses the same estimated hazard ratio, 0.98, and the same 95.47% confidence interval of 0.84 to 1.14. The registry reports a superiority p-value of 0.3813.

The p-value is a measure of compatibility with the superiority hypothesis under the specified testing framework; it is not an effect-size measure. The HR and confidence interval provide the information about the magnitude and precision of the estimated relative hazard.

The registry explicitly identifies this as the second step in the pre-defined hierarchical testing approach. Therefore, the result should be understood within that sequential testing framework rather than treated as an isolated hypothesis test.

The confidence interval remains important: it extends from 0.84 to 1.14, encompassing a range of plausible relative hazard values on both sides of 1. As with the non-inferiority analysis, the HR describes relative instantaneous event rates under the Cox model rather than a simple difference in cumulative risks.

Non-inferiority is not the same as superiority. A successful non-inferiority test asks whether the new treatment's effect is not unacceptably worse than the comparator according to a prespecified margin. A superiority test asks whether the data provide evidence that the treatments differ in the specified direction. CAROLINA's registry explicitly records these as sequential first and second steps of a hierarchical testing approach.

8. Secondary Results: Hierarchical Testing Pathway

The registry identifies several subsequent analyses as steps in the pre-defined hierarchical testing approach. These analyses cover cardiovascular outcomes, composite glycaemic/weight/hypoglycaemia outcomes, glycaemic measures, laboratory outcomes, insulin secretion, and cognition.

4-point MACE

Hazard ratio for first 4-point MACE

0.99

95.47% CI: 0.86–1.14   ·   P = 0.4334

Two-sided confidence interval. Treated set. Hypothesis type: superiority.

The first 4-point MACE was analyzed as a time-to-event endpoint using a Cox proportional-hazards model with treatment as a factor. The registry identifies this as the third step in the pre-defined hierarchical testing approach.

Clinical Biostats interpretation

The HR of 0.99 is close to 1, indicating that the estimated hazards were very similar under this analysis. The 95.47% CI of 0.86 to 1.14 quantifies the uncertainty around that estimate. The p-value of 0.4334 is a hypothesis-test result and should not be confused with a measure of effect magnitude.

Glycaemic control without rescue, excessive weight gain, and moderate/severe hypoglycaemia

Odds ratio

1.68

95.47% CI: 1.43–1.96   ·   P < 0.0001

Logistic regression; two-sided confidence interval; TS without duplicates with non-completers considered failure.

This binary endpoint measured the percentage of participants taking trial medication at trial end who maintained HbA1c ≤7.0% without rescue medication, without >2% weight gain, and without moderate/severe hypoglycaemic episodes during the maintenance phase from Visit 6 (Week 16) to the final visit (Week 432).

Clinical Biostats interpretation

An odds ratio of 1.68 means the estimated odds of meeting the composite criterion were 1.68 times as high for linagliptin relative to glimepiride under the logistic model. Odds are not the same quantity as probabilities or risk ratios.

The 95.47% CI of 1.43 to 1.96 lies above 1. The p-value of <0.0001 addresses the statistical hypothesis associated with the comparison; it does not say that there is a 0.01% or smaller probability that the observed effect is due to chance, nor does it quantify clinical importance.

The analysis also has an important population definition: patients who were off-drug or died before regular study stop were handled as non-completers and considered failures. That rule directly affects the binary outcome being analyzed.

Glycaemic control without rescue medication and without >2% weight gain

Odds ratio

1.29

95.47% CI: 1.11–1.48   ·   P = 0.0004

Logistic regression; two-sided confidence interval; TS without duplicates with non-completers considered failure.

This endpoint used the same maintenance-phase framework but did not include the requirement concerning moderate/severe hypoglycaemic episodes. The registry identifies it as the fifth step in the pre-defined hierarchical testing approach and notes multiplicity adjustment in the analysis text.

Clinical Biostats interpretation

The OR of 1.29 indicates higher estimated odds of meeting this composite criterion with linagliptin under the fitted logistic model. The 95.47% CI of 1.11 to 1.48 provides the stated precision around the odds-ratio estimate.

The p-value of 0.0004 is evidence from the specified hypothesis test; it is not a measure of how large or clinically important the difference is. The hierarchical sequence and multiplicity adjustment matter because multiple related outcomes were examined.

9. Additional Cardiovascular and Clinical Outcomes

OutcomeEstimate95% CIP-valueMethod
Time to first occurrence of any components of the composite endpoint of all adjudication-confirmed events HR 0.96 0.85–1.09 0.5249 Cox proportional-hazards model
Accelerated cognitive decline at end of follow-up OR 1.01 0.86–1.18 0.9112 Logistic regression

The adjudication-confirmed time-to-event endpoint was measured from the start of treatment until 7 days after the end of treatment, up to 433 weeks. The cognitive endpoint was assessed at 433 weeks. The registry identifies multiplicity adjustment in both analysis pathways and intention-to-treat analysis among the concepts associated with the cognitive analysis.

Clinical Biostats interpretation

For the adjudication-confirmed time-to-event endpoint, the HR of 0.96 is close to 1, with a 95% CI of 0.85 to 1.09. For accelerated cognitive decline, the OR of 1.01 is also close to 1, with a 95% CI of 0.86 to 1.18. These are estimates of different statistical quantities and should not be directly compared as though they represented the same endpoint scale.

10. Glycaemic Outcomes

HbA1c

Mean difference in final HbA1c

-0.09

95% CI: -0.15 to -0.03   ·   P = 0.0023

Mean difference = Linagliptin mean − Glimepiride mean.

The analysis evaluated change from baseline to the final visit in HbA1c at baseline and Week 432. The ANCOVA model included treatment as a fixed categorical effect and baseline HbA1c as a continuous covariate. The analysis used the TS without duplicates considering all available data.

Clinical Biostats interpretation

The negative mean difference means that the final HbA1c mean was estimated to be 0.09 percentage points lower with linagliptin than with glimepiride under the stated ANCOVA model. The 95% CI extends from -0.15 to -0.03.

The important statistical feature is the use of baseline covariate adjustment. ANCOVA compares treatment groups while accounting for the continuous baseline HbA1c measurement in the model, rather than simply comparing raw final means.

Fasting Plasma Glucose

Mean difference in final FPG

-7.3

95% CI: -9.7 to -4.8 mg/dL   ·   P < 0.0001

Mean difference = Linagliptin mean − Glimepiride mean.

Fasting plasma glucose was assessed at baseline and Week 432. The ANCOVA model included treatment as a fixed categorical effect and baseline FPG as a continuous covariate.

Clinical Biostats interpretation

The estimate of -7.3 mg/dL indicates a lower estimated final FPG value for linagliptin relative to glimepiride under the stated model. The confidence interval, -9.7 to -4.8 mg/dL, gives the registry-reported precision of that estimate.

As with HbA1c, the p-value is not an effect-size measure. The numerical effect estimate and confidence interval are the appropriate quantities for describing magnitude and uncertainty.

11. Lipid and Renal Outcomes

OutcomeEstimate95% CIP-valueMethod
LDL cholesterolMean difference 0.4 mg/dL-1.3 to 2.10.6400ANCOVA
HDL cholesterolMean difference 0.5 mg/dL0.0 to 1.00.0497ANCOVA
Total cholesterolMean difference -0.4 mg/dL-2.4 to 1.60.6823ANCOVA
TriglyceridesMean difference -3.5 mg/dL-9.6 to 2.70.2678ANCOVA
CreatinineMean difference -0.01 mg/dL-0.03 to 0.010.5165ANCOVA
eGFRMean difference 1.0 mL/minute/1.73 meter20.2 to 1.80.5165ANCOVA
UACRRatio of geometric means 0.970.91 to 1.030.2921ANCOVA

These outcomes were generally analyzed using ANCOVA with treatment as a fixed categorical effect and the relevant baseline measurement as a continuous covariate. The registry identifies multiplicity adjustment in these analysis records.

Why a geometric-mean ratio for UACR?

The UACR analysis used a ratio of geometric means rather than an arithmetic mean difference. The registry defines the ratio as:

gMean ratio = Linagliptin mean ÷ Glimepiride mean

An estimate of 0.97 therefore represents a multiplicative comparison: the estimated geometric mean for linagliptin was 0.97 times that for glimepiride under the stated analysis. Its 95% CI was 0.91 to 1.03.

12. Insulin Secretion and Equivalence Testing

Mean difference in insulin secretion rate

4.13

95% CI: -36.46 to 44.71   ·   P = 0.8402

ANCOVA; equivalence hypothesis; mean difference = Linagliptin mean − Glimepiride mean.

The insulin secretion rate (ISR) outcome was measured at a fixed glucose concentration at 208 weeks. The analysis used the Meal Tolerance Test last-observation-carried-forward set, comprising randomized and treated patients with one dose of study drug and signed sub-study informed consent with the required assessments described in the registry.

Clinical Biostats interpretation

The estimated net mean difference was 4.13 pmol/min/m², with a 95% CI of -36.46 to 44.71. The p-value was 0.8402.

This analysis is explicitly labelled equivalence in the registry. Equivalence is not established merely because a conventional p-value is large. Equivalence requires a prespecified equivalence interval or margins and demonstration that the confidence interval lies within those bounds. The ClinicalTrials.gov record identifies the hypothesis type but does not provide the numerical equivalence margins, so no additional equivalence conclusion is added here.

13. Statistical Methods Explained

Why was a Cox proportional-hazards model used for 3P-MACE?

3P-MACE is a time-to-event endpoint: participants can experience the first event at different times, and some observations may be censored. A Cox model uses the timing information rather than reducing follow-up to a simple yes/no event indicator. Its principal treatment effect measure is the hazard ratio.

What does an HR of 0.98 mean?

For the primary endpoint, an HR of 0.98 means the estimated hazard under linagliptin was 0.98 times the estimated hazard under glimepiride. Equivalently, the estimated relative hazard was about 2% lower. This does not mean that the probability of having a MACE was exactly 2% lower.

Why is non-inferiority judged against a margin?

Non-inferiority asks whether any disadvantage of the new treatment is sufficiently small according to a prespecified clinical/statistical margin. The key comparison is therefore between the confidence bound and that margin, rather than simply asking whether a conventional superiority p-value is below 0.05. CAROLINA's registry explicitly states that the upper confidence bound was compared with the non-inferiority margin.

Why was ANCOVA used for HbA1c and other continuous outcomes?

ANCOVA provides a model-based comparison of treatment groups while adjusting for a prespecified baseline continuous covariate. In CAROLINA, the registry states that baseline HbA1c, FPG, LDL cholesterol, HDL cholesterol, total cholesterol, creatinine, eGFR, UACR, or ISR was included as the corresponding continuous baseline covariate in the relevant analyses.

What does an odds ratio of 1.68 mean?

An OR of 1.68 means that the estimated odds of meeting the specified binary composite criterion were 1.68 times as high with linagliptin as with glimepiride under the logistic regression model. Odds are different from probabilities, so an OR should not automatically be described as a 68% increase in probability.

Why does the hierarchy matter?

The registry identifies a pre-defined hierarchical testing approach. The primary 3P-MACE analysis was the first step for non-inferiority and the second step for superiority. Later endpoints are also identified as subsequent steps. A hierarchy can control how confirmatory claims are sequenced across multiple hypotheses rather than treating every p-value as an independent primary test.

14. Multiplicity, Missing Data, and Analysis Populations

The posted analysis records explicitly identify multiplicity adjustment for several secondary analyses. This is statistically important because CAROLINA contains numerous secondary endpoints and multiple hypothesis-testing steps. Without accounting for multiplicity, interpreting every nominal p-value as though it came from a single prespecified hypothesis could overstate the strength of evidence.

The analysis populations were not identical across endpoints. The primary 3P-MACE analysis used the treated set, defined as all patients treated with at least one dose of trial drug. Several continuous and binary analyses used the treated set without duplicates. The insulin secretion analysis used a specialized Meal Tolerance Test LOCF set, while the cognitive analysis used a Full Analysis Set cognition population.

The registry also explicitly identifies multiple imputation / missing data as a concept for the insulin secretion analysis. For the binary maintenance-phase outcomes, patients who were off-drug or died before regular study stop were handled as non-completers and considered failures. These rules are part of the estimand being analyzed: changing the missing-data or intercurrent-event strategy can change the meaning of the resulting treatment comparison.

15. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.

TreatmentSerious adverse eventsAt risk
Linagliptin14033023
Glimepiride14483010
Serious adverse events by arm
Linagliptin
1403/3023
Glimepiride
1448/3010

The ClinicalTrials.gov record provides affected and at-risk counts but do not provide a formal statistical comparison for these serious adverse events. Accordingly, this page does not add an unreported p-value, confidence interval, risk ratio, or odds ratio.

16. Limitations

Registry-derived scope

This analysis is restricted to the ClinicalTrials.gov record. It does not import numerical results from publications or other sources.

Non-inferiority margin

The registry states that the primary HR confidence bound was compared with a prespecified margin, but the ClinicalTrials.gov record does not provide the numerical margin.

Different analysis populations

Different endpoints use different analysis sets, including TS, TS without duplicates, a cognitive full analysis set, and a specialized ISR LOCF set.

Multiple endpoints

The registry reports 20 outcome measures and 17 statistical analyses. Several analyses are explicitly identified as later steps in a hierarchical testing approach.

A second limitation is that the registry extract reported here does not provide every design detail that might be needed for a complete independent reconstruction of the statistical analysis plan. In particular, the numerical non-inferiority and equivalence margins are not available in the ClinicalTrials.gov record. They are therefore not inferred.

The primary endpoint is a composite cardiovascular endpoint. Composite outcomes can be statistically efficient because several event types contribute information, but the component events may differ in clinical frequency and importance. The ClinicalTrials.gov record defines the components but do not provide component-specific estimates, so no component-level interpretation is added.

17. Why This Trial Matters Statistically

CAROLINA is statistically instructive because it combines several core ideas in clinical-trial methodology within one randomized comparison.

First, the primary cardiovascular outcome is a classic time-to-event endpoint. The Cox model allows the analysis to incorporate both whether and when the first qualifying event occurred, while the hazard ratio provides a relative measure of the instantaneous event rate under the model.

Second, the primary objective illustrates the distinction between non-inferiority and superiority. The registry records non-inferiority as the first testing step and superiority as the second. That sequence is fundamentally different from simply asking whether one treatment has a statistically significant advantage over another.

Third, the trial demonstrates how hierarchical testing and multiplicity adjustment can structure a large family of endpoints. Several later analyses are explicitly described as steps in the pre-defined hierarchy. This matters because a collection of individually small p-values cannot automatically be interpreted as though each were the sole prespecified hypothesis.

Fourth, CAROLINA illustrates why analysis population definitions matter. The treated set is used for the primary cardiovascular analysis, while other endpoints use different populations and rules for missing observations or intercurrent events. The numerical result is inseparable from the population and analysis rule that generated it.

Finally, the trial provides examples of three major regression frameworks: Cox regression for time-to-event outcomes, logistic regression for binary outcomes, and ANCOVA for continuous outcomes. Understanding which estimand and outcome structure each model addresses is more informative than treating all reported p-values as interchangeable measures of evidence.

18. Key Statistical Takeaways

Primary endpoint

First 3-point MACE was analyzed with a Cox proportional-hazards model in the treated set.

Non-inferiority

The primary HR was 0.98 with a 95.47% CI of 0.84–1.14; the registry states that the upper bound was compared with a prespecified non-inferiority margin.

Superiority

The same HR of 0.98 had a superiority p-value of 0.3813 in the second step of the hierarchy.

Model diversity

Cox regression, logistic regression, and ANCOVA were all used, matched to different endpoint structures.

Covariate adjustment

ANCOVA analyses included the relevant baseline measurement as a continuous covariate.

Multiplicity

Several secondary analyses explicitly identify multiplicity adjustment and hierarchical testing.

Explore the statistical methods behind CAROLINA

Use the related tutorials and calculators to study the trial's core concepts, including Cox models, non-inferiority, ANCOVA, odds ratios, confidence intervals, and multiplicity.