← Clinical Trials
Type 2 Diabetes Mellitus Phase 3 Non-Inferiority NCT01986881

VERTIS CV: Complete Statistical Analysis of Ertugliflozin in Type 2 Diabetes With Vascular Disease

A statistical review of the randomized, double-blind, placebo-controlled phase 3 VERTIS CV trial of ertugliflozin 5 mg and 15 mg in adults with type 2 diabetes mellitus and vascular disease: a cardiovascular non-inferiority question answered with a stratified Cox model, glycemic sub-studies analysed with constrained longitudinal data analysis, and a large family of secondary cardiovascular and renal endpoints.

Registry: NCT01986881  ·  Start: 2013-11-04  ·  Primary completion: 2019-12-27  ·  Status: Completed
About this analysis

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

VERTIS CV (MK-8835-004) was a randomized, double-blind, parallel-group phase 3 trial sponsored by Merck Sharp & Dohme LLC. Its central statistical task was to show that ertugliflozin did not increase the hazard of major adverse cardiovascular events (MACE) by more than a prespecified margin relative to placebo in people with type 2 diabetes mellitus and vascular disease. Embedded glycemic sub-studies separately tested whether ertugliflozin lowered hemoglobin A1C over 18 weeks when added to specific background therapies.

8,246
Enrollment
Three randomized arms
0.97
MACE HR
All ertugliflozin vs placebo; 95.6% CI 0.848–1.114
1.3
NI margin
Hazard ratio for MACE
<0.001
NI p-value
One-sided, pooled doses
FeatureVERTIS CV
PhasePhase 3
ConditionType 2 diabetes mellitus in participants with vascular disease
DesignRandomized, parallel-group, double-masked, placebo-controlled
ArmsErtugliflozin 5 mg; ertugliflozin 15 mg; placebo (glycemic rescue medication available)
Enrollment8,246
Primary cardiovascular endpointTime to first MACE (CV death, non-fatal MI, non-fatal stroke), on-treatment + 365-day approach
Primary glycemic endpointsChange from baseline in A1C at Week 18 (excluding rescue) in three add-on glycemic sub-studies
Primary hypothesis (MACE)Non-inferiority, hazard-ratio margin 1.3
Study datesStart 2013-11-04; primary completion 2019-12-27
ClinicalTrials.govNCT01986881
SponsorMerck Sharp & Dohme LLC (industry)

2. Clinical Question

The trial asked a safety-oriented efficacy question typical of cardiovascular outcomes trials for glucose-lowering drugs: can an increase in cardiovascular risk of a clinically unacceptable size be ruled out? The primary cardiovascular hypothesis was non-inferiority, not superiority. Separately, the glycemic sub-studies asked the more conventional superiority question of whether ertugliflozin lowered A1C compared with placebo.

Population

Adults with type 2 diabetes mellitus and vascular disease, with glycemic sub-studies in participants on insulin with or without metformin, sulfonylurea monotherapy, or metformin with a sulfonylurea.

Intervention

Ertugliflozin 5 mg or ertugliflozin 15 mg, analysed separately and pooled as "All Ertugliflozin" for the primary cardiovascular comparison.

Comparator

Matching placebo. Participants in any arm who met glycemic rescue criteria received glycemic rescue medication.

Primary question

Is the hazard ratio for first MACE with ertugliflozin versus placebo credibly below 1.3? And, within each glycemic sub-study, does ertugliflozin reduce A1C at Week 18 more than placebo?

3. Trial Design

01
Randomize8,246 enrolled, 3 arms
02
Double-blind treatmentErtugliflozin 5 mg, 15 mg or placebo
03
Week 18A1C endpoints in glycemic sub-studies
04
Event follow-upAdjudicated CV and renal events
05
Final analysisUp to approximately 6 years
ARM 1 · SAE population n = 2746

Ertugliflozin 5 mg

  • Ertugliflozin 5 mg, double-blind
  • Compared with placebo individually
  • Contributes to the pooled "All Ertugliflozin" group
ARM 2 · SAE population n = 2747

Ertugliflozin 15 mg

  • Ertugliflozin 15 mg, double-blind
  • Compared with placebo individually
  • Contributes to the pooled "All Ertugliflozin" group
ARM 3 · SAE population n = 2745

Placebo

  • Matching placebo, double-blind
  • Common control for both doses
ALL ARMS

Glycemic rescue

  • Rescue medication for participants meeting glycemic rescue criteria
  • Post-rescue A1C data excluded from the primary glycemic analyses

The arm sizes above are the numbers at risk in the overall cardiovascular study's serious-adverse-event tables, which count treated participants. The ClinicalTrials.gov record does not report the randomization ratio or the randomization stratification scheme, but the Cox analyses state that "cohort category" was used as a stratification factor.

Three arms, two kinds of comparison. Because the two ertugliflozin doses share one placebo group, the registry reports each dose versus placebo and a pooled "All Ertugliflozin" versus placebo comparison. The pooled comparison is the one that carries the primary non-inferiority claim; the dose-specific hazard ratios describe consistency across doses and are less precise because each uses roughly two-thirds as many participants.

4. Analysis Populations and Event-Counting Approaches

A distinctive feature of VERTIS CV is that different endpoints use different analysis populations and different rules for which events count. Readers comparing hazard ratios across endpoints need to keep these differences in view.

AnalysisPopulation (registry definition)Event window
Primary MACE (non-inferiority)All randomized participants who received at least 1 dose of study medicationOn-treatment + 365-day approach: events between first dose and the on-treatment censor date
Secondary CV, renal and mortality endpointsAll randomized participantsOn-study approach: events between randomization and the on-study censor date
Glycemic sub-study A1C (primary)Randomized participants in the sub-study who received at least 1 dose and had at least 1 measurement of the endpoint from Baseline to Week 18Baseline to Week 18, excluding data after initiation of rescue
Overall-study A1C (secondary)Randomized participants who received at least one dose of blinded study medication and had at least one assessment at or after baselineBaseline to Week 18, excluding rescue

For the primary endpoint, person-years were calculated as the sum of participants' time to first event or time to censoring, with censoring at the earliest of the end-of-study date, death date, last contact date, or 365 days after the last dose. Arm-level results are expressed as events per 100 person-years.

Why the choice of window matters for non-inferiority: in a superiority trial, an intention-to-treat style analysis is conservative because non-adherence and post-discontinuation follow-up dilute differences toward the null. In a non-inferiority trial, that same dilution makes the arms look more alike and therefore makes non-inferiority easier to claim. An on-treatment window, extended here by 365 days after the last dose, limits dilution from time spent off study drug while still capturing events that occur after stopping treatment. The secondary endpoints, which test superiority, use the broader on-study window in all randomized participants.

5. Endpoints

The registry lists seven primary outcome measures: one cardiovascular endpoint for the overall study, and a baseline A1C measure plus a Week 18 change in A1C for each of the three glycemic sub-studies.

Primary endpoint (registry wording, shortened)DefinitionTime frame
Time to first occurrence of MACE (on-treatment + 365-day approach), overall cardiovascular studyTime to the first occurrence of any adjudicated component of 3-point MACE: CV death (including fatal stroke and fatal MI), non-fatal MI, and non-fatal strokeUp to approximately 6 years
Baseline A1C, insulin with or without metformin add-on glycemic sub-studyWeek 0 A1C, reported as a percentage of glycated hemoglobinBaseline
Change from baseline in A1C at Week 18, excluding rescue approach, insulin with or without metformin sub-studyWeek 18 A1C minus Week 0 A1C; negative values indicate reduction; data after rescue initiation excludedBaseline and Week 18
Baseline A1C, sulfonylurea monotherapy add-on glycemic sub-studyWeek 0 A1CBaseline
Change from baseline in A1C at Week 18, excluding rescue approach, sulfonylurea monotherapy sub-studyWeek 18 A1C minus Week 0 A1C; data after rescue initiation excludedBaseline and Week 18
Baseline A1C, metformin with sulfonylurea add-on glycemic sub-studyWeek 0 A1CBaseline
Change from baseline in A1C at Week 18, excluding rescue approach, metformin with sulfonylurea sub-studyWeek 18 A1C minus Week 0 A1C; data after rescue initiation excludedBaseline and Week 18

The baseline A1C measures are descriptive: they characterise each sub-study population and anchor the change-from-baseline endpoints, but they are not treatment comparisons and carry no statistical test.

Secondary endpoints with posted analyses

All secondary cardiovascular and renal endpoints below were analysed in the overall cardiovascular study using the on-study approach, with an assessment window of up to approximately 6 years.

6. Primary Result: MACE Non-Inferiority

The primary cardiovascular endpoint was analysed with a Cox proportional hazards model including treatment as an explanatory factor and cohort category as a stratification factor. Non-inferiority was assessed against a hazard-ratio margin of 1.3 using a one-sided p-value, with a two-sided 95.6% confidence interval reported for the hazard ratio.

Hazard ratio for first MACE, all ertugliflozin vs placebo

0.97

95.6% CI: 0.848–1.114   ·   One-sided non-inferiority p < 0.001

Non-inferiority margin: HR 1.3   ·   Upper confidence bound 1.114 lies below the margin

ComparisonHR95.6% CIOne-sided NI p-valueUpper bound vs 1.3 margin
All ertugliflozin vs placebo0.970.848–1.114<0.001Below margin
Ertugliflozin 15 mg vs placebo1.040.887–1.2110.002Below margin
Ertugliflozin 5 mg vs placebo0.910.773–1.065<0.001Below margin
Hazard ratios with 95.6% CIs (axis 0.6 to 1.4; solid line at HR 1.0, gold line at the 1.3 margin)
All ertugliflozin
0.97
Ertugliflozin 15 mg
1.04
Ertugliflozin 5 mg
0.91
Clinical Biostats interpretation

What the estimate means. A hazard ratio of 0.97 means that, under the stratified Cox model and over the on-treatment + 365-day window, the estimated instantaneous rate of a first MACE event with ertugliflozin was about 3% lower than with placebo. That is a point estimate very close to 1, i.e. essentially no difference.

What it does not mean. It does not show that ertugliflozin reduced MACE. The trial was designed to rule out harm, and the confidence interval of 0.848 to 1.114 includes 1, so the data are compatible with a modest reduction, no effect, or a modest increase in hazard. Nor is the hazard ratio a relative risk or a statement about the proportion of participants who had an event.

What the confidence interval says. The key quantity for non-inferiority is the upper bound. At 1.114, it corresponds to at most an 11.4% higher estimated hazard, comfortably inside the 1.3 margin. The interval is fairly narrow because the pooled comparison uses both ertugliflozin arms and a long event-driven follow-up. The dose-specific intervals are wider, with the 15 mg upper bound at 1.211, closer to but still below the margin.

Why the p-value is not an effect size. The one-sided p < 0.001 tests the null hypothesis that the true HR is at least 1.3. A very small p-value says that a hazard ratio of 1.3 or worse is implausible given the data; it says nothing about whether the true HR is 0.97, 1.0 or 1.1. The 15 mg p-value of 0.002 is larger only because its interval is wider and sits slightly closer to the margin, not because it shows "more harm" in any graded sense.

Cautions. (1) The CI level is 95.6%, not 95%, which indicates that the confirmatory alpha for this test was adjusted; the registry does not describe the allocation. (2) A single hazard ratio summarises about six years of follow-up and assumes approximately proportional hazards; the registry does not report a check of that assumption. (3) The analysis is in treated participants with an on-treatment + 365-day window, which is appropriate for non-inferiority but is not an intention-to-treat analysis of all randomized participants. (4) Non-inferiority conclusions depend on the margin; a margin of 1.3 excludes a 30% increase in hazard, not every possible increase.

7. Primary Results: Glycemic Sub-Studies

Each glycemic sub-study compared each ertugliflozin dose with placebo for change from baseline in A1C at Week 18, excluding data after initiation of glycemic rescue. The effect measure is the difference in least squares (LS) means from a constrained longitudinal data analysis (cLDA) model, reported with a two-sided 95% CI and a superiority hypothesis. Negative differences favour ertugliflozin.

7a. Insulin with or without metformin add-on sub-study

ComparisonLS mean difference (A1C %)95% CIP-value
Ertugliflozin 15 mg vs placebo-0.65-0.78 to -0.51<0.001
Ertugliflozin 5 mg vs placebo-0.58-0.71 to -0.44<0.001

Numbers at risk in this sub-study's serious-adverse-event tables were 348 (5 mg), 370 (15 mg) and 347 (placebo).

Clinical Biostats interpretation

What the estimate means. After adjustment within the cLDA model, mean A1C at Week 18 was estimated to fall 0.65 percentage points more with ertugliflozin 15 mg than with placebo, and 0.58 percentage points more with 5 mg. These are differences in absolute A1C percentage points, not relative percentage changes.

What it does not mean. An LS mean difference is a population-average contrast; it does not say that every participant's A1C fell by that amount, and it does not compare the two doses with each other. The overlapping intervals give no basis for claiming one dose is better.

Precision. Both intervals are narrow (about a quarter of a percentage point wide) and lie entirely below zero, so the direction of effect is well established in this sub-study.

P-value. P < 0.001 indicates that a zero difference is highly incompatible with the data; the size of the benefit is conveyed by the estimate and interval, not by the p-value.

Cautions. The "excluding rescue" approach estimates the effect had rescue not been used, a hypothetical estimand. If rescue was more common with placebo, excluding post-rescue data removes the participants with the poorest control disproportionately from one arm, and the cLDA model then relies on a missing-at-random assumption for those removed values. Each dose is also tested separately, so multiplicity across doses and sub-studies needs to be handled by the prespecified testing strategy.

7b. Sulfonylurea monotherapy add-on sub-study

ComparisonLS mean difference (A1C %)95% CIP-value
Ertugliflozin 15 mg vs placebo-0.22-0.60 to 0.160.247
Ertugliflozin 5 mg vs placebo-0.35-0.72 to 0.020.063

This was the smallest sub-study: numbers at risk in its serious-adverse-event tables were 55 (5 mg), 54 (15 mg) and 48 (placebo).

Clinical Biostats interpretation

What the estimate means. Point estimates favour ertugliflozin (-0.22 and -0.35 percentage points), but neither comparison reached conventional statistical significance.

What it does not mean. A non-significant result is not evidence that ertugliflozin has no glycemic effect in this setting. "Absence of evidence is not evidence of absence": the intervals span values from a meaningful reduction (-0.60, -0.72) to no reduction or a slight increase (0.16, 0.02).

Precision. The intervals are nearly three times as wide as in the insulin sub-study, which is what one expects from about a sixth of the sample size. The sub-study was underpowered to detect differences of the size observed.

P-value. The 5 mg p-value of 0.063 is not "almost significant" evidence of a larger effect than the 15 mg comparison (p = 0.247). The two doses' intervals overlap almost completely, and the ordering of the point estimates (5 mg larger than 15 mg) is consistent with random variation in small samples.

Cautions. With so few participants, the cLDA results are sensitive to a handful of missing or rescue-excluded values, and any imbalance in rescue between arms matters more here than in the larger sub-studies.

7c. Metformin with sulfonylurea add-on sub-study

ComparisonLS mean difference (A1C %)95% CIP-value
Ertugliflozin 15 mg vs placebo-0.75-0.98 to -0.53<0.001
Ertugliflozin 5 mg vs placebo-0.66-0.89 to -0.43<0.001

Numbers at risk in this sub-study's serious-adverse-event tables were 100 (5 mg), 113 (15 mg) and 117 (placebo).

Clinical Biostats interpretation

What the estimate means. Both doses produced a larger estimated mean reduction in A1C than placebo at Week 18: 0.75 percentage points with 15 mg and 0.66 with 5 mg.

What it does not mean. The larger point estimates here than in the insulin sub-study do not demonstrate that ertugliflozin works better on a metformin-plus-sulfonylurea background. The sub-studies enrolled different populations with different baseline A1C, and no formal between-sub-study comparison is reported.

Precision. The intervals are roughly 0.45 percentage points wide, wider than in the insulin sub-study because of the smaller sample, but both lie well below zero.

P-value. P < 0.001 rejects a zero difference; the clinical magnitude comes from the estimate and interval.

Cautions. As in the other sub-studies, the analysis population required at least one post-baseline measurement, and the rescue-exclusion estimand assumes post-rescue values are missing at random given the observed data.

8. Secondary Cardiovascular and Renal Endpoints

The registry posts stratified Cox analyses for a large family of secondary time-to-event endpoints, all in all randomized participants with the on-study approach and all tested for superiority. Three of them (CV death or HHF, CV death, and the renal composite) are reported with 95.8% intervals; the others use 95% intervals. The ClinicalTrials.gov record also posts many further secondary and other analyses; the principal hazard-ratio comparisons are summarised below.

First-event analyses (stratified Cox model)

EndpointComparisonHRCI (level)P-value
CV death or HHFAll ertugliflozin vs placebo0.880.750–1.034 (95.8%)0.108
15 mg vs placebo0.880.725–1.057 (95.8%)0.150
5 mg vs placebo0.890.735–1.068 (95.8%)0.188
CV deathAll ertugliflozin vs placebo0.920.767–1.113 (95.8%)0.385
15 mg vs placebo0.920.739–1.139 (95.8%)0.417
5 mg vs placebo0.930.750–1.154 (95.8%)0.494
Renal compositeAll ertugliflozin vs placebo0.810.630–1.036 (95.8%)0.081
15 mg vs placebo0.850.638–1.137 (95.8%)0.258
5 mg vs placebo0.760.568–1.028 (95.8%)0.065
MACE plusAll ertugliflozin vs placebo0.920.823–1.038 (95%)0.183
15 mg vs placebo0.950.831–1.086 (95%)0.450
5 mg vs placebo0.900.785–1.029 (95%)0.123
Fatal or non-fatal MIAll ertugliflozin vs placebo1.040.861–1.259 (95%)0.676
15 mg vs placebo1.170.949–1.451 (95%)0.139
5 mg vs placebo0.910.727–1.141 (95%)0.416
Fatal or non-fatal strokeAll ertugliflozin vs placebo1.060.820–1.365 (95%)0.663
15 mg vs placebo1.130.845–1.505 (95%)0.415
5 mg vs placebo0.990.736–1.334 (95%)0.953
Hospitalization for heart failureAll ertugliflozin vs placebo0.700.539–0.902 (95%)0.006
15 mg vs placebo0.680.502–0.932 (95%)0.016
5 mg vs placebo0.710.524–0.964 (95%)0.028
Death from any causeAll ertugliflozin vs placebo0.930.797–1.081 (95%)0.340
15 mg vs placebo0.940.784–1.117 (95%)0.463
5 mg vs placebo0.920.771–1.100 (95%)0.363

Total-event analyses (Andersen-Gill model)

For total MACE and total CV death or HHF, the registry reports hazard ratios from the Andersen-Gill model for recurrent events at the end of study, with 95% CIs. No p-values are posted for these analyses.

EndpointComparisonHR95% CI
Total MACEAll ertugliflozin vs placebo1.010.898–1.131
15 mg vs placebo1.070.937–1.219
5 mg vs placebo0.950.828–1.085
Total CV death or HHFAll ertugliflozin vs placebo0.820.716–0.945
15 mg vs placebo0.790.673–0.935
5 mg vs placebo0.850.725–1.001

Hospitalization for heart failure, all ertugliflozin vs placebo

0.70

95% CI: 0.539–0.902   ·   P = 0.006 (superiority, two-sided CI)

A 30% lower estimated hazard of first HHF; consistent across doses (0.68 and 0.71)

Reading the secondary results

The secondary results form a coherent pattern: point estimates close to 1 for atherosclerotic outcomes (MI, stroke, MACE plus, total MACE), and lower estimates for heart-failure-related outcomes. HHF alone shows an estimated hazard ratio of 0.70 with an interval excluding 1, and the recurrent-event analysis of CV death or HHF gives 0.82 (0.716–0.945).

However, the composite of CV death or HHF, reported with a 95.8% interval, had p = 0.108 for the pooled comparison, and its interval (0.750–1.034) includes 1. The use of a 95.8% level for this endpoint, CV death and the renal composite suggests they belonged to a formally alpha-controlled set of key secondary hypotheses. The registry does not describe the testing sequence, but in a fixed-sequence or gatekeeping strategy a non-significant result on an earlier hypothesis means later hypotheses, including HHF alone, cannot be claimed as confirmatory, however small their nominal p-values. The HHF finding is best read as a nominally significant, hypothesis-supporting result rather than a confirmed treatment effect.

The renal composite shows a similar picture: an estimate of 0.81 with an interval (0.630–1.036) that includes 1, and dose-specific estimates of 0.85 and 0.76 whose ordering should not be over-interpreted given their wide, overlapping intervals.

9. Secondary Glycemic Result: Overall Cardiovascular Study

In the overall cardiovascular study, change from baseline in A1C at Week 18 (excluding rescue) for ertugliflozin 15 mg versus placebo was analysed with a cLDA model with fixed effects for treatment, time, baseline eGFR (continuous) and the time-by-treatment interaction.

ComparisonLS mean difference (A1C %)95% CIP-value
Ertugliflozin 15 mg vs placebo-0.50-0.55 to -0.46<0.001

The very narrow interval reflects the size of the overall study. Adjusting for baseline eGFR is statistically sensible for an SGLT2 inhibitor, whose glucose-lowering action depends on renal filtration: eGFR is a strong prognostic covariate for A1C response, and including it improves precision. Because it is a baseline covariate, adjustment does not compromise the randomized comparison.

10. Safety: Serious Adverse Events

The registry reports the number of participants with at least one serious adverse event (SAE) by arm, separately for the overall cardiovascular study and for each glycemic sub-study. No formal statistical comparison of SAE frequencies is posted.

Study populationErtugliflozin 5 mg
(affected / at risk)
Ertugliflozin 15 mg
(affected / at risk)
Placebo
(affected / at risk)
Overall cardiovascular study958 / 2746937 / 2747990 / 2745
Insulin with or without metformin sub-study33 / 34827 / 37037 / 347
Sulfonylurea monotherapy sub-study4 / 551 / 542 / 48
Metformin with sulfonylurea sub-study7 / 1008 / 1136 / 117

In the overall study, the counts of participants with an SAE are similar across the three arms, with the placebo arm having the highest count against nearly identical denominators. Two cautions apply. First, SAE counts in a trial lasting up to approximately six years are strongly influenced by exposure time; a crude proportion does not account for differences in time on study, which is why efficacy outcomes in this trial are expressed per 100 person-years. Second, in a cardiovascular outcomes trial many serious adverse events are themselves cardiovascular events, so SAE totals partly overlap the efficacy endpoints. In the small sulfonylurea monotherapy sub-study, single-digit counts cannot support any comparison between arms.

11. Statistical Methodology

Stratified Cox proportional hazards model

All time-to-event comparisons used a Cox proportional hazards model with treatment as an explanatory factor and cohort category as a stratification factor. Stratification allows each cohort to have its own baseline hazard function while estimating a single common treatment hazard ratio across cohorts.

Conceptual form
h(t | treatment, cohort s) = h0s(t) · exp(β · treatment)   →   HR = exp(β)

where h0s(t) is an unspecified baseline hazard for stratum s. The model assumes the treatment hazard ratio is constant over time and the same in every stratum.

For the pooled comparisons, the registry labels the groups as "Placebo vs All Ertugliflozin"; the accompanying analysis descriptions state that the hazard ratios compare ertugliflozin with placebo, and they are read that way throughout this page, so values below 1 indicate a lower estimated hazard with ertugliflozin.

Non-inferiority testing against a hazard-ratio margin

Hypotheses for the primary MACE endpoint
H0: HR ≥ 1.3    versus    H1: HR < 1.3    (one-sided)

Non-inferiority is concluded when the upper limit of the two-sided confidence interval for the HR lies below 1.3, which is equivalent to a one-sided test at half the complementary alpha. For VERTIS CV, the upper limit of the 95.6% CI was 1.114.

Non-standard confidence levels and alpha control

The primary MACE analyses use 95.6% intervals and three key secondary endpoints use 95.8% intervals, while other endpoints use 95%. Confidence levels other than 95% usually arise when part of the overall type I error has been spent elsewhere, for example at interim analyses, or when alpha is split or recycled among several hypotheses. The registry does not document the multiplicity strategy, but these levels signal that the confirmatory analyses were conducted under a prespecified error-control framework and that endpoints reported with plain 95% intervals should be interpreted as nominal.

Andersen-Gill model for recurrent events

The total MACE and total CV death or HHF analyses used the Andersen-Gill model, an extension of the Cox model in which each participant can contribute multiple events. It estimates a common hazard ratio for all events, first and subsequent, under the assumption that events within a participant are conditionally independent given covariates; robust (sandwich) variance estimation is conventionally used to account for within-participant correlation.

Andersen-Gill intensity
λi(t) = Yi(t) · λ0(t) · exp(β · treatmenti)

where Yi(t) indicates that participant i is still under observation at time t. Unlike a first-event analysis, participants remain at risk after an event.

Constrained longitudinal data analysis (cLDA)

The A1C endpoints were analysed with cLDA, a mixed model for repeated measures in which baseline A1C is treated as part of the outcome vector rather than as a covariate, and the baseline means are constrained to be equal across randomized groups. This reflects randomization (groups have the same expected baseline) and allows participants with a baseline value but missing post-baseline values to contribute to the estimation of the baseline mean and covariance structure.

cLDA structure (overall-study model as posted)
A1Cij = timej + treatment × timej (post-baseline only) + baseline eGFR + εij

The treatment effect at Week 18 is the difference in LS means for the treatment-by-time term at that visit. An unstructured within-participant covariance is typical for this model; valid inference relies on data being missing at random.

Handling of glycemic rescue

Under the "excluding rescue" approach, all A1C data after initiation of rescue therapy are set aside. Combined with cLDA, this targets the treatment effect in the absence of rescue, with post-rescue values implicitly imputed from the model under missing at random. It is a clear estimand, but a hypothetical one; a treatment-policy estimand that includes post-rescue data would answer a different question.

12. Statistical Methods Explained

Why is non-inferiority judged against the 1.3 margin rather than the p-value alone?

The margin defines what "not unacceptably worse" means. The p-value of <0.001 is computed relative to that margin, so it is only meaningful alongside it. Reading the confidence interval is more informative: an upper bound of 1.114 tells you how large an excess hazard remains plausible, and you can compare that to any margin you personally consider relevant, not only 1.3.

Why is the primary interval 95.6% rather than 95%?

A 95.6% interval corresponds to a two-sided alpha of 0.044, which is less than the conventional 0.05. That reduction typically means some alpha was reserved for other looks at the data or other hypotheses. It makes the interval slightly wider than a 95% interval, which is the price of controlling the overall false-positive rate.

Why does the primary MACE analysis use an on-treatment + 365-day window while secondary endpoints use on-study?

For non-inferiority, including large amounts of off-drug follow-up pushes the arms toward the same outcome and makes non-inferiority easier to show. Limiting follow-up to the treatment period plus 365 days reduces that bias. For superiority testing of secondary endpoints, the on-study approach in all randomized participants is the conservative choice. The two windows answer slightly different questions, so hazard ratios from them should not be compared as if they came from the same analysis.

Why do total-event and first-event hazard ratios differ for CV death or HHF?

The first-event hazard ratio was 0.88 (pooled), while the Andersen-Gill total-event hazard ratio was 0.82. Recurrent-event analyses use every hospitalization, not just the first, which adds information and usually narrows the interval. They also weight participants with multiple events more heavily. A difference between the two does not mean either is wrong; they estimate different quantities.

What does a difference in LS means of -0.65 mean?

It is the model-estimated difference between the ertugliflozin 15 mg and placebo groups in mean change in A1C at Week 18, in percentage points of A1C. If placebo participants' A1C changed by some amount, the estimated change with ertugliflozin 15 mg was 0.65 percentage points more negative. "Least squares" indicates the mean is adjusted for the other terms in the model rather than being a raw arithmetic mean.

Why use cLDA instead of an ANCOVA on Week 18 values?

An ANCOVA uses only participants with a Week 18 value, adjusting for baseline. cLDA uses all available measurements, keeps participants with incomplete follow-up in the model, and constrains baseline means to be equal as randomization implies. Under missing at random, this gives valid and typically more efficient estimates than a completers-only analysis.

Why were the two doses pooled for the primary comparison?

Pooling both ertugliflozin doses doubles the number of treated participants contrasted with placebo, narrowing the interval for the safety-oriented question "does ertugliflozin increase cardiovascular risk?". The dose-specific estimates (1.04 and 0.91) are reported to check that neither dose alone drives an unfavourable result; the dose-specific upper bounds (1.211 and 1.065) also sit below 1.3.

13. Limitations

14. Why This Trial Matters Statistically

VERTIS CV is a useful teaching case because it combines a regulatory-style non-inferiority design with superiority-tested secondary and glycemic endpoints, several analysis windows, recurrent-event methods and longitudinal models for continuous outcomes, all in one trial.

ConceptHow it appears in VERTIS CV
Non-inferiority designPrimary MACE comparison against a hazard-ratio margin of 1.3 with a one-sided test
Hazard ratioRelative measure for MACE, CV death, HHF, renal and mortality endpoints
Stratified Cox modelTreatment as explanatory factor, cohort category as stratification factor
Confidence intervals95.6%, 95.8% and 95% levels reflecting alpha allocation
MultiplicityMany endpoints, two doses and three sub-studies; nominal vs confirmatory p-values
Recurrent eventsAndersen-Gill model for total MACE and total CV death or HHF
Analysis windowsOn-treatment + 365-day vs on-study approaches
cLDALongitudinal A1C analysis with constrained baseline means
Missing data / estimandsExclusion of post-rescue data and the missing-at-random assumption
Covariate adjustmentBaseline eGFR in the overall-study A1C model
Sample size and precisionContrast between the small sulfonylurea sub-study and the larger sub-studies

15. Related Tutorials

Learn more about the methods used in this trial:

16. Related Calculators

17. Sources

Explore the methods behind the results

Follow the statistical ideas in this trial, from non-inferiority margins to longitudinal models, through our tutorials and calculators.

18. Record Summary

VERTIS CV met its primary cardiovascular objective of non-inferiority: the pooled ertugliflozin-versus-placebo hazard ratio for first MACE was 0.97, with a 95.6% CI upper bound of 1.114, below the 1.3 margin. Ertugliflozin lowered A1C relative to placebo in the insulin and metformin-with-sulfonylurea sub-studies, while the small sulfonylurea monotherapy sub-study was inconclusive. Secondary analyses showed a lower estimated hazard of hospitalization for heart failure, but the composite of CV death or HHF did not reach statistical significance. The most accurate reading of the trial therefore combines the logic of non-inferiority, the distinction between confirmatory and nominal p-values, the analysis window behind each estimate, and the estimand implied by excluding rescue data.