← Clinical Trials
Type 2 Diabetes Phase 3 Cardiovascular Outcomes NCT01131676

EMPA-REG OUTCOME: Complete Statistical Analysis of Empagliflozin in Type 2 Diabetes

An independent statistical analysis of the randomized, double-blind phase 3 EMPA-REG OUTCOME trial evaluating empagliflozin versus placebo in patients with type 2 diabetes mellitus, with emphasis on its time-to-event endpoints, non-inferiority framework, Cox proportional-hazards analyses, and reported cardiovascular and microvascular outcomes.

Trial start: 2010-07  ·  Primary completion: 2015-04  ·  Status: Completed
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

EMPA-REG OUTCOME was a randomized, double-blind, parallel phase 3 trial in patients with type 2 diabetes mellitus. The registry reports an enrollment of 7064 participants and evaluates empagliflozin against placebo using time-to-event cardiovascular and microvascular endpoints.

7064
Enrollment
Registered participants
3
Arms
Placebo, 10 mg, 25 mg
0.86
Primary MACE HR
95.02% CI 0.74–0.99
0.65
HF Hospitalisation HR
95% CI 0.50–0.85
FeatureEMPA-REG OUTCOME
Trial nameEMPA-REG OUTCOME
NCT identifierNCT01131676
PhasePhase 3
ConditionDiabetes Mellitus, Type 2
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment7064.0
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry
Trial statusCompleted
Registry record: The official trial record is available at ClinicalTrials.gov NCT01131676.

2. Clinical Question

The central statistical question was whether empagliflozin was non-inferior to placebo for the time to first occurrence of the registered 3-point major adverse cardiovascular event (MACE) composite, followed within the prespecified testing strategy by a superiority assessment.

Population

Patients with Diabetes Mellitus, Type 2 enrolled in the EMPA-REG OUTCOME trial.

Intervention

BI 10773 (empagliflozin), represented by low-dose and high-dose study arms.

Comparator

Placebo corresponding to the empagliflozin study arms.

Primary question

Is the time to the first occurrence of 3-point MACE non-inferior to placebo, and does the subsequent superiority test support a lower hazard?

3. Trial Design

01
Enroll7064.0 participants
02
RandomizeRandomized allocation
03
MaskDouble-masked design
04
ObserveTime-to-event follow-up
05
AnalyzeCox proportional hazards
STUDY ARM · EMPAGLIFLOZIN LOW DOSE

Empagliflozin 10 mg

  • BI 10773 low-dose study intervention
  • Serious adverse events: 876/2345 affected/at risk
STUDY ARM · EMPAGLIFLOZIN HIGH DOSE

Empagliflozin 25 mg

  • BI 10773 high-dose study intervention
  • Serious adverse events: 913/2342 affected/at risk
COMPARATOR

Placebo

  • Placebo study intervention
  • Serious adverse events: 988/2333 affected/at risk
Important arm-level distinction: the primary and secondary statistical analyses in the ClinicalTrials.gov record compare Placebo vs All Empagliflozin. The registry's safety data, however, are reported separately for placebo, empagliflozin 10 mg, and empagliflozin 25 mg.

4. Trial Conduct and Registry Details

Allocation
Randomized
Design model
Parallel
Masking
Double
Primary purpose
Treatment
Start
2010-07
Primary completion
2015-04

The ClinicalTrials.gov record identifies randomization, parallel allocation, and double masking. They do not provide a crossover description, factorial structure, interim-analysis specification, missing-data or imputation strategy, or Bayesian analysis. Those design features are therefore not incorporated into the statistical interpretation of this page.

5. Primary Endpoint

EndpointRegistry definitionTime frameEndpoint type
3-point MACE Time to the first occurrence of any of the following adjudicated components of the primary composite endpoint: cardiovascular (CV) death (including fatal stroke and fatal myocardial infarction (MI)), non-fatal MI (excluding silent MI), and non-fatal stroke. Percentage of patients with the event are presented. From randomisation to individual end of observation, up to 4.6 years Time-to-event

The primary endpoint is a composite time-to-event endpoint. A participant contributes an event when the first qualifying component occurs. Because the endpoint combines CV death, non-fatal MI, and non-fatal stroke, its treatment effect summarizes the time until the first occurrence of any of those adjudicated events rather than estimating a separate effect for each component.

6. Statistical Methodology

Cox proportional-hazards model

The registry reports the Cox proportional-hazards model as the primary analytical method. The effect measure is the hazard ratio (HR), comparing All Empagliflozin with Placebo.

Hazard-ratio interpretation
HR = estimated hazard in All Empagliflozin / estimated hazard in Placebo

An HR below 1 indicates a lower estimated instantaneous event hazard in the empagliflozin group under the fitted Cox model. It is not an absolute risk difference, a probability of an event, or a statement that every participant experiences the same proportional reduction.

Analysis population

The primary analyses use the TS analysis population. Secondary analyses use either TS or TS (evaluable cases), as specified for each endpoint. This distinction matters because an analysis restricted to evaluable cases does not necessarily represent exactly the same population as the full TS analysis.

Covariate adjustment

For the reported superiority analysis of the primary endpoint and the secondary endpoints, the registry analysis notes specify a Cox proportional-hazards model with factors for treatment, age, gender, categorised BMI, HbA1c, eGFR and geographical region.

Non-inferiority framework

The primary objective was to establish non-inferiority of All Empagliflozin relative to placebo for time to first 3-point MACE. The ClinicalTrials.gov record states that the non-inferiority margin was 1.3, chosen based on FDA guidance for cardiovascular-risk evaluation of new antidiabetic therapies.

Non-inferiority logic
Upper confidence limit < 1.3  →  non-inferiority criterion satisfied

For a hazard ratio where values above 1 represent greater hazard with empagliflozin, the non-inferiority question asks whether the plausible upper boundary remains below the prespecified margin of 1.3. This is different from simply asking whether a conventional superiority p-value is below a chosen threshold.

Multiplicity and hierarchical testing

The registry states that a 4-step hierarchical testing strategy was followed: non-inferiority testing of the primary endpoint, followed by non-inferiority testing of the key secondary endpoint, followed by superiority testing of the primary endpoint and then the key secondary endpoint. The non-inferiority margin for the primary and key secondary endpoints was 1.3.

Testing stepEndpoint roleHypothesis typeMargin / comparison
Step 1Primary 3-point MACENon-inferiorityMargin 1.3
Step 2Key secondary 4-point MACENon-inferiorityMargin 1.3
Step 3Primary 3-point MACESuperiorityHR comparison with 1
Step 4Key secondary 4-point MACESuperiorityHR comparison with 1

7. Primary Results: 3-Point MACE

Non-inferiority analysis

Hazard ratio for first 3-point MACE

0.86

95.02% CI: 0.74–0.99   ·   P < 0.0001

All Empagliflozin vs Placebo  ·  TS population

The reported Cox proportional-hazards model estimated an HR of 0.86 for All Empagliflozin versus Placebo. The two-sided 95.02% confidence interval was 0.74–0.99, and the registry reports P < 0.0001 for the non-inferiority hypothesis.

Clinical Biostats interpretation

An HR of 0.86 corresponds to a 14% lower estimated hazard of the first 3-point MACE event for All Empagliflozin relative to Placebo, using the usual interpretation of the hazard ratio. This is a relative time-to-event measure; it does not mean that exactly 14% of participants avoided an event, nor does it directly provide an absolute reduction in event probability.

The confidence interval of 0.74–0.99 indicates the precision of the estimated HR under the model and statistical framework. It does not describe the range of effects that individual patients experienced.

The reported P < 0.0001 is evidence against the non-inferiority null under the prespecified testing framework; it is not a measure of the size of the treatment effect. The key non-inferiority comparison is with the prespecified margin of 1.3, not merely with 1.0.

Superiority analysis

Superiority hazard ratio for first 3-point MACE

0.86

95.02% CI: 0.74–0.99   ·   P = 0.0382

All Empagliflozin divided by Placebo

The same estimated HR of 0.86 was reported for the subsequent superiority analysis, with a two-sided 95.02% confidence interval of 0.74–0.99 and a reported P-value of 0.0382. The superiority Cox model included treatment, age, gender, categorised BMI, HbA1c, eGFR and geographical region.

Clinical Biostats interpretation

The HR of 0.86 indicates a lower estimated hazard for the first 3-point MACE event in the All Empagliflozin group relative to Placebo. The estimate describes the relative event hazard over the analyzed observation period; it is not a direct estimate of absolute event probability or individual benefit.

The 0.74–0.99 confidence interval is relatively close to 1 at its upper boundary, illustrating why the precision of the estimate matters in addition to the point estimate. A confidence interval describes statistical uncertainty around the estimated treatment effect; it does not establish that the true effect for every patient lies within that interval.

The superiority P = 0.0382 quantifies evidence against the null hypothesis used for this superiority test under the specified analysis. It does not quantify the magnitude or clinical importance of the effect. The reported analysis also sits within the trial's hierarchical testing strategy, so its interpretation should not be separated from that prespecified sequence.

8. Key Secondary Endpoint: 4-Point MACE

The registry reports a key secondary time-to-event endpoint defined as the first occurrence of all events adjudicated in the 4-point MACE composite: CV death, non-fatal MI, non-fatal stroke, and hospitalization for unstable angina pectoris.

EndpointAnalysis populationMethodEffect measureEstimate95.02% CIP-value
4-point MACE TS Cox proportional-hazards model Hazard ratio 0.89 0.78–1.01 <0.0001 for non-inferiority

Non-inferiority analysis

Hazard ratio for 4-point MACE

0.89

95.02% CI: 0.78–1.01   ·   P < 0.0001

All Empagliflozin vs Placebo  ·  TS population

The registry reports an HR of 0.89 with a two-sided 95.02% confidence interval of 0.78–1.01. The non-inferiority P-value is reported as <0.0001, with the non-inferiority margin set at 1.3.

Clinical Biostats interpretation

An HR of 0.89 corresponds to an 11% lower estimated hazard of the first 4-point MACE event for All Empagliflozin versus Placebo. Again, this is not an 11-percentage-point reduction in event probability and does not mean that each participant experiences an identical reduction.

The confidence interval of 0.78–1.01 extends slightly above 1.0, so the interval does not by itself exclude equal hazards. For the non-inferiority question, however, the relevant boundary is 1.3, not 1.0. The registry reports P < 0.0001 for that non-inferiority assessment.

The P-value addresses the hypothesis being tested; it does not measure the magnitude or practical importance of the HR. The composite endpoint also contains an additional component—hospitalization for unstable angina pectoris—relative to the 3-point MACE definition.

Superiority analysis

Superiority hazard ratio for 4-point MACE

0.89

95.02% CI: 0.78–1.01   ·   P = 0.0795

All Empagliflozin divided by Placebo

For the subsequent superiority analysis, the registry reports the same HR of 0.89 and 95.02% CI of 0.78–1.01, with a superiority P-value of 0.0795.

Clinical Biostats interpretation

The point estimate remains below 1, corresponding to an estimated 11% lower hazard in the All Empagliflozin group. The confidence interval, however, includes 1.0, indicating uncertainty that encompasses equal hazards.

The reported P = 0.0795 is a hypothesis-test result for superiority and should not be interpreted as the probability that the treatment has no effect. It also should not be treated as a measure of effect size. The distinction between the non-inferiority and superiority questions is essential: the same HR can support non-inferiority against a margin of 1.3 while not providing the same evidence for superiority against 1.0.

Because this superiority test appears as a later step in the stated hierarchical sequence, its interpretation also depends on the trial's prespecified multiplicity strategy.

9. Secondary Time-to-Event Results

The ClinicalTrials.gov record contains five additional secondary time-to-event analyses. All compare All Empagliflozin with Placebo using Cox proportional-hazards models.

Secondary endpointAnalysis populationHR95% CIP-value
Percentage of Participants With Silent MI TS (evaluable cases) 1.28 0.70–2.33 0.4172
Percentage of Participants With Heart Failure Requiring Hospitalisation (Adjudicated) TS 0.65 0.50–0.85 0.0017
Percentage of Participants With New Onset Albuminuria TS (evaluable cases) 0.95 0.87–1.04 0.2547
Percentage of Participants With New Onset Macroalbuminuria TS (evaluable cases) 0.62 0.54–0.72 <0.0001
Percentage of Participants With the Composite Microvascular Outcome TS (evaluable cases) 0.62 0.54–0.70 <0.0001

Heart failure requiring hospitalisation

Hazard ratio

0.65

95% CI: 0.50–0.85   ·   P = 0.0017

All Empagliflozin vs Placebo  ·  TS population

The HR of 0.65 corresponds to a 35% lower estimated hazard of adjudicated heart failure requiring hospitalisation in the All Empagliflozin group relative to Placebo, using the conventional interpretation of the reported HR.

Clinical Biostats interpretation

The 95% confidence interval of 0.50–0.85 indicates the statistical uncertainty around the estimated HR. The interval remains below 1.0, while the reported P-value is 0.0017. The P-value does not measure the size of the treatment effect, and the HR does not directly state an absolute difference in hospitalization probability.

Because this is a time-to-event endpoint, censoring and the assumptions of the Cox model remain relevant. The registry supplies the Cox method but does not provide enough information in the ClinicalTrials.gov record to independently assess the proportional-hazards assumption.

New onset macroalbuminuria

Hazard ratio

0.62

95% CI: 0.54–0.72   ·   P < 0.0001

All Empagliflozin vs Placebo  ·  TS (evaluable cases)

The reported HR of 0.62 represents a 38% lower estimated hazard of new onset macroalbuminuria for All Empagliflozin relative to Placebo. The confidence interval of 0.54–0.72 provides the reported uncertainty around this estimate.

Composite microvascular outcome

Hazard ratio

0.62

95% CI: 0.54–0.70   ·   P < 0.0001

All Empagliflozin vs Placebo  ·  TS (evaluable cases)

The composite microvascular endpoint has the same point estimate as new onset macroalbuminuria, 0.62, but a different confidence interval of 0.54–0.70. This illustrates why the point estimate alone is insufficient: the endpoint definition and uncertainty interval both matter.

Silent myocardial infarction

EndpointHR95% CIP-value
Silent MI1.280.70–2.330.4172

The silent-MI analysis produced an HR of 1.28 with a 95% CI of 0.70–2.33. The interval is wide and includes 1.0. The registry reports P = 0.4172 for the superiority hypothesis. The analysis population was TS (evaluable cases).

New onset albuminuria

EndpointHR95% CIP-value
New onset albuminuria0.950.87–1.040.2547

The HR of 0.95 indicates an estimated hazard close to that of the comparator, while the 95% CI of 0.87–1.04 includes 1.0. The reported superiority P-value is 0.2547.

Multiplicity matters: these secondary endpoints should not be interpreted as though each were an isolated, independently powered confirmatory experiment. The ClinicalTrials.gov record explicitly identify multiplicity adjustment in the analysis text for the primary 3-point MACE and key secondary 4-point MACE analyses, while the broader set of secondary results should be interpreted in the context of the trial's stated hierarchical testing framework.

10. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

The registered endpoints are time-to-event outcomes, meaning that both whether an event occurred and when it occurred contribute information. A Cox model is designed for this setting and produces a hazard ratio that compares the instantaneous event hazards between treatment groups while allowing for censored observations.

What does an HR of 0.86 mean?

An HR of 0.86 means that the estimated instantaneous hazard in the All Empagliflozin group is 0.86 times the estimated hazard in the Placebo group under the fitted model. Expressed as a relative difference, this is a 14% lower estimated hazard. It does not mean that 14% of participants benefited or that event probability was reduced by exactly 14 percentage points.

Why was the non-inferiority margin 1.3 important?

Non-inferiority asks whether the treatment is not unacceptably worse than the comparator by more than a prespecified amount. For this trial, the ClinicalTrials.gov record specifies a margin of 1.3. Therefore, the relevant question is whether the confidence bound remains below 1.3—not simply whether the confidence interval excludes 1.0.

How can non-inferiority and superiority have different P-values for the same HR?

They are different hypotheses. The non-inferiority analysis tests against a margin of 1.3, whereas superiority tests whether the hazard ratio is compatible with the null value of 1.0. In EMPA-REG OUTCOME, the primary HR is 0.86, with P < 0.0001 for the reported non-inferiority test and P = 0.0382 for the reported superiority test.

Why does the 4-point MACE result illustrate this distinction particularly well?

The 4-point MACE HR is 0.89 with a 95.02% CI of 0.78–1.01. The upper confidence limit is well below the non-inferiority margin of 1.3, supporting the reported non-inferiority result. But the interval includes 1.0, and the reported superiority P-value is 0.0795. Thus, non-inferiority and superiority are not interchangeable claims.

Why does the analysis population matter?

The primary MACE analyses use the TS population, while silent MI, new onset albuminuria, new onset macroalbuminuria, and the composite microvascular outcome use TS (evaluable cases). Restricting an analysis to evaluable cases can change the population contributing data, so the analysis population should always be read alongside the estimate.

What does a confidence interval tell us that a P-value does not?

A confidence interval communicates the precision and plausible range of the estimated treatment effect under the specified statistical framework. A P-value summarizes evidence against a particular null hypothesis. Neither one directly measures clinical importance, and neither should be confused with the probability that a treatment effect is true.

11. Confidence Intervals and Effect Size

The reported results illustrate several distinct patterns of statistical uncertainty.

EndpointHRConfidence intervalWhat the interval indicates
3-point MACE0.860.74–0.99Relatively narrow interval ending below 1.0
4-point MACE0.890.78–1.01Interval includes 1.0 but remains below 1.3
Silent MI1.280.70–2.33Wide interval spanning both below and above 1.0
Heart failure requiring hospitalisation0.650.50–0.85Interval below 1.0
New onset albuminuria0.950.87–1.04Interval includes 1.0 and is relatively close to it
New onset macroalbuminuria0.620.54–0.72Interval below 1.0
Composite microvascular outcome0.620.54–0.70Interval below 1.0

A confidence interval should be read together with the endpoint definition, analysis population, model, and hypothesis being tested. In particular, an interval crossing 1.0 is not equivalent to proof of no effect; it indicates that the data are compatible with equal hazards within the stated interval and statistical framework.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm as the number affected divided by the number at risk.

ArmSerious adverse events affected / at risk
Placebo988 / 2333
Empagliflozin 10 mg876 / 2345
Empagliflozin 25 mg913 / 2342
Serious adverse events: affected / at risk
Placebo
988/2333
Empagliflozin 10 mg
876/2345
Empagliflozin 25 mg
913/2342

The figures above are reported as affected participants divided by participants at risk. They should not be transformed here into calculated percentages because the page is restricted to the registry-reported trial numbers and the instruction not to recompute or round results.

Safety interpretation: the serious-adverse-event counts are descriptive arm-level safety data. They are not the same statistical endpoint as the time-to-event efficacy analyses, and the ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse events between the three arms.

13. Registry Limitations and Data Integrity Notes

The ClinicalTrials.gov record contains several operational caveats that matter when interpreting the reported enrollment and analyses.

Why this matters statistically: the registry's enrollment figure and its analysis populations are not interchangeable concepts. Enrollment describes the trial's registered participant count, whereas the reported statistical analyses specify their own analysis population, including TS and TS (evaluable cases).

14. Time-to-Event Endpoints and Censoring

The primary and secondary efficacy outcomes in the ClinicalTrials.gov record are time-to-event endpoints. The time frame is consistently described as from randomisation to individual end of observation, up to 4.6 years.

Why time-to-event analysis is different
Time-to-event analysis = event occurrence + event timing + censoring information

A simple comparison of percentages at a single time point can discard information about when events occurred and how long participants were followed. Cox regression instead uses the observed event times and censoring structure to estimate a relative hazard.

Because follow-up ends at each participant's individual end of observation, some observations can be censored rather than observed through an event. The ClinicalTrials.gov record identifies the Cox model but do not provide enough information to reconstruct individual censoring histories or independently evaluate censoring assumptions.

15. Non-Inferiority vs Superiority

EMPA-REG OUTCOME is particularly useful for teaching the distinction between two related but different clinical-trial questions.

Non-inferiority

Asks whether the treatment's hazard is sufficiently below the prespecified upper margin of 1.3 to rule out an unacceptably large disadvantage relative to placebo.

Superiority

Asks whether the treatment hazard is statistically distinguishable from the null value of 1.0 in the favorable direction.

EndpointHR95.02% CINI P-valueSuperiority P-value
3-point MACE0.860.74–0.99<0.00010.0382
4-point MACE0.890.78–1.01<0.00010.0795

The two rows show why a trial can provide evidence for non-inferiority without the same endpoint necessarily satisfying a superiority test. The relevant null value changes from the non-inferiority margin of 1.3 to the superiority value of 1.0.

16. Multiplicity and Hierarchical Testing

The registry-reported analysis notes explicitly identify multiplicity adjustment for the primary endpoint and the key secondary endpoint. The four-step hierarchical sequence is therefore an important part of the statistical interpretation rather than a peripheral technical detail.

StepEndpointQuestionReported result
13-point MACENon-inferiorityHR 0.86; P < 0.0001
24-point MACENon-inferiorityHR 0.89; P < 0.0001
33-point MACESuperiorityHR 0.86; P = 0.0382
44-point MACESuperiorityHR 0.89; P = 0.0795

Hierarchical testing is designed to control how confirmatory conclusions are sequenced across multiple hypotheses. The important statistical lesson is that individual P-values should be interpreted in the context of the prespecified testing hierarchy rather than as isolated numbers.

17. What the Hazard Ratio Does — and Does Not — Mean

Relative treatment effect

An HR of 0.65 for heart failure requiring hospitalisation means that the fitted Cox model estimates a lower instantaneous hazard in the All Empagliflozin group than in the Placebo group. Expressed relatively, the estimated hazard is 35% lower.

It does not mean that 35% of participants avoided hospitalization, that the absolute hospitalization risk fell by 35 percentage points, or that every individual had exactly the same proportional reduction.

Confidence interval

The 95% confidence interval of 0.50–0.85 for this endpoint describes uncertainty around the estimated HR under the model and sampling framework. It does not describe the range of treatment effects among individual patients.

P-value

The P-value addresses evidence against the specified null hypothesis. A small P-value does not imply a large treatment effect, while a larger P-value does not prove that the treatment has no effect. Effect size and uncertainty should be read from the HR and confidence interval.

18. Composite Endpoints: Why the Definition Matters

EMPA-REG OUTCOME's primary endpoint is a composite of three adjudicated events: CV death, non-fatal MI excluding silent MI, and non-fatal stroke. The key secondary composite adds hospitalization for unstable angina pectoris.

3-point MACE

CV death, including fatal stroke and fatal MI; non-fatal MI excluding silent MI; and non-fatal stroke.

4-point MACE

CV death, non-fatal MI, non-fatal stroke, and hospitalization for unstable angina pectoris.

A composite endpoint allows several clinically relevant events to contribute to one time-to-event analysis, but the resulting HR applies to the first occurrence of the composite. It should not automatically be interpreted as though it represents the treatment effect on each individual component separately.

19. Analysis Population and Evaluable Cases

EndpointAnalysis population
Primary 3-point MACETS
Key secondary 4-point MACETS
Silent MITS (evaluable cases)
Heart failure requiring hospitalisationTS
New onset albuminuriaTS (evaluable cases)
New onset macroalbuminuriaTS (evaluable cases)
Composite microvascular outcomeTS (evaluable cases)

The use of TS (evaluable cases) for several secondary endpoints means that those analyses are not necessarily based on precisely the same analytic population as the primary MACE analysis. That distinction is particularly important when comparing estimates across endpoints.

20. Important Limitations and Interpretation Issues

21. Why This Trial Matters Statistically

EMPA-REG OUTCOME is a useful teaching case because the registry data bring together several core concepts in clinical-trial survival analysis: randomized treatment allocation, double masking, composite time-to-event endpoints, Cox proportional-hazards modeling, hazard ratios, confidence intervals, non-inferiority margins, superiority testing, and hierarchical multiplicity control.

ConceptHow it appears in EMPA-REG OUTCOME
RandomizationThe trial is registered as randomized with a parallel design.
BlindingThe trial is registered as double-masked.
Time-to-event endpointThe primary endpoint measures time to first 3-point MACE event from randomisation to individual end of observation, up to 4.6 years.
Cox modelThe statistical analyses posted on ClinicalTrials.gov use Cox proportional-hazards models.
Hazard ratioHR is the reported effect measure for the primary and secondary time-to-event analyses.
Confidence intervalPrimary analyses report 95.02% confidence intervals; other registry-reported analyses report 95% confidence intervals.
Non-inferiorityThe primary and key secondary analyses use a prespecified margin of 1.3.
SuperioritySuperiority is tested after the non-inferiority steps in the stated 4-step hierarchy.
MultiplicityThe registry analysis notes identify multiplicity adjustment and a hierarchical testing strategy.
Analysis populationsAnalyses use TS or TS (evaluable cases), depending on endpoint.
SafetySerious adverse events are reported as affected/at risk for each study arm.

22. A Statistical Reading of the Overall Results

The primary 3-point MACE analysis produced an HR of 0.86 with a 95.02% CI of 0.74–0.99. In the registry analysis, this result supported the prespecified non-inferiority objective and was subsequently evaluated for superiority, with a reported P-value of 0.0382.

The key secondary 4-point MACE analysis produced an HR of 0.89 with a 95.02% CI of 0.78–1.01. Its reported non-inferiority P-value was <0.0001, while the subsequent superiority P-value was 0.0795. The difference between these two hypothesis tests is an important statistical feature of the trial.

Among the additional secondary endpoints reported, the HR was 0.65 for heart failure requiring hospitalisation, 0.62 for new onset macroalbuminuria, and 0.62 for the composite microvascular outcome. New onset albuminuria had an HR of 0.95, while silent MI had an HR of 1.28. These estimates come from different endpoints and, for several outcomes, different analysis populations, so they should not be collapsed into a single generalized treatment effect.

Statistical takeaway: the most informative way to read EMPA-REG OUTCOME is to keep the endpoint, hypothesis, effect measure, confidence interval, analysis population, and multiplicity framework together. The same numerical HR can answer different questions depending on whether it is being evaluated against a non-inferiority margin of 1.3 or the superiority null value of 1.0.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue with the statistical methods

Explore the survival-analysis, trial-design, inference, and multiplicity concepts that provide the statistical framework for interpreting EMPA-REG OUTCOME.

26. Record Summary

EMPA-REG OUTCOME provides a compact teaching example of how a randomized clinical trial can combine a time-to-event primary endpoint with a formal non-inferiority framework and a subsequent superiority assessment. The primary 3-point MACE analysis reported an HR of 0.86, with a 95.02% CI of 0.74–0.99, a non-inferiority P-value of <0.0001, and a superiority P-value of 0.0382. The key secondary 4-point MACE analysis reported an HR of 0.89, with a 95.02% CI of 0.78–1.01, a non-inferiority P-value of <0.0001, and a superiority P-value of 0.0795.

The additional registry-posted analyses demonstrate why statistical interpretation must remain endpoint-specific. Heart failure requiring hospitalisation had an HR of 0.65, while new onset macroalbuminuria and the composite microvascular outcome each had an HR of 0.62. Silent MI and new onset albuminuria had HRs of 1.28 and 0.95, respectively. These results use the stated Cox model but differ in endpoint definition and, for several outcomes, analysis population.

Clinical Biostats methodology: A trial-results page should not merely repeat numerical results. The goal is to connect each estimate to its endpoint, analysis population, hypothesis, confidence interval, and statistical model while keeping reported evidence separate from educational interpretation.