This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
EMPA-REG OUTCOME was a randomized, double-blind, parallel phase 3 trial in patients with type 2 diabetes mellitus. The registry reports an enrollment of 7064 participants and evaluates empagliflozin against placebo using time-to-event cardiovascular and microvascular endpoints.
| Feature | EMPA-REG OUTCOME |
|---|---|
| Trial name | EMPA-REG OUTCOME |
| NCT identifier | NCT01131676 |
| Phase | Phase 3 |
| Condition | Diabetes Mellitus, Type 2 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 7064.0 |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | Industry |
| Trial status | Completed |
2. Clinical Question
The central statistical question was whether empagliflozin was non-inferior to placebo for the time to first occurrence of the registered 3-point major adverse cardiovascular event (MACE) composite, followed within the prespecified testing strategy by a superiority assessment.
Population
Patients with Diabetes Mellitus, Type 2 enrolled in the EMPA-REG OUTCOME trial.
Intervention
BI 10773 (empagliflozin), represented by low-dose and high-dose study arms.
Comparator
Placebo corresponding to the empagliflozin study arms.
Primary question
Is the time to the first occurrence of 3-point MACE non-inferior to placebo, and does the subsequent superiority test support a lower hazard?
3. Trial Design
Empagliflozin 10 mg
- BI 10773 low-dose study intervention
- Serious adverse events: 876/2345 affected/at risk
Empagliflozin 25 mg
- BI 10773 high-dose study intervention
- Serious adverse events: 913/2342 affected/at risk
Placebo
- Placebo study intervention
- Serious adverse events: 988/2333 affected/at risk
4. Trial Conduct and Registry Details
The ClinicalTrials.gov record identifies randomization, parallel allocation, and double masking. They do not provide a crossover description, factorial structure, interim-analysis specification, missing-data or imputation strategy, or Bayesian analysis. Those design features are therefore not incorporated into the statistical interpretation of this page.
5. Primary Endpoint
| Endpoint | Registry definition | Time frame | Endpoint type |
|---|---|---|---|
| 3-point MACE | Time to the first occurrence of any of the following adjudicated components of the primary composite endpoint: cardiovascular (CV) death (including fatal stroke and fatal myocardial infarction (MI)), non-fatal MI (excluding silent MI), and non-fatal stroke. Percentage of patients with the event are presented. | From randomisation to individual end of observation, up to 4.6 years | Time-to-event |
The primary endpoint is a composite time-to-event endpoint. A participant contributes an event when the first qualifying component occurs. Because the endpoint combines CV death, non-fatal MI, and non-fatal stroke, its treatment effect summarizes the time until the first occurrence of any of those adjudicated events rather than estimating a separate effect for each component.
6. Statistical Methodology
Cox proportional-hazards model
The registry reports the Cox proportional-hazards model as the primary analytical method. The effect measure is the hazard ratio (HR), comparing All Empagliflozin with Placebo.
An HR below 1 indicates a lower estimated instantaneous event hazard in the empagliflozin group under the fitted Cox model. It is not an absolute risk difference, a probability of an event, or a statement that every participant experiences the same proportional reduction.
Analysis population
The primary analyses use the TS analysis population. Secondary analyses use either TS or TS (evaluable cases), as specified for each endpoint. This distinction matters because an analysis restricted to evaluable cases does not necessarily represent exactly the same population as the full TS analysis.
Covariate adjustment
For the reported superiority analysis of the primary endpoint and the secondary endpoints, the registry analysis notes specify a Cox proportional-hazards model with factors for treatment, age, gender, categorised BMI, HbA1c, eGFR and geographical region.
Non-inferiority framework
The primary objective was to establish non-inferiority of All Empagliflozin relative to placebo for time to first 3-point MACE. The ClinicalTrials.gov record states that the non-inferiority margin was 1.3, chosen based on FDA guidance for cardiovascular-risk evaluation of new antidiabetic therapies.
For a hazard ratio where values above 1 represent greater hazard with empagliflozin, the non-inferiority question asks whether the plausible upper boundary remains below the prespecified margin of 1.3. This is different from simply asking whether a conventional superiority p-value is below a chosen threshold.
Multiplicity and hierarchical testing
The registry states that a 4-step hierarchical testing strategy was followed: non-inferiority testing of the primary endpoint, followed by non-inferiority testing of the key secondary endpoint, followed by superiority testing of the primary endpoint and then the key secondary endpoint. The non-inferiority margin for the primary and key secondary endpoints was 1.3.
| Testing step | Endpoint role | Hypothesis type | Margin / comparison |
|---|---|---|---|
| Step 1 | Primary 3-point MACE | Non-inferiority | Margin 1.3 |
| Step 2 | Key secondary 4-point MACE | Non-inferiority | Margin 1.3 |
| Step 3 | Primary 3-point MACE | Superiority | HR comparison with 1 |
| Step 4 | Key secondary 4-point MACE | Superiority | HR comparison with 1 |
7. Primary Results: 3-Point MACE
Non-inferiority analysis
Hazard ratio for first 3-point MACE
95.02% CI: 0.74–0.99 · P < 0.0001
All Empagliflozin vs Placebo · TS population
The reported Cox proportional-hazards model estimated an HR of 0.86 for All Empagliflozin versus Placebo. The two-sided 95.02% confidence interval was 0.74–0.99, and the registry reports P < 0.0001 for the non-inferiority hypothesis.
An HR of 0.86 corresponds to a 14% lower estimated hazard of the first 3-point MACE event for All Empagliflozin relative to Placebo, using the usual interpretation of the hazard ratio. This is a relative time-to-event measure; it does not mean that exactly 14% of participants avoided an event, nor does it directly provide an absolute reduction in event probability.
The confidence interval of 0.74–0.99 indicates the precision of the estimated HR under the model and statistical framework. It does not describe the range of effects that individual patients experienced.
The reported P < 0.0001 is evidence against the non-inferiority null under the prespecified testing framework; it is not a measure of the size of the treatment effect. The key non-inferiority comparison is with the prespecified margin of 1.3, not merely with 1.0.
Superiority analysis
Superiority hazard ratio for first 3-point MACE
95.02% CI: 0.74–0.99 · P = 0.0382
All Empagliflozin divided by Placebo
The same estimated HR of 0.86 was reported for the subsequent superiority analysis, with a two-sided 95.02% confidence interval of 0.74–0.99 and a reported P-value of 0.0382. The superiority Cox model included treatment, age, gender, categorised BMI, HbA1c, eGFR and geographical region.
The HR of 0.86 indicates a lower estimated hazard for the first 3-point MACE event in the All Empagliflozin group relative to Placebo. The estimate describes the relative event hazard over the analyzed observation period; it is not a direct estimate of absolute event probability or individual benefit.
The 0.74–0.99 confidence interval is relatively close to 1 at its upper boundary, illustrating why the precision of the estimate matters in addition to the point estimate. A confidence interval describes statistical uncertainty around the estimated treatment effect; it does not establish that the true effect for every patient lies within that interval.
The superiority P = 0.0382 quantifies evidence against the null hypothesis used for this superiority test under the specified analysis. It does not quantify the magnitude or clinical importance of the effect. The reported analysis also sits within the trial's hierarchical testing strategy, so its interpretation should not be separated from that prespecified sequence.
8. Key Secondary Endpoint: 4-Point MACE
The registry reports a key secondary time-to-event endpoint defined as the first occurrence of all events adjudicated in the 4-point MACE composite: CV death, non-fatal MI, non-fatal stroke, and hospitalization for unstable angina pectoris.
| Endpoint | Analysis population | Method | Effect measure | Estimate | 95.02% CI | P-value |
|---|---|---|---|---|---|---|
| 4-point MACE | TS | Cox proportional-hazards model | Hazard ratio | 0.89 | 0.78–1.01 | <0.0001 for non-inferiority |
Non-inferiority analysis
Hazard ratio for 4-point MACE
95.02% CI: 0.78–1.01 · P < 0.0001
All Empagliflozin vs Placebo · TS population
The registry reports an HR of 0.89 with a two-sided 95.02% confidence interval of 0.78–1.01. The non-inferiority P-value is reported as <0.0001, with the non-inferiority margin set at 1.3.
An HR of 0.89 corresponds to an 11% lower estimated hazard of the first 4-point MACE event for All Empagliflozin versus Placebo. Again, this is not an 11-percentage-point reduction in event probability and does not mean that each participant experiences an identical reduction.
The confidence interval of 0.78–1.01 extends slightly above 1.0, so the interval does not by itself exclude equal hazards. For the non-inferiority question, however, the relevant boundary is 1.3, not 1.0. The registry reports P < 0.0001 for that non-inferiority assessment.
The P-value addresses the hypothesis being tested; it does not measure the magnitude or practical importance of the HR. The composite endpoint also contains an additional component—hospitalization for unstable angina pectoris—relative to the 3-point MACE definition.
Superiority analysis
Superiority hazard ratio for 4-point MACE
95.02% CI: 0.78–1.01 · P = 0.0795
All Empagliflozin divided by Placebo
For the subsequent superiority analysis, the registry reports the same HR of 0.89 and 95.02% CI of 0.78–1.01, with a superiority P-value of 0.0795.
The point estimate remains below 1, corresponding to an estimated 11% lower hazard in the All Empagliflozin group. The confidence interval, however, includes 1.0, indicating uncertainty that encompasses equal hazards.
The reported P = 0.0795 is a hypothesis-test result for superiority and should not be interpreted as the probability that the treatment has no effect. It also should not be treated as a measure of effect size. The distinction between the non-inferiority and superiority questions is essential: the same HR can support non-inferiority against a margin of 1.3 while not providing the same evidence for superiority against 1.0.
Because this superiority test appears as a later step in the stated hierarchical sequence, its interpretation also depends on the trial's prespecified multiplicity strategy.
9. Secondary Time-to-Event Results
The ClinicalTrials.gov record contains five additional secondary time-to-event analyses. All compare All Empagliflozin with Placebo using Cox proportional-hazards models.
| Secondary endpoint | Analysis population | HR | 95% CI | P-value |
|---|---|---|---|---|
| Percentage of Participants With Silent MI | TS (evaluable cases) | 1.28 | 0.70–2.33 | 0.4172 |
| Percentage of Participants With Heart Failure Requiring Hospitalisation (Adjudicated) | TS | 0.65 | 0.50–0.85 | 0.0017 |
| Percentage of Participants With New Onset Albuminuria | TS (evaluable cases) | 0.95 | 0.87–1.04 | 0.2547 |
| Percentage of Participants With New Onset Macroalbuminuria | TS (evaluable cases) | 0.62 | 0.54–0.72 | <0.0001 |
| Percentage of Participants With the Composite Microvascular Outcome | TS (evaluable cases) | 0.62 | 0.54–0.70 | <0.0001 |
Heart failure requiring hospitalisation
Hazard ratio
95% CI: 0.50–0.85 · P = 0.0017
All Empagliflozin vs Placebo · TS population
The HR of 0.65 corresponds to a 35% lower estimated hazard of adjudicated heart failure requiring hospitalisation in the All Empagliflozin group relative to Placebo, using the conventional interpretation of the reported HR.
The 95% confidence interval of 0.50–0.85 indicates the statistical uncertainty around the estimated HR. The interval remains below 1.0, while the reported P-value is 0.0017. The P-value does not measure the size of the treatment effect, and the HR does not directly state an absolute difference in hospitalization probability.
Because this is a time-to-event endpoint, censoring and the assumptions of the Cox model remain relevant. The registry supplies the Cox method but does not provide enough information in the ClinicalTrials.gov record to independently assess the proportional-hazards assumption.
New onset macroalbuminuria
Hazard ratio
95% CI: 0.54–0.72 · P < 0.0001
All Empagliflozin vs Placebo · TS (evaluable cases)
The reported HR of 0.62 represents a 38% lower estimated hazard of new onset macroalbuminuria for All Empagliflozin relative to Placebo. The confidence interval of 0.54–0.72 provides the reported uncertainty around this estimate.
Composite microvascular outcome
Hazard ratio
95% CI: 0.54–0.70 · P < 0.0001
All Empagliflozin vs Placebo · TS (evaluable cases)
The composite microvascular endpoint has the same point estimate as new onset macroalbuminuria, 0.62, but a different confidence interval of 0.54–0.70. This illustrates why the point estimate alone is insufficient: the endpoint definition and uncertainty interval both matter.
Silent myocardial infarction
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Silent MI | 1.28 | 0.70–2.33 | 0.4172 |
The silent-MI analysis produced an HR of 1.28 with a 95% CI of 0.70–2.33. The interval is wide and includes 1.0. The registry reports P = 0.4172 for the superiority hypothesis. The analysis population was TS (evaluable cases).
New onset albuminuria
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| New onset albuminuria | 0.95 | 0.87–1.04 | 0.2547 |
The HR of 0.95 indicates an estimated hazard close to that of the comparator, while the 95% CI of 0.87–1.04 includes 1.0. The reported superiority P-value is 0.2547.
10. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
The registered endpoints are time-to-event outcomes, meaning that both whether an event occurred and when it occurred contribute information. A Cox model is designed for this setting and produces a hazard ratio that compares the instantaneous event hazards between treatment groups while allowing for censored observations.
What does an HR of 0.86 mean?
An HR of 0.86 means that the estimated instantaneous hazard in the All Empagliflozin group is 0.86 times the estimated hazard in the Placebo group under the fitted model. Expressed as a relative difference, this is a 14% lower estimated hazard. It does not mean that 14% of participants benefited or that event probability was reduced by exactly 14 percentage points.
Why was the non-inferiority margin 1.3 important?
Non-inferiority asks whether the treatment is not unacceptably worse than the comparator by more than a prespecified amount. For this trial, the ClinicalTrials.gov record specifies a margin of 1.3. Therefore, the relevant question is whether the confidence bound remains below 1.3—not simply whether the confidence interval excludes 1.0.
How can non-inferiority and superiority have different P-values for the same HR?
They are different hypotheses. The non-inferiority analysis tests against a margin of 1.3, whereas superiority tests whether the hazard ratio is compatible with the null value of 1.0. In EMPA-REG OUTCOME, the primary HR is 0.86, with P < 0.0001 for the reported non-inferiority test and P = 0.0382 for the reported superiority test.
Why does the 4-point MACE result illustrate this distinction particularly well?
The 4-point MACE HR is 0.89 with a 95.02% CI of 0.78–1.01. The upper confidence limit is well below the non-inferiority margin of 1.3, supporting the reported non-inferiority result. But the interval includes 1.0, and the reported superiority P-value is 0.0795. Thus, non-inferiority and superiority are not interchangeable claims.
Why does the analysis population matter?
The primary MACE analyses use the TS population, while silent MI, new onset albuminuria, new onset macroalbuminuria, and the composite microvascular outcome use TS (evaluable cases). Restricting an analysis to evaluable cases can change the population contributing data, so the analysis population should always be read alongside the estimate.
What does a confidence interval tell us that a P-value does not?
A confidence interval communicates the precision and plausible range of the estimated treatment effect under the specified statistical framework. A P-value summarizes evidence against a particular null hypothesis. Neither one directly measures clinical importance, and neither should be confused with the probability that a treatment effect is true.
11. Confidence Intervals and Effect Size
The reported results illustrate several distinct patterns of statistical uncertainty.
| Endpoint | HR | Confidence interval | What the interval indicates |
|---|---|---|---|
| 3-point MACE | 0.86 | 0.74–0.99 | Relatively narrow interval ending below 1.0 |
| 4-point MACE | 0.89 | 0.78–1.01 | Interval includes 1.0 but remains below 1.3 |
| Silent MI | 1.28 | 0.70–2.33 | Wide interval spanning both below and above 1.0 |
| Heart failure requiring hospitalisation | 0.65 | 0.50–0.85 | Interval below 1.0 |
| New onset albuminuria | 0.95 | 0.87–1.04 | Interval includes 1.0 and is relatively close to it |
| New onset macroalbuminuria | 0.62 | 0.54–0.72 | Interval below 1.0 |
| Composite microvascular outcome | 0.62 | 0.54–0.70 | Interval below 1.0 |
A confidence interval should be read together with the endpoint definition, analysis population, model, and hypothesis being tested. In particular, an interval crossing 1.0 is not equivalent to proof of no effect; it indicates that the data are compatible with equal hazards within the stated interval and statistical framework.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized arm as the number affected divided by the number at risk.
| Arm | Serious adverse events affected / at risk |
|---|---|
| Placebo | 988 / 2333 |
| Empagliflozin 10 mg | 876 / 2345 |
| Empagliflozin 25 mg | 913 / 2342 |
The figures above are reported as affected participants divided by participants at risk. They should not be transformed here into calculated percentages because the page is restricted to the registry-reported trial numbers and the instruction not to recompute or round results.
13. Registry Limitations and Data Integrity Notes
The ClinicalTrials.gov record contains several operational caveats that matter when interpreting the reported enrollment and analyses.
- Duplicate records: 2 sets of duplicates involving 4 patients were counted only once.
- Sites: 652 sites were initiated and 611 enrolled.
- Site transfer: one site transferred all patients to other sites.
- Exclusions: 27 patients randomized and 13 patients screened were excluded from analyses due to serious non-compliance.
14. Time-to-Event Endpoints and Censoring
The primary and secondary efficacy outcomes in the ClinicalTrials.gov record are time-to-event endpoints. The time frame is consistently described as from randomisation to individual end of observation, up to 4.6 years.
A simple comparison of percentages at a single time point can discard information about when events occurred and how long participants were followed. Cox regression instead uses the observed event times and censoring structure to estimate a relative hazard.
Because follow-up ends at each participant's individual end of observation, some observations can be censored rather than observed through an event. The ClinicalTrials.gov record identifies the Cox model but do not provide enough information to reconstruct individual censoring histories or independently evaluate censoring assumptions.
15. Non-Inferiority vs Superiority
EMPA-REG OUTCOME is particularly useful for teaching the distinction between two related but different clinical-trial questions.
Non-inferiority
Asks whether the treatment's hazard is sufficiently below the prespecified upper margin of 1.3 to rule out an unacceptably large disadvantage relative to placebo.
Superiority
Asks whether the treatment hazard is statistically distinguishable from the null value of 1.0 in the favorable direction.
| Endpoint | HR | 95.02% CI | NI P-value | Superiority P-value |
|---|---|---|---|---|
| 3-point MACE | 0.86 | 0.74–0.99 | <0.0001 | 0.0382 |
| 4-point MACE | 0.89 | 0.78–1.01 | <0.0001 | 0.0795 |
The two rows show why a trial can provide evidence for non-inferiority without the same endpoint necessarily satisfying a superiority test. The relevant null value changes from the non-inferiority margin of 1.3 to the superiority value of 1.0.
16. Multiplicity and Hierarchical Testing
The registry-reported analysis notes explicitly identify multiplicity adjustment for the primary endpoint and the key secondary endpoint. The four-step hierarchical sequence is therefore an important part of the statistical interpretation rather than a peripheral technical detail.
| Step | Endpoint | Question | Reported result |
|---|---|---|---|
| 1 | 3-point MACE | Non-inferiority | HR 0.86; P < 0.0001 |
| 2 | 4-point MACE | Non-inferiority | HR 0.89; P < 0.0001 |
| 3 | 3-point MACE | Superiority | HR 0.86; P = 0.0382 |
| 4 | 4-point MACE | Superiority | HR 0.89; P = 0.0795 |
Hierarchical testing is designed to control how confirmatory conclusions are sequenced across multiple hypotheses. The important statistical lesson is that individual P-values should be interpreted in the context of the prespecified testing hierarchy rather than as isolated numbers.
17. What the Hazard Ratio Does — and Does Not — Mean
An HR of 0.65 for heart failure requiring hospitalisation means that the fitted Cox model estimates a lower instantaneous hazard in the All Empagliflozin group than in the Placebo group. Expressed relatively, the estimated hazard is 35% lower.
It does not mean that 35% of participants avoided hospitalization, that the absolute hospitalization risk fell by 35 percentage points, or that every individual had exactly the same proportional reduction.
The 95% confidence interval of 0.50–0.85 for this endpoint describes uncertainty around the estimated HR under the model and sampling framework. It does not describe the range of treatment effects among individual patients.
The P-value addresses evidence against the specified null hypothesis. A small P-value does not imply a large treatment effect, while a larger P-value does not prove that the treatment has no effect. Effect size and uncertainty should be read from the HR and confidence interval.
18. Composite Endpoints: Why the Definition Matters
EMPA-REG OUTCOME's primary endpoint is a composite of three adjudicated events: CV death, non-fatal MI excluding silent MI, and non-fatal stroke. The key secondary composite adds hospitalization for unstable angina pectoris.
3-point MACE
CV death, including fatal stroke and fatal MI; non-fatal MI excluding silent MI; and non-fatal stroke.
4-point MACE
CV death, non-fatal MI, non-fatal stroke, and hospitalization for unstable angina pectoris.
A composite endpoint allows several clinically relevant events to contribute to one time-to-event analysis, but the resulting HR applies to the first occurrence of the composite. It should not automatically be interpreted as though it represents the treatment effect on each individual component separately.
19. Analysis Population and Evaluable Cases
| Endpoint | Analysis population |
|---|---|
| Primary 3-point MACE | TS |
| Key secondary 4-point MACE | TS |
| Silent MI | TS (evaluable cases) |
| Heart failure requiring hospitalisation | TS |
| New onset albuminuria | TS (evaluable cases) |
| New onset macroalbuminuria | TS (evaluable cases) |
| Composite microvascular outcome | TS (evaluable cases) |
The use of TS (evaluable cases) for several secondary endpoints means that those analyses are not necessarily based on precisely the same analytic population as the primary MACE analysis. That distinction is particularly important when comparing estimates across endpoints.
20. Important Limitations and Interpretation Issues
- Registry-level information: this page is limited to the ClinicalTrials.gov record. Details not contained in those data are not added from external publications or memory.
- Proportional-hazards assumption: the Cox model produces a hazard ratio whose usual interpretation relies on the model's proportional-hazards structure. The ClinicalTrials.gov record does not provide diagnostics to assess that assumption.
- Analysis population: primary MACE analyses use TS, whereas several secondary analyses use TS (evaluable cases). Estimates should therefore be interpreted with their stated population.
- Composite endpoints: the HR applies to the first qualifying event in the composite, not automatically to every component individually.
- Non-inferiority versus superiority: the margin of 1.3 and the null value of 1.0 answer different questions. A non-inferiority result should not automatically be described as evidence of superiority.
- Multiplicity: the primary and key secondary endpoint analyses were embedded in a 4-step hierarchical testing strategy. Secondary P-values should not be treated as though they were all independent confirmatory tests.
- Safety comparison: serious adverse-event counts are reported by arm, but the ClinicalTrials.gov record does not include a formal statistical comparison of those safety counts.
- Operational exclusions: the registry notes that 27 randomized patients and 13 screened patients were excluded from analyses because of serious non-compliance.
- Missing-data details: the ClinicalTrials.gov record does not specify a missing-data or imputation strategy, so none is inferred.
- Interim analysis: the ClinicalTrials.gov record does not provide an interim-analysis specification, so no interim-monitoring interpretation is added.
- Bayesian methods: no Bayesian method is reported in the statistical analyses posted on ClinicalTrials.gov.
21. Why This Trial Matters Statistically
EMPA-REG OUTCOME is a useful teaching case because the registry data bring together several core concepts in clinical-trial survival analysis: randomized treatment allocation, double masking, composite time-to-event endpoints, Cox proportional-hazards modeling, hazard ratios, confidence intervals, non-inferiority margins, superiority testing, and hierarchical multiplicity control.
| Concept | How it appears in EMPA-REG OUTCOME |
|---|---|
| Randomization | The trial is registered as randomized with a parallel design. |
| Blinding | The trial is registered as double-masked. |
| Time-to-event endpoint | The primary endpoint measures time to first 3-point MACE event from randomisation to individual end of observation, up to 4.6 years. |
| Cox model | The statistical analyses posted on ClinicalTrials.gov use Cox proportional-hazards models. |
| Hazard ratio | HR is the reported effect measure for the primary and secondary time-to-event analyses. |
| Confidence interval | Primary analyses report 95.02% confidence intervals; other registry-reported analyses report 95% confidence intervals. |
| Non-inferiority | The primary and key secondary analyses use a prespecified margin of 1.3. |
| Superiority | Superiority is tested after the non-inferiority steps in the stated 4-step hierarchy. |
| Multiplicity | The registry analysis notes identify multiplicity adjustment and a hierarchical testing strategy. |
| Analysis populations | Analyses use TS or TS (evaluable cases), depending on endpoint. |
| Safety | Serious adverse events are reported as affected/at risk for each study arm. |
22. A Statistical Reading of the Overall Results
The primary 3-point MACE analysis produced an HR of 0.86 with a 95.02% CI of 0.74–0.99. In the registry analysis, this result supported the prespecified non-inferiority objective and was subsequently evaluated for superiority, with a reported P-value of 0.0382.
The key secondary 4-point MACE analysis produced an HR of 0.89 with a 95.02% CI of 0.78–1.01. Its reported non-inferiority P-value was <0.0001, while the subsequent superiority P-value was 0.0795. The difference between these two hypothesis tests is an important statistical feature of the trial.
Among the additional secondary endpoints reported, the HR was 0.65 for heart failure requiring hospitalisation, 0.62 for new onset macroalbuminuria, and 0.62 for the composite microvascular outcome. New onset albuminuria had an HR of 0.95, while silent MI had an HR of 1.28. These estimates come from different endpoints and, for several outcomes, different analysis populations, so they should not be collapsed into a single generalized treatment effect.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: EMPA-REG OUTCOME — NCT01131676. The trial registry is the source for the trial design, endpoint definitions, statistical analyses, reported estimates, confidence intervals, P-values, safety counts, and registry limitations represented on this page.
- Linked publication: PubMed record — PMID 41198855.
- Linked publication: PubMed record — PMID 39026628.
- Linked publication: PubMed record — PMID 38992713.
- Linked publication: PubMed record — PMID 38770818.
- Linked publication: PubMed record — PMID 38419787.
Continue with the statistical methods
Explore the survival-analysis, trial-design, inference, and multiplicity concepts that provide the statistical framework for interpreting EMPA-REG OUTCOME.
26. Record Summary
EMPA-REG OUTCOME provides a compact teaching example of how a randomized clinical trial can combine a time-to-event primary endpoint with a formal non-inferiority framework and a subsequent superiority assessment. The primary 3-point MACE analysis reported an HR of 0.86, with a 95.02% CI of 0.74–0.99, a non-inferiority P-value of <0.0001, and a superiority P-value of 0.0382. The key secondary 4-point MACE analysis reported an HR of 0.89, with a 95.02% CI of 0.78–1.01, a non-inferiority P-value of <0.0001, and a superiority P-value of 0.0795.
The additional registry-posted analyses demonstrate why statistical interpretation must remain endpoint-specific. Heart failure requiring hospitalisation had an HR of 0.65, while new onset macroalbuminuria and the composite microvascular outcome each had an HR of 0.62. Silent MI and new onset albuminuria had HRs of 1.28 and 0.95, respectively. These results use the stated Cox model but differ in endpoint definition and, for several outcomes, analysis population.