← Clinical Trials
Heart Failure Phase 3 Time-to-Event NCT03057977

EMPEROR-Reduced: Complete Statistical Analysis of Empagliflozin in Chronic Heart Failure With Reduced Ejection Fraction

An independent statistical review of the randomized phase 3 EMPEROR-Reduced trial comparing 10 mg empagliflozin with placebo in patients with heart failure, focusing on the registered cardiovascular death or hospitalization for heart failure endpoint and the trial's reported survival, recurrent-event, renal, longitudinal, and patient-reported outcome analyses.

Trial status: COMPLETED  ·  Enrollment: 3730  ·  Primary completion: 2020-05-01
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are taken from the ClinicalTrials.gov record. The registry reports 10 outcome measures and 10 statistical analyses, including one primary-endpoint analysis.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

EMPEROR-Reduced was a randomized, double-blind, parallel phase 3 trial evaluating empagliflozin versus placebo in heart failure. The registry identifies one primary time-to-event endpoint: time to the first event of adjudicated cardiovascular death or adjudicated hospitalization for heart failure.

3730
Enrolled
Randomized patients
2
Arms
Empagliflozin vs placebo
0.75
Primary HR
95.04% CI 0.65–0.86
<0.0001
Primary P-value
Two-sided analysis
FeatureEMPEROR-Reduced
PhasePhase 3
ConditionHeart Failure
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Enrollment3730
InterventionsEmpagliflozin; Placebo
Primary endpoint typeTime-to-event
Results postedYes
Statistical analyses posted10
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry
ClinicalTrials.govNCT03057977

2. Clinical Question

The trial asks whether treatment with 10 mg empagliflozin, compared with placebo, changes the time to the first event of adjudicated cardiovascular death or adjudicated hospitalization for heart failure during the planned treatment period.

Population

Patients enrolled in the phase 3 EMPEROR-Reduced trial for the condition recorded as heart failure.

Intervention

10 mg empagliflozin.

Comparator

Placebo.

Primary question

Does empagliflozin change the hazard of the first adjudicated cardiovascular death or adjudicated hospitalization for heart failure?

3. Trial Design

01
Randomize 3730 patients
02
Double-blind Empagliflozin or placebo
03
Follow Planned treatment period
04
Assess Time-to-event and longitudinal outcomes
05
Analyze Prespecified analysis populations
ARM A

10 mg Empagliflozin

  • Empagliflozin 10 mg
  • Randomized treatment assignment
  • Double-blind trial design
ARM B

Placebo

  • Placebo
  • Randomized treatment assignment
  • Double-blind trial design
Allocation
Randomized allocation was used to create the treatment comparison.
Masking
The registry describes the study as double-blind.
Design model
Parallel-group design with two treatment arms.
Study period
The trial started on 2017-03-06 and had a primary completion date of 2020-05-01.

4. Analysis Populations

The registry distinguishes the analysis population according to the question being analyzed. This distinction is important because a randomized-set analysis and a treated-set analysis answer somewhat different statistical questions.

PopulationDefinition in registry analysisRole
Randomised Set (RS) All randomised patients. Used for the primary endpoint and most time-to-event secondary endpoints.
Treated Set (TS) All patients treated with at least one dose of the study medication and at least one on-treatment measurement of eGFR. Used for the eGFR slope analysis.
Randomised set with available endpoint data Patients in the randomized set with available data for the KCCQ Clinical Summary Score endpoint, including values obtained on treatment or post-treatment. Used for the KCCQ longitudinal analysis.
Randomised set with pre-DM Patients in the randomised set with pre-diabetes. Used for the time-to-onset of diabetes mellitus analysis.

5. Endpoints

The registry identifies one primary endpoint and nine additional posted statistical analyses. The primary endpoint is a composite time-to-event outcome. The secondary endpoints include recurrent hospitalization outcomes, renal outcomes, mortality outcomes, diabetes onset, a patient-reported outcome, and an eGFR slope measure.

RoleEndpointTime frameType
Primary Time to the First Event of Adjudicated Cardiovascular (CV) Death or Adjudicated Hospitalisation for Heart Failure (HHF) From randomisation until completion of the planned treatment period, up to 1040 days. Time-to-event
Secondary Occurrence of Adjudicated Hospitalisation for Heart Failure (HHF) (First and Recurrent) From randomisation until completion of the planned treatment phase, up to 1040 days. Time-to-event
Secondary eGFR (CKD-EPI) cr Slope of Change From Baseline Assessed at baseline, week 4, 12, 32, 52, 76, 100, 124, 148 and at end of treatment (EOT), up to 1040 days. Continuous
Secondary Time to First Event in Composite Renal Endpoint: Chronic Dialysis, Renal Transplant or Sustained Reduction of eGFR(CKD-EPI)cr From randomisation until completion of the planned treatment period, up to 1040 days. Time-to-event
Secondary Time to First Adjudicated Hospitalisation for Heart Failure (HHF) From randomisation until completion of the planned treatment period, up to 1040 days. Time-to-event
Secondary Time to Adjudicated Cardiovascular (CV) Death From randomisation until completion of the planned treatment period, up to 1040 days. Time-to-event
Secondary Time to All-cause Mortality From randomisation until completion of the planned treatment period, up to 1040 days. Time-to-event
Secondary Time to Onset of Diabetes Mellitus (DM) From randomisation until completion of the planned treatment period, up to 1040 days. Time-to-event
Secondary Change From Baseline in KCCQ (Kansas City Cardiomyopathy Questionnaire) Clinical Summary Score at Week 52 Assessed at baseline, week 12, week 32 and week 52. Continuous
Secondary Number of All-cause Hospitalizations (First and Recurrent) From randomisation until completion of the planned treatment phase, up to 1040 days. Time-to-event

6. Statistical Methodology

Cox proportional-hazards model

The primary endpoint was analyzed with a Cox proportional-hazards model. The analysis population was the Randomised Set, defined as all randomized patients. The comparison was placebo versus 10 mg empagliflozin.

The primary model included terms for age, baseline eGFR (CKD-EPI)cr, region, baseline diabetes status, sex, baseline LVEF, and treatment. The registry describes the resulting effect measure as a hazard ratio.

Conceptual hazard-ratio model
h(t | X) = h0(t) exp(β1X1 + ··· + βpXp)

The Cox model relates covariates to the instantaneous event hazard. The treatment coefficient is transformed into a hazard ratio, which summarizes the relative event hazard between treatment groups under the fitted model.

Joint frailty models

Two recurrent-event analyses used a joint frailty model. The first evaluated first and recurrent adjudicated hospitalizations for heart failure. The model explicitly accounted for dependence between recurrent HHF and cardiovascular death.

The all-cause hospitalization analysis likewise used a joint frailty model, accounting for dependence between recurrent all-cause hospitalizations and all-cause mortality. This is statistically different from treating every hospitalization as an independent observation.

Random coefficient model for eGFR slope

The eGFR analysis used a random coefficient model allowing for a random intercept and random slope per patient. The model included the same major factors used for the primary endpoint, together with time, treatment-by-time interaction, and baseline eGFR-by-time interaction.

Mixed model for repeated KCCQ measurements

The KCCQ Clinical Summary Score analysis used a mixed model, specifically described in the registry analysis notes as a mixed model for repeated measures. Fixed effects included age, baseline eGFR as linear covariate(s), region, baseline diabetes status, sex, baseline LVEF, week reachable, treatment-by-visit interaction, and baseline KCCQ Clinical Summary Score-by-visit interaction. An unstructured covariance structure was used.

Covariate adjustment

Several models adjust for baseline prognostic variables rather than estimating the treatment effect from treatment assignment alone. In the primary analysis, the adjustment variables were age, baseline eGFR, region, baseline diabetes status, sex, baseline LVEF, and treatment.

Important distinction: covariate adjustment does not change the randomized treatment assignment. It changes the statistical model used to estimate the treatment effect and can improve precision when the included baseline variables explain outcome variation.

7. Primary Endpoint Result

The primary endpoint was time to the first event of adjudicated cardiovascular death or adjudicated hospitalization for heart failure, measured from randomization until completion of the planned treatment period, up to 1040 days.

Hazard ratio for the primary composite endpoint

0.75

95.04% CI: 0.65–0.86   ·   P < 0.0001

Analysis: Cox proportional-hazards model  ·  Randomised Set

Estimated relative hazard
Placebo
1.00 reference
10 mg Empagliflozin
HR 0.75
Clinical Biostats interpretation

An HR of 0.75 means that the fitted Cox model estimated the instantaneous hazard of the first adjudicated cardiovascular death or adjudicated hospitalization for heart failure to be about 25% lower with 10 mg empagliflozin than with placebo over the analyzed follow-up.

The HR does not mean that 25% of patients avoided the endpoint, that an individual patient's probability was reduced by exactly 25%, or that the absolute difference in event probability was 25 percentage points. A hazard ratio is a relative, model-based time-to-event measure.

The two-sided 95.04% confidence interval of 0.65–0.86 describes uncertainty around the estimated hazard ratio under the model and statistical framework used for the analysis. It does not describe the range of individual patient effects.

The P < 0.0001 value addresses evidence against the null hypothesis specified for the comparison. It is not a measure of the size or clinical importance of the treatment effect. The HR and its confidence interval provide the effect-size information.

Because this is a Cox analysis, interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative hazard under the model rather than describing the entire time-varying event experience of every patient.

The registry analysis notes also identify interim analysis / alpha spending, with alpha reported as 0.0496 resulting from the interim analysis. That design feature matters when interpreting the formal evidence threshold because the analysis was not simply an unplanned single look at the data.

8. Secondary Endpoint Results

First and Recurrent Hospitalisation for Heart Failure

The occurrence of adjudicated hospitalization for heart failure, including first and recurrent events, was analyzed using a joint frailty model. The model accounted for dependence between recurrent HHF and cardiovascular death.

Hazard ratio for first and recurrent HHF

0.70

95.04% CI: 0.58–0.85   ·   P = 0.0003

Analysis: Joint frailty model  ·  Randomised Set

An HR of 0.70 corresponds to an estimated 30% lower hazard in the empagliflozin group relative to placebo under the fitted model. Because this analysis concerns recurrent hospitalizations and mortality jointly, the estimate should not be interpreted as though it came from a simple comparison of the proportion of patients hospitalized.

eGFR Slope of Change From Baseline

The eGFR slope endpoint was assessed repeatedly at baseline, weeks 4, 12, 32, 52, 76, 100, 124, 148 and end of treatment, up to 1040 days. The analysis used a random coefficient model with a random intercept and random slope per patient.

Treatment-by-time interaction

1.733

99.9% CI: 0.669–2.796   ·   P < 0.0001

Analysis: Random intercept random coefficient model  ·  Treated Set

The reported effect measure is a treatment-by-time interaction, not a hazard ratio. It therefore should not be described as a percentage reduction in risk. Its meaning is tied to the fitted longitudinal model and the difference in modeled eGFR trajectories over time.

Composite Renal Endpoint

The composite renal endpoint was time to the first event of chronic dialysis, renal transplant, or sustained reduction of eGFR(CKD-EPI)cr. It was analyzed with a Cox proportional-hazards model in the Randomised Set.

Hazard ratio for the composite renal endpoint

0.50

95% CI: 0.32–0.77   ·   P = 0.0019

Analysis: Cox proportional-hazards model  ·  Randomised Set

An HR of 0.50 represents a 50% lower estimated hazard under the fitted model. The confidence interval indicates substantial statistical uncertainty around the point estimate, but the entire reported interval is below 1.

First Adjudicated Hospitalisation for Heart Failure

Hazard ratio for first HHF

0.69

95% CI: 0.59–0.81   ·   P < 0.0001

Analysis: Cox proportional-hazards model  ·  Randomised Set

The first-HHF analysis estimated a 31% lower hazard with 10 mg empagliflozin relative to placebo. This endpoint is narrower than the primary composite because it considers the first adjudicated hospitalization for heart failure rather than the combined first occurrence of cardiovascular death or HHF.

Cardiovascular Death

Hazard ratio for adjudicated cardiovascular death

0.92

95% CI: 0.75–1.12   ·   P = 0.4133

Analysis: Cox proportional-hazards model  ·  Randomised Set

The point estimate of 0.92 corresponds to an estimated 8% lower hazard, but the confidence interval extends on both sides of 1. The reported P-value is 0.4133. This illustrates why a point estimate alone is insufficient: the confidence interval communicates the uncertainty surrounding that estimate.

All-cause Mortality

Hazard ratio for all-cause mortality

0.92

95% CI: 0.77–1.10   ·   P = 0.3536

Analysis: Cox proportional-hazards model  ·  Randomised Set

The all-cause mortality estimate was also below 1, with an HR of 0.92, but its 95% confidence interval includes 1. The result therefore should be read as an estimate with uncertainty rather than as proof of a specific mortality reduction.

Time to Onset of Diabetes Mellitus

Hazard ratio for onset of diabetes mellitus

0.86

95% CI: 0.62–1.19   ·   P = 0.3576

Analysis: Cox proportional-hazards model  ·  Randomised Set with pre-DM

This analysis was restricted to patients in the randomized set with pre-diabetes. The point estimate corresponds to an estimated 14% lower hazard, while the confidence interval spans values below and above 1. The result should therefore be interpreted primarily through the complete interval and the specified analysis population rather than through the point estimate alone.

KCCQ Clinical Summary Score at Week 52

Difference in adjusted means

2.06

95% CI: 0.16–3.96   ·   P = 0.0340

Analysis: Mixed model for repeated measures  ·  Randomised Set with available endpoint data

The reported effect measure is a difference of adjusted means. Unlike the hazard ratios above, this estimate is expressed on the KCCQ Clinical Summary Score scale. The model uses repeated measurements and adjusts for baseline and visit-related covariates.

First and Recurrent All-cause Hospitalizations

Hazard ratio for all-cause hospitalizations

0.85

95% CI: 0.75–0.95   ·   P = 0.0065

Analysis: Joint frailty model  ·  Randomised Set

The HR of 0.85 corresponds to a 15% lower estimated hazard under the joint frailty model. Because the endpoint includes first and recurrent hospitalizations and the model accounts for dependence with all-cause mortality, the result is not equivalent to a simple risk ratio for "ever hospitalized."

9. Results Overview

EndpointEffect estimateConfidence intervalP-valueModel
Primary: CV death or HHF HR 0.75 95.04% CI 0.65–0.86 <0.0001 Cox proportional-hazards
First and recurrent HHF HR 0.70 95.04% CI 0.58–0.85 0.0003 Joint frailty
eGFR slope 1.733 99.9% CI 0.669–2.796 <0.0001 Random coefficient
Composite renal endpoint HR 0.50 95% CI 0.32–0.77 0.0019 Cox proportional-hazards
First HHF HR 0.69 95% CI 0.59–0.81 <0.0001 Cox proportional-hazards
CV death HR 0.92 95% CI 0.75–1.12 0.4133 Cox proportional-hazards
All-cause mortality HR 0.92 95% CI 0.77–1.10 0.3536 Cox proportional-hazards
Onset of diabetes mellitus HR 0.86 95% CI 0.62–1.19 0.3576 Cox proportional-hazards
KCCQ Clinical Summary Score at week 52 Mean difference 2.06 95% CI 0.16–3.96 0.0340 Mixed model
First and recurrent all-cause hospitalizations HR 0.85 95% CI 0.75–0.95 0.0065 Joint frailty

The results illustrate an important statistical principle: endpoints that appear related clinically can have different estimands and therefore different statistical interpretations. A composite cardiovascular endpoint, a recurrent hospitalization endpoint, an eGFR slope, a mortality endpoint, and a patient-reported score cannot be reduced to one common effect measure.

10. Statistical Methods Explained

Why was a Cox proportional-hazards model used for the primary endpoint?

The primary endpoint is a time-to-event outcome. Patients can experience the first qualifying event at different times, and some observations can be censored before the event occurs. A Cox model uses the timing information rather than reducing the analysis to a simple event/no-event comparison.

Its principal effect measure is the hazard ratio. In EMPEROR-Reduced, the primary HR was 0.75, estimated after adjustment for the baseline variables specified in the registry analysis.

What does a hazard ratio of 0.75 mean?

An HR of 0.75 means the fitted model estimates a hazard 25% lower in the empagliflozin group relative to placebo. It is not a statement that 25% of patients benefited or that the probability of the endpoint fell by 25 percentage points.

The distinction matters because hazards are instantaneous rates conditional on having remained event-free to a particular time. Risk, cumulative incidence, and hazard are related but different quantities.

Why does the primary model adjust for baseline eGFR, LVEF, age, and other covariates?

Adjustment can account for prognostic baseline variables while estimating the treatment effect. The registry specifies age, baseline eGFR, region, baseline diabetes status, sex, baseline LVEF, and treatment in the primary model.

Because treatment assignment was randomized, these covariates are not required to create randomization. Their role is within the statistical model used to estimate the treatment comparison.

Why use a joint frailty model for recurrent hospitalizations?

Recurrent hospitalizations create a statistical dependence problem: the same patient can contribute more than one hospitalization, and patients also differ in their underlying susceptibility to hospitalization and death. The registry states that the joint frailty analysis accounts for dependence between recurrent HHF and cardiovascular death, and separately between recurrent all-cause hospitalizations and all-cause mortality.

A conventional analysis that treated every hospitalization as an independent observation would fail to represent this within-patient dependence appropriately.

What does the eGFR treatment-by-time interaction measure?

The eGFR analysis is longitudinal. Rather than asking only whether two groups differ at one time point, the random coefficient model estimates patient-specific trajectories using repeated measurements. The treatment-by-time interaction captures the difference in modeled change over time between treatment groups.

Because the reported estimate is 1.733 rather than a hazard ratio, it should not be translated into a percentage reduction in an event hazard.

Why was a mixed-effects model used for KCCQ?

KCCQ measurements were obtained repeatedly at baseline, week 12, week 32 and week 52. A mixed model for repeated measures can use the longitudinal structure of these observations and account for within-patient correlation. The registry specifies an unstructured covariance structure and includes treatment-by-visit and baseline-score-by-visit interactions.

Why does alpha spending matter?

The primary analysis notes identify an interim analysis and alpha spending, with alpha reported as 0.0496 resulting from the interim analysis. Interim monitoring creates an important statistical issue: repeatedly looking at accumulating data can affect the probability of a false-positive finding unless the testing procedure accounts for those looks.

Alpha spending is designed to allocate the available type I error across the planned analysis framework rather than treating every interim look as though it were the only analysis.

11. Confidence Intervals and P-values

The EMPEROR-Reduced results illustrate why a confidence interval and a P-value should be interpreted together with the effect estimate rather than treated as interchangeable quantities.

EndpointEstimateCI width / rangeP-value
Primary CV death or HHFHR 0.750.65–0.86<0.0001
First and recurrent HHFHR 0.700.58–0.850.0003
CV deathHR 0.920.75–1.120.4133
All-cause mortalityHR 0.920.77–1.100.3536
KCCQ at week 52Mean difference 2.060.16–3.960.0340
How to read the interval

A confidence interval describes uncertainty around the estimated parameter under the statistical model and sampling framework. A relatively narrow interval indicates greater precision than a much wider interval, all else equal. It does not describe the distribution of treatment effects among individual patients.

How to read the P-value

A P-value quantifies how incompatible the observed data and analysis statistic are with the specified null hypothesis, under the assumptions of the test. It does not tell us the probability that the null hypothesis is true, nor does it measure the magnitude of an effect.

Why both matter

The primary HR of 0.75 communicates the estimated relative treatment effect, the 0.65–0.86 confidence interval communicates its statistical precision, and the P-value of <0.0001 communicates evidence relative to the null hypothesis. These are three different pieces of information.

12. Time-to-Event Analysis

Most of the reported EMPEROR-Reduced endpoints are time-to-event outcomes. These analyses use the time from randomization to an event rather than simply asking whether an event occurred during the study.

Time-to-event structure
Randomisation → follow-up time → event or censoring

The analysis can use the timing of events and information from patients who have not experienced the endpoint by the time their follow-up ends.

The primary endpoint combines two clinically distinct event types: adjudicated cardiovascular death and adjudicated hospitalization for heart failure. The registry specifies that the endpoint is the first event of either component.

Several secondary endpoints then separate aspects of this composite, including first HHF and cardiovascular death, while other analyses explicitly address recurrent hospitalizations. This provides different statistical views of the same broad clinical domain.

Kaplan-Meier context: Kaplan-Meier estimation is a standard descriptive method for time-to-event data, but the ClinicalTrials.gov record identifies the formal primary method as Cox regression and do not state that a Kaplan-Meier method was the formal statistical method for the posted primary analysis. The page therefore does not attribute a Kaplan-Meier analysis to the registry beyond this general methodological context.

13. Covariate Adjustment and Model Structure

The primary Cox model included seven types of terms: age, baseline eGFR (CKD-EPI)cr, region, baseline diabetes status, sex, baseline LVEF, and treatment. Several secondary Cox analyses used closely related adjustment sets.

Variable / componentRole in analysis
AgeBaseline covariate in the primary model.
Baseline eGFR (CKD-EPI)crBaseline renal-function covariate.
RegionGeographic covariate in the model.
Baseline diabetes statusBaseline disease-status covariate.
SexBaseline demographic covariate.
Baseline LVEFBaseline cardiac-function covariate.
TreatmentRandomized treatment comparison.

The important statistical point is that the treatment effect remains the target parameter even when other baseline variables are included. The covariates help define the fitted model; they do not turn the randomized trial into an observational comparison.

14. Interim Analysis and Alpha Spending

The primary analysis notes identify interim analysis / alpha spending. The reported analysis specifies alpha = 0.0496, resulting from the interim analysis.

Why interim looks matter

If investigators repeatedly test accumulating data without accounting for those looks, the nominal type I error can no longer represent the overall false-positive probability of the testing strategy.

What alpha spending does

An alpha-spending framework allocates the available type I error across the planned analysis sequence so that interim monitoring can occur within a prespecified inferential framework.

The ClinicalTrials.gov record does not specify the complete spending function or the full sequence of interim information fractions. Accordingly, this page does not attribute a particular boundary or spending function to EMPEROR-Reduced beyond the registry's stated interim-analysis and alpha-spending information.

Interpretation caution: the primary P-value should be interpreted in the context of the interim-analysis framework and the reported alpha of 0.0496. It should not be treated as though the trial necessarily involved one completely unplanned, single final look at the data.

15. Multiplicity and Hypothesis Testing

The trial data identify the primary endpoint as a superiority analysis and provide statistical analyses for nine secondary endpoints. The primary analysis states the null hypothesis as: there is no difference between the effect of placebo and the effect of empagliflozin.

The presence of multiple endpoints creates an important interpretive distinction. A statistically small P-value for one endpoint does not automatically establish that every other endpoint is confirmatory, nor does the statistical evidence for one endpoint determine the effect estimate for another.

Analysis levelStatistical roleInterpretive issue
Primary endpoint Superiority comparison Formal analysis uses Cox regression with the reported interim-adjusted alpha.
Secondary endpoints Additional treatment-effect analyses Each has its own estimand, analysis population, model, confidence interval and P-value.
Recurrent-event endpoints Repeated outcome analysis Dependence between events and mortality is handled through joint frailty models.
Longitudinal endpoints Repeated measurements eGFR and KCCQ use models that explicitly represent longitudinal observations.

The ClinicalTrials.gov record does not provide a complete multiplicity-adjustment hierarchy for all 10 posted analyses. Therefore, the secondary P-values should be reported as the registry reports them rather than assigning them a confirmatory status that is not documented in the ClinicalTrials.gov record.

16. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected patients divided by patients at risk.

Safety measurePlacebo10 mg Empagliflozin
Serious adverse events 896/1863 772/1863
Serious adverse events: affected / at risk
Placebo
896 / 1863
10 mg Empagliflozin
772 / 1863

The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events. Accordingly, this page reports the affected/at-risk counts without constructing an additional hypothesis test.

Safety interpretation: safety and efficacy use different statistical perspectives. The randomized comparison of the primary efficacy endpoint is based on the Randomised Set, while safety summaries concern exposure to treatment. A safety count should therefore not be treated as another estimate of the primary treatment effect.

17. Interpreting the Primary Hazard Ratio

What the HR means

The primary HR of 0.75 means the fitted Cox model estimated the instantaneous hazard of the first adjudicated cardiovascular death or adjudicated hospitalization for heart failure to be approximately 75% of the placebo-group hazard, or approximately 25% lower.

What the HR does not mean

It does not mean that 25% of patients avoided the endpoint, that every patient's risk fell by 25%, or that the absolute event probability was reduced by 25 percentage points.

Why the confidence interval matters

The 95.04% CI of 0.65–0.86 indicates the statistical uncertainty around the estimated HR under the specified model. The interval is an uncertainty statement about the parameter estimate, not a range of individual patient responses.

Why the P-value does not measure effect size

The P-value of <0.0001 is evidence against the specified null hypothesis within the trial's testing framework. It does not say that the treatment effect is "99.99% effective," nor does it quantify the size of the effect. The HR and confidence interval perform that role.

Composite endpoints require component-level reading

The primary endpoint combines cardiovascular death and hospitalization for heart failure. Its HR therefore summarizes the composite endpoint rather than either component individually. The separately reported cardiovascular-death and first-HHF analyses show why the composite should not be interpreted as though it were a single type of event.

18. Comparing the Different Estimands

One of the most useful statistical lessons from EMPEROR-Reduced is that the word "effect" can refer to different estimands depending on the endpoint and model.

EndpointEstimand / effect measureInterpretation
Primary CV death or HHF Hazard ratio Relative hazard of the first composite event under a Cox model.
First and recurrent HHF Hazard ratio from joint frailty model Relative event hazard accounting for recurrent HHF and cardiovascular death dependence.
eGFR slope Treatment-by-time interaction Difference in modeled longitudinal trajectories rather than an event hazard.
KCCQ at week 52 Difference of adjusted means Difference between adjusted treatment-group means in the longitudinal model.
All-cause hospitalizations Hazard ratio from joint frailty model Relative event measure accounting for recurrent hospitalizations and mortality dependence.

This distinction prevents a common statistical error: treating every number below 1 as though it represented the same kind of "risk reduction." The number's meaning depends on the endpoint, model, population, and estimand.

19. Limitations

20. Why This Trial Matters Statistically

EMPEROR-Reduced is a useful statistical teaching case because it combines several important clinical-trial methods within one randomized comparison. The primary outcome is a time-to-event composite analyzed with a covariate-adjusted Cox model, while secondary analyses extend the framework to recurrent events, longitudinal renal measurements, patient-reported outcomes, mortality, and diabetes onset.

ConceptHow it appears in EMPEROR-Reduced
Randomization3730 patients were randomized to two parallel treatment arms.
Double blindingThe registry describes the study as double-blind.
Time-to-event analysisThe primary endpoint and multiple secondary endpoints are time-to-event outcomes.
Hazard ratioThe primary and several secondary time-to-event analyses report HRs.
Cox proportional-hazards modelUsed for the primary endpoint and several secondary time-to-event outcomes.
Covariate adjustmentAge, baseline eGFR, region, diabetes status, sex, baseline LVEF, and treatment appear in the primary model.
Recurrent eventsJoint frailty models address first and recurrent HHF and all-cause hospitalization outcomes.
Mixed-effects modelingKCCQ uses a mixed model for repeated measures.
Random coefficientsThe eGFR slope model allows a random intercept and random slope per patient.
Interim analysisThe primary analysis notes identify interim analysis / alpha spending.
Confidence intervalsReported alongside the principal treatment-effect estimates.
P-valuesReported for all 10 posted statistical analyses.

21. A Practical Reading Sequence for This Trial

When reading the EMPEROR-Reduced statistical results, a disciplined sequence helps avoid common interpretation errors.

01
Define Identify the endpoint and time frame
02
Identify Check the analysis population
03
Read Identify the statistical model
04
Estimate Read the effect and CI together
05
Interpret Use the P-value within its testing framework

For the primary endpoint, this sequence produces: a time-to-event composite, analyzed in the Randomised Set, using a Cox proportional-hazards model, producing HR 0.75 with a 95.04% CI of 0.65–0.86 and P < 0.0001. The interim-analysis and alpha-spending framework then provides important context for the formal inference.

22. Related Statistical Concepts

Learn more about the methods used in this trial:

23. Related Statistical Calculators

Use the same concepts in hands-on statistical calculations:

24. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts behind randomized trials, time-to-event endpoints, regression models, longitudinal data, confidence intervals, and clinical-trial inference.

25. Record Summary

EMPEROR-Reduced provides a compact example of how modern randomized-trial analysis can combine several statistical frameworks. Its primary endpoint is a time-to-event composite analyzed in the Randomised Set using a covariate-adjusted Cox proportional-hazards model, with a reported HR of 0.75, 95.04% CI 0.65–0.86, and P < 0.0001. Secondary analyses extend the statistical framework to recurrent hospitalization through joint frailty models, eGFR trajectories through a random coefficient model, and repeated KCCQ measurements through a mixed model.

The broader statistical lesson is that treatment effects should always be interpreted in the context of the endpoint, estimand, analysis population, model, confidence interval, and testing framework. An HR is not a mean difference; a treatment-by-time interaction is not a hazard ratio; a recurrent-event analysis is not a simple first-event analysis; and a P-value does not replace the effect estimate or its uncertainty.

Clinical Biostats methodology: The goal of a trial-results page is not merely to repeat reported numbers. It is to explain what each estimate represents, identify the statistical model that produced it, distinguish primary from secondary analyses, and make clear where the available registry data do and do not support additional interpretation.