← Clinical Trials
Heart Failure Phase 3 Completed NCT04435626

FINEARTS-HF: Complete Statistical Analysis of Finerenone in Heart Failure

An independent statistical review of the randomized phase 3 FINEARTS-HF trial evaluating finerenone versus placebo in participants with heart failure and left ventricular ejection fraction greater than or equal to 40%, with emphasis on recurrent heart failure events, cardiovascular death, symptom scores, functional status, renal events, and mortality.

Phase 3  ·  Randomized  ·  Parallel  ·  Quadruple-masked  ·  Enrollment 6016
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics are restricted to the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

FINEARTS-HF was a randomized, parallel, quadruple-masked phase 3 trial evaluating the efficacy and safety of finerenone in participants with heart failure and left ventricular ejection fraction greater than or equal to 40%. The trial enrolled 6016 participants and compared finerenone with placebo.

6016
Enrolled
Phase 3 trial
2
Arms
Finerenone vs placebo
0.84
Primary rate ratio
95% CI 0.74–0.95
0.0072
Primary P-value
Superiority analysis
FeatureFINEARTS-HF
Trial nameFINEARTS-HF
NCT identifierNCT04435626
PhasePhase 3
ConditionHeart Failure
PopulationParticipants with heart failure and left ventricular ejection fraction greater than or equal to 40%
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment6016
InterventionsFinerenone (BAY94-8862) and placebo
Lead sponsorBayer
Sponsor typeIndustry
Study datesStart: 2020-09-14; primary completion: 2024-05-15
StatusCompleted

2. Clinical Question

The central statistical question was whether randomized assignment to finerenone, compared with placebo, was associated with a difference in the occurrence of a composite endpoint consisting of cardiovascular death and total first and recurrent heart failure events in participants with heart failure and left ventricular ejection fraction greater than or equal to 40%.

Population

Participants with heart failure and left ventricular ejection fraction greater than or equal to 40%.

Intervention

Finerenone (BAY94-8862).

Comparator

Placebo.

Primary question

Does finerenone reduce the occurrence of cardiovascular death and total first and recurrent heart failure events relative to placebo?

3. Trial Design

01
Randomize6016 participants
02
Two armsFinerenone vs placebo
03
ParallelRandomized treatment comparison
04
Follow-upFrom randomization to end of study
05
AnalysisRecurrent events and other endpoints
ARM A

Finerenone

  • Finerenone (BAY94-8862)
  • Compared with placebo in the randomized parallel design
  • Included in the full analysis set for the reported efficacy analyses
ARM B

Placebo

  • Placebo
  • Randomized comparator arm
  • Included in the full analysis set for the reported efficacy analyses

The trial was quadruple-masked. In statistical terms, masking is intended to reduce the potential for knowledge of treatment assignment to influence participant behavior, investigator decisions, outcome assessment, or other trial processes. The registry data identify the masking level but do not provide additional masking-role details.

Allocation
Randomized. Randomization creates the primary framework for comparing the treatment groups while reducing systematic differences in assignment.
Design model
Parallel. Participants were assigned to one of two treatment groups rather than moving between treatment sequences.
Masking
Quadruple. The registry classifies the study as quadruple-masked.
Primary purpose
Treatment. The study was designed to evaluate treatment efficacy and safety.

4. Endpoints

The registry contains two entries for the primary endpoint, one describing the number of composite endpoint events and one describing the number of participants with composite endpoint events. Both use the same endpoint wording and time frame.

EndpointRegistry definitionTime frame
Occurrence of the Composite Endpoint of Cardiovascular Death and Total (First and Recurrent) Heart Failure Events Number of composite endpoint events of cardiovascular death and total (first and recurrent) heart failure events, defined as hospitalization for heart failure or urgent HF visit. From randomization up until the end of study, with an average study duration of 32 months.
Occurrence of the Composite Endpoint of Cardiovascular Death and Total (First and Recurrent) Heart Failure Events Number of participants with composite endpoint events of cardiovascular death and total (first and recurrent) heart failure events, defined as hospitalization for heart failure or urgent HF visit. From randomization up until the end of study, with an average study duration of 32 months.

Secondary endpoints with reported analyses

EndpointTime frameReported methodEffect measure
Occurrence of Total (First and Recurrent) Heart Failure Events From randomization up until the end of study, with an average study duration of 32 months Stratified Andersen-Gill model Rate ratio
Change From Baseline in Total Symptom Score (TSS) of the Kansas City Cardiomyopathy Questionnaire (KCCQ) From baseline to Month 6, 9 and 12 Mixed Models Analysis / MMRM Difference in LS Mean
Proportion of Participants Who Showed Improvement in NYHA Class From baseline to Month 12 Logistic regression Odds ratio
First Occurrence of Renal Composite Events From randomization up to end of study visit, with an average study duration of 32 months Stratified log-rank test; stratified Cox proportional hazards model Cause-specific hazard ratio
Occurrence of All-cause Mortality From randomization until end of study, with an average of 32 months Stratified log-rank; stratified Cox proportional hazards model Cause-specific hazard ratio

5. Statistical Methodology

Recurrent-event analysis: the Andersen-Gill model

The most distinctive feature of the primary analysis is that the endpoint includes total first and recurrent heart failure events. A participant can therefore contribute more than one event during follow-up. A conventional time-to-first-event analysis would discard information about subsequent heart failure events after the first event; the Andersen-Gill framework instead models recurrent event occurrences over time.

Primary effect measure
Rate Ratio = Finerenone event rate / Placebo event rate

The reported primary estimate was a rate ratio of 0.84, with a two-sided 95% confidence interval of 0.74–0.95.

The registry specifies a stratified Andersen-Gill model with a robust sandwich estimate for the covariance matrix of recurrent events. The robust covariance estimator is important because repeated events from the same participant are correlated; treating those observations as independent would generally understate the uncertainty.

Intention-to-treat analysis

The reported Andersen-Gill analyses were performed in the full analysis set and identify intention-to-treat analysis as an associated concept. The basic intention-to-treat principle is to preserve the randomized comparison rather than redefining treatment groups according to what participants subsequently received.

Mixed-effects model for repeated measurements

The KCCQ total symptom score was assessed longitudinally from baseline to Months 6, 9 and 12. The registry reports a mixed-effects model for repeated measures (MMRM). This approach is appropriate for repeated measurements because observations from the same participant are correlated and because the scientific question concerns change over multiple follow-up visits rather than a single isolated measurement.

Logistic regression

Improvement in NYHA class is a binary outcome. The registry reports logistic regression and an odds ratio comparing finerenone with placebo. Logistic regression models the log odds of the binary outcome as a function of treatment and any specified covariates or adjustment structure.

Stratified log-rank and Cox analysis

For the renal composite and all-cause mortality endpoints, the registry reports stratified log-rank testing together with stratified Cox proportional-hazards models. These methods are designed for time-to-event outcomes, where participants can have different follow-up times and some observations can be right-censored.

Hazard-ratio interpretation
HR < 1  →  lower estimated instantaneous event rate in the finerenone group

A hazard ratio is a relative, model-based time-to-event measure. It is not a probability, an absolute risk reduction, or the percentage of participants who experience benefit.

6. Primary Result: Cardiovascular Death and Total First and Recurrent Heart Failure Events

The primary endpoint was analyzed in the full analysis set using a stratified Andersen-Gill model. The reported effect measure was a rate ratio comparing finerenone with placebo.

Primary rate ratio

0.84

95% CI: 0.74–0.95   ·   P = 0.0072

Analysis: stratified Andersen-Gill model with robust sandwich covariance estimate for recurrent events.

Primary endpointFinerenone vs placebo95% CIP-value
Cardiovascular death and total first and recurrent heart failure events Rate ratio 0.84 0.74–0.95 0.0072
Clinical Biostats interpretation

The rate ratio of 0.84 means that the estimated event rate under the reported recurrent-event model was 0.84 times the corresponding rate in the placebo group. Expressed as a simple relative interpretation, this corresponds to an estimated 16% lower event rate for finerenone relative to placebo under the model.

This does not mean that 16% of participants avoided an event, that every participant had a 16% reduction in risk, or that the absolute difference in the number of events was 16 percentage points. The endpoint is a recurrent-event composite, so the rate ratio summarizes event occurrence rather than simply comparing the proportion of participants with at least one event.

The two-sided 95% confidence interval of 0.74–0.95 describes uncertainty around the estimated rate ratio under the specified statistical framework. It does not describe the range of effects that must occur in individual participants.

The P-value of 0.0072 addresses statistical evidence against the relevant null hypothesis within the reported analysis. It is not a measure of effect size: a P-value does not tell us that the treatment effect is "0.0072 large," nor does it describe clinical importance. Effect size and uncertainty are better represented by the rate ratio and its confidence interval.

Because this is an Andersen-Gill recurrent-event analysis, interpretation also depends on the model's handling of repeated events and within-participant dependence. The robust sandwich covariance estimate addresses correlation in the variance calculation, while the rate ratio itself remains a model-based summary of recurrent event occurrence.

Why a recurrent-event endpoint changes the interpretation

Suppose two participants each experience several heart failure events. A time-to-first-event analysis would focus only on the first event for each participant. The FINEARTS-HF primary endpoint instead counts total first and recurrent heart failure events alongside cardiovascular death. That makes the analysis sensitive not only to whether a participant experiences an initial event, but also to the subsequent burden of recurrent events.

This distinction is statistically important. A rate ratio from an Andersen-Gill model should not be translated into a conventional risk ratio without additional information. The estimand is tied to the recurrent-event framework used in the registry analysis.

7. Secondary Result: Total First and Recurrent Heart Failure Events

The secondary analysis isolated total first and recurrent heart failure events. It used the same general recurrent-event framework: a stratified Andersen-Gill model with robust sandwich covariance estimation.

Heart failure event rate ratio

0.82

95% CI: 0.71–0.94   ·   P = 0.0062

Full analysis set; finerenone vs placebo.

Clinical Biostats interpretation

The rate ratio of 0.82 corresponds to an estimated event rate equal to 82% of the placebo-group rate under the reported model, or an estimated 18% lower event rate for finerenone relative to placebo.

Again, this is not an 18-percentage-point reduction in the probability of experiencing heart failure. The endpoint counts first and recurrent events, so the measure incorporates repeated events rather than only the first event per participant.

The 95% confidence interval of 0.71–0.94 indicates the statistical precision of the estimated rate ratio. The interval remains below 1, but its width shows that the point estimate should not be treated as an exact or fixed treatment effect.

The P-value of 0.0062 is evidence against the null hypothesis under the reported statistical test. It does not quantify clinical magnitude or tell us the probability that the treatment "works."

8. Secondary Result: KCCQ Total Symptom Score

The registry reports change from baseline in the Total Symptom Score of the Kansas City Cardiomyopathy Questionnaire at Months 6, 9 and 12. The analysis used a mixed-effects model for repeated measures in the full analysis set with available KCCQ measurements.

Difference in least-squares mean

1.56

95% CI: 0.79–2.34   ·   P < .0001

Mixed-effects model for repeated measures; finerenone vs placebo.

Clinical Biostats interpretation

The reported difference in LS mean of 1.56 summarizes the adjusted difference between treatment groups from the mixed-effects repeated-measures analysis. The direction of the reported estimate is the finerenone group relative to placebo.

The confidence interval of 0.79–2.34 represents uncertainty around the estimated difference under the model. It does not describe the distribution of individual participant-level changes and should not be interpreted as a prediction interval for individual patients.

The P-value < .0001 indicates strong statistical evidence against the null hypothesis used for this reported comparison. It does not indicate that the treatment effect is large in clinical terms. Statistical significance and clinical importance are distinct questions.

Because the endpoint includes repeated measurements at Months 6, 9 and 12, the mixed-effects framework accounts for the longitudinal structure rather than treating every measurement as an independent observation. The registry specifically describes the analysis as an MMRM.

9. Secondary Result: Improvement in NYHA Class

The proportion of participants who showed improvement in NYHA class from baseline to Month 12 was analyzed using logistic regression in the full analysis set with an available NYHA measurement at baseline.

Odds ratio for improvement

1.01

95% CI: 0.88–1.15   ·   P = 0.9295

Logistic regression; odds ratio defined as finerenone / placebo.

Clinical Biostats interpretation

An odds ratio of 1.01 indicates that the estimated odds of improvement in NYHA class were approximately the same between the finerenone and placebo groups under the reported logistic regression model.

The 95% confidence interval of 0.88–1.15 spans 1.00. Thus, the data are compatible with modestly lower or modestly higher odds under the interval represented by this estimate and its confidence interval.

The P-value of 0.9295 does not measure the magnitude of the odds ratio. It indicates that the reported data provide little statistical evidence against the null comparison in this particular analysis. A nonsignificant P-value should not be converted into a claim that the two groups are proven identical.

An odds ratio is also not the same as a risk ratio. When interpreting a binary endpoint, the distinction matters because odds and probabilities have different mathematical definitions.

10. Secondary Result: First Occurrence of Renal Composite Events

The first occurrence of renal composite events was analyzed as a time-to-event endpoint from randomization through the end-of-study visit, with an average study duration of 32 months. The registry reports a stratified log-rank test and a stratified Cox proportional-hazards model.

Cause-specific hazard ratio

1.33

95% CI: 0.94–1.89   ·   P = 0.1071

Stratified log-rank test with stratified Cox proportional-hazards model.

Clinical Biostats interpretation

The reported hazard ratio of 1.33 means that the estimated cause-specific hazard of the first renal composite event was 1.33 times the corresponding hazard in the placebo group under the fitted Cox model.

A hazard ratio above 1 should not automatically be interpreted as evidence of harm. The confidence interval is 0.94–1.89, which includes 1.00, and the P-value is 0.1071. The reported analysis therefore does not provide conventional statistical evidence of a difference in this endpoint.

The confidence interval is important because it is relatively broad compared with the point estimate. The data are compatible with effects on either side of the null value represented by 1.00 within the reported interval.

As with other Cox-model results, the hazard ratio is a relative time-to-event measure rather than an absolute difference in event probability. Its interpretation also relies on the model framework and the handling of censoring.

11. Secondary Result: All-cause Mortality

All-cause mortality was analyzed from randomization until the end of study, with an average study duration of 32 months. The registry reports a stratified log-rank analysis and a stratified Cox proportional-hazards model in the full analysis set.

Cause-specific hazard ratio

0.93

95% CI: 0.83–1.06   ·   P = 0.2794

Stratified log-rank test with stratified Cox proportional-hazards model.

Clinical Biostats interpretation

The hazard ratio of 0.93 corresponds to an estimated hazard approximately 93% of the placebo-group hazard under the reported model. As a descriptive relative interpretation, the point estimate is consistent with an estimated 7% lower instantaneous hazard of all-cause mortality.

That point estimate should not be treated as proof of a mortality reduction. The 95% confidence interval of 0.83–1.06 includes 1.00, and the P-value is 0.2794. The reported analysis therefore does not establish a statistically detectable difference in all-cause mortality.

The confidence interval also illustrates why the point estimate alone is insufficient. A point estimate of 0.93 may look directionally favorable, but uncertainty around the estimate is what determines how precisely the treatment effect has been estimated.

12. Comparative View of the Reported Results

EndpointAnalysisEffect estimate95% CIP-value
Cardiovascular death + total first and recurrent HF events Stratified Andersen-Gill Rate ratio 0.84 0.74–0.95 0.0072
Total first and recurrent HF events Stratified Andersen-Gill Rate ratio 0.82 0.71–0.94 0.0062
KCCQ TSS change MMRM LS mean difference 1.56 0.79–2.34 <.0001
Improvement in NYHA class Logistic regression Odds ratio 1.01 0.88–1.15 0.9295
First renal composite event Stratified log-rank / Cox Hazard ratio 1.33 0.94–1.89 0.1071
All-cause mortality Stratified log-rank / Cox Hazard ratio 0.93 0.83–1.06 0.2794

This table illustrates why trial interpretation cannot be reduced to a single P-value. FINEARTS-HF used different estimands and statistical models for different endpoint types. The primary recurrent-event analysis produced a rate ratio; the KCCQ analysis estimated a longitudinal mean difference; NYHA improvement used an odds ratio; and renal and mortality endpoints used hazard ratios.

13. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. The affected and at-risk counts were reported as follows.

GroupParticipants with serious adverse eventsParticipants at risk
Finerenone (BAY94-8862)11572993
Placebo12132993
Never Received Treatment015
Clinical Biostats interpretation

The registry reports 1157/2993 participants with serious adverse events in the finerenone group and 1213/2993 in the placebo group. These are counts of affected participants among the reported at-risk populations.

The safety data should be kept conceptually separate from the efficacy estimands. A serious adverse event count is not directly comparable with the primary recurrent heart failure event rate ratio because the two summaries answer different questions and use different denominators and analytical frameworks.

The ClinicalTrials.gov record also identify 0/15 serious adverse events among participants classified as never having received treatment. That group is presented separately rather than being combined with either randomized treatment arm.

14. Statistical Methods Explained

Why was an Andersen-Gill model used for the primary endpoint?

The primary endpoint includes first and recurrent heart failure events. Participants can therefore contribute multiple heart failure events during follow-up. The Andersen-Gill approach is a recurrent-event extension of survival-analysis methodology that allows multiple event times to contribute to the analysis rather than stopping after the first event.

What does a rate ratio of 0.84 mean?

A rate ratio of 0.84 means that the estimated event rate for finerenone was 0.84 times the placebo event rate under the reported model. A convenient descriptive translation is a 16% lower estimated event rate. It does not mean a 16-percentage-point reduction in risk and does not imply that every participant experiences the same proportional reduction.

Why was a robust sandwich estimate used?

Repeated events from the same participant are correlated. The robust sandwich covariance estimator provides a way to obtain variance estimates that account for this within-participant dependence. This matters because the primary endpoint intentionally includes recurrent events rather than treating each event as an independent observation.

Why was an MMRM used for the KCCQ endpoint?

The KCCQ total symptom score was measured repeatedly at Months 6, 9 and 12. A mixed-effects model for repeated measures is designed for this longitudinal structure. It allows the analysis to model repeated observations from the same participant while estimating treatment-group differences across the scheduled assessment period.

What does an odds ratio of 1.01 mean?

The NYHA endpoint is binary: participants either showed improvement or did not under the registered definition. An odds ratio of 1.01 means the estimated odds of improvement were approximately the same between the finerenone and placebo groups. It is important not to call 1.01 a 1% increase in probability; an odds ratio and a risk ratio are different quantities.

Why use a hazard ratio for renal events and mortality?

Renal composite events and all-cause mortality are time-to-event outcomes. Participants can be followed for different lengths of time, and some may reach the end of observation without experiencing the event. Stratified log-rank testing and Cox proportional-hazards modeling are designed for this type of censored time-to-event data.

What does a confidence interval add beyond the P-value?

The P-value summarizes evidence against a null hypothesis under the specified statistical test. The confidence interval instead shows the range of effect estimates compatible with the statistical model and sampling uncertainty at the stated confidence level. For FINEARTS-HF, the difference is especially visible when comparing a point estimate such as the renal hazard ratio of 1.33 with its 95% CI of 0.94–1.89.

15. Interpreting Hazard Ratios Correctly

Two of the secondary endpoints use hazard ratios: first occurrence of renal composite events and all-cause mortality. These estimates should be interpreted differently from the primary rate ratio.

Hazard ratio

A relative measure of the instantaneous event rate between treatment groups under a time-to-event model.

Rate ratio

A relative comparison of event rates, particularly relevant here because the primary endpoint includes recurrent events.

Odds ratio

A relative comparison of odds for a binary outcome, used here for improvement in NYHA class.

Mean difference

A difference between modeled treatment-group means, used here for longitudinal KCCQ symptom scores.

These measures cannot simply be placed on one common scale. A rate ratio of 0.84, odds ratio of 1.01, and hazard ratio of 0.93 represent different estimands. The correct interpretation begins with the endpoint and statistical model, not with the numerical value alone.

16. Intention-to-Treat and Analysis Populations

The registry identifies the full analysis set for the primary recurrent-event analysis and for several secondary analyses. The KCCQ analysis uses the full analysis set with available KCCQ measurements, while the NYHA analysis uses the full analysis set with an available NYHA measurement at baseline.

AnalysisReported population
Primary composite recurrent-event endpointFull analysis set
Total first and recurrent heart failure eventsFull analysis set
KCCQ total symptom scoreFull analysis set with available KCCQ measurements
Improvement in NYHA classFull analysis set with available NYHA measurement at baseline
First renal composite eventFull analysis set
All-cause mortalityFull analysis set

The distinction between the overall full analysis set and endpoint-specific availability requirements is important. A longitudinal questionnaire analysis necessarily depends on the availability of the relevant measurements, whereas a mortality endpoint can be analyzed without a patient-reported questionnaire measurement.

Analysis-population caution: The ClinicalTrials.gov record does not provide detailed missing-data rules, imputation procedures, treatment-discontinuation rules, or censoring algorithms. Those details should not be inferred from the reported effect estimates.

17. Stratification and Survival Analysis

The registry identifies stratified analysis for the Andersen-Gill primary endpoint and the secondary recurrent heart failure endpoint. It also reports stratified log-rank testing and stratified Cox proportional-hazards models for renal composite events and all-cause mortality.

Stratification allows the treatment comparison to account for prespecified or otherwise specified strata without necessarily estimating a separate treatment effect for every stratum. In a stratified Cox analysis, for example, the baseline hazard can differ between strata while the treatment effect is summarized across the strata.

Conceptual survival comparison
HR = instantaneous event rate under finerenone / instantaneous event rate under placebo

For the renal endpoint, the reported HR was 1.33. For all-cause mortality, the reported HR was 0.93.

The ClinicalTrials.gov record does not identify the actual stratification factors. Consequently, this page does not assign clinical variables to the strata or reconstruct a stratification scheme beyond what the registry data explicitly support.

18. Multiplicity and Interpretation of Multiple Endpoints

FINEARTS-HF contains multiple registered endpoints and multiple reported statistical analyses. The primary endpoint has a reported superiority analysis, while the ClinicalTrials.gov record also contain secondary analyses for recurrent heart failure events, KCCQ symptoms, NYHA improvement, renal composite events, and all-cause mortality.

Endpoint roleReported endpointStatistical interpretation
Primary Cardiovascular death and total first and recurrent heart failure events Reported superiority analysis with rate ratio 0.84, 95% CI 0.74–0.95, P = 0.0072
Secondary Total first and recurrent heart failure events Rate ratio 0.82, 95% CI 0.71–0.94, P = 0.0062
Secondary KCCQ total symptom score LS mean difference 1.56, 95% CI 0.79–2.34, P < .0001
Secondary NYHA class improvement Odds ratio 1.01, 95% CI 0.88–1.15, P = 0.9295
Secondary First renal composite event Hazard ratio 1.33, 95% CI 0.94–1.89, P = 0.1071
Secondary All-cause mortality Hazard ratio 0.93, 95% CI 0.83–1.06, P = 0.2794

The ClinicalTrials.gov record does not specify a multiplicity adjustment procedure, alpha hierarchy, alpha-spending scheme, or formal gatekeeping strategy. Therefore, the secondary P-values are reported as posted but should not automatically be interpreted as though every secondary endpoint were an independent confirmatory hypothesis test with its own unrestricted type I error budget.

19. Interim Analysis, Crossover, and Bayesian Methods

The ClinicalTrials.gov record does not report an interim analysis procedure, alpha-spending method, crossover plan, or Bayesian analysis. These design topics are therefore not characterized here.

Why this matters: statistical trial pages should distinguish between methods that are documented and methods that merely might have been used in a trial of this type. The registry-reported FINEARTS-HF data support recurrent-event, mixed-effects, logistic-regression, and stratified survival analyses; they do not support adding an unreported interim, crossover, or Bayesian framework.

20. Confidence Intervals and Statistical Precision

Confidence intervals are particularly informative when the estimates point in different directions across endpoints.

EndpointEstimate95% CIDoes the CI include the null value?
Primary compositeRate ratio 0.840.74–0.95No; null for a ratio is 1
Total HF eventsRate ratio 0.820.71–0.94No; null for a ratio is 1
KCCQ TSSLS mean difference 1.560.79–2.34No; null for a difference is 0
NYHA improvementOdds ratio 1.010.88–1.15Yes; null for a ratio is 1
Renal compositeHazard ratio 1.330.94–1.89Yes; null for a ratio is 1
All-cause mortalityHazard ratio 0.930.83–1.06Yes; null for a ratio is 1

The null value depends on the effect measure. For rate ratios, odds ratios, and hazard ratios, the null value is 1. For a mean difference, the null value is 0. This simple distinction prevents a common interpretive error: treating every effect estimate as though it uses the same scale.

21. Limitations

22. Why This Trial Matters Statistically

FINEARTS-HF is a useful statistical teaching case because it combines several distinct clinical-trial estimands within one randomized study. The primary endpoint is not simply a conventional time-to-first-event outcome: it includes cardiovascular death and total first and recurrent heart failure events. That requires a recurrent-event framework and makes the interpretation of the rate ratio different from the interpretation of the hazard ratios reported for renal events and mortality.

ConceptHow it appears in FINEARTS-HF
RandomizationRandomized comparison of finerenone and placebo
BlindingQuadruple-masked design
Intention-to-treat analysisIdentified as an analysis concept for the primary and several secondary analyses
Recurrent eventsPrimary endpoint includes total first and recurrent heart failure events
Andersen-Gill modelPrimary and total heart failure event analyses
Robust covarianceRobust sandwich estimate for recurrent-event covariance
Mixed-effects modelLongitudinal KCCQ total symptom score analysis
Logistic regressionNYHA class improvement analysis
Stratified log-rank testRenal composite and all-cause mortality analyses
Hazard ratioCause-specific effect measure for renal composite events and all-cause mortality
Rate ratioPrimary effect measure for recurrent-event endpoints
Odds ratioEffect measure for binary NYHA improvement
Confidence interval95% two-sided intervals reported for all six statistical analyses
MultiplicityMultiple primary and secondary analyses require endpoint-specific interpretation

23. Overall Statistical Interpretation

The primary statistical story

The primary analysis reported a rate ratio of 0.84 with a 95% CI of 0.74–0.95 and P = 0.0072 for cardiovascular death and total first and recurrent heart failure events. Under the reported stratified Andersen-Gill model, this corresponds to an estimated 16% lower event rate for finerenone relative to placebo.

The important qualification is that the endpoint is a recurrent-event composite. The rate ratio therefore describes event rates within the recurrent-event analysis rather than a simple proportion of participants experiencing an outcome.

What the secondary analyses add

The total first and recurrent heart failure event analysis produced a rate ratio of 0.82. The KCCQ analysis produced a least-squares mean difference of 1.56. In contrast, the NYHA improvement odds ratio was 1.01, the renal composite hazard ratio was 1.33, and the all-cause mortality hazard ratio was 0.93.

These results should not be collapsed into a single "overall effect." Each endpoint measures a different aspect of disease burden and uses a different estimand. A statistically meaningful result for one endpoint does not automatically establish the same result for another endpoint.

Why the confidence intervals matter

The reported intervals show that the precision of the estimates differs across endpoints. The primary rate ratio has a 95% CI of 0.74–0.95, while the renal composite hazard ratio has a wider interval of 0.94–1.89. Looking at the point estimate without its confidence interval would conceal this difference in statistical precision.

Why P-values are not effect sizes

The P-values range from <.0001 for the KCCQ analysis to 0.9295 for NYHA improvement. These values summarize evidence under their respective statistical tests; they do not provide a common scale for comparing clinical importance. The appropriate effect measure depends on the endpoint: rate ratio, mean difference, odds ratio, or hazard ratio.

24. Clinical Biostats Interpretation: Reading the Trial as a Statistician

The most important lesson from FINEARTS-HF is that the statistical estimand follows the clinical question. When the question concerns repeated heart failure events, an event-history model can use information beyond the first event. When the question concerns a repeatedly measured symptom score, a longitudinal mixed-effects model is appropriate to the repeated structure. When the outcome is a binary change in NYHA class, logistic regression produces an odds ratio. When the endpoint is the time to a first renal event or death, survival methods produce a hazard ratio.

This is more than a collection of statistical techniques. Each model determines what the resulting number means. A rate ratio of 0.84 and a hazard ratio of 0.93 cannot be interpreted as interchangeable measures simply because both are below 1. Similarly, an odds ratio of 1.01 is not a "1% probability difference."

The primary result is therefore best understood in its full context: a randomized, quadruple-masked, parallel phase 3 comparison with a recurrent-event primary endpoint, analyzed using a stratified Andersen-Gill model with robust sandwich covariance estimation. That framework is central to understanding what the reported 0.84 rate ratio actually represents.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Calculators

27. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore recurrent-event analysis, survival methods, longitudinal models, logistic regression, effect measures, and confidence intervals.

28. Record Summary

FINEARTS-HF provides a particularly useful example of endpoint-specific clinical-trial statistics. The randomized, parallel, quadruple-masked phase 3 design compared finerenone with placebo in 6016 participants with heart failure and left ventricular ejection fraction greater than or equal to 40%. The primary endpoint combined cardiovascular death with total first and recurrent heart failure events and was analyzed using a stratified Andersen-Gill model with robust sandwich covariance estimation. The resulting rate ratio was 0.84 with a 95% CI of 0.74–0.95 and P = 0.0072.

The secondary analyses demonstrate why statistical interpretation must remain endpoint-specific. Recurrent heart failure events used the same recurrent-event framework; KCCQ symptom scores used an MMRM; NYHA improvement used logistic regression; and renal composite events and all-cause mortality used stratified log-rank and Cox methods. The reported estimates therefore represent different quantities and should not be treated as though they were measurements of a single common effect.

Clinical Biostats methodology: A trial-results page should not merely repeat reported numbers. The objective is to explain the estimand, statistical model, uncertainty, and interpretation of each major result while preserving the distinction between documented trial methods and educational statistical explanation.