← Clinical Trials
Chronic Kidney Disease Phase 3 Finerenone NCT02540993

FIDELIO-DKD: Complete Statistical Analysis of Finerenone in Diabetic Kidney Disease

An independent statistical analysis of the randomized phase 3 FIDELIO-DKD trial evaluating finerenone versus placebo in subjects with type 2 diabetes mellitus and diabetic kidney disease, with emphasis on the primary renal composite endpoint, cardiovascular outcomes, mortality, hospitalization, UACR, and the statistical methods used to analyze them.

Phase 3  ·  Randomized  ·  Parallel  ·  Quadruple masking  ·  Completed
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information contained in the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.

Important source note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

FIDELIO-DKD was a randomized, parallel-group, phase 3 trial evaluating the efficacy and safety of finerenone in subjects with type 2 diabetes mellitus and diabetic kidney disease. The registry identifies chronic kidney disease as the condition and reports 5734 participants allocated to two treatment arms.

5734
Enrollment
Randomized trial
2
Treatment arms
Finerenone vs placebo
0.825
Primary renal HR
95% CI 0.732–0.928
0.0014
Primary P-value
Two-sided analysis
FeatureFIDELIO-DKD
Brief titleEfficacy and Safety of Finerenone in Subjects With Type 2 Diabetes Mellitus and Diabetic Kidney Disease
PhasePhase 3
StatusCompleted
ConditionChronic Kidney Disease
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment5734
InterventionsFinerenone (BAY94-8862); placebo
Lead sponsorBayer
Sponsor typeIndustry
ClinicalTrials.govNCT02540993

2. Clinical Question

The central question was whether treatment with finerenone, compared with placebo, affected the time to the first occurrence of a prespecified composite renal outcome in subjects with type 2 diabetes mellitus and diabetic kidney disease.

Population

Subjects with type 2 diabetes mellitus and diabetic kidney disease, with chronic kidney disease identified as the trial condition.

Intervention

Finerenone (BAY94-8862).

Comparator

Placebo.

Primary question

Does finerenone change the hazard of experiencing the first occurrence of kidney failure, a sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death?

3. Trial Design

01
Randomize 5734 participants
02
Parallel groups Finerenone vs placebo
03
Follow Time-to-event outcomes
04
Analyze Log-rank and Cox model
05
Assess Efficacy and safety
ARM A · FINERENONE

Finerenone

  • Finerenone (BAY94-8862)
  • Randomized treatment assignment
  • Evaluated against placebo
  • Included in the full analysis set for reported efficacy analyses
ARM B · PLACEBO

Placebo

  • Placebo
  • Randomized treatment assignment
  • Comparator for finerenone
  • Included in the full analysis set for reported efficacy analyses
Allocation
Randomized parallel-group design.
Masking
Quadruple masking.
Primary purpose
Treatment.
Trial status
Completed.

4. Trial Timeline

2015-09-17

Trial start

The registry reports a start date of September 17, 2015.

2020-04-14

Primary completion

The registry reports April 14, 2020 as the primary completion date.

Completed

Results available

The trial is listed as completed, with six outcome measures and six statistical analyses posted in the ClinicalTrials.gov record.

5. Endpoints

The registered primary endpoint is a composite time-to-event outcome. The registry describes assessment from randomization until the first occurrence of the primary renal composite endpoint or censoring at the end of the specified follow-up period.

EndpointRegistry definition / time frameAnalysis
Primary renal composite The first occurrence of the composite endpoint of onset of kidney failure, a sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death. From randomization up until the first occurrence of the primary renal composite endpoint, or censoring at the end of the specified follow-up. Stratified log-rank test; stratified Cox proportional hazards regression for HR and two-sided 95% CI.
Key secondary cardiovascular composite The first occurrence of cardiovascular death, non-fatal myocardial infarction, non-fatal stroke, or hospitalization for heart failure. From randomization up until the first occurrence of the key secondary CV composite endpoint, or censoring at the end of the specified follow-up. Stratified log-rank test; stratified Cox proportional hazards regression.
All-cause mortality From randomization up until death due to any cause, or censoring at the end of the study visit, with an average of 32 months. Stratified log-rank test; stratified Cox proportional hazards regression.
All-cause hospitalization From randomization up until the first occurrence of hospitalization due to any cause, or censoring at the end of the study. Stratified log-rank test; stratified Cox proportional hazards regression.
Change in UACR Change in urinary albumin-to-creatinine ratio (UACR) from baseline up until Month 4. ANCOVA among subjects in the full analysis set with measurements available within the Month 4 time window.
Secondary renal composite The first occurrence of the composite endpoint of onset of kidney failure, a sustained decrease in eGFR of ≥57% from baseline over at least 4 weeks, or renal death. From randomization up until the first occurrence of the composite endpoint, or censoring at the end of the study. Stratified log-rank test; stratified Cox proportional hazards regression.
Endpoint structure matters. The primary renal endpoint is not a single laboratory measurement. It is a composite time-to-event outcome in which the participant is counted when the first qualifying component occurs. This makes both the definition of the components and the time-to-event analysis important to interpretation.

6. Statistical Methodology

Primary analysis: stratified log-rank test

The primary endpoint was analyzed with a stratified log-rank test. The registry analysis text also identifies intention-to-treat analysis and stratified analysis as concepts used in the primary comparison.

The log-rank test compares the observed pattern of event occurrence between treatment groups across follow-up. It is particularly suited to time-to-event endpoints because it uses information about both the timing of events and the number of participants remaining under observation.

Hazard-ratio estimation

The registry states that a stratified Cox proportional hazards regression model was used to provide the point estimate of the hazard ratio and its corresponding two-sided 95% confidence interval.

Conceptual hazard-ratio interpretation
HR = estimated hazard in finerenone group ÷ estimated hazard in placebo group

An HR below 1 indicates a lower estimated instantaneous event rate in the finerenone group under the fitted model. It is a relative time-to-event measure, not an absolute probability and not a statement that every participant experiences the same proportional change.

ANCOVA for UACR

The change in UACR from baseline to Month 4 was analyzed using ANCOVA. The analysis population was subjects in the full analysis set with measurements available within the Month 4 time window.

The reported effect measure was the ratio of least squares means. This is important because the reported value of 0.688 is a ratio rather than a conventional arithmetic mean difference. Interpreting it as a difference of 0.688 units would therefore be incorrect.

Analysis population

The posted statistical analyses identify the full analysis set for the time-to-event endpoints. The UACR analysis likewise begins with the full analysis set but restricts the analysis to subjects with measurements available within the Month 4 time window.

Time-to-event analyses

Five posted analyses use a stratified log-rank framework and a stratified Cox proportional hazards model for the reported hazard ratio and confidence interval.

Continuous endpoint

The Month 4 UACR analysis uses ANCOVA and reports a ratio of least squares means.

7. Primary Result: Renal Composite

The primary endpoint was the first occurrence of kidney failure, a sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death. The registry reports a formal superiority analysis in the full analysis set using a stratified log-rank test, with the hazard ratio estimated using a stratified Cox proportional hazards regression model.

Primary renal composite hazard ratio

0.825

95% CI: 0.732–0.928   ·   P = 0.0014

Finerenone vs placebo  ·  two-sided 95% confidence interval

Primary endpointFinerenone vs placeboAnalysis
First occurrence of kidney failure, sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death HR 0.825
95% CI 0.732–0.928
P = 0.0014
Full analysis set; stratified log-rank test; stratified Cox proportional hazards regression
Clinical Biostats interpretation

What the estimate means: An HR of 0.825 means that the estimated instantaneous hazard of the first primary renal composite event in the finerenone group was 82.5% of the estimated hazard in the placebo group under the fitted Cox model. Equivalently, the point estimate corresponds to a 17.5% lower estimated hazard because 1 − 0.825 = 0.175.

What it does not mean: It does not mean that 17.5% of participants avoided the event, that each participant had exactly a 17.5% reduction in risk, or that the absolute probability of an event was reduced by 17.5 percentage points. Those are different quantities.

What the confidence interval says: The two-sided 95% CI of 0.732–0.928 quantifies uncertainty around the estimated hazard ratio under the statistical model and analysis framework. Its width also shows that the estimate is not known with unlimited precision. The entire interval lies below 1, which is consistent with a lower estimated hazard in the finerenone group under the prespecified superiority comparison.

Why the P-value is not an effect size: P = 0.0014 addresses evidence against the null hypothesis in the specified testing framework. It does not say that the treatment effect is 0.0014 in magnitude, nor does it measure clinical importance. The HR and its confidence interval provide the information about relative effect size and precision.

Model caution: The HR comes from a Cox proportional hazards model. Its usual interpretation relies on a proportional-hazards framework in which the relative hazard is adequately summarized by a single HR over the analyzed period. The ClinicalTrials.gov record does not provide a separate assessment of that assumption, so the HR should be interpreted as the model-based summary reported by the trial.

Censoring matters: Participants who do not experience the first composite event before the end of their available follow-up contribute information until censoring. The validity of a survival analysis therefore depends in part on how censoring relates to the event process.

8. Key Secondary Cardiovascular Composite

The key secondary cardiovascular endpoint was the first occurrence of cardiovascular death, non-fatal myocardial infarction, non-fatal stroke, or hospitalization for heart failure. It was analyzed from randomization until the first occurrence of the endpoint or censoring at the end of the specified follow-up.

Key secondary cardiovascular composite

HR 0.860

95% CI: 0.747–0.989   ·   P = 0.0339

Finerenone vs placebo  ·  stratified log-rank test

Clinical Biostats interpretation

An HR of 0.860 corresponds to a 14.0% lower estimated hazard because 1 − 0.860 = 0.140. This is a relative hazard interpretation, not a 14.0-percentage-point reduction in the probability of experiencing the composite endpoint.

The 95% CI of 0.747–0.989 is relatively close to the null value of 1 at its upper boundary. That makes the precision of the estimate particularly relevant: the point estimate is below 1, but the interval indicates uncertainty about the magnitude of the relative effect.

The P-value of 0.0339 provides evidence against the null hypothesis within the reported superiority testing framework. It should not be interpreted as a measure of how large the cardiovascular effect is. Effect magnitude is described by the HR and its confidence interval.

The endpoint is also a composite. An overall HR for the composite does not establish that every individual component has the same treatment effect. The registry analysis reports the composite result, but it does not provide component-specific statistical estimates in the data used for this page.

9. All-Cause Mortality

All-cause mortality was defined as time from randomization until death due to any cause, with censoring at the end of the study visit. The registry gives an average follow-up of 32 mo for this endpoint.

All-cause mortality

HR 0.895

95% CI: 0.746–1.075   ·   P = 0.2348

Clinical Biostats interpretation

The point estimate of 0.895 corresponds to a 10.5% lower estimated hazard for death from any cause in the finerenone group under the fitted model. However, the confidence interval of 0.746–1.075 crosses 1, so the estimate is compatible with both a lower and a higher hazard under the uncertainty represented by the interval.

The P-value of 0.2348 is not a measure of the magnitude of the observed HR. Rather, it summarizes the evidence against the specified null hypothesis under the reported testing procedure. A P-value above a conventional threshold should not be converted into a statement that the treatment has exactly no effect.

For an all-cause mortality endpoint, the time-to-event framework is important because participants enter and leave observation at different times. The analysis therefore incorporates event timing and censoring rather than reducing the outcome to a simple proportion.

10. All-Cause Hospitalization

All-cause hospitalization was evaluated from randomization until the first occurrence of hospitalization due to any cause or censoring at the end of the study.

All-cause hospitalization

HR 0.946

95% CI: 0.876–1.022   ·   P = 0.1623

Clinical Biostats interpretation

The HR of 0.946 corresponds to a 5.4% lower estimated hazard for the first all-cause hospitalization under the fitted model. The effect is modest on the relative-hazard scale compared with the primary renal composite estimate.

The 95% CI of 0.876–1.022 crosses 1. Therefore, the uncertainty interval includes the null value as well as values representing a lower estimated hazard. The P-value of 0.1623 likewise should be read as a statement about statistical evidence under the specified hypothesis test, not as a direct measurement of treatment effect size.

Because the endpoint is the first occurrence of hospitalization, the reported analysis does not automatically answer questions about the total number of hospitalizations per participant. Recurrent-event analyses require different statistical formulations when repeated hospitalizations are the outcome of interest.

11. Change in UACR From Baseline to Month 4

The urinary albumin-to-creatinine ratio (UACR) endpoint was analyzed from baseline up until Month 4. Unlike the time-to-event endpoints, this is a continuous endpoint and was analyzed using ANCOVA.

Ratio of least squares means

0.688

95% CI: 0.662–0.715   ·   P < 0.0001

Finerenone vs placebo  ·  ANCOVA

Clinical Biostats interpretation

The reported effect measure is a ratio of least squares means. A value of 0.688 indicates that the adjusted UACR measure summarized by the model was estimated at 68.8% of the corresponding value in the placebo group. On this ratio scale, the point estimate is therefore 31.2% below 1.

This should not be described as a mean difference of 0.688 units. A ratio and an arithmetic difference have different interpretations and different units. The registry explicitly reports the effect as a ratio of least squares means.

The 95% CI of 0.662–0.715 is relatively narrow around the point estimate, indicating greater numerical precision than would be conveyed by the point estimate alone. The P-value of <0.0001 indicates strong evidence against the specified null hypothesis within the reported ANCOVA framework, but it does not itself quantify the magnitude of the UACR effect.

The analysis population is also important: subjects were included if they were in the full analysis set and had measurements available within the Month 4 time window. That differs from the structure of the time-to-event analyses, which use follow-up until an event or censoring.

12. Secondary Renal Composite Using eGFR ≥57%

A second renal composite used a more stringent eGFR decline threshold: onset of kidney failure, a sustained decrease in eGFR of ≥57% from baseline over at least 4 weeks, or renal death. The registry specifies analysis from randomization until the first occurrence of the composite endpoint or censoring at the end of the study.

Secondary renal composite

HR 0.763

95% CI: 0.648–0.900   ·   P = 0.0012

Clinical Biostats interpretation

The HR of 0.763 corresponds to a 23.7% lower estimated hazard for the secondary renal composite under the fitted model. The direction of the estimate is the same as for the primary renal composite, although the endpoint definition uses the more stringent eGFR decline threshold of ≥57%.

The 95% CI of 0.648–0.900 lies below 1. The interval therefore supports a lower estimated hazard under the reported model while also showing that the exact magnitude is uncertain.

The P-value of 0.0012 provides evidence against the specified null hypothesis under the reported statistical framework. It should not be treated as a 0.12% probability that the treatment effect is real, nor as a direct measure of the size or clinical importance of the treatment effect.

13. Results Summary

EndpointEffect measure95% CIP-valueMethod
Primary renal composite: kidney failure, sustained eGFR decrease ≥40%, or renal death HR 0.825 0.732–0.928 = 0.0014 Stratified log-rank; stratified Cox model
Key secondary CV composite HR 0.860 0.747–0.989 = 0.0339 Stratified log-rank; stratified Cox model
All-cause mortality HR 0.895 0.746–1.075 = 0.2348 Stratified log-rank; stratified Cox model
All-cause hospitalization HR 0.946 0.876–1.022 = 0.1623 Stratified log-rank; stratified Cox model
UACR change baseline to Month 4 Ratio of LS means 0.688 0.662–0.715 < 0.0001 ANCOVA
Secondary renal composite: kidney failure, sustained eGFR decrease ≥57%, or renal death HR 0.763 0.648–0.900 = 0.0012 Stratified log-rank; stratified Cox model
Read the table horizontally. Each row combines an endpoint definition, an effect measure, an uncertainty interval, a P-value, and an analysis method. Comparing P-values alone would discard much of the statistical information. For time-to-event outcomes, the HR and its confidence interval describe the estimated relative treatment effect; for UACR, the ratio of least squares means is the relevant effect measure.

14. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.

Safety measureFinerenonePlacebo
Serious adverse events 902 / 2827 971 / 2831
Serious adverse events: affected / at risk
Finerenone
902 / 2827
Placebo
971 / 2831

The ClinicalTrials.gov record does not provide a formal statistical analysis, confidence interval, or P-value for serious adverse events. Accordingly, this page does not infer a comparative treatment effect from the affected/at-risk counts alone.

Clinical Biostats interpretation

The numerator and denominator provide descriptive information about serious adverse events in the two groups. A formal comparative safety analysis would need to specify the estimand and statistical method—for example, whether the outcome is treated as a binary occurrence over a defined period, a time-to-first-event outcome, or another safety estimand.

It is also important not to confuse the descriptive affected/at-risk counts with the denominator of the full randomized population. The registry-reported serious-adverse-event figures have their own reported at-risk denominators and should be presented exactly as reported rather than reconstructed into a different analysis population.

15. Statistical Methods Explained

Why was a log-rank test used?

The principal efficacy outcomes are time-to-event endpoints. A log-rank test is designed to compare the event-time experience of two groups while using the timing of events and accommodating right-censored observations. This makes it more appropriate for these endpoints than a simple comparison of the proportion of participants who eventually experienced an event.

Why was the log-rank test stratified?

The registry explicitly describes the primary analysis as a stratified log-rank test and identifies stratified analysis as an analysis concept. Stratification allows the comparison to account for prespecified strata rather than treating all participants as if they came from one homogeneous risk set. The ClinicalTrials.gov record does not identify the individual stratification variables, so none are added here.

Why was a Cox model used in addition to the log-rank test?

The log-rank test provides a hypothesis test for the time-to-event comparison, while the Cox proportional hazards model supplies a quantitative estimate of the relative hazard in the form of a hazard ratio and its confidence interval. Using both therefore provides complementary information: evidence for a treatment difference and an estimate of its magnitude.

What does HR 0.825 mean?

An HR of 0.825 means that the fitted model estimates the instantaneous event hazard in the finerenone group at 82.5% of the corresponding hazard in the placebo group. The associated relative reduction in estimated hazard is 17.5%. It does not mean that 17.5% of patients benefited or that every participant had a 17.5% reduction in absolute risk.

Why is the UACR analysis different?

UACR change from baseline to Month 4 is a continuous outcome measured at a defined follow-up time rather than an event that occurs at an uncertain time. The registry therefore reports ANCOVA rather than a survival-analysis method. The effect measure is a ratio of least squares means, which must be interpreted on a ratio scale rather than as an arithmetic mean difference.

What does the 95% confidence interval add?

A point estimate alone does not show how precisely the treatment effect has been estimated. The 95% confidence interval supplies an uncertainty range under the model and sampling framework. For example, the primary renal HR of 0.825 is accompanied by a 95% CI of 0.732–0.928. The interval gives substantially more information than the HR alone because it shows the plausible statistical precision of the estimate.

Why should P-values and effect sizes be read separately?

A P-value answers a hypothesis-testing question, whereas an effect estimate describes magnitude. A very small P-value does not imply a large effect, and a larger P-value does not prove that the true effect is exactly zero. The most informative interpretation considers the estimate, confidence interval, P-value, endpoint definition, analysis population, and study design together.

16. Intention-to-Treat and Analysis Populations

The registry's statistical analysis text identifies intention-to-treat analysis as a concept for the posted efficacy analyses, and the analysis population is identified as the full analysis set.

Randomized comparison
The treatment comparison begins with randomized assignment rather than redefining groups based on later outcomes.
Full analysis set
The posted time-to-event analyses are identified as using the full analysis set.
UACR population
The UACR ANCOVA uses subjects in the full analysis set with measurements available within the Month 4 time window.
Interpretive consequence
Different endpoints can legitimately have different analysis-population specifications because their data requirements differ.

For randomized clinical trials, analyzing participants according to randomized assignment helps preserve the comparison created by randomization. The important statistical distinction is that the analysis population must be identified before interpreting an estimate: an HR from the full analysis set and an ANCOVA restricted to participants with an available Month 4 measurement do not represent precisely the same analytic population.

17. Stratified Survival Analysis

All five posted time-to-event statistical analyses use the same broad structure: a stratified log-rank test together with a stratified Cox proportional hazards regression model. This consistency is useful because the primary renal endpoint, cardiovascular composite, mortality, hospitalization, and secondary renal composite are all fundamentally time-to-event questions.

Time-to-event endpointLog-rankCox HR model
Primary renal compositeStratifiedStratified Cox model
Key secondary CV compositeStratifiedStratified Cox model
All-cause mortalityStratifiedStratified Cox model
All-cause hospitalizationStratifiedStratified Cox model
Secondary renal compositeStratifiedStratified Cox model

This structure also illustrates why a trial's reported HR should not be interpreted as a simple ratio of event percentages. Survival analysis uses the entire observed event-time structure, including censoring, rather than collapsing follow-up into a single binary outcome.

18. Understanding Composite Endpoints

The primary endpoint combines three possible first events: kidney failure, sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death. The secondary renal endpoint similarly combines kidney failure, sustained decrease in eGFR of ≥57% from baseline over at least 4 weeks, or renal death.

Why use a composite?

A composite can capture several clinically relevant ways in which the disease course may reach the prespecified endpoint. Statistically, the participant enters the event analysis at the time the first qualifying component occurs.

What the composite HR means

The reported HR summarizes the time to the first occurrence of any qualifying component. It is not automatically the HR for each individual component.

Component interpretation

A composite result should not be decomposed into individual component treatment effects unless those component-specific analyses are separately reported.

Different eGFR thresholds

The primary endpoint uses a sustained eGFR decrease of ≥40%, while the secondary renal composite uses ≥57%. These are distinct endpoint definitions and should not be conflated.

For statistical interpretation, the exact endpoint definition is therefore inseparable from the HR. Saying simply that "the renal HR was 0.825" loses the information that the estimate concerns the first occurrence of a particular composite endpoint.

19. Confidence Intervals and Statistical Precision

The six posted statistical analyses provide a useful range of confidence intervals. They illustrate why the point estimate alone is insufficient for interpretation.

EndpointEstimate95% CIRelationship to HR = 1
Primary renal compositeHR 0.8250.732–0.928Entire interval below 1
Key secondary CV compositeHR 0.8600.747–0.989Entire interval below 1
All-cause mortalityHR 0.8950.746–1.075Interval crosses 1
All-cause hospitalizationHR 0.9460.876–1.022Interval crosses 1
UACRRatio 0.6880.662–0.715Entire interval below 1
Secondary renal compositeHR 0.7630.648–0.900Entire interval below 1

For the four hazard-ratio endpoints, the null value is HR = 1. For the UACR ratio, a ratio of 1 represents equal adjusted values between groups. These reference values are different because the effect measures are different.

A practical reading rule

Start with the endpoint definition, then identify the effect measure, then inspect the confidence interval, and only then consider the P-value. This sequence prevents a common error in which the P-value becomes the entire interpretation of the trial.

20. Multiplicity and Hierarchical Testing

The ClinicalTrials.gov record includes a specific hierarchy statement for the secondary endpoints. They state that if treatment effect on both the primary and key secondary endpoint is significant, other secondary efficacy endpoints—including all-cause mortality, all-cause hospitalization, change in UACR from baseline to Month 4, and the secondary renal composite endpoint—will be tested hierarchically, starting with all-cause mortality.

EndpointRole in the ClinicalTrials.gov recordReported estimate
Primary renal compositePrimary endpointHR 0.825; 95% CI 0.732–0.928; P = 0.0014
Key secondary CV compositeKey secondary endpointHR 0.860; 95% CI 0.747–0.989; P = 0.0339
All-cause mortalitySecondary efficacy endpointHR 0.895; 95% CI 0.746–1.075; P = 0.2348
All-cause hospitalizationSecondary efficacy endpointHR 0.946; 95% CI 0.876–1.022; P = 0.1623
UACR changeSecondary efficacy endpointRatio of LS means 0.688; 95% CI 0.662–0.715; P < 0.0001
Secondary renal compositeSecondary efficacy endpointHR 0.763; 95% CI 0.648–0.900; P = 0.0012
Why hierarchy matters: Multiple endpoint tests can increase the chance of observing a small P-value by chance alone. The ClinicalTrials.gov record explicitly describe hierarchical testing of additional secondary efficacy endpoints conditional on the preceding results. Therefore, the order and testing framework are part of the statistical interpretation rather than an afterthought.

The ClinicalTrials.gov record does not provide the complete alpha-allocation details or a numerical multiplicity-adjustment calculation. Accordingly, this page does not reconstruct an unreported familywise error calculation.

21. Proportional-Hazards Assumption and Censoring

The Cox model supplies a compact hazard-ratio summary, but that summary should be interpreted in the context of the model's assumptions. In particular, the conventional Cox proportional-hazards interpretation assumes that the relative hazard can be represented adequately by a common treatment effect over time.

Why the assumption matters
hfinerenone(t) / hplacebo(t) = HR

A single HR is most naturally interpreted when the relative hazard is reasonably represented as stable over time. If hazards vary substantially in their relative separation, one HR can become a less complete description of the treatment effect.

The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic. The appropriate conclusion is therefore limited: the trial reported a Cox-model HR, and the HR should be understood as the model-based summary rather than as a universal statement about every point in follow-up.

Censoring is also fundamental. A participant who has not experienced the first qualifying event when follow-up ends contributes information up to the censoring time. Survival-analysis validity depends on assumptions concerning censoring and the relationship between follow-up availability and the event process.

22. What the P-values Do — and Do Not — Say

P = 0.0014

For the primary renal composite, this is the reported P-value for the superiority comparison using the stratified log-rank framework.

P = 0.0339

For the key secondary cardiovascular composite, this is the reported P-value under the corresponding time-to-event analysis.

P = 0.2348

For all-cause mortality, this P-value accompanies an HR of 0.895 and a 95% CI that crosses 1.

P = 0.1623

For all-cause hospitalization, this P-value accompanies an HR of 0.946 and a 95% CI that crosses 1.

A P-value is calculated under a specified null hypothesis and statistical model. It does not measure the probability that the null hypothesis is true, the probability that the treatment works, or the clinical importance of the result.

For example, the primary HR of 0.825 and its CI of 0.732–0.928 describe the estimated effect and its precision. The P-value of 0.0014 adds information about statistical evidence against the null hypothesis. Neither number substitutes for the other.

23. Design Features the Registry Does Not Quantify Here

The ClinicalTrials.gov record identifies several design characteristics directly, but do not provide numerical details for every possible design topic. The distinction is important because an educational analysis should not fill gaps with assumptions.

Design topicWhat is supported by the ClinicalTrials.gov recordInterpretation
RandomizationAllocation is randomized.The treatment comparison is based on randomized groups.
Parallel designDesign model is parallel.The two interventions form parallel treatment groups.
BlindingMasking is quadruple.The registry identifies a quadruple-masked trial.
SuperiorityHypothesis type is superiority.The reported efficacy tests are framed as superiority comparisons.
Non-inferiority marginNot reported.No non-inferiority margin is applicable to the reported superiority analyses presented here.
CrossoverNot reported.No crossover analysis is added.
Factorial designNot reported; design model is parallel.No factorial analysis is added.
Bayesian methodsNot reported.No Bayesian analysis is added.
Interim analysisNot reported in the ClinicalTrials.gov record.No interim-analysis schedule or alpha-spending method is inferred.
Missing-data imputationNot reported.No specific imputation method is attributed to the trial.
Methodological discipline: absence of a detail from the ClinicalTrials.gov record is not evidence that the trial did not use such a procedure. It means only that the detail is not available in the dataset used for this page, so it is not reconstructed here.

24. Why the Primary Result Is a Time-to-Event Result

The primary renal endpoint is defined by the first occurrence of a qualifying event after randomization. This creates a time-to-event estimand rather than a simple binary endpoint at a fixed calendar date.

Time-to-event structure
Randomization → follow-up → first qualifying renal event or censoring

The statistical analysis uses the timing of the first qualifying event. Participants without an event by the end of their observed follow-up contribute information until censoring.

This distinction explains why the trial reports hazard ratios rather than simply reporting the difference in the number of participants with an event. A participant who experiences an event earlier and a participant who experiences the same event much later do not provide identical time-to-event information.

The same logic applies to the key secondary cardiovascular composite, all-cause mortality, all-cause hospitalization, and the secondary renal composite. Each is defined around time from randomization to the first qualifying event or censoring.

25. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary renal composite produced an HR of 0.825 with a two-sided 95% CI of 0.732–0.928 and P = 0.0014 under the reported stratified survival-analysis framework. The key secondary cardiovascular composite produced an HR of 0.860 with a 95% CI of 0.747–0.989 and P = 0.0339.

Effect-size interpretation

The HRs quantify relative differences in estimated hazard, while the confidence intervals quantify statistical uncertainty. The primary point estimate corresponds to a 17.5% lower estimated hazard, and the cardiovascular point estimate corresponds to a 14.0% lower estimated hazard.

Continuous endpoint

The UACR result uses a different estimand: a ratio of least squares means of 0.688 with a 95% CI of 0.662–0.715 and P < 0.0001. It should not be interpreted as a hazard ratio.

Safety interpretation

Serious adverse events are reported descriptively as 902/2827 for finerenone and 971/2831 for placebo. The ClinicalTrials.gov record does not provide a formal comparative safety test for this outcome.

26. Important Limitations and Interpretation Issues

27. Why This Trial Matters Statistically

FIDELIO-DKD is a useful teaching case because it places several recurring clinical-trial methods in one analysis: randomized parallel-group comparison, quadruple masking, composite time-to-event endpoints, stratified log-rank testing, Cox hazard-ratio estimation, ANCOVA for a continuous endpoint, intention-to-treat concepts, confidence intervals, P-values, and hierarchical testing.

ConceptHow it appears in FIDELIO-DKD
RandomizationThe registry identifies randomized allocation with two treatment arms.
Parallel designThe design model is parallel.
BlindingThe trial is quadruple masked.
Time-to-event endpointsThe primary, cardiovascular, mortality, hospitalization, and secondary renal outcomes are analyzed from randomization to an event or censoring.
Log-rank testAll five posted time-to-event analyses use a stratified log-rank test.
Hazard ratioStratified Cox regression supplies the reported HRs and two-sided 95% confidence intervals.
ANCOVAChange in UACR from baseline to Month 4 is analyzed using ANCOVA.
Ratio of least squares meansThe UACR effect is reported as a ratio of least squares means rather than an HR.
Intention-to-treatIntention-to-treat analysis is identified as a concept in the posted efficacy analyses.
Stratified analysisThe time-to-event analyses are described as stratified.
MultiplicityThe registry-reported analysis notes describe hierarchical testing of additional secondary efficacy endpoints.
Composite endpointThe primary and secondary renal endpoints combine several possible first events.

The most important statistical lesson is that a trial result is not a single number. The HR, confidence interval, P-value, endpoint definition, analysis population, statistical model, censoring structure, and multiplicity framework jointly define what the result means.

28. A Worked Reading of the Primary Analysis

Suppose the primary result is encountered in abbreviated form as: HR 0.825, 95% CI 0.732–0.928, P = 0.0014. A statistically literate reading proceeds in several steps.

01
Identify endpoint Primary renal composite
02
Identify estimand Relative hazard
03
Read magnitude HR 0.825
04
Read precision 95% CI 0.732–0.928
05
Read evidence P = 0.0014

First, the endpoint must be identified: this is not all-cause mortality or hospitalization but the first occurrence of the specified renal composite. Second, the effect measure is a hazard ratio. Third, the point estimate indicates a lower estimated hazard in the finerenone group. Fourth, the confidence interval describes uncertainty around that estimate. Fifth, the P-value describes the statistical evidence under the specified hypothesis-testing framework.

This approach prevents two common errors. The first is treating an HR as if it were an absolute risk reduction. The second is treating the P-value as if it were the treatment effect. Neither interpretation is statistically correct.

29. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the methods behind randomized clinical-trial analysis, survival endpoints, confidence intervals, hypothesis testing, and continuous-outcome models.