This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information contained in the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
FIDELIO-DKD was a randomized, parallel-group, phase 3 trial evaluating the efficacy and safety of finerenone in subjects with type 2 diabetes mellitus and diabetic kidney disease. The registry identifies chronic kidney disease as the condition and reports 5734 participants allocated to two treatment arms.
| Feature | FIDELIO-DKD |
|---|---|
| Brief title | Efficacy and Safety of Finerenone in Subjects With Type 2 Diabetes Mellitus and Diabetic Kidney Disease |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Chronic Kidney Disease |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 5734 |
| Interventions | Finerenone (BAY94-8862); placebo |
| Lead sponsor | Bayer |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT02540993 |
2. Clinical Question
The central question was whether treatment with finerenone, compared with placebo, affected the time to the first occurrence of a prespecified composite renal outcome in subjects with type 2 diabetes mellitus and diabetic kidney disease.
Population
Subjects with type 2 diabetes mellitus and diabetic kidney disease, with chronic kidney disease identified as the trial condition.
Intervention
Finerenone (BAY94-8862).
Comparator
Placebo.
Primary question
Does finerenone change the hazard of experiencing the first occurrence of kidney failure, a sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death?
3. Trial Design
Finerenone
- Finerenone (BAY94-8862)
- Randomized treatment assignment
- Evaluated against placebo
- Included in the full analysis set for reported efficacy analyses
Placebo
- Placebo
- Randomized treatment assignment
- Comparator for finerenone
- Included in the full analysis set for reported efficacy analyses
4. Trial Timeline
Trial start
The registry reports a start date of September 17, 2015.
Primary completion
The registry reports April 14, 2020 as the primary completion date.
Results available
The trial is listed as completed, with six outcome measures and six statistical analyses posted in the ClinicalTrials.gov record.
5. Endpoints
The registered primary endpoint is a composite time-to-event outcome. The registry describes assessment from randomization until the first occurrence of the primary renal composite endpoint or censoring at the end of the specified follow-up period.
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Primary renal composite | The first occurrence of the composite endpoint of onset of kidney failure, a sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death. From randomization up until the first occurrence of the primary renal composite endpoint, or censoring at the end of the specified follow-up. | Stratified log-rank test; stratified Cox proportional hazards regression for HR and two-sided 95% CI. |
| Key secondary cardiovascular composite | The first occurrence of cardiovascular death, non-fatal myocardial infarction, non-fatal stroke, or hospitalization for heart failure. From randomization up until the first occurrence of the key secondary CV composite endpoint, or censoring at the end of the specified follow-up. | Stratified log-rank test; stratified Cox proportional hazards regression. |
| All-cause mortality | From randomization up until death due to any cause, or censoring at the end of the study visit, with an average of 32 months. | Stratified log-rank test; stratified Cox proportional hazards regression. |
| All-cause hospitalization | From randomization up until the first occurrence of hospitalization due to any cause, or censoring at the end of the study. | Stratified log-rank test; stratified Cox proportional hazards regression. |
| Change in UACR | Change in urinary albumin-to-creatinine ratio (UACR) from baseline up until Month 4. | ANCOVA among subjects in the full analysis set with measurements available within the Month 4 time window. |
| Secondary renal composite | The first occurrence of the composite endpoint of onset of kidney failure, a sustained decrease in eGFR of ≥57% from baseline over at least 4 weeks, or renal death. From randomization up until the first occurrence of the composite endpoint, or censoring at the end of the study. | Stratified log-rank test; stratified Cox proportional hazards regression. |
6. Statistical Methodology
Primary analysis: stratified log-rank test
The primary endpoint was analyzed with a stratified log-rank test. The registry analysis text also identifies intention-to-treat analysis and stratified analysis as concepts used in the primary comparison.
The log-rank test compares the observed pattern of event occurrence between treatment groups across follow-up. It is particularly suited to time-to-event endpoints because it uses information about both the timing of events and the number of participants remaining under observation.
Hazard-ratio estimation
The registry states that a stratified Cox proportional hazards regression model was used to provide the point estimate of the hazard ratio and its corresponding two-sided 95% confidence interval.
An HR below 1 indicates a lower estimated instantaneous event rate in the finerenone group under the fitted model. It is a relative time-to-event measure, not an absolute probability and not a statement that every participant experiences the same proportional change.
ANCOVA for UACR
The change in UACR from baseline to Month 4 was analyzed using ANCOVA. The analysis population was subjects in the full analysis set with measurements available within the Month 4 time window.
The reported effect measure was the ratio of least squares means. This is important because the reported value of 0.688 is a ratio rather than a conventional arithmetic mean difference. Interpreting it as a difference of 0.688 units would therefore be incorrect.
Analysis population
The posted statistical analyses identify the full analysis set for the time-to-event endpoints. The UACR analysis likewise begins with the full analysis set but restricts the analysis to subjects with measurements available within the Month 4 time window.
Time-to-event analyses
Five posted analyses use a stratified log-rank framework and a stratified Cox proportional hazards model for the reported hazard ratio and confidence interval.
Continuous endpoint
The Month 4 UACR analysis uses ANCOVA and reports a ratio of least squares means.
7. Primary Result: Renal Composite
The primary endpoint was the first occurrence of kidney failure, a sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death. The registry reports a formal superiority analysis in the full analysis set using a stratified log-rank test, with the hazard ratio estimated using a stratified Cox proportional hazards regression model.
Primary renal composite hazard ratio
95% CI: 0.732–0.928 · P = 0.0014
Finerenone vs placebo · two-sided 95% confidence interval
| Primary endpoint | Finerenone vs placebo | Analysis |
|---|---|---|
| First occurrence of kidney failure, sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death | HR 0.825 95% CI 0.732–0.928 P = 0.0014 |
Full analysis set; stratified log-rank test; stratified Cox proportional hazards regression |
What the estimate means: An HR of 0.825 means that the estimated instantaneous hazard of the first primary renal composite event in the finerenone group was 82.5% of the estimated hazard in the placebo group under the fitted Cox model. Equivalently, the point estimate corresponds to a 17.5% lower estimated hazard because 1 − 0.825 = 0.175.
What it does not mean: It does not mean that 17.5% of participants avoided the event, that each participant had exactly a 17.5% reduction in risk, or that the absolute probability of an event was reduced by 17.5 percentage points. Those are different quantities.
What the confidence interval says: The two-sided 95% CI of 0.732–0.928 quantifies uncertainty around the estimated hazard ratio under the statistical model and analysis framework. Its width also shows that the estimate is not known with unlimited precision. The entire interval lies below 1, which is consistent with a lower estimated hazard in the finerenone group under the prespecified superiority comparison.
Why the P-value is not an effect size: P = 0.0014 addresses evidence against the null hypothesis in the specified testing framework. It does not say that the treatment effect is 0.0014 in magnitude, nor does it measure clinical importance. The HR and its confidence interval provide the information about relative effect size and precision.
Model caution: The HR comes from a Cox proportional hazards model. Its usual interpretation relies on a proportional-hazards framework in which the relative hazard is adequately summarized by a single HR over the analyzed period. The ClinicalTrials.gov record does not provide a separate assessment of that assumption, so the HR should be interpreted as the model-based summary reported by the trial.
Censoring matters: Participants who do not experience the first composite event before the end of their available follow-up contribute information until censoring. The validity of a survival analysis therefore depends in part on how censoring relates to the event process.
8. Key Secondary Cardiovascular Composite
The key secondary cardiovascular endpoint was the first occurrence of cardiovascular death, non-fatal myocardial infarction, non-fatal stroke, or hospitalization for heart failure. It was analyzed from randomization until the first occurrence of the endpoint or censoring at the end of the specified follow-up.
Key secondary cardiovascular composite
95% CI: 0.747–0.989 · P = 0.0339
Finerenone vs placebo · stratified log-rank test
An HR of 0.860 corresponds to a 14.0% lower estimated hazard because 1 − 0.860 = 0.140. This is a relative hazard interpretation, not a 14.0-percentage-point reduction in the probability of experiencing the composite endpoint.
The 95% CI of 0.747–0.989 is relatively close to the null value of 1 at its upper boundary. That makes the precision of the estimate particularly relevant: the point estimate is below 1, but the interval indicates uncertainty about the magnitude of the relative effect.
The P-value of 0.0339 provides evidence against the null hypothesis within the reported superiority testing framework. It should not be interpreted as a measure of how large the cardiovascular effect is. Effect magnitude is described by the HR and its confidence interval.
The endpoint is also a composite. An overall HR for the composite does not establish that every individual component has the same treatment effect. The registry analysis reports the composite result, but it does not provide component-specific statistical estimates in the data used for this page.
9. All-Cause Mortality
All-cause mortality was defined as time from randomization until death due to any cause, with censoring at the end of the study visit. The registry gives an average follow-up of 32 mo for this endpoint.
All-cause mortality
95% CI: 0.746–1.075 · P = 0.2348
The point estimate of 0.895 corresponds to a 10.5% lower estimated hazard for death from any cause in the finerenone group under the fitted model. However, the confidence interval of 0.746–1.075 crosses 1, so the estimate is compatible with both a lower and a higher hazard under the uncertainty represented by the interval.
The P-value of 0.2348 is not a measure of the magnitude of the observed HR. Rather, it summarizes the evidence against the specified null hypothesis under the reported testing procedure. A P-value above a conventional threshold should not be converted into a statement that the treatment has exactly no effect.
For an all-cause mortality endpoint, the time-to-event framework is important because participants enter and leave observation at different times. The analysis therefore incorporates event timing and censoring rather than reducing the outcome to a simple proportion.
10. All-Cause Hospitalization
All-cause hospitalization was evaluated from randomization until the first occurrence of hospitalization due to any cause or censoring at the end of the study.
All-cause hospitalization
95% CI: 0.876–1.022 · P = 0.1623
The HR of 0.946 corresponds to a 5.4% lower estimated hazard for the first all-cause hospitalization under the fitted model. The effect is modest on the relative-hazard scale compared with the primary renal composite estimate.
The 95% CI of 0.876–1.022 crosses 1. Therefore, the uncertainty interval includes the null value as well as values representing a lower estimated hazard. The P-value of 0.1623 likewise should be read as a statement about statistical evidence under the specified hypothesis test, not as a direct measurement of treatment effect size.
Because the endpoint is the first occurrence of hospitalization, the reported analysis does not automatically answer questions about the total number of hospitalizations per participant. Recurrent-event analyses require different statistical formulations when repeated hospitalizations are the outcome of interest.
11. Change in UACR From Baseline to Month 4
The urinary albumin-to-creatinine ratio (UACR) endpoint was analyzed from baseline up until Month 4. Unlike the time-to-event endpoints, this is a continuous endpoint and was analyzed using ANCOVA.
Ratio of least squares means
95% CI: 0.662–0.715 · P < 0.0001
Finerenone vs placebo · ANCOVA
The reported effect measure is a ratio of least squares means. A value of 0.688 indicates that the adjusted UACR measure summarized by the model was estimated at 68.8% of the corresponding value in the placebo group. On this ratio scale, the point estimate is therefore 31.2% below 1.
This should not be described as a mean difference of 0.688 units. A ratio and an arithmetic difference have different interpretations and different units. The registry explicitly reports the effect as a ratio of least squares means.
The 95% CI of 0.662–0.715 is relatively narrow around the point estimate, indicating greater numerical precision than would be conveyed by the point estimate alone. The P-value of <0.0001 indicates strong evidence against the specified null hypothesis within the reported ANCOVA framework, but it does not itself quantify the magnitude of the UACR effect.
The analysis population is also important: subjects were included if they were in the full analysis set and had measurements available within the Month 4 time window. That differs from the structure of the time-to-event analyses, which use follow-up until an event or censoring.
12. Secondary Renal Composite Using eGFR ≥57%
A second renal composite used a more stringent eGFR decline threshold: onset of kidney failure, a sustained decrease in eGFR of ≥57% from baseline over at least 4 weeks, or renal death. The registry specifies analysis from randomization until the first occurrence of the composite endpoint or censoring at the end of the study.
Secondary renal composite
95% CI: 0.648–0.900 · P = 0.0012
The HR of 0.763 corresponds to a 23.7% lower estimated hazard for the secondary renal composite under the fitted model. The direction of the estimate is the same as for the primary renal composite, although the endpoint definition uses the more stringent eGFR decline threshold of ≥57%.
The 95% CI of 0.648–0.900 lies below 1. The interval therefore supports a lower estimated hazard under the reported model while also showing that the exact magnitude is uncertain.
The P-value of 0.0012 provides evidence against the specified null hypothesis under the reported statistical framework. It should not be treated as a 0.12% probability that the treatment effect is real, nor as a direct measure of the size or clinical importance of the treatment effect.
13. Results Summary
| Endpoint | Effect measure | 95% CI | P-value | Method |
|---|---|---|---|---|
| Primary renal composite: kidney failure, sustained eGFR decrease ≥40%, or renal death | HR 0.825 | 0.732–0.928 | = 0.0014 | Stratified log-rank; stratified Cox model |
| Key secondary CV composite | HR 0.860 | 0.747–0.989 | = 0.0339 | Stratified log-rank; stratified Cox model |
| All-cause mortality | HR 0.895 | 0.746–1.075 | = 0.2348 | Stratified log-rank; stratified Cox model |
| All-cause hospitalization | HR 0.946 | 0.876–1.022 | = 0.1623 | Stratified log-rank; stratified Cox model |
| UACR change baseline to Month 4 | Ratio of LS means 0.688 | 0.662–0.715 | < 0.0001 | ANCOVA |
| Secondary renal composite: kidney failure, sustained eGFR decrease ≥57%, or renal death | HR 0.763 | 0.648–0.900 | = 0.0012 | Stratified log-rank; stratified Cox model |
14. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.
| Safety measure | Finerenone | Placebo |
|---|---|---|
| Serious adverse events | 902 / 2827 | 971 / 2831 |
The ClinicalTrials.gov record does not provide a formal statistical analysis, confidence interval, or P-value for serious adverse events. Accordingly, this page does not infer a comparative treatment effect from the affected/at-risk counts alone.
The numerator and denominator provide descriptive information about serious adverse events in the two groups. A formal comparative safety analysis would need to specify the estimand and statistical method—for example, whether the outcome is treated as a binary occurrence over a defined period, a time-to-first-event outcome, or another safety estimand.
It is also important not to confuse the descriptive affected/at-risk counts with the denominator of the full randomized population. The registry-reported serious-adverse-event figures have their own reported at-risk denominators and should be presented exactly as reported rather than reconstructed into a different analysis population.
15. Statistical Methods Explained
Why was a log-rank test used?
The principal efficacy outcomes are time-to-event endpoints. A log-rank test is designed to compare the event-time experience of two groups while using the timing of events and accommodating right-censored observations. This makes it more appropriate for these endpoints than a simple comparison of the proportion of participants who eventually experienced an event.
Why was the log-rank test stratified?
The registry explicitly describes the primary analysis as a stratified log-rank test and identifies stratified analysis as an analysis concept. Stratification allows the comparison to account for prespecified strata rather than treating all participants as if they came from one homogeneous risk set. The ClinicalTrials.gov record does not identify the individual stratification variables, so none are added here.
Why was a Cox model used in addition to the log-rank test?
The log-rank test provides a hypothesis test for the time-to-event comparison, while the Cox proportional hazards model supplies a quantitative estimate of the relative hazard in the form of a hazard ratio and its confidence interval. Using both therefore provides complementary information: evidence for a treatment difference and an estimate of its magnitude.
What does HR 0.825 mean?
An HR of 0.825 means that the fitted model estimates the instantaneous event hazard in the finerenone group at 82.5% of the corresponding hazard in the placebo group. The associated relative reduction in estimated hazard is 17.5%. It does not mean that 17.5% of patients benefited or that every participant had a 17.5% reduction in absolute risk.
Why is the UACR analysis different?
UACR change from baseline to Month 4 is a continuous outcome measured at a defined follow-up time rather than an event that occurs at an uncertain time. The registry therefore reports ANCOVA rather than a survival-analysis method. The effect measure is a ratio of least squares means, which must be interpreted on a ratio scale rather than as an arithmetic mean difference.
What does the 95% confidence interval add?
A point estimate alone does not show how precisely the treatment effect has been estimated. The 95% confidence interval supplies an uncertainty range under the model and sampling framework. For example, the primary renal HR of 0.825 is accompanied by a 95% CI of 0.732–0.928. The interval gives substantially more information than the HR alone because it shows the plausible statistical precision of the estimate.
Why should P-values and effect sizes be read separately?
A P-value answers a hypothesis-testing question, whereas an effect estimate describes magnitude. A very small P-value does not imply a large effect, and a larger P-value does not prove that the true effect is exactly zero. The most informative interpretation considers the estimate, confidence interval, P-value, endpoint definition, analysis population, and study design together.
16. Intention-to-Treat and Analysis Populations
The registry's statistical analysis text identifies intention-to-treat analysis as a concept for the posted efficacy analyses, and the analysis population is identified as the full analysis set.
For randomized clinical trials, analyzing participants according to randomized assignment helps preserve the comparison created by randomization. The important statistical distinction is that the analysis population must be identified before interpreting an estimate: an HR from the full analysis set and an ANCOVA restricted to participants with an available Month 4 measurement do not represent precisely the same analytic population.
17. Stratified Survival Analysis
All five posted time-to-event statistical analyses use the same broad structure: a stratified log-rank test together with a stratified Cox proportional hazards regression model. This consistency is useful because the primary renal endpoint, cardiovascular composite, mortality, hospitalization, and secondary renal composite are all fundamentally time-to-event questions.
| Time-to-event endpoint | Log-rank | Cox HR model |
|---|---|---|
| Primary renal composite | Stratified | Stratified Cox model |
| Key secondary CV composite | Stratified | Stratified Cox model |
| All-cause mortality | Stratified | Stratified Cox model |
| All-cause hospitalization | Stratified | Stratified Cox model |
| Secondary renal composite | Stratified | Stratified Cox model |
This structure also illustrates why a trial's reported HR should not be interpreted as a simple ratio of event percentages. Survival analysis uses the entire observed event-time structure, including censoring, rather than collapsing follow-up into a single binary outcome.
18. Understanding Composite Endpoints
The primary endpoint combines three possible first events: kidney failure, sustained decrease of eGFR ≥40% from baseline over at least 4 weeks, or renal death. The secondary renal endpoint similarly combines kidney failure, sustained decrease in eGFR of ≥57% from baseline over at least 4 weeks, or renal death.
Why use a composite?
A composite can capture several clinically relevant ways in which the disease course may reach the prespecified endpoint. Statistically, the participant enters the event analysis at the time the first qualifying component occurs.
What the composite HR means
The reported HR summarizes the time to the first occurrence of any qualifying component. It is not automatically the HR for each individual component.
Component interpretation
A composite result should not be decomposed into individual component treatment effects unless those component-specific analyses are separately reported.
Different eGFR thresholds
The primary endpoint uses a sustained eGFR decrease of ≥40%, while the secondary renal composite uses ≥57%. These are distinct endpoint definitions and should not be conflated.
For statistical interpretation, the exact endpoint definition is therefore inseparable from the HR. Saying simply that "the renal HR was 0.825" loses the information that the estimate concerns the first occurrence of a particular composite endpoint.
19. Confidence Intervals and Statistical Precision
The six posted statistical analyses provide a useful range of confidence intervals. They illustrate why the point estimate alone is insufficient for interpretation.
| Endpoint | Estimate | 95% CI | Relationship to HR = 1 |
|---|---|---|---|
| Primary renal composite | HR 0.825 | 0.732–0.928 | Entire interval below 1 |
| Key secondary CV composite | HR 0.860 | 0.747–0.989 | Entire interval below 1 |
| All-cause mortality | HR 0.895 | 0.746–1.075 | Interval crosses 1 |
| All-cause hospitalization | HR 0.946 | 0.876–1.022 | Interval crosses 1 |
| UACR | Ratio 0.688 | 0.662–0.715 | Entire interval below 1 |
| Secondary renal composite | HR 0.763 | 0.648–0.900 | Entire interval below 1 |
For the four hazard-ratio endpoints, the null value is HR = 1. For the UACR ratio, a ratio of 1 represents equal adjusted values between groups. These reference values are different because the effect measures are different.
Start with the endpoint definition, then identify the effect measure, then inspect the confidence interval, and only then consider the P-value. This sequence prevents a common error in which the P-value becomes the entire interpretation of the trial.
20. Multiplicity and Hierarchical Testing
The ClinicalTrials.gov record includes a specific hierarchy statement for the secondary endpoints. They state that if treatment effect on both the primary and key secondary endpoint is significant, other secondary efficacy endpoints—including all-cause mortality, all-cause hospitalization, change in UACR from baseline to Month 4, and the secondary renal composite endpoint—will be tested hierarchically, starting with all-cause mortality.
| Endpoint | Role in the ClinicalTrials.gov record | Reported estimate |
|---|---|---|
| Primary renal composite | Primary endpoint | HR 0.825; 95% CI 0.732–0.928; P = 0.0014 |
| Key secondary CV composite | Key secondary endpoint | HR 0.860; 95% CI 0.747–0.989; P = 0.0339 |
| All-cause mortality | Secondary efficacy endpoint | HR 0.895; 95% CI 0.746–1.075; P = 0.2348 |
| All-cause hospitalization | Secondary efficacy endpoint | HR 0.946; 95% CI 0.876–1.022; P = 0.1623 |
| UACR change | Secondary efficacy endpoint | Ratio of LS means 0.688; 95% CI 0.662–0.715; P < 0.0001 |
| Secondary renal composite | Secondary efficacy endpoint | HR 0.763; 95% CI 0.648–0.900; P = 0.0012 |
The ClinicalTrials.gov record does not provide the complete alpha-allocation details or a numerical multiplicity-adjustment calculation. Accordingly, this page does not reconstruct an unreported familywise error calculation.
21. Proportional-Hazards Assumption and Censoring
The Cox model supplies a compact hazard-ratio summary, but that summary should be interpreted in the context of the model's assumptions. In particular, the conventional Cox proportional-hazards interpretation assumes that the relative hazard can be represented adequately by a common treatment effect over time.
A single HR is most naturally interpreted when the relative hazard is reasonably represented as stable over time. If hazards vary substantially in their relative separation, one HR can become a less complete description of the treatment effect.
The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic. The appropriate conclusion is therefore limited: the trial reported a Cox-model HR, and the HR should be understood as the model-based summary rather than as a universal statement about every point in follow-up.
Censoring is also fundamental. A participant who has not experienced the first qualifying event when follow-up ends contributes information up to the censoring time. Survival-analysis validity depends on assumptions concerning censoring and the relationship between follow-up availability and the event process.
22. What the P-values Do — and Do Not — Say
P = 0.0014
For the primary renal composite, this is the reported P-value for the superiority comparison using the stratified log-rank framework.
P = 0.0339
For the key secondary cardiovascular composite, this is the reported P-value under the corresponding time-to-event analysis.
P = 0.2348
For all-cause mortality, this P-value accompanies an HR of 0.895 and a 95% CI that crosses 1.
P = 0.1623
For all-cause hospitalization, this P-value accompanies an HR of 0.946 and a 95% CI that crosses 1.
A P-value is calculated under a specified null hypothesis and statistical model. It does not measure the probability that the null hypothesis is true, the probability that the treatment works, or the clinical importance of the result.
For example, the primary HR of 0.825 and its CI of 0.732–0.928 describe the estimated effect and its precision. The P-value of 0.0014 adds information about statistical evidence against the null hypothesis. Neither number substitutes for the other.
23. Design Features the Registry Does Not Quantify Here
The ClinicalTrials.gov record identifies several design characteristics directly, but do not provide numerical details for every possible design topic. The distinction is important because an educational analysis should not fill gaps with assumptions.
| Design topic | What is supported by the ClinicalTrials.gov record | Interpretation |
|---|---|---|
| Randomization | Allocation is randomized. | The treatment comparison is based on randomized groups. |
| Parallel design | Design model is parallel. | The two interventions form parallel treatment groups. |
| Blinding | Masking is quadruple. | The registry identifies a quadruple-masked trial. |
| Superiority | Hypothesis type is superiority. | The reported efficacy tests are framed as superiority comparisons. |
| Non-inferiority margin | Not reported. | No non-inferiority margin is applicable to the reported superiority analyses presented here. |
| Crossover | Not reported. | No crossover analysis is added. |
| Factorial design | Not reported; design model is parallel. | No factorial analysis is added. |
| Bayesian methods | Not reported. | No Bayesian analysis is added. |
| Interim analysis | Not reported in the ClinicalTrials.gov record. | No interim-analysis schedule or alpha-spending method is inferred. |
| Missing-data imputation | Not reported. | No specific imputation method is attributed to the trial. |
24. Why the Primary Result Is a Time-to-Event Result
The primary renal endpoint is defined by the first occurrence of a qualifying event after randomization. This creates a time-to-event estimand rather than a simple binary endpoint at a fixed calendar date.
The statistical analysis uses the timing of the first qualifying event. Participants without an event by the end of their observed follow-up contribute information until censoring.
This distinction explains why the trial reports hazard ratios rather than simply reporting the difference in the number of participants with an event. A participant who experiences an event earlier and a participant who experiences the same event much later do not provide identical time-to-event information.
The same logic applies to the key secondary cardiovascular composite, all-cause mortality, all-cause hospitalization, and the secondary renal composite. Each is defined around time from randomization to the first qualifying event or censoring.
25. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary renal composite produced an HR of 0.825 with a two-sided 95% CI of 0.732–0.928 and P = 0.0014 under the reported stratified survival-analysis framework. The key secondary cardiovascular composite produced an HR of 0.860 with a 95% CI of 0.747–0.989 and P = 0.0339.
Effect-size interpretation
The HRs quantify relative differences in estimated hazard, while the confidence intervals quantify statistical uncertainty. The primary point estimate corresponds to a 17.5% lower estimated hazard, and the cardiovascular point estimate corresponds to a 14.0% lower estimated hazard.
Continuous endpoint
The UACR result uses a different estimand: a ratio of least squares means of 0.688 with a 95% CI of 0.662–0.715 and P < 0.0001. It should not be interpreted as a hazard ratio.
Safety interpretation
Serious adverse events are reported descriptively as 902/2827 for finerenone and 971/2831 for placebo. The ClinicalTrials.gov record does not provide a formal comparative safety test for this outcome.
26. Important Limitations and Interpretation Issues
- Composite endpoint: the primary HR summarizes time to the first occurrence of any qualifying component, not a separately estimated effect for each component.
- Hazard ratio interpretation: HRs are model-based relative measures and should not be converted into absolute risk differences without the underlying event-time and survival information.
- Proportional hazards: the registry reports a Cox proportional hazards model, but the ClinicalTrials.gov record does not report a formal diagnostic of the proportional-hazards assumption.
- Censoring: time-to-event analyses depend on the handling and assumptions associated with censoring.
- Analysis populations: time-to-event analyses use the full analysis set, while the UACR analysis requires measurements available within the Month 4 time window.
- Multiplicity: the registry-reported analysis notes describe a hierarchical testing strategy for secondary efficacy endpoints. P-values should therefore be interpreted in the context of that hierarchy rather than as a collection of unrelated tests.
- Missing-data methods: the ClinicalTrials.gov record does not specify a missing-data or imputation method, so no particular method is attributed to the trial.
- Component-specific effects: the ClinicalTrials.gov record does not provide statistical estimates for individual components of the composite endpoints, so component-level treatment effects are not inferred.
- Safety comparison: serious adverse events are provided as affected/at-risk counts without a formal comparative analysis in the ClinicalTrials.gov record.
- Generalizability: the trial population is specifically described as subjects with type 2 diabetes mellitus and diabetic kidney disease, and the registry condition is chronic kidney disease. Application outside the studied population requires separate clinical judgment.
27. Why This Trial Matters Statistically
FIDELIO-DKD is a useful teaching case because it places several recurring clinical-trial methods in one analysis: randomized parallel-group comparison, quadruple masking, composite time-to-event endpoints, stratified log-rank testing, Cox hazard-ratio estimation, ANCOVA for a continuous endpoint, intention-to-treat concepts, confidence intervals, P-values, and hierarchical testing.
| Concept | How it appears in FIDELIO-DKD |
|---|---|
| Randomization | The registry identifies randomized allocation with two treatment arms. |
| Parallel design | The design model is parallel. |
| Blinding | The trial is quadruple masked. |
| Time-to-event endpoints | The primary, cardiovascular, mortality, hospitalization, and secondary renal outcomes are analyzed from randomization to an event or censoring. |
| Log-rank test | All five posted time-to-event analyses use a stratified log-rank test. |
| Hazard ratio | Stratified Cox regression supplies the reported HRs and two-sided 95% confidence intervals. |
| ANCOVA | Change in UACR from baseline to Month 4 is analyzed using ANCOVA. |
| Ratio of least squares means | The UACR effect is reported as a ratio of least squares means rather than an HR. |
| Intention-to-treat | Intention-to-treat analysis is identified as a concept in the posted efficacy analyses. |
| Stratified analysis | The time-to-event analyses are described as stratified. |
| Multiplicity | The registry-reported analysis notes describe hierarchical testing of additional secondary efficacy endpoints. |
| Composite endpoint | The primary and secondary renal endpoints combine several possible first events. |
The most important statistical lesson is that a trial result is not a single number. The HR, confidence interval, P-value, endpoint definition, analysis population, statistical model, censoring structure, and multiplicity framework jointly define what the result means.
28. A Worked Reading of the Primary Analysis
Suppose the primary result is encountered in abbreviated form as: HR 0.825, 95% CI 0.732–0.928, P = 0.0014. A statistically literate reading proceeds in several steps.
First, the endpoint must be identified: this is not all-cause mortality or hospitalization but the first occurrence of the specified renal composite. Second, the effect measure is a hazard ratio. Third, the point estimate indicates a lower estimated hazard in the finerenone group. Fourth, the confidence interval describes uncertainty around that estimate. Fifth, the P-value describes the statistical evidence under the specified hypothesis-testing framework.
This approach prevents two common errors. The first is treating an HR as if it were an absolute risk reduction. The second is treating the P-value as if it were the treatment effect. Neither interpretation is statistically correct.
29. Sources
- ClinicalTrials.gov: NCT02540993 — FIDELIO-DKD.
- PubMed record: PMID 33264825.
- PubMed record: PMID 33198491.
- PubMed record: PMID 42234437.
- PubMed record: PMID 41827014.
- PubMed record: PMID 41679125.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the methods behind randomized clinical-trial analysis, survival endpoints, confidence intervals, hypothesis testing, and continuous-outcome models.