← Clinical Trials
Heart Failure Phase 2 Randomized NCT01807221

ARTS-HF: Complete Statistical Analysis of Finerenone in Heart Failure

An independent statistical analysis of the randomized phase 2 ARTS-HF study evaluating different oral doses of finerenone (BAY94-8862) in subjects with worsening chronic heart failure and left ventricular systolic dysfunction and either type 2 diabetes mellitus with or without chronic kidney disease or chronic kidney disease alone.

Completed  ·  Enrollment 1066  ·  Six-arm parallel design  ·  Quadruple masking
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics are restricted to the ClinicalTrials.gov data posted on ClinicalTrials.gov for ARTS-HF. The registry provides formal primary-endpoint analyses using the full analysis set and a chi-squared method.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ARTS-HF was a randomized, parallel-group, quadruple-masked phase 2 trial evaluating different oral doses of finerenone against eplerenone and placebo in subjects with worsening chronic heart failure and left ventricular systolic dysfunction and selected diabetes or chronic kidney disease characteristics.

1066
Enrollment
Participants
6
Arms
Parallel design
5
Primary analyses
All with estimate + CI
90%
Confidence interval
Two-sided
FeatureARTS-HF
Trial nameARTS-HF
NCT IDNCT01807221
PhasePhase 2
StatusCompleted
Start2013-06-17
Primary completion2014-11-11
Enrollment1066
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Lead sponsorBayer
Sponsor typeIndustry
ConditionHeart Failure
InterventionsFinerenone (BAY94-8862), placebo, and Inspra (eplerenone)

2. Clinical Question

The registered primary question was whether different oral doses of finerenone could produce a relative decrease in NT-proBNP of more than 30% from baseline to Day 90 in subjects with worsening chronic heart failure and left ventricular systolic dysfunction and either type 2 diabetes mellitus with or without chronic kidney disease or chronic kidney disease alone.

Population

Subjects with worsening chronic heart failure and left ventricular systolic dysfunction and either type 2 diabetes mellitus with or without chronic kidney disease or chronic kidney disease alone.

Intervention

Different oral doses of finerenone (BAY94-8862), including 2.5-5 mg OD, 5-10 mg OD, 7.5-15 mg OD, 10-20 mg OD, and 15-20 mg OD.

Comparator

Eplerenone (INSPRA®), with placebo also included as a trial intervention.

Primary question

What percentage of participants had a relative decrease in NT-proBNP of more than 30% from baseline to Day 90, and how did the finerenone dose groups compare with eplerenone?

3. Trial Design

01
Enroll1066 participants
02
RandomizeSix parallel arms
03
MaskQuadruple masking
04
AssessNT-proBNP
05
AnalyzeDay 90 endpoint
COMPARATOR · n = 221

Eplerenone (INSPRA®)

  • Eplerenone (INSPRA®)
  • Serious adverse events: 77/221
FINERENONE · 2.5-5 mg OD · n = 172

Finerenone (BAY94-8862)

  • Finerenone 2.5-5 mg OD
  • Serious adverse events: 72/172
FINERENONE · 5-10 mg OD · n = 163

Finerenone (BAY94-8862)

  • Finerenone 5-10 mg OD
  • Serious adverse events: 47/163
FINERENONE · 7.5-15 mg OD · n = 167

Finerenone (BAY94-8862)

  • Finerenone 7.5-15 mg OD
  • Serious adverse events: 52/167
FINERENONE · 10-20 mg OD · n = 169

Finerenone (BAY94-8862)

  • Finerenone 10-20 mg OD
  • Serious adverse events: 46/169
FINERENONE · 15-20 mg OD · n = 163

Finerenone (BAY94-8862)

  • Finerenone 15-20 mg OD
  • Serious adverse events: 57/163
Six-arm structure matters. The registry identifies six study arms, including eplerenone, five finerenone dose ranges, and placebo as trial interventions. The five formal primary analyses reported in the ClinicalTrials.gov record are comparisons of eplerenone with the five finerenone dose groups. The ClinicalTrials.gov record does not provide a formal primary-endpoint estimate for placebo.

4. Endpoints

EndpointRegistry definition / time frameEndpoint type
Primary endpoint Percentage of Participants With a Relative Decrease in NT-proBNP of More Than 30% From Baseline to Day 90 Binary
Time frame Baseline and Day 90

The registry describes NT-proBNP as a blood measurement used for screening and diagnosis of acute and chronic heart failure and that may be useful for establishing prognosis in heart failure. The registered endpoint converts the continuous biomarker measurement into a binary outcome: whether a participant achieved a relative decrease of more than 30% between baseline and Day 90.

Binary endpoint definition
Responder = relative decrease in NT-proBNP > 30% from baseline to Day 90

The analysis therefore concerns the percentage of participants meeting the registered threshold, rather than the mean change in NT-proBNP itself.

5. Statistical Methodology

Full analysis set

The primary analyses used the full analysis set (FAS). The registry defines this as all participants who received the study drug and had baseline and at least one post-baseline NT-proBNP value, or who died or experienced permanent withdrawal of study treatment lasting at least 5 consecutive days.

This population definition is important because it does not simply say "all randomized participants." Eligibility for the FAS depends on treatment exposure and the availability of relevant post-baseline information, with specific handling of death or permanent withdrawal.

Chi-squared test

The registry reports a chi-squared method for each of the five primary comparisons. Because the primary endpoint is binary, each participant can be classified according to whether the more-than-30% relative NT-proBNP decrease criterion was met. The chi-squared framework evaluates the distribution of this categorical outcome across the compared treatment groups.

Conceptual 2 × 2 comparison
Treatment group × { NT-proBNP decrease > 30%,  NT-proBNP decrease ≤ 30% }

The central statistical object is the difference between the proportions meeting the binary response criterion in the two treatment groups.

Reported effect measure: mean difference

Although the endpoint itself is binary and is reported as a percentage of participants, the registry records the effect measure as Mean Difference (Final Values), normalized here as mean difference. The analysis notes state that treatment differences were represented as πBi - πC.

For interpretation, a negative percentage-point difference means the first quantity in that treatment-difference representation is lower than the comparator quantity, while a positive value means it is higher. This is a difference in percentages, not a relative percentage change.

Confidence intervals

The registry reports two-sided 90% confidence intervals. Clopper-Pearson confidence intervals were calculated for each treatment group, while exact unconditional confidence limits were calculated for treatment differences.

This distinction is useful statistically. The confidence interval for an individual treatment-group percentage addresses uncertainty around that proportion. The confidence interval for the treatment difference addresses uncertainty around the contrast between groups.

Intention-to-treat concept

The statistical-analysis records also identify intention-to-treat analysis as an analysis concept. At the same time, the registry-reported definition of the FAS is more specific than simply saying that every randomized participant was analyzed. A careful reading therefore distinguishes the registry's stated ITT concept from the operational FAS definition used for the posted primary analyses.

6. Primary Results: Eplerenone vs Finerenone

The registry posts five formal primary-endpoint analyses. All use the same binary endpoint, the same Baseline-to-Day-90 time frame, the same FAS framework, and the chi-squared method. What changes is the finerenone dose range being compared with eplerenone.

ComparisonMean differenceTwo-sided 90% CIP-value
Eplerenone vs finerenone 2.5-5 mg OD-6.3-14.9 to 2.30.8771
Eplerenone vs finerenone 5-10 mg OD-4.7-13.4 to 40.7945
Eplerenone vs finerenone 7.5-15 mg OD0.1-8.5 to 8.80.5
Eplerenone vs finerenone 10-20 mg OD1.6-7.1 to 10.20.4225
Eplerenone vs finerenone 15-20 mg OD-3-11.7 to 5.70.6865

Finerenone 2.5-5 mg OD vs eplerenone

Reported treatment difference

-6.3

Two-sided 90% CI: -14.9 to 2.3   ·   P = 0.8771

Effect measure reported by the registry: Mean Difference (Final Values).

Clinical Biostats interpretation

The reported treatment difference is -6.3 percentage points in the registry's πBi - πC representation. In practical terms, the estimated percentage meeting the more-than-30% NT-proBNP decrease criterion was lower for the finerenone 2.5-5 mg OD comparison quantity than for its eplerenone comparator quantity by 6.3 percentage points.

The estimate does not mean that NT-proBNP itself fell by 6.3%, and it does not describe an individual participant's biomarker response. It is a between-group difference in the percentage meeting a binary response threshold.

The two-sided 90% CI extends from -14.9 to 2.3. Because this interval includes zero, the data represented by the registry are compatible with a negative difference, little difference, or a positive difference within that interval. The interval therefore communicates substantially more than the point estimate alone.

The P-value of 0.8771 is evidence against the particular null hypothesis tested under the reported analysis framework; it is not a measure of effect size, clinical importance, or the probability that one treatment is better than another.

Finerenone 5-10 mg OD vs eplerenone

Reported treatment difference

-4.7

Two-sided 90% CI: -13.4 to 4   ·   P = 0.7945

Effect measure reported by the registry: Mean Difference (Final Values).

Clinical Biostats interpretation

The estimated treatment difference for finerenone 5-10 mg OD relative to eplerenone is -4.7 percentage points under the registry's πBi - πC representation. The quantity describes the difference in the percentage of participants achieving the registered NT-proBNP response criterion.

It does not mean that every participant had a 4.7% lower biomarker response, nor does it establish a proportional reduction in NT-proBNP concentrations. The endpoint is a threshold-based binary classification.

The two-sided 90% CI is -13.4 to 4. Since zero lies inside the interval, the reported estimate is not separated from no difference by this confidence interval. The width of the interval also shows that the point estimate should not be treated as an exact measurement of the underlying treatment difference.

The P-value of 0.7945 should be read as a hypothesis-testing quantity under the specified analysis, not as a probability that the observed difference is correct. It also should not be compared directly across dose groups as if smaller P-values represented larger treatment effects.

Finerenone 7.5-15 mg OD vs eplerenone

Reported treatment difference

0.1

Two-sided 90% CI: -8.5 to 8.8   ·   P = 0.5

Effect measure reported by the registry: Mean Difference (Final Values).

Clinical Biostats interpretation

The estimated difference is 0.1 percentage points, which is very close to zero as a point estimate. The appropriate interpretation is that the two compared percentages were estimated to be nearly the same under this analysis, not that the treatments are proven to be identical.

The two-sided 90% CI ranges from -8.5 to 8.8. That interval includes zero and spans both directions of possible treatment difference. It therefore indicates uncertainty around the near-zero point estimate rather than establishing equivalence.

The P-value of 0.5 does not measure the size of the observed 0.1-point difference. A P-value and an effect estimate answer different questions: the estimate describes the observed contrast, while the P-value addresses compatibility with the hypothesis being tested.

Because the endpoint is binary, the result should also be understood in terms of the underlying responder proportions. The ClinicalTrials.gov record does not provide those arm-specific percentages, so the treatment-difference estimate should not be reverse-engineered into unreported response rates.

Finerenone 10-20 mg OD vs eplerenone

Reported treatment difference

1.6

Two-sided 90% CI: -7.1 to 10.2   ·   P = 0.4225

Effect measure reported by the registry: Mean Difference (Final Values).

Clinical Biostats interpretation

The estimated treatment difference is 1.6 percentage points in the registry's treatment-difference representation. This is a small positive point estimate, meaning the percentage represented by the finerenone treatment quantity was estimated to exceed the comparator quantity by 1.6 percentage points.

The estimate does not establish that this positive difference is a reproducible treatment effect. The two-sided 90% CI extends from -7.1 to 10.2, crossing zero and allowing for differences in either direction.

The P-value of 0.4225 is not an effect-size statistic. It should not be interpreted as saying that there is a 42.25% probability of benefit, nor as quantifying how clinically important the 1.6-point estimate might be.

For a binary endpoint, the confidence interval is particularly useful because it places the observed treatment difference in context. The registry's registry-reported interval is compatible with a modest negative difference as well as a larger positive difference.

Finerenone 15-20 mg OD vs eplerenone

Reported treatment difference

-3

Two-sided 90% CI: -11.7 to 5.7   ·   P = 0.6865

Effect measure reported by the registry: Mean Difference (Final Values).

Clinical Biostats interpretation

The estimated treatment difference is -3 percentage points in the registry's πBi - πC representation. This describes a three-percentage-point lower estimated response percentage for the finerenone treatment quantity relative to its comparator quantity.

It does not mean that NT-proBNP concentrations were three percentage points lower, and it does not imply that three out of every one hundred individual participants experienced a particular biomarker change.

The two-sided 90% CI of -11.7 to 5.7 crosses zero. The confidence interval therefore does not isolate a single direction of treatment difference. Its breadth also emphasizes why the point estimate should be interpreted together with uncertainty.

The P-value of 0.6865 is a hypothesis-testing result under the reported chi-squared framework. It does not quantify treatment benefit or harm and should not be treated as a substitute for the effect estimate and its confidence interval.

7. Interpreting the Five Primary Comparisons Together

The five posted analyses form a dose-ranging comparison structure rather than five unrelated endpoints. Every analysis uses the same registered primary endpoint and Baseline-to-Day-90 time frame, while the finerenone dose range changes from one comparison to the next.

Finerenone dose rangeEstimate90% CIP-valueWhat the interval indicates
2.5-5 mg OD-6.3-14.9 to 2.30.8771Includes zero
5-10 mg OD-4.7-13.4 to 40.7945Includes zero
7.5-15 mg OD0.1-8.5 to 8.80.5Includes zero
10-20 mg OD1.6-7.1 to 10.20.4225Includes zero
15-20 mg OD-3-11.7 to 5.70.6865Includes zero

A useful statistical observation is that all five registry-reported confidence intervals include zero. This means none of the five posted interval estimates excludes no treatment difference at the stated 90% confidence level. That is a statement about the reported statistical uncertainty; it is not equivalent to proving that the treatments have identical effects.

The point estimates themselves range from -6.3 to 1.6. That range should not be interpreted as evidence of a monotonic dose-response relationship. With only these summary estimates, formal conclusions about dose-response require an appropriate prespecified model or comparison strategy rather than visually ordering the point estimates.

Multiplicity matters: the registry supplies five primary-endpoint treatment comparisons. Because several comparisons are made, the overall interpretation of the collection of P-values depends on the trial's multiplicity strategy. The ClinicalTrials.gov record does not specify an adjustment procedure or an overall familywise error allocation, so the five nominal P-values should not be treated as though a multiplicity procedure were known when it is not reported here.

8. Statistical Methods Explained

Why was a chi-squared test used?

The primary endpoint is binary: a participant either had a relative decrease in NT-proBNP of more than 30% from baseline to Day 90 or did not. A chi-squared test is a standard categorical-data method for comparing distributions of such outcomes between treatment groups. The registry specifically reports chi-squared for all five primary analyses.

Why is the endpoint analyzed as a percentage?

The registered outcome is explicitly the percentage of participants meeting the more-than-30% reduction criterion. Converting the continuous biomarker change into a responder/non-responder endpoint gives the trial a clinically interpretable threshold, but it also discards information about how far individual participants moved beyond or below that threshold.

What does a mean difference of -6.3 mean here?

For the 2.5-5 mg OD comparison, the reported effect estimate is -6.3. Because the endpoint is a percentage of participants and the registry represents treatment differences as πBi - πC, the quantity is interpreted as a percentage-point difference between the two treatment groups. It is not a hazard ratio, odds ratio, risk ratio, or mean change in NT-proBNP.

Why does the confidence interval matter?

A point estimate alone can give a false impression of precision. For the 2.5-5 mg OD comparison, for example, the estimate is -6.3 but the two-sided 90% confidence interval extends from -14.9 to 2.3. The interval communicates that the estimated treatment difference is uncertain and that zero is within the interval.

What does a P-value of 0.8771 mean?

It is a hypothesis-testing quantity produced by the reported analysis. It does not say that there is an 87.71% probability that the null hypothesis is true, and it does not say that the treatment effect is 87.71% absent. Effect size and statistical evidence against a null hypothesis are different concepts.

What is the role of the full analysis set?

The FAS defines which participants contribute to the primary analyses. Here, the registry specifies receipt of study drug plus baseline and post-baseline NT-proBNP information, with additional inclusion through death or permanent withdrawal lasting at least 5 consecutive days. Analysis-population definitions can materially affect which observations enter a treatment comparison, so they should be read alongside the statistical method.

Why should the five dose comparisons not be treated as five independent experiments?

They arise from the same randomized trial and use the same primary endpoint. Multiple comparisons can increase the chance of obtaining a small P-value somewhere in a collection of tests even when there is no true effect. Appropriate interpretation therefore depends on whether and how multiplicity was handled. The ClinicalTrials.gov record does not specify such an adjustment.

9. Confidence Intervals and the Meaning of Zero

All five posted treatment-difference confidence intervals cross zero. This feature deserves careful interpretation because "the confidence interval includes zero" is often oversimplified as "there is no effect."

What crossing zero tells us

The registry-reported interval does not exclude a treatment difference of zero at the stated confidence level. Both positive and negative differences remain compatible with the interval.

What it does not tell us

It does not prove the treatments are equivalent, nor does it establish that any true difference is clinically unimportant.

Why the point estimate still matters

The estimate identifies the observed direction and magnitude of the treatment contrast. It should be interpreted together with the interval rather than discarded when the interval crosses zero.

Why the 90% level matters

The registry specifically reports two-sided 90% confidence intervals. Confidence-level interpretation should therefore not silently substitute a different confidence level.

Treatment-difference framework
Difference = πBi − πC

The registry's analysis notes use πBi - πC for treatment differences. The ClinicalTrials.gov record therefore describe the contrast between the finerenone treatment group and its comparator under that representation.

10. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk. These figures provide an arm-level safety summary, separate from the primary NT-proBNP efficacy endpoint.

ArmSerious adverse eventsAffected / at risk
Eplerenone (INSPRA®)7777/221
Finerenone (BAY94-8862) 2.5-5 mg OD7272/172
Finerenone (BAY94-8862) 5-10 mg OD4747/163
Finerenone (BAY94-8862) 7.5-15 mg OD5252/167
Finerenone (BAY94-8862) 10-20 mg OD4646/169
Finerenone (BAY94-8862) 15-20 mg OD5757/163

The numerator and denominator should be kept together when reading these safety data. The raw counts are not directly comparable without considering the different numbers of participants at risk in each arm. Conversely, converting the registry-reported fractions into new percentages would constitute a derived calculation that is unnecessary for the registry-based presentation and is therefore not used here.

Safety interpretation: serious adverse events and the NT-proBNP endpoint answer different questions. The safety counts describe participants affected by serious adverse events, whereas the primary efficacy endpoint describes the percentage meeting a biomarker-response threshold. Neither should be substituted for the other when interpreting the trial.

11. Analysis Population and Missing Information

The FAS definition explicitly addresses post-baseline information and permanent withdrawal. Participants were included if they had baseline and at least one post-baseline NT-proBNP value, or if they died or experienced permanent withdrawal of study lasting at least 5 consecutive days after receiving study drug.

This is statistically important because a binary endpoint measured at Day 90 can be difficult to define for participants who discontinue early or do not have the scheduled biomarker measurement. The registry's reported FAS definition provides a rule for determining who enters the analysis, but the ClinicalTrials.gov record does not provide a detailed imputation algorithm for missing Day 90 NT-proBNP values.

Observed endpoint

The primary outcome is defined around the change from baseline to Day 90 and classifies participants according to a greater-than-30% relative decrease.

FAS rule

The registry includes participants with relevant post-baseline NT-proBNP information and specifies additional inclusion through death or permanent withdrawal.

Missing-data limitation

The ClinicalTrials.gov record does not specify a separate imputation model for missing Day 90 values.

Interpretive consequence

The exact construction of the analysis population matters because the reported treatment differences apply to the defined FAS rather than automatically to every enrolled participant.

12. Randomization and Blinding

ARTS-HF was randomized, parallel, and quadruple-masked. Randomization is the principal design feature that allows treatment groups to be compared without deliberately assigning treatment according to participants' baseline characteristics. Masking is a separate protection against knowledge of treatment assignment influencing trial conduct or assessment.

Allocation
Randomized
Participants were allocated through a randomized trial design.
Design
Parallel
The registry identifies a parallel-group design with six arms.
Masking
Quadruple
The registry classifies the study as quadruple-masked.
Purpose
Treatment
The primary purpose is classified as treatment.

For statistical interpretation, randomization and masking should not be conflated. Randomization concerns allocation and comparability of treatment groups; masking concerns knowledge of assignment. Both are design features, but they address different sources of bias.

13. What the Primary Results Do — and Do Not — Establish

What the estimates establish descriptively

The registry reports five treatment differences for the same binary NT-proBNP endpoint, ranging from -6.3 to 1.6, each with a two-sided 90% confidence interval and a P-value.

What the confidence intervals add

All five registry-reported intervals include zero. Thus, at the reported 90% confidence level, none of the five intervals excludes a zero treatment difference. The intervals also demonstrate that the point estimates are uncertain rather than exact measurements of a population effect.

What the P-values do not measure

The P-values do not measure the size of the treatment effect, the clinical importance of the biomarker threshold, or the probability that a treatment is effective. Those questions require interpretation of the effect estimate, uncertainty interval, endpoint definition, and trial design.

Why this is not an equivalence conclusion

A confidence interval that includes zero is not by itself an equivalence test. Demonstrating equivalence requires a prespecified equivalence margin and an analysis explicitly designed to test whether the treatment difference lies within that margin. No such margin is reported in the ClinicalTrials.gov record.

14. Important Limitations and Interpretation Issues

15. Why This Trial Matters Statistically

ARTS-HF is a useful teaching case because it illustrates how a randomized multi-arm clinical trial can combine a binary biomarker endpoint, categorical-data testing, treatment-difference estimates, exact confidence intervals, and a defined analysis population.

ConceptHow it appears in ARTS-HF
RandomizationThe trial uses randomized allocation across six arms.
Parallel designThe registry classifies the design model as parallel.
Quadruple maskingThe registry classifies the study as quadruple-masked.
Binary endpointParticipants are classified according to whether NT-proBNP decreased by more than 30% from baseline to Day 90.
Chi-squared testReported statistical method for all five formal primary analyses.
Mean differenceThe registry reports Mean Difference (Final Values) as the effect measure.
Confidence intervalTwo-sided 90% intervals are reported for treatment differences.
Exact methodsClopper-Pearson intervals are reported for treatment groups and exact unconditional limits for treatment differences.
Analysis populationPrimary analyses use a registry-defined full analysis set.
MultiplicityFive primary treatment comparisons are posted for the same endpoint.
Safety analysisSerious adverse events are reported separately for each treatment arm.

16. Statistical Interpretation of the Dose Comparisons

The most important analytical distinction is between describing the observed pattern and claiming a dose-response relationship. The five estimates are -6.3, -4.7, 0.1, 1.6, and -3 across the listed finerenone dose ranges. These values do not form a monotonic sequence.

That observation alone does not establish that there is no dose-response relationship. A formal dose-response analysis would require an appropriate statistical model, prespecified contrasts, or another methodology designed to test the relationship between dose and response. The ClinicalTrials.gov record reports pairwise treatment comparisons rather than such a dose-response model.

Similarly, comparing P-values across dose groups is not a substitute for comparing effect estimates. A P-value can change because of sampling variability and precision even when two effect estimates have similar magnitude. The confidence interval and treatment difference should remain central to interpretation.

A useful reading order
Endpoint definition → analysis population → effect estimate → confidence interval → P-value → multiplicity → clinical context

Reading results in this order helps prevent a P-value from becoming the only statistic used to describe a randomized treatment comparison.

17. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The five reported finerenone-versus-eplerenone treatment differences range from -6.3 to 1.6, and every registry-reported two-sided 90% confidence interval includes zero. The registry reports chi-squared analyses for these binary endpoint comparisons.

Clinical interpretation

The primary endpoint measures a specific biomarker-response threshold at Day 90. The ClinicalTrials.gov record does not establish that the observed treatment differences translate directly into differences in clinical outcomes such as survival or hospitalization.

This separation is important. Statistical evidence describes uncertainty in the observed treatment comparison. Clinical interpretation asks what the endpoint itself represents and how meaningful that endpoint is for patients and disease management. The two levels of interpretation should not be collapsed into a single P-value.

18. Related Statistical Tutorials

Learn more about the methods used in this trial:

19. Related Statistical Calculators

20. Sources

Continue through the Clinical Biostats statistical library

Use the related tutorials and calculators to explore the categorical-data methods, confidence intervals, P-values, randomization, and analysis-population concepts illustrated by ARTS-HF.

21. Record Summary

ARTS-HF provides a compact example of statistical analysis for a randomized, six-arm, quadruple-masked phase 2 trial. Its registered primary endpoint is a binary biomarker-response measure: the percentage of participants with a relative decrease in NT-proBNP of more than 30% from baseline to Day 90. The five posted primary analyses compare eplerenone with five finerenone dose ranges using a chi-squared method in the registry-defined full analysis set, with mean differences and two-sided 90% confidence intervals.

The five treatment differences are -6.3, -4.7, 0.1, 1.6, and -3, respectively. Every registry-reported confidence interval includes zero. The appropriate statistical reading is therefore to describe the observed treatment contrasts and their uncertainty without converting those results into an equivalence claim or treating the P-values as measures of effect size.

Clinical Biostats methodology: A trial-results page should not merely repeat registry fields. The goal is to reconstruct the statistical story of the trial while keeping reported evidence, uncertainty, analysis-population definitions, and educational interpretation clearly separated.