← Clinical Trials
Septic Shock Phase 3 Completed NCT03434028

CLOVERS: Complete Statistical Analysis of Early Fluid Restriction in Septic Shock

An independent statistical review of the randomized phase 3 CLOVERS trial comparing restrictive-fluid and liberal-fluid resuscitation strategies in patients with septic shock, with emphasis on the registered day-90 mortality endpoint and the statistical interpretation of its confidence interval and Z-test.

Trial start: March 7, 2018  ·  Primary completion: May 10, 2022  ·  Enrollment: 1563
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the ClinicalTrials.gov record. The registry reports 14 outcome measures and 14 statistical analyses.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

CLOVERS was a randomized, parallel, unmasked, phase 3 treatment trial evaluating early vasopressors and early fluids in patients with septic shock. The trial enrolled 1563 participants and reported a primary endpoint of death before discharge home by day 90.

1563
Enrollment
Total participants
2
Arms
Parallel randomized design
−0.9
Primary difference
Restrictive − liberal mortality rate
0.61
Primary P-value
Two-sided Z-test
FeatureCLOVERS
Trial nameCLOVERS
Brief titleCrystalloid Liberal or Vasopressors Early Resuscitation in Sepsis
PhasePhase 3
ConditionSeptic Shock
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment1563
InterventionsEarly Vasopressors; Early Fluids
Lead sponsorMassachusetts General Hospital
StatusCompleted
ClinicalTrials.govNCT03434028

2. Clinical Question

The central statistical question was whether the two randomized resuscitation strategies differed in the probability of death from any cause before discharge home by day 90.

Population

Participants in the CLOVERS trial with the registered condition of septic shock.

Intervention

Early vasopressors, identified in the ClinicalTrials.gov record as an intervention and represented in the analysis data by the restrictive-fluid group.

Comparator

Early fluids, represented in the statistical analyses by the liberal-fluid group.

Primary question

Is there evidence of a difference between restrictive fluids and liberal fluids in death before discharge home by day 90?

The important point statistically is that the question is framed as a randomized comparison. The analysis therefore estimates a between-group difference rather than attempting to explain mortality from observational associations alone.

3. Trial Design

01
Randomize 1563 participants
02
Two arms Early vasopressors / early fluids
03
Follow Day 90 primary window
04
Time-to-event Kaplan-Meier estimation
05
Compare Difference and Z-test
Allocation
Randomized allocation to two parallel study arms.
Masking
None.
Primary hypothesis
Superiority.
Statistical methods
Kaplan-Meier estimation and Wald / Z-test methodology were identified in the registry data.
ANALYSIS GROUP · RESTRICTIVE FLUIDS

Restrictive-fluid strategy

  • Compared directly with the liberal-fluid group.
  • 109 deaths were reported for the primary endpoint.
  • 5 patients had censored data for the primary endpoint.
  • The primary point estimate was obtained from the Kaplan-Meier framework.
ANALYSIS GROUP · LIBERAL FLUIDS

Liberal-fluid strategy

  • Compared directly with the restrictive-fluid group.
  • 116 deaths were reported for the primary endpoint.
  • 4 patients had censored data for the primary endpoint.
  • The primary point estimate was obtained from the Kaplan-Meier framework.
Terminology: the ClinicalTrials.gov record lists the interventions as Early Vasopressors and Early Fluids, while the posted statistical analyses compare Restrictive Fluids versus Liberal Fluids. This page retains those registry terms rather than assigning additional treatment details that are not present in the ClinicalTrials.gov record.

4. Trial Timeline

2018-03-07

Trial start

The registered trial start date was March 7, 2018.

2022-05-10

Primary completion

The registered primary completion date was May 10, 2022.

Completed

Registry status

The trial is recorded as completed, with results posted.

5. Primary Endpoint

EndpointRegistered definitionAnalysis
Death Before Discharge Home by Day 90 From randomization to discharge home up to and including day 90. The primary outcome was death from any cause before discharge home by day 90. Kaplan-Meier point estimates with a reported Z-test comparing restrictive fluids with liberal fluids.

The registry further defines home as the same setting or a setting similar to the setting where the patient resided before becoming ill. The registry text reports the primary endpoint using a time-to-event framework: there were 109 deaths and 5 patients with censored data in the restrictive-fluid group, and 116 deaths and 4 patients with censored data in the liberal-fluid group.

This distinction matters. Although the endpoint can be described in binary terms as whether death occurred by day 90, the posted point estimates were obtained from Kaplan-Meier curves. That means the analysis retained information about when an event or censoring occurred rather than treating every participant simply as an uncensored yes/no observation at day 90.

6. Primary Result: Death Before Discharge Home by Day 90

Difference in mortality rate

−0.9

Restrictive Fluids − Liberal Fluids

95% CI: −4.4 to 2.6   ·   P = 0.61

Two-sided Z-test; superiority hypothesis

Primary endpointRestrictive FluidsLiberal FluidsDifference95% CIP-value
Death Before Discharge Home by Day 90 109 deaths; 5 censored 116 deaths; 4 censored −0.9 −4.4 to 2.6 0.61
Clinical Biostats interpretation

The estimated difference in mortality rate was −0.9, defined as restrictive fluids minus liberal fluids. Thus, the point estimate is in the direction of a lower mortality rate in the restrictive-fluid group, but the estimate itself should not be treated as evidence that the two strategies necessarily differ.

The 95% confidence interval of −4.4 to 2.6 is important because it spans zero. Under the statistical framework used for the reported estimate, the data are compatible with a range of differences extending from a larger negative difference to a positive difference. The interval therefore communicates substantially more information than the point estimate alone.

The P-value of 0.61 is a measure of compatibility with the null hypothesis under the specified test; it is not a measure of the magnitude of the treatment effect, the probability that the null hypothesis is true, or the probability that either strategy is clinically preferable.

Because the point estimates came from Kaplan-Meier curves and some participants were censored, the comparison should also be understood as a time-to-event analysis rather than as a simple raw proportion comparison. Censoring allows participants to contribute information up to the time at which their outcome information becomes censored.

Finally, the posted analysis is explicitly identified as a superiority analysis. The appropriate interpretation is therefore about evidence for a difference, not about proving that the strategies are equivalent or identical.

7. Understanding the Primary Analysis

Why use Kaplan-Meier estimation for a day-90 mortality endpoint?

Kaplan-Meier estimation is appropriate when the outcome is defined over time and some participants may not have a fully observed event time. Instead of assigning every participant a simple event/no-event label without regard to follow-up timing, the method uses the observed event and censoring times to estimate the survival function.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at event time ti, while ni represents the number at risk immediately before that time.

For CLOVERS, the ClinicalTrials.gov record explicitly states that the primary point estimates were from Kaplan-Meier curves. The presence of 5 censored observations in the restrictive-fluid group and 4 in the liberal-fluid group is therefore directly relevant to understanding why a time-to-event framework was used.

What does the reported Z-test do?

The posted primary analysis identifies a Z-test, normalized here as a Wald / Z-test. In general, a Z-type comparison standardizes an estimated treatment difference by its estimated standard error. Conceptually, the statistic has the form:

Wald / Z-test concept
Z = estimate / standard error

The resulting statistic measures how far the estimated effect lies from the null value relative to its estimated sampling uncertainty.

The registry reports the resulting two-sided P-value as 0.61. The P-value should therefore be interpreted in the context of the prespecified superiority comparison and its corresponding null hypothesis, rather than as a standalone measure of clinical importance.

8. Confidence Interval Interpretation

Point estimate

The reported difference was −0.9. It is the single best point estimate reported by the posted analysis.

Interval estimate

The 95% confidence interval was −4.4 to 2.6. It expresses uncertainty around the estimated difference.

Zero as the reference

For a difference measure, zero represents no difference between the two groups on that scale.

P-value

The two-sided P-value was 0.61. It does not quantify the size of the observed difference.

A useful way to read the result is to start with the scale of the effect. The estimate is a difference, not a hazard ratio. Therefore, an estimate of −0.9 is interpreted as the restrictive-minus-liberal difference on the reported mortality-rate scale. It should not be translated into a relative percentage reduction unless such a relative measure is explicitly reported.

The confidence interval also prevents an overly narrow interpretation. A point estimate close to zero can occur even when the underlying uncertainty is broader. Here, the interval extends on both sides of zero, so the estimate alone does not establish the direction of the underlying treatment difference.

Do not confuse the confidence interval with the range of individual patient outcomes. The 95% CI of −4.4 to 2.6 describes uncertainty around the estimated group-level effect under the analysis framework. It does not mean that individual patients have treatment effects confined to that interval.

9. Secondary Endpoint Results

The registry data contain statistical analyses for 13 secondary endpoints in addition to the primary endpoint. The posted estimates below are reproduced exactly as reported in the registry. Unless a P-value is explicitly reported in the trial data, no formal P-value is presented here.

Secondary endpoint Time frame Effect 95% CI
Organ Support Free Days 28 days after randomization Mean difference: 0.3 days −0.5 to 1.2
Ventilator Free Days (VFD) 28 days after randomization Mean difference: 0.6 days −0.4 to 1.6
Renal Replacement Free Days 28 days after randomization Mean difference: 0.2 days −0.8 to 1.2
Vasopressor Free Days From study day 2 through day 28 Mean difference: 0.4 days −0.5 to 1.3
ICU Free Days 28 days after randomization Mean difference: 0.1 days −0.8 to 1.0
Hospital Free Days to Discharge Home 28 days after randomization Mean difference: 0.8 days −0.3 to 1.9
New Intubation With Invasive Mechanical Ventilation by 28 Days 28 days after randomization Risk difference: −1.7 −5.1 to 1.7
Initiation of Renal Replacement Therapy 28 days after randomization Risk difference: 0.0 −1.8 to 1.8
Kidney Disease: Change in Creatinine-based Global Outcomes (KIDGO) Score Between Baseline and 72 Hours 72 hours after randomization Mean difference: 0.0 −0.1 to 0.1
Change in SOFA (Sepsis Related Organ Failure Assessment) Score 72 hours after randomization Mean difference: 0.1 −0.3 to 0.4
Development of ARDS 7 days after randomization Risk difference: −0.1 −1.7 to 1.5
New Onset Atrial or Ventricular Arrhythmia 28 days after randomization Risk difference: −1.0 −3.7 to 1.7
Death From Any Cause at Any Location by Day 90 From randomization to and including day 90 Risk difference: 0.5 −3.6 to 4.7

All of the listed confidence intervals are two-sided 95% intervals. The posted hypothesis type for these analyses is superiority.

10. Secondary Time-to-Event Analysis: Organ Support Free Days

Organ Support Free Days

0.3 days

Mean difference, Restrictive Fluids − Liberal Fluids

95% CI: −0.5 to 1.2

28 days after randomization

The registry identifies Kaplan-Meier estimation for this endpoint and notes that the reported mean is an estimate from the Kaplan-Meier curve. That is an important methodological detail: although the outcome is expressed in days, the posted analysis is not simply described as an ordinary comparison of arithmetic means from complete observations.

Clinical Biostats interpretation

The estimated difference was 0.3 days, meaning the restrictive-fluid group had an estimated mean that was 0.3 days higher on the reported scale. The 95% CI extends from −0.5 to 1.2 days, so the interval includes both negative and positive differences.

The registry explicitly states that available data were used without imputation for primary and secondary outcomes. Consequently, the interpretation of this endpoint should recognize that the denominator may vary for selected outcomes because of missing data.

No P-value is posted on ClinicalTrials.gov for this secondary analysis in the ClinicalTrials.gov record. It would therefore be inappropriate to manufacture a significance assessment from the confidence interval or to infer an unreported formal hypothesis test.

11. Secondary Continuous Outcomes

EndpointEstimate95% CIStatistical note
Ventilator Free Days (VFD) 0.6 −0.4 to 1.6 Mean difference; method not reported
Renal Replacement Free Days 0.2 −0.8 to 1.2 Mean difference; method not reported
Vasopressor Free Days 0.4 −0.5 to 1.3 Mean difference; method not reported
ICU Free Days 0.1 −0.8 to 1.0 Mean difference; method not reported
Hospital Free Days to Discharge Home 0.8 −0.3 to 1.9 Mean difference; method not reported
KIDGO Score Change 0.0 −0.1 to 0.1 Mean difference; method not reported
Change in SOFA Score 0.1 −0.3 to 0.4 Mean difference; method not reported

The common feature of these estimates is that the registry-reported 95% confidence intervals cross zero. That observation describes the interval estimates; it does not by itself establish the absence of an effect. A confidence interval that includes zero can reflect uncertainty that is consistent with both directions of the treatment difference.

For the endpoints where the registry does not report the analysis method, the appropriate statistical discipline is to preserve that limitation. Different methods for free-day outcomes can make different assumptions about distributions, censoring, death, and the relationship between an event and the number of observed free days. The ClinicalTrials.gov record does not provide enough information to select one particular unreported method as the actual CLOVERS method.

12. Secondary Binary Outcomes

EndpointEffect measureEstimate95% CI
New Intubation With Invasive Mechanical Ventilation by 28 Days Risk difference −1.7 −5.1 to 1.7
Initiation of Renal Replacement Therapy Risk difference 0.0 −1.8 to 1.8
Development of ARDS Risk difference −0.1 −1.7 to 1.5
New Onset Atrial or Ventricular Arrhythmia Risk difference −1.0 −3.7 to 1.7
Death From Any Cause at Any Location by Day 90 Risk difference 0.5 −3.6 to 4.7

Why risk difference is useful

A risk difference is an absolute comparison. Conceptually, it can be written as:

Risk difference
RD = RiskRestrictive − RiskLiberal

A negative value indicates a lower observed risk in the restrictive group on the reported scale; a positive value indicates a higher observed risk.

This is different from a risk ratio, odds ratio, or hazard ratio. For example, the reported risk difference of 0.0 for initiation of renal replacement therapy is a difference on an absolute scale. It should not be re-expressed as a relative effect without the underlying group-specific risks.

How to read the secondary binary estimates

The intubation estimate of −1.7 has a 95% CI of −5.1 to 1.7. The ARDS estimate of −0.1 has a 95% CI of −1.7 to 1.5. The arrhythmia estimate of −1.0 has a 95% CI of −3.7 to 1.7. In each case, the interval crosses zero, so the direction suggested by the point estimate is not established by the interval alone.

The death-from-any-cause-at-any-location endpoint has an estimate of 0.5 with a 95% CI of −3.6 to 4.7. It is a different endpoint definition from the primary endpoint because the primary endpoint specifically concerns death before discharge home by day 90.

13. Primary Endpoint vs Other Day-90 Mortality Endpoint

The registry contains two mortality measures that should not be silently combined.

FeaturePrimary endpointSecondary mortality endpoint
Endpoint Death Before Discharge Home by Day 90 Death From Any Cause at Any Location by Day 90
Time frame From randomization to discharge home up to and including day 90 From randomization to and including day 90
Effect measure Difference in mortality rate Risk difference
Estimate −0.9 0.5
95% CI −4.4 to 2.6 −3.6 to 4.7
Formal method reported Z-test Not reported

The distinction illustrates why endpoint definitions matter in clinical-trial statistics. Two outcomes can both concern death by day 90 while differing in where the event must occur or how discharge status enters the definition. The analysis should therefore preserve the registry wording rather than treating the endpoints as interchangeable.

14. Missing Data and Censoring

The registry-reported analysis text states that primary and secondary outcomes were reported using available data without imputation. It also notes that, because of missing data, the denominator may vary for selected outcomes as reported in the primary publication.

No imputation

The registry analysis states that available data were used without imputation for primary and secondary outcomes.

Variable denominators

Missing data can cause the denominator to vary for selected secondary outcomes.

Censoring

The primary endpoint explicitly includes censored observations within its Kaplan-Meier analysis.

Different concepts

Missing data and right-censoring are related to incomplete information but are not statistically identical concepts.

Missing data is not the same as censoring

For a time-to-event endpoint, censoring means that the event time is not observed beyond a known point of follow-up, while the participant can still contribute information before that point. Missing data can instead mean that a variable or outcome measurement is unavailable without necessarily defining an observed follow-up endpoint.

The distinction is especially important when interpreting secondary endpoints. The ClinicalTrials.gov record explicitly identifies missing-data considerations and the absence of imputation, but it does not provide enough information here to specify a missing-data mechanism such as missing completely at random, missing at random, or missing not at random.

Do not infer an imputation strategy. The ClinicalTrials.gov record explicitly says that primary and secondary outcomes were reported using available data without imputation. No additional imputation method should therefore be attributed to the trial.

15. Statistical Methodology

Kaplan-Meier estimation

Kaplan-Meier estimation was the principal time-to-event method identified in the ClinicalTrials.gov record. For the primary endpoint, the registry specifically states that point estimates were from Kaplan-Meier curves.

The Kaplan-Meier estimator constructs a stepwise estimate of the probability of remaining event-free over time. At each observed event time, the estimated survival probability is adjusted according to the number of events and the number of participants at risk.

Core Kaplan-Meier idea
Each event time updates the estimated survival probability using the risk set immediately before that time.

Participants who are censored contribute information while they remain under observation, but are not counted as events after their censoring time.

Wald / Z-test

The primary statistical analysis was reported as a Z-test and normalized as a Wald / Z-test. A Wald-type test compares an estimated effect with its null value after accounting for its estimated standard error.

For a difference measure, the null value is generally zero. The resulting test statistic is then compared with a reference normal distribution to obtain a two-sided P-value when a two-sided test is specified.

Effect measures

The registry-reported CLOVERS analyses use several effect-measure families: mean difference, other difference, and risk difference. These measures answer different statistical questions.

Effect measureInterpretationCLOVERS examples
Difference in mortality rate Absolute difference between the randomized groups on the reported mortality-rate scale. Primary endpoint
Mean difference Difference between group-level mean outcome values. Free-day and score outcomes
Risk difference Absolute difference in risk between groups. Intubation, renal replacement therapy, ARDS, arrhythmia, and day-90 mortality at any location

Superiority hypothesis

The trial data identify the hypothesis type as superiority. A superiority framework asks whether the randomized groups differ, rather than asking whether one treatment is sufficiently close to another to establish non-inferiority.

That distinction is central to interpreting the primary confidence interval. A confidence interval that includes zero does not establish equivalence. It indicates that the interval estimate includes the null value for the difference.

16. Why Randomization Matters

Randomization is the core design feature that supports a causal comparison between the two strategies. By randomly assigning participants to the study arms, the trial is designed so that systematic differences in baseline factors are not intentionally allocated to one strategy.

The causal comparison
Observed treatment contrast = outcome difference between randomized groups

The strength of the randomized design is that treatment assignment is determined independently of the participant's expected outcome, subject to the trial's actual implementation.

Randomization does not guarantee identical groups on every characteristic in a finite sample. Its statistical purpose is different: it provides a principled basis for comparing outcomes under the assigned strategies without relying on observational adjustment alone.

The parallel design also means that participants belong to one randomized arm rather than switching between randomized strategies as part of the primary comparison. The ClinicalTrials.gov record does not identify a crossover analysis, factorial design, or non-inferiority margin, so those concepts are not imposed on this trial.

17. Multiplicity and Multiple Endpoints

The registry contains one primary endpoint and multiple secondary endpoints. The trial data identify 14 outcome measures and 14 statistical analyses in total.

ComponentRegistry informationStatistical implication
Primary endpoint Death Before Discharge Home by Day 90 Primary confirmatory question
Secondary endpoints 13 additional posted analyses Supportive and additional outcome comparisons
Hypothesis type Superiority Differences are evaluated rather than non-inferiority margins
Multiplicity adjustment Not specified in the ClinicalTrials.gov record No additional multiplicity procedure should be attributed to the trial

Multiple secondary endpoints create a statistical interpretation issue even when every estimate is reported correctly. The more comparisons a study examines, the more opportunities there are for apparently notable findings to occur by chance. The ClinicalTrials.gov record does not specify a multiplicity-adjustment strategy, so this page does not infer one.

The primary endpoint should therefore remain conceptually distinct from the collection of secondary outcomes. A secondary confidence interval crossing zero does not become more or less informative because another secondary endpoint has a different point estimate.

Registry limitation: the ClinicalTrials.gov record does not report a multiplicity-adjustment procedure, alpha-spending plan, or interim-analysis rule. These methods should not be reconstructed from the observed P-value of 0.61.

18. Safety Results

The ClinicalTrials.gov record reports serious adverse events by analysis group. The counts are presented as affected participants divided by participants at risk.

Safety measureRestrictive FluidsLiberal Fluids
Serious adverse events 18 / 782 19 / 781

These figures should be interpreted as counts over the corresponding at-risk populations reported in the ClinicalTrials.gov record. The trial data do not provide a formal statistical comparison, confidence interval, or P-value for this safety measure, so none is added here.

Safety interpretation

The numbers show that serious adverse events were reported in 18 of 782 participants in the restrictive-fluid group and 19 of 781 participants in the liberal-fluid group. Descriptively, the affected-participant counts are close.

However, a descriptive comparison is not the same as a formal statistical test. Without an explicitly reported confidence interval or hypothesis test for this safety endpoint, it would be inappropriate to claim a statistically demonstrated difference or no difference.

Registry caveat concerning all-cause mortality

All-cause mortality clarification: the registry notes that death was not classified as anticipated or not anticipated, and that cause of death was not collected for this trial.

This caveat is relevant to interpretation of mortality as a safety-related outcome. The primary efficacy endpoint is explicitly defined as death from any cause before discharge home by day 90, while the registry separately states that cause of death was not collected.

19. Statistical Methods Explained

Why was a Kaplan-Meier analysis used for the primary endpoint?

The primary endpoint is defined from randomization through discharge home up to and including day 90. Because the registry reports censored observations and states that point estimates came from Kaplan-Meier curves, the analysis preserves information about the timing of events and censoring. This is more informative than treating every participant as though their outcome were observed at exactly the same time.

What does a difference of −0.9 mean?

The reported estimate is the restrictive-fluid mortality rate minus the liberal-fluid mortality rate. A value of −0.9 therefore points toward a lower mortality rate in the restrictive group on the reported scale. It does not mean a 0.9-fold hazard, a 0.9% relative reduction, or a 0.9% probability that an individual patient benefits.

Why is zero important for the confidence interval?

Zero is the null value for a difference. The primary 95% CI runs from −4.4 to 2.6 and therefore includes zero. This means the interval estimate includes the possibility of no difference as well as differences in either direction on the reported scale.

What does P = 0.61 mean?

The two-sided P-value of 0.61 is generated by the reported Z-test under the superiority testing framework. It quantifies how compatible the observed test statistic is with the null hypothesis under that framework. It does not measure effect size, clinical importance, or the probability that the treatment strategies are equivalent.

Why should a secondary endpoint without a P-value not be given one?

Because a P-value is an output of a particular statistical test. The ClinicalTrials.gov record reports confidence intervals and estimates for many secondary outcomes but do not provide formal P-values for them. Reconstructing an unreported P-value would require assumptions about the exact test, standard error, distribution, and analysis population.

What is the difference between a risk difference and a mean difference?

A risk difference compares probabilities of a binary event. A mean difference compares average values on a quantitative outcome scale. In CLOVERS, risk differences are reported for binary outcomes such as new intubation, while mean differences are reported for several free-day and score outcomes.

Does a confidence interval crossing zero prove that there is no treatment effect?

No. It means that zero is contained within the reported interval estimate. The interval reflects statistical uncertainty around the estimated effect; it does not prove that the true effect is exactly zero. The distinction is particularly important when the interval covers a range of potentially meaningful positive and negative differences.

20. Interpreting the Primary Result Without Overstating It

What the estimate says

The estimated restrictive-minus-liberal mortality difference was −0.9.

What the CI says

The 95% CI was −4.4 to 2.6, spanning zero and therefore spanning both directions of difference.

What the P-value says

The two-sided Z-test produced P = 0.61 under the superiority framework.

What it does not say

It does not establish equivalence, quantify individual benefit, or provide a relative hazard measure.

There is a useful hierarchy for interpreting the result. First identify the estimand: here, a difference in mortality rate. Next identify the direction and magnitude of the point estimate. Then inspect the confidence interval to understand uncertainty. Only after that should the P-value be considered as evidence relative to the prespecified null hypothesis.

This order prevents a common statistical error: reducing the entire result to a binary "significant/not significant" label. A P-value is only one component of the evidence. The effect measure tells us what was estimated, while the confidence interval tells us how precisely it was estimated under the analysis framework.

21. What the Primary Result Does — and Does Not — Mean

Effect estimate

The estimate of −0.9 means that the estimated mortality-rate difference was negative when calculated as restrictive fluids minus liberal fluids. The estimate therefore points toward a lower rate in the restrictive group on that scale.

It does not mean that patients assigned to restrictive fluids had exactly a 0.9% lower individual probability of death, nor does it imply a relative reduction of 0.9%.

Confidence interval

The 95% CI of −4.4 to 2.6 quantifies uncertainty around the estimated difference. Because it includes zero, the interval is compatible with no difference as well as with differences in either direction.

The confidence interval is not a prediction interval for individual patients and should not be interpreted as the range of mortality effects experienced by patients.

P-value

The P-value of 0.61 belongs to the reported two-sided Z-test. It describes the evidence against the null hypothesis under that test; it does not measure how large the treatment difference is.

Superiority framework

The registry identifies the hypothesis type as superiority. Therefore, a result that does not establish a difference should not automatically be translated into a claim of equivalence or identical outcomes.

22. Endpoint-Specific Interpretation

The secondary results illustrate why clinical-trial interpretation should remain endpoint-specific. The estimates are not all expressed on the same statistical scale.

Endpoint familyExamples in CLOVERSHow to interpret the effect
Mortality difference Death Before Discharge Home by Day 90 Absolute difference in the reported mortality-rate scale.
Mean difference Organ Support Free Days; VFD; ICU Free Days Difference in average outcome values between groups.
Risk difference New intubation; ARDS; arrhythmia Absolute difference in event risk between groups.

For example, the estimate of 0.8 for Hospital Free Days to Discharge Home is not directly comparable in meaning to the risk difference of −1.7 for new intubation. One is expressed as a mean difference in days; the other is expressed as an absolute risk difference.

Similarly, the primary estimate of −0.9 should not be compared numerically with the secondary risk difference of 0.5 as though they were interchangeable endpoints. The definitions, time frames, and analysis methods differ.

23. Important Limitations and Interpretation Issues

24. Why This Trial Matters Statistically

CLOVERS is a useful statistical teaching case because the primary endpoint combines several fundamental ideas in clinical-trial analysis: randomization, time-to-event follow-up, censoring, Kaplan-Meier estimation, an absolute treatment contrast, confidence intervals, and a two-sided Z-test.

ConceptHow it appears in CLOVERS
Randomization The trial used randomized allocation to two parallel groups.
Time-to-event endpoint The primary outcome was followed from randomization through discharge home up to and including day 90.
Kaplan-Meier estimation Primary point estimates were obtained from Kaplan-Meier curves.
Censoring 5 censored observations were reported in the restrictive-fluid group and 4 in the liberal-fluid group for the primary endpoint.
Absolute effect The primary effect was reported as a difference in mortality rate.
Risk difference Several secondary binary outcomes used risk differences.
Mean difference Several secondary continuous outcomes used mean differences.
Confidence interval The primary estimate had a two-sided 95% CI of −4.4 to 2.6.
Wald / Z-test The primary comparison used a reported Z-test.
P-value The primary two-sided P-value was 0.61.
Missing data The registry-reported analysis states that outcomes were reported using available data without imputation.
Multiple endpoints 14 outcome measures and 14 statistical analyses were posted.

25. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

The PubMed links above identify publications associated with the trial. The numerical trial results and methodological facts presented on this page are restricted to the ClinicalTrials.gov-derived trial data.

Continue through the Clinical Biostats statistical library

Use the trial's endpoints and methods as a starting point for deeper study of confidence intervals, survival analysis, missing data, risk differences, and hypothesis testing.

28. Record Summary

CLOVERS provides a clear example of how a randomized clinical trial can combine an absolute treatment contrast with time-to-event estimation. The primary endpoint was death before discharge home by day 90, with point estimates obtained from Kaplan-Meier curves and a reported two-sided Z-test. The estimated mortality-rate difference was −0.9, with a 95% CI of −4.4 to 2.6 and P = 0.61.

The most important statistical lesson is that these three quantities answer different questions. The estimate describes the observed direction and magnitude on the reported scale. The confidence interval describes uncertainty around that estimate. The P-value evaluates the observed test statistic relative to the null hypothesis under the specified testing framework. None of the three should be used as a substitute for the others.

The secondary outcomes reinforce the same principle. Some are reported as mean differences, others as risk differences, and Organ Support Free Days is described as a Kaplan-Meier-derived mean estimate. The ClinicalTrials.gov record also states that available data were used without imputation and that denominators may vary for selected outcomes. These details are not peripheral: they determine what statistical quantity is actually being estimated.

Finally, the registry's limitations should remain part of the statistical story. The ClinicalTrials.gov record does not specify methods for many secondary analyses, do not provide secondary P-values, do not report a multiplicity-adjustment strategy, and do not provide a non-inferiority margin, crossover analysis, factorial structure, Bayesian analysis, or proportional-hazards model for this record. A rigorous trial-results page should preserve those boundaries rather than filling them with assumptions.

Clinical Biostats methodology: The purpose of an independent statistical analysis is not simply to reproduce a trial record. It is to identify the estimand, understand the analysis method, distinguish effect size from statistical evidence, preserve endpoint definitions, and make uncertainty explicit without attributing methods or results that are not supported by the underlying record.