Tutorials › Biostatistics › Confidence Intervals Explained for Clinical Trials

Clinical Trial Statistics

Confidence Intervals Explained for Clinical Trials

A practical guide to confidence intervals in clinical research, including interpretation, standard errors, treatment effects, hazard ratios, risk differences, odds ratios, sample size, precision, and common reporting mistakes.

Intermediate16 min read

What You'll Learn

  • What a confidence interval represents in a clinical trial
  • How confidence intervals are calculated from estimates and standard errors
  • How CIs work for means, proportions, risk differences, and ratios
  • Why interval width matters for clinical interpretation
  • How CIs relate to hypothesis tests and p-values
  • How CIs are reported and QC'd in clinical-trial outputs

Introduction

A clinical trial rarely produces a treatment effect that is known with perfect certainty. Even when the underlying effect is fixed, the estimate observed in a particular trial depends on the patients enrolled, the measurements collected, and random sampling variation.

A confidence interval (CI) quantifies uncertainty around an estimated treatment effect or other population parameter. Instead of reporting only a point estimate, a confidence interval provides a range produced by a statistical procedure designed to have a stated long-run coverage probability under its assumptions.

Key idea: A point estimate tells you the estimated effect. A confidence interval tells you how precisely that effect has been estimated.

What Is a Confidence Interval?

Suppose a clinical trial estimates a treatment effect with point estimate \(\hat{\theta}\) and standard error \(SE(\hat{\theta})\). A commonly used two-sided confidence interval is:

\[ \hat{\theta}\pm c\times SE(\hat{\theta}) \]

where \(c\) is a critical value determined by the desired confidence level and reference distribution. For a large-sample 95% normal approximation, \(c\) is approximately 1.96.

\[ CI_{95\%}=\hat{\theta}\pm1.96\times SE(\hat{\theta}) \]

The interval can therefore be viewed as the point estimate plus or minus a margin of error:

\[ \text{Margin of Error}=c\times SE(\hat{\theta}) \]

What Does "95% Confidence" Mean?

The standard frequentist interpretation concerns the performance of the interval-generating procedure over repeated samples. If the same design and analysis procedure were repeated many times, approximately 95% of the resulting 95% confidence intervals would contain the true parameter, assuming the procedure's underlying assumptions hold.

Do not say: "There is a 95% probability that the true treatment effect is inside this particular interval." That is not the standard frequentist interpretation of a completed 95% CI.

For clinical communication, it is reasonable to describe a CI as a range of values compatible with the observed data and the statistical model, while remembering that the confidence level describes the long-run behavior of the procedure.

Point Estimate Versus Confidence Interval

Suppose a trial estimates a treatment difference of 6 points with a 95% CI from 1 to 11 points.

QuantityExampleWhat it tells you
Point estimate6 pointsBest single estimate of the treatment difference
Lower limit1 pointLower end of the interval
Upper limit11 pointsUpper end of the interval
CI width10 pointsReflects statistical precision

The point estimate is useful, but the interval adds information about how much uncertainty surrounds that estimate.

How a Confidence Interval Is Calculated

1
Calculate the point estimate of the treatment effect or parameter.
2
Estimate the standard error.
3
Select the appropriate critical value or reference distribution.
4
Calculate the margin of error.
5
Construct the lower and upper confidence limits.
6
Interpret the interval using the endpoint's effect scale, direction, and clinical context.

Worked Example: Mean Treatment Difference

Suppose the estimated treatment difference is:

\[ \hat{\theta}=8 \]

with:

\[ SE(\hat{\theta})=3 \]

Using a large-sample 95% normal approximation:

\[ 8\pm1.96(3) \]

The margin of error is:

\[ 1.96\times3=5.88 \]

The confidence limits are:

\[ 8-5.88=2.12 \qquad 8+5.88=13.88 \]

Therefore:

\[ CI_{95\%}=(2.12,\;13.88) \]
Clinical interpretation: The estimated treatment difference is 8 units, with a 95% CI from approximately 2.1 to 13.9 units. The interval communicates both the estimated magnitude and the uncertainty around it.

Confidence Intervals and Precision

A narrow confidence interval indicates greater statistical precision, whereas a wide confidence interval indicates less precision. For a fixed confidence level, interval width is driven strongly by the standard error.

\[ \text{CI Width}\propto SE(\hat{\theta}) \]

All else being equal, increasing sample size generally reduces standard error:

\[ SE(\hat{\theta})\propto\frac{1}{\sqrt{n}} \]

Consequently, larger trials often produce narrower intervals.

Same Estimate, Different Precision

Estimate95% CIInterpretation
5(4, 6)Very precise estimate
5(1, 9)Moderately precise estimate
5(−8, 18)Much greater uncertainty

Confidence Intervals for Treatment Differences

For a difference in mean outcomes, a common treatment contrast is:

\[ \hat{\Delta}=\bar{X}_T-\bar{X}_C \]

where \(\bar{X}_T\) and \(\bar{X}_C\) are the treatment and control means. The confidence interval is constructed around the estimated contrast using the appropriate standard error and inferential method.

Confidence Intervals for Risk Differences

For a binary endpoint, let \(p_T\) be the event probability in treatment and \(p_C\) the event probability in control. The risk difference is:

\[ RD=p_T-p_C \]

If the event rates are 60% and 45%:

\[ RD=0.60-0.45=0.15 \]

This corresponds to a 15-percentage-point difference.

Interpret the scale correctly: A risk difference of 0.15 is a 15-percentage-point difference, not a 15% relative increase.

Confidence Intervals for Risk Ratios

A risk ratio compares event probabilities:

\[ RR=\frac{p_T}{p_C} \]

If \(p_T=0.60\) and \(p_C=0.50\):

\[ RR=\frac{0.60}{0.50}=1.20 \]

For ratio measures, the null value is 1 rather than 0.

Effect measureNull valueTypical interpretation
Mean difference0No mean difference
Risk difference0No risk difference
Risk ratio1Equal risks
Odds ratio1Equal odds
Hazard ratio1Equal hazards

Confidence Intervals for Odds Ratios

Odds ratios are common in logistic regression and binary clinical endpoints. Suppose:

\[ OR=0.70 \]

with:

\[ CI_{95\%}=(0.52,\;0.94) \]

Because the entire interval is below 1, the estimated odds are lower in the numerator treatment group, assuming the ratio has been defined as treatment relative to control. Whether this is clinically favorable depends on what the event represents.

Confidence Intervals for Hazard Ratios

Time-to-event analyses commonly report hazard ratios from a Cox proportional hazards model. Suppose:

\[ HR=0.78 \]

with:

\[ CI_{95\%}=(0.64,\;0.95) \]

A hazard ratio below 1 suggests a lower instantaneous event hazard for treatment relative to control under the model and endpoint definition.

Important: A hazard ratio is not the same thing as a risk ratio. It summarizes a relative hazard under a time-to-event model and should not automatically be described as a percentage reduction in risk.

Ratio Confidence Intervals and the Log Scale

For ratio measures, statistical models often operate naturally on the logarithmic scale:

\[ \log(\hat{\theta})\pm c\times SE\{\log(\hat{\theta})\} \]

The limits are then transformed back to the ratio scale:

\[ CI=\left(\exp(L),\;\exp(U)\right) \]

This produces asymmetric intervals on the original ratio scale while preserving positivity.

Confidence Intervals and the Null Value

For many compatible two-sided procedures, whether a 95% CI includes the null value corresponds closely to the result of a two-sided hypothesis test at the 5% level.

\[ 0\in CI_{95\%} \]

is relevant for difference measures, while:

\[ 1\in CI_{95\%} \]

is relevant for ratio measures.

Important: This relationship assumes compatible models and inferential procedures. Multiplicity adjustments, one-sided testing, alternative interval constructions, or other design features can change the relationship.

Confidence Intervals Are More Informative Than "Significant / Not Significant"

A p-value helps address a hypothesis-testing question, but it does not by itself communicate effect magnitude or precision.

StudyEstimate95% CIp-value
A2.0(1.8, 2.2)<0.001
B8.0(−1.0, 17.0)0.08

Study A provides a highly precise estimate of a smaller effect. Study B has a larger point estimate but much more uncertainty. A simple significant/not-significant label would obscure that distinction.

Statistical Significance Versus Clinical Importance

A confidence interval can help separate statistical evidence from clinical relevance. Suppose a clinically important difference has been prespecified as 5 units.

\[ CI_{95\%}=(6,\;10) \]

If the treatment effect is 8 units and the entire interval exceeds 5, the data are consistent with an effect exceeding that threshold.

By contrast:

\[ CI_{95\%}=(2,\;14) \]

includes values below and above the clinically important threshold. The point estimate alone would not reveal this uncertainty.

Clinical interpretation should ask: What effects are compatible with the data, and are those effects clinically meaningful?

Confidence Intervals and Non-Inferiority

Confidence intervals are central to non-inferiority trials. Suppose the treatment effect is defined as treatment minus control and the non-inferiority margin is \(-5\) units. Consider:

\[ CI_{95\%}=(-2,\;7) \]

The lower confidence limit is above the non-inferiority margin. Under the prespecified estimand, analysis, and inferential framework, this can support a non-inferiority conclusion.

Margin matters: Non-inferiority cannot be determined from whether a CI crosses zero alone. The relevant comparison is against the prespecified non-inferiority margin and the correct effect direction.

Confidence Intervals and Equivalence

Equivalence testing also relies on prespecified margins. If the acceptable difference is between \(-\delta\) and \(+\delta\), the confidence interval must lie entirely within that interval for an equivalence conclusion under the applicable procedure:

\[ -\delta

This is different from simply failing to demonstrate a statistically significant difference.

Confidence Intervals for Subgroup Analyses

Clinical trials often display treatment effects by subgroup. Forest plots commonly show the point estimate and confidence interval for each subgroup.

SubgroupEstimate95% CI
Overall0.78(0.66, 0.92)
Age <650.74(0.59, 0.93)
Age ≥650.84(0.66, 1.08)

The subgroup intervals describe uncertainty within each subgroup. They do not, by themselves, establish that treatment effects differ between subgroups.

Common mistake: "Significant in one subgroup but not significant in another" does not establish a treatment-by-subgroup interaction. An appropriate interaction analysis addresses that question.

Confidence Intervals and Interaction Tests

If the scientific question is whether treatment effects differ between subgroups, the formal hypothesis concerns the interaction term:

\[ H_0:\theta_{\text{interaction}}=0 \]

Subgroup-specific CIs remain useful for describing individual estimates, but the interaction analysis addresses the comparison of effects.

Confidence Intervals in Regression Models

Confidence intervals are routinely produced for regression coefficients and model-derived effects. For a regression coefficient:

\[ \hat{\beta}\pm c\times SE(\hat{\beta}) \]

In logistic regression and Cox regression, exponentiating a coefficient gives an odds ratio or hazard ratio:

\[ OR=\exp(\hat{\beta}) \qquad HR=\exp(\hat{\beta}) \]

Confidence Intervals for Kaplan-Meier Estimates

Time-to-event analyses can also report confidence intervals around estimated survival probabilities. For example:

\[ \hat{S}(24)=0.72 \qquad CI_{95\%}=(0.64,\;0.79) \]

This communicates an estimated 72% survival probability at Week 24 together with uncertainty calculated using the specified survival-analysis method.

Confidence Intervals Are Not Prediction Intervals

IntervalMain question
Confidence intervalWhat values of the population parameter are supported by the analysis?
Prediction intervalWhat range might be expected for a future individual observation or outcome?

Prediction intervals generally incorporate additional individual-level variability and are therefore often wider.

Confidence Level: 90%, 95%, and 99%

Confidence levelTypical normal critical valueRelative width
90%1.645Narrower
95%1.96Common default
99%2.576Wider
\[ \text{Higher confidence level} \Rightarrow \text{larger critical value} \Rightarrow \text{wider CI} \]

Small Samples and the t Distribution

For some small-sample mean-based analyses, a t distribution is used rather than a standard normal critical value:

\[ CI=\bar{X}\pm t_{1-\alpha/2,\;df}\times SE(\bar{X}) \]

The degrees of freedom depend on the analysis. As sample size increases, the t distribution approaches the standard normal distribution.

Confidence Intervals Are Method-Dependent

There is no single universal CI method for every clinical-trial endpoint. The appropriate method depends on the estimand, endpoint, model, sample size, distributional assumptions, censoring, stratification, and analysis strategy.

Examples include Wald-type intervals, score-based or exact methods for proportions, transformed intervals for ratios, profile-likelihood intervals, and model-based intervals.

Best practice: Do not assume that two software procedures producing the same point estimate will necessarily produce identical confidence intervals. The interval method and variance estimator matter.

What Makes a Clinical-Trial CI Wide?

  • Small sample size
  • High variability in the endpoint
  • Few events in binary or time-to-event analyses
  • Small subgroup populations
  • Large residual variance
  • Limited information for the model parameter

A wide interval indicates imprecision. It is not automatically evidence that the treatment has no effect.

Precision Is Not the Same as Validity

A narrow interval can still surround a biased estimate. Confidence intervals quantify statistical uncertainty under the specified analysis; they do not automatically correct for bias, confounding, measurement error, or model misspecification.

Key distinction: Precision concerns how tightly the analysis estimates a parameter. Validity concerns whether the estimate appropriately represents the target quantity.

Confidence Intervals and Missing Data

The treatment-effect estimate and CI depend on how missing data are handled. Mixed models, multiple imputation, tipping-point analyses, and other sensitivity analyses can produce different intervals because they rely on different assumptions or estimands.

Confidence Intervals and Multiplicity

Clinical trials may involve multiple endpoints, treatment comparisons, subgroups, or interim analyses. When multiplicity adjustments are part of the inferential strategy, confidence intervals may also need to be adjusted so that the interval and testing procedures remain coherent.

Important: A nominal 95% CI should not automatically be interpreted as an adjusted 95% CI in a multiplicity-controlled analysis.

Confidence Intervals in Interim Analyses

Repeated looks at accumulating data can affect the operating characteristics of naive inferential procedures. In group-sequential or other adaptive designs, confidence intervals should follow the prespecified inferential method when adjusted inference is required.

Confidence Intervals in Clinical Study Reports

Confidence intervals are commonly included in efficacy and safety tables, forest plots, Kaplan-Meier displays, regression summaries, and key treatment-effect outputs.

ComponentExample
EndpointChange from baseline at Week 12
Effect measureTreatment minus control
Point estimate−4.2
95% CI(−7.1, −1.3)
p-value0.004

The direction of the treatment contrast should always be clear because the clinical meaning of a positive or negative estimate depends on how the endpoint and contrast are defined.

Confidence Intervals in Forest Plots

A forest plot displays treatment-effect estimates and their confidence intervals across subgroups, endpoints, or other analysis strata.

1
Point estimate: location of the treatment-effect marker.
2
Confidence interval: horizontal uncertainty range.
3
Null line: 0 for differences or 1 for ratios.
4
Clinical interpretation: consider precision and clinical thresholds, not only whether the interval crosses the null.

Confidence Intervals in SAS

Many SAS procedures provide confidence limits for model parameters and effect measures. For example:

proc logistic data=analysis;
    class treatment(ref="Control") / param=ref;
    model response(event="1") = treatment age baseline;
    oddsratio treatment / cl=wald;
run;

The CI method should match the statistical analysis plan and the properties of the endpoint.

Confidence Intervals in R

For a simple large-sample calculation:

estimate <- 8
se <- 3
z <- qnorm(0.975)

lower <- estimate - z * se
upper <- estimate + z * se

c(lower = lower, upper = upper)

For fitted models, the appropriate model-specific function can return confidence intervals using the selected method.

A Practical Clinical-Trial Interpretation Workflow

1
Identify the endpoint and estimand.
2
Identify the effect measure and direction of comparison.
3
Read the point estimate.
4
Read the lower and upper confidence limits.
5
Identify the null value: usually 0 for differences and 1 for ratios.
6
Assess interval width and statistical precision.
7
Compare the interval with clinically important thresholds or decision margins.
8
Check for multiplicity, interim-analysis, or other prespecified adjustments.

Common Mistakes

  1. Interpreting a 95% CI as a 95% probability statement about a fixed parameter. The standard frequentist interpretation concerns repeated use of the interval procedure.
  2. Ignoring the effect scale. Difference measures generally use 0 as the null; ratio measures generally use 1.
  3. Assuming a CI crossing the null proves there is no effect. It may indicate insufficient precision to distinguish the effect from the null.
  4. Equating statistical significance with clinical importance. A small effect can be estimated very precisely.
  5. Assuming a narrow CI guarantees an unbiased estimate. Precision does not eliminate systematic bias.
  6. Comparing subgroup significance instead of testing interaction. A formal interaction analysis is needed to compare subgroup effects.
  7. Using an unadjusted CI after a multiplicity-adjusted analysis. Inferential intervals should match the analysis strategy.
  8. Reporting a ratio without identifying the comparison direction. The numerator and denominator determine interpretation.
  9. Assuming every software procedure uses the same CI construction. Variance estimators and interval methods can differ.

Confidence Interval Quality Control

Confidence intervals should be checked as part of statistical-programming QC.

  • Confirm the analysis population.
  • Confirm treatment-group definitions and contrasts.
  • Verify the endpoint and estimand.
  • Verify the point estimate and standard error.
  • Verify the confidence level.
  • Verify the CI construction method.
  • Confirm lower and upper limits are correctly ordered.
  • Check the null value and treatment-effect direction.
  • Reproduce selected intervals independently.
  • Confirm multiplicity or interim-analysis adjustments.

Manual Validation Example

Suppose a treatment difference is 8 with standard error 3 and a 95% normal critical value of 1.96.

ComponentValue
Estimate8
Standard error3
Critical value1.96
Margin of error5.88
Lower limit2.12
Upper limit13.88
\[ 8-1.96(3)=2.12 \]
\[ 8+1.96(3)=13.88 \]

This independent check can catch errors in effect direction, standard-error selection, critical values, or limit formatting.

How to Report a Confidence Interval

Example: The estimated difference between treatment and control in mean change from baseline at Week 12 was −4.2 units (95% CI: −7.1 to −1.3). The negative value indicates a lower mean change in the treatment group when the treatment effect is defined as treatment minus control.
Ratio example: The estimated hazard ratio for progression or death was 0.78 (95% CI: 0.64 to 0.95), with treatment compared with control as specified in the analysis.

Confidence Intervals in Decision Making

A useful interpretation considers the entire interval rather than focusing only on its midpoint.

  • What is the estimated treatment effect?
  • How precise is the estimate?
  • Does the interval include the null value?
  • Does it include clinically trivial effects?
  • Does it include clinically important effects?
  • Is it consistent with the prespecified decision margin?
  • Were multiplicity and other design features handled appropriately?

Key Takeaways

1
CIs quantify uncertainty around an estimated parameter under the specified statistical procedure.
2
Interpret the point estimate and CI together. The estimate gives magnitude; the interval communicates precision.
3
Know the null value. Differences generally use 0; ratios generally use 1.
4
Wide intervals indicate imprecision, not necessarily no effect.
5
Statistical significance is not clinical importance. Compare the interval with meaningful clinical thresholds.
6
The CI method matters. Use the method prespecified for the endpoint, estimand, model, and trial design.
Clinical Trials

See these methods in real clinical trials

See the method applied to published trial results, with the estimates, confidence intervals and interpretation explained.

REVEAL
Independent statistical analysis of the phase 3 REVEAL trial of anacetrapib versus placebo in atherosclerotic cardiovascular disease, including trial design, time-to-event endpoints,…
Phase 3 · n = 30,449
FOURIER
Independent statistical analysis of the FOURIER phase 3 trial of evolocumab versus placebo in subjects with elevated cardiovascular risk and dyslipidemia, including…
Phase 3 · n = 27,564
TAILORx
Independent statistical analysis of the TAILORx phase 3 trial, focusing on 5-year disease-free survival, the Cox proportional-hazards model, hazard ratios, confidence intervals,…
Phase 3 · n = 10,273
TEAM
Independent statistical analysis of TEAM, the randomized phase 3 trial comparing exemestane with tamoxifen followed by exemestane in postmenopausal patients with receptor-positive…
Phase 3 · n = 9,779
COMPASS
Independent statistical analysis of the COMPASS phase 3 trial of rivaroxaban-based antithrombotic treatment in coronary or peripheral artery disease, including trial design,…
Phase 3 · n = 27,395
TRA 2P-TIMI 50
Independent statistical analysis of TRA 2P-TIMI 50 (NCT00526474), including randomized trial design, time-to-event endpoints, Cox proportional-hazards methods, efficacy results, bleeding outcomes, post-hoc…
Phase 3 · n = 26,449
See all 407 trials using confidence intervals →