Tutorials › Biostatistics › Equivalence Trial Design Principles

Clinical Trial Design & Hypothesis Testing

Equivalence Trial Design Principles

A practical guide to designing equivalence trials, including equivalence margins, the two one-sided tests procedure, confidence intervals, sample size, power, Type I error, analysis populations, and a complete worked example for a continuous endpoint.

Advanced 18 min read

What You'll Learn

  • Why equivalence trials use a fundamentally different hypothesis-testing framework
  • How the equivalence margin defines the clinically acceptable difference
  • How the Two One-Sided Tests (TOST) procedure establishes equivalence
  • How confidence intervals provide an intuitive equivalence decision
  • How sample size, variability, margin, power, and Type I error interact
  • How equivalence differs from superiority and non-inferiority trials

Introduction

Many clinical trials are designed to determine whether a new treatment is better than a control treatment.

Equivalence trials ask a different question: Are the two treatments sufficiently similar that any difference is clinically unimportant?

This distinction is fundamental. A failure to demonstrate superiority does not demonstrate equivalence. Likewise, observing a small difference between treatments does not by itself establish equivalence.

An equivalence trial must be designed specifically to demonstrate that the difference between treatments lies within a prespecified range of clinically acceptable values.

Key idea: Equivalence is demonstrated when the entire confidence interval for the treatment difference falls inside the prespecified equivalence margins. The margins therefore play a central role in both the design and the interpretation of the study.

What Is an Equivalence Trial?

An equivalence trial is designed to determine whether the difference between two treatments is sufficiently small to be considered clinically acceptable.

Let the treatment effect be defined as:

$$ \Delta=\mu_T-\mu_R $$

where:

  • \(\mu_T\) is the mean outcome under the test treatment.
  • \(\mu_R\) is the mean outcome under the reference treatment.
  • \(\Delta\) is the treatment difference.

Suppose differences between \(-\Delta_E\) and \(+\Delta_E\) are considered clinically acceptable. Then equivalence is defined by:

$$ -\Delta_E < \Delta < +\Delta_E $$

The quantity \(\Delta_E\) is the equivalence margin.

Equivalence Is Not "No Difference"

This is one of the most important concepts in equivalence trial design.

A conventional statistical test of:

$$ H_0:\Delta=0 $$

asks whether there is evidence of a difference from exactly zero.

That is not the same question as whether the treatments are sufficiently similar.

A study with a very large sample size could find a statistically significant difference of only 0.2 units even when that difference has no meaningful clinical consequence.

Conversely, a small study might fail to detect a statistically significant difference of 5 units even though such a difference would be clinically important.

Equivalence asks a clinical question, not simply a significance question: Is the treatment difference small enough that the remaining uncertainty is entirely contained within a clinically acceptable region?

Superiority, Non-Inferiority, and Equivalence

These three designs have different objectives.

Design Primary Question Typical Alternative
Superiority Is one treatment better? Difference in one favorable direction
Non-inferiority Is the test treatment not unacceptably worse? Difference is above a one-sided margin
Equivalence Are the treatments sufficiently similar? Difference lies within two-sided margins

The statistical structure follows directly from these questions.

The Equivalence Region

Suppose the clinically acceptable difference is ±5 units. Then:

$$ -\Delta_E=-5 \qquad\text{and}\qquad +\Delta_E=+5 $$

The equivalence region is therefore:

$$ -5<\Delta<5 $$

A treatment difference of:

  • −3 units is within the equivalence region.
  • 0 units is within the equivalence region.
  • +4 units is within the equivalence region.
  • +7 units is outside the equivalence region.
  • −6 units is outside the equivalence region.

The boundaries are determined by clinical judgment rather than by the observed data.

Choosing the Equivalence Margin

The equivalence margin is arguably the most important design parameter in an equivalence trial.

It should represent the largest difference between treatments that would still be considered clinically acceptable.

A margin should not be selected simply because it produces a convenient sample size.

For an efficacy endpoint, investigators may consider:

  • The established effect of the reference treatment
  • The variability of the endpoint
  • Historical randomized trial evidence
  • Clinical relevance of different effect sizes
  • Patient perspective
  • Regulatory considerations
  • The intended use of the new treatment
Margin selection comes before sample-size optimization. If the equivalence margin is chosen primarily to make the trial smaller, the statistical design may no longer provide a clinically meaningful demonstration of equivalence.

Symmetric vs. Asymmetric Margins

Not every equivalence margin must be symmetric.

A symmetric margin has:

$$ -\Delta_L=-\Delta \qquad\text{and}\qquad \Delta_U=+\Delta $$

But some clinical settings may justify different lower and upper limits:

$$ \Delta_L<\Delta<\Delta_U $$

For example, suppose a difference between −2 and +4 units is considered clinically acceptable. Then:

$$ -2<\Delta<4 $$

The statistical framework can accommodate asymmetric equivalence margins.

The Equivalence Hypotheses

This is where equivalence trials differ most clearly from conventional hypothesis tests.

Suppose the equivalence limits are:

$$ -\Delta_E \qquad\text{and}\qquad +\Delta_E $$

The null hypothesis is:

$$ H_0: \Delta\le-\Delta_E \quad\text{or}\quad \Delta\ge+\Delta_E $$

The alternative hypothesis is:

$$ H_A: -\Delta_E<\Delta<+\Delta_E $$

In words: the null says the treatments are not equivalent, while the alternative says they are equivalent.

This reverses the usual intuition. In a superiority trial, the null usually represents no meaningful treatment difference. In an equivalence trial, the null represents differences that are too large to be considered clinically acceptable.

The Two One-Sided Tests Procedure

The standard method for demonstrating equivalence is the Two One-Sided Tests (TOST) procedure.

Instead of testing one two-sided null hypothesis directly, equivalence is demonstrated by conducting two one-sided tests.

The first test evaluates the lower equivalence boundary:

$$ H_{01}:\Delta\le-\Delta_E $$ versus
$$ H_{A1}:\Delta>-\Delta_E $$

The second test evaluates the upper equivalence boundary:

$$ H_{02}:\Delta\ge+\Delta_E $$ versus
$$ H_{A2}:\Delta<+\Delta_E $$

Both null hypotheses must be rejected for equivalence to be demonstrated.

Why Two Tests?

The two tests address two different ways the treatments could fail to be equivalent.

1
Lower boundary: rule out a treatment difference that is too unfavorable in the negative direction.
2
Upper boundary: rule out a treatment difference that is too favorable in the positive direction.
3
Both pass: the entire clinically relevant region has been established as plausible.

The Confidence Interval Interpretation

The TOST procedure has an equivalent and often more intuitive confidence interval interpretation.

For a two-sided equivalence test conducted at significance level \(\alpha\), calculate a:

$$ 100(1-2\alpha)\% $$

confidence interval for the treatment difference.

If:

$$ CI_{1-2\alpha} \subset (-\Delta_E,+\Delta_E) $$

then equivalence is demonstrated.

For the common choice:

$$ \alpha=0.05 $$

the corresponding confidence interval is:

$$ 90\%\ CI $$

Thus, equivalence at the 5% significance level is commonly expressed as the entire 90% confidence interval falling inside the equivalence margins.

Practical rule: For a two-sided equivalence test with α = 0.05, calculate the 90% confidence interval for the treatment difference. Equivalence is demonstrated only if both the lower and upper confidence limits lie inside the prespecified equivalence margins.

A Visual Way to Think About Equivalence

Confidence Interval Conclusion
Entirely inside the margins Equivalence demonstrated
Crosses either margin Equivalence not demonstrated
Entirely outside one margin Clearly not equivalent under the specified margins
Very wide interval crossing both margins Data are too imprecise to establish equivalence

This makes clear why simply observing a small point estimate is insufficient.

A point estimate of zero with a confidence interval from −10 to +10 would not establish equivalence if the margin were ±5.

Worked Example: Continuous Endpoint

Suppose a randomized, parallel-group clinical trial compares a test treatment with a reference treatment using a continuous efficacy endpoint.

Assume that larger values represent better outcomes and that the clinically acceptable difference has been established as ±5 units.

Design Parameter Value
Endpoint Continuous
Equivalence margin ±5 units
Type I error 5%
Target power 90%
Expected true difference 0 units
Common SD 12 units

The treatment effect is defined as:

$$ \Delta=\mu_T-\mu_R $$

The equivalence objective is:

$$ -5<\Delta<5 $$

Step 1: Define the Hypotheses

The first one-sided test is:

$$ H_{01}:\Delta\le-5 \qquad \text{vs.} \qquad H_{A1}:\Delta>-5 $$

The second is:

$$ H_{02}:\Delta\ge5 \qquad \text{vs.} \qquad H_{A2}:\Delta<5 $$

Both tests must reject their respective null hypotheses.

Step 2: Suppose the Trial Produces This Result

Assume the estimated treatment difference is:

$$ \hat{\Delta}=1.2 $$

and the standard error is:

$$ SE(\hat{\Delta})=2.5 $$

The 90% confidence interval uses:

$$ \hat{\Delta} \pm z_{0.95}SE $$

where:

$$ z_{0.95}\approx1.645 $$

Therefore:

$$ 1.2 \pm (1.645)(2.5) $$

The margin of error is:

$$ 1.645(2.5)=4.1125 $$

Therefore:

$$ CI_{90\%} = 1.2\pm4.1125 $$
$$ CI_{90\%} = (-2.9125,\;5.3125) $$

The upper confidence limit exceeds +5. Therefore, the entire confidence interval does not fall inside the equivalence region.

The trial therefore does not demonstrate equivalence.

Important: The point estimate of +1.2 units is comfortably inside the equivalence margin. Nevertheless, equivalence is not demonstrated because the uncertainty around the estimate extends beyond the upper equivalence limit.

A Second Example: A Narrower Confidence Interval

Suppose the estimated difference remains:

$$ \hat{\Delta}=1.2 $$

but the standard error is only:

$$ SE=1.5 $$

Then:

$$ 1.2 \pm 1.645(1.5) $$

giving:

$$ 1.2\pm2.4675 $$
$$ CI_{90\%} = (-1.2675,\;3.6675) $$

The entire interval lies between −5 and +5.

Therefore, equivalence is demonstrated.

What the Example Demonstrates

The two examples have exactly the same point estimate.

Scenario Estimate 90% CI Equivalence?
Less precise 1.2 (−2.91, 5.31) No
More precise 1.2 (−1.27, 3.67) Yes

The difference is precision.

This is why sample size is especially important in equivalence trials: the study must be sufficiently precise to exclude clinically meaningful differences.

Sample Size for Two-Group Equivalence Trials

For a simple two-group parallel trial with a continuous endpoint and equal allocation, a useful planning approximation is:

$$ n_{\text{per group}} \approx \frac{2\sigma^2 (z_{1-\alpha}+z_{1-\beta})^2} {\Delta_E^2} $$

where:

  • \(\sigma\) = assumed common standard deviation
  • \(\Delta_E\) = equivalence margin
  • \(\alpha\) = one-sided TOST significance level
  • \(\beta\) = Type II error
  • \(n\) = sample size per group

The factor of 2 arises because the variance of the difference between two independent equal-sized group means is:

$$ \operatorname{Var}(\bar X_T-\bar X_R) = \frac{\sigma^2}{n} + \frac{\sigma^2}{n} = \frac{2\sigma^2}{n} $$

Worked Sample Size Calculation

Return to the example:

  • Equivalence margin = 5
  • Common SD = 12
  • One-sided α = 0.05
  • Power = 90%

The relevant normal quantiles are approximately:

$$ z_{1-\alpha}=1.645 $$

and:

$$ z_{1-\beta}=1.282 $$

Therefore:

$$ n \approx \frac{ 2(12^2)(1.645+1.282)^2 }{ 5^2 } $$

First calculate:

$$ 1.645+1.282=2.927 $$

Then:

$$ 2(144)(2.927)^2 \approx2466 $$

and:

$$ \frac{2466}{25} \approx98.6 $$

Thus the approximate required sample size is:

$$ \boxed{99\text{ patients per group}} $$

In practice, the final sample size would generally be rounded upward and adjusted for the chosen analysis method, dropout, allocation ratio, and potential use of a t-distribution rather than a normal approximation.

Why the Equivalence Margin Has Such a Large Effect on Sample Size

The sample size is approximately inversely proportional to the square of the equivalence margin:

$$ n\propto\frac{1}{\Delta_E^2} $$

Therefore, making the margin half as large can require approximately four times the sample size, all else equal.

Margin Relative Sample Size
\(\Delta_E\)
\(0.75\Delta_E\) Approximately 1.78×
\(0.50\Delta_E\) Approximately 4×
\(0.25\Delta_E\) Approximately 16×

This mathematical relationship illustrates why margin selection and trial feasibility must be considered together.

Effect of Variability on Sample Size

Sample size also increases with the square of the assumed standard deviation:

$$ n\propto\sigma^2 $$

If the true variability is greater than expected, the study may have less precision than anticipated.

For example, increasing the SD from 10 to 12 increases the variance by:

$$ \frac{12^2}{10^2} = 1.44 $$

or 44%.

This can have a substantial impact on the required sample size.

Effect of Power on Sample Size

Increasing power increases sample size.

Power \(z_{1-\beta}\) Relative Planning Burden
80% 0.842 Lower
90% 1.282 Higher
95% 1.645 Higher still

The appropriate power depends on the clinical and regulatory context.

Equivalence and Type I Error

In an equivalence study, Type I error concerns incorrectly concluding that the treatments are equivalent when the true difference lies outside the acceptable margins.

The TOST framework uses one-sided tests.

For an overall equivalence conclusion at the conventional 5% level, each component test is typically performed at:

$$ \alpha=0.05 $$

The confidence interval equivalent is therefore a 90% two-sided interval.

Do not automatically divide α by 2. The TOST procedure consists of two one-sided tests, each conducted at the specified one-sided significance level. The familiar 90% confidence interval arises from this structure.

Equivalence vs. Failure to Demonstrate Superiority

Consider a superiority study comparing two treatments.

Suppose the study finds:

$$ p=0.30 $$

for a test of whether the treatment difference is statistically significant, and the null hypothesis is not rejected.

That result does not demonstrate equivalence.

The study may simply be too small to rule out clinically meaningful differences.

For example, if the estimated difference is 1 unit but the confidence interval is:

$$ (-8,\;10) $$

a difference of ±5 or more remains entirely plausible.

The correct conclusion is: equivalence was not established.

It is not: the treatments are equivalent.

Equivalence vs. Non-Inferiority

Non-inferiority typically focuses on one direction.

Suppose negative values indicate worse efficacy for the test treatment. A non-inferiority margin might be:

$$ -\Delta_{NI} $$

The question is whether the treatment effect is sufficiently above that boundary.

Equivalence requires both:

$$ \Delta>-\Delta_E $$

and:

$$ \Delta<+\Delta_E $$

Therefore, equivalence is a two-sided requirement while non-inferiority is normally one-sided.

Characteristic Non-Inferiority Equivalence
Number of margins One Two
Direction of concern Usually one direction Both directions
Goal Rule out unacceptable loss Rule out unacceptable differences in either direction
Typical CI criterion CI excludes the non-inferiority margin Entire CI lies within both equivalence margins

Parallel-Group Equivalence Designs

The simplest equivalence design is a randomized parallel-group trial.

1
Randomize participants to test or reference treatment.
2
Collect the prespecified primary endpoint.
3
Estimate the treatment difference.
4
Calculate the appropriate confidence interval.
5
Compare the entire interval with the equivalence margins.

Crossover Equivalence Trials

Equivalence studies are also common in crossover designs, particularly when within-subject comparisons are scientifically appropriate.

For example, each participant might receive both:

  • The test treatment
  • The reference treatment

The treatment difference is then estimated using within-subject information.

Because the same participant contributes information under both treatments, within-subject variability can be substantially smaller than between-subject variability.

This can make crossover designs efficient when the endpoint and treatment effects are sufficiently stable over time.

Crossover designs require additional assumptions. Period effects, sequence effects, carryover, washout adequacy, and treatment stability must be considered when determining whether a crossover equivalence design is appropriate.

Equivalence for Binary Endpoints

The same conceptual framework applies to binary endpoints, although the treatment effect and statistical model differ.

Suppose:

$$ p_T=\text{response probability under test} $$

and:

$$ p_R=\text{response probability under reference} $$

An absolute-risk-difference equivalence criterion might be:

$$ -\Delta_E < p_T-p_R < +\Delta_E $$

Other effect measures may be appropriate depending on the clinical question, including risk ratios or odds ratios.

Ratio-Based Equivalence

Some equivalence problems are more naturally expressed as ratios rather than differences.

Suppose the relevant parameter is:

$$ \theta=\frac{\mu_T}{\mu_R} $$

An equivalence interval might be:

$$ L<\theta

For example, an equivalence criterion could be expressed as:

$$ 0.80<\frac{\mu_T}{\mu_R}<1.25 $$

This type of ratio-based framework is particularly familiar in bioequivalence studies.

For multiplicative effects, logarithmic transformation is often useful:

$$ \log(\theta) = \log(\mu_T)-\log(\mu_R) $$

which converts a ratio problem into a difference on the log scale.

Bioequivalence vs. Clinical Equivalence

Bioequivalence and clinical equivalence are related but distinct concepts.

Bioequivalence typically concerns pharmacokinetic measures such as:

  • AUC
  • Cmax
  • Other prespecified exposure parameters

Clinical equivalence concerns whether treatments produce sufficiently similar clinical outcomes according to a prespecified margin.

The statistical logic of confidence intervals and equivalence margins is related, but the estimands, endpoints, transformations, and regulatory requirements can differ substantially.

Estimand Considerations

An equivalence trial should define the treatment effect being estimated before the analysis begins.

Important estimand attributes include:

  • Population
  • Treatment conditions
  • Endpoint
  • Summary measure
  • Handling of intercurrent events

This matters because different analysis strategies can answer different questions about equivalence.

Analysis Populations in Equivalence Trials

Equivalence trials require particularly careful consideration of analysis populations.

Common analysis sets include:

  • Intention-to-treat population
  • Per-protocol population
  • Safety population

In superiority trials, deviations and treatment nonadherence can sometimes make treatment groups appear more similar and therefore bias against finding a difference.

In equivalence trials, that property can be problematic because reduced separation between groups can make equivalence easier to conclude.

Key principle: In equivalence and non-inferiority studies, protocol deviations, treatment switching, missing data, and nonadherence require especially careful consideration because they can bias results toward similarity.

Why Both ITT and Per-Protocol Analyses Can Matter

Because different analysis populations can have different biases in equivalence settings, investigators may prespecify complementary analyses.

For example:

Analysis Purpose
Intent-to-treat Reflects outcomes according to randomized treatment assignment
Per-protocol Assesses treatment differences among participants sufficiently adherent to the protocol
Safety Characterizes treatment-emergent safety outcomes

The exact strategy should be prespecified and justified for the trial rather than selected after seeing which analysis supports equivalence.

Missing Data

Missing observations can have a particularly important effect in equivalence trials.

A missing-data strategy should therefore be established before database lock.

Potential approaches depend on the endpoint and estimand and may include:

  • Model-based methods
  • Multiple imputation
  • Mixed-effects models
  • Prespecified sensitivity analyses
  • Pattern-mixture or tipping-point analyses

The appropriate approach depends on the mechanism and amount of missingness and the scientific estimand.

Assay Sensitivity

An important concern in equivalence and non-inferiority studies is assay sensitivity.

Assay sensitivity refers to the ability of the trial to distinguish effective treatments from ineffective treatments under the conditions of the study.

If the trial is poorly designed or the endpoint is insufficiently sensitive, two ineffective treatments could appear similar.

Therefore, demonstrating equivalence requires more than obtaining a narrow confidence interval. The study must also be scientifically capable of detecting meaningful differences if they exist.

Constancy and Historical Evidence

Historical evidence may be important when selecting equivalence margins and interpreting the reference treatment.

Investigators may need to establish that the reference treatment's historical effect remains reasonably applicable to the current trial setting.

Changes in:

  • Patient population
  • Background therapy
  • Endpoint definitions
  • Standard of care
  • Comparator implementation
  • Trial conduct

can affect the interpretation of historical treatment effects.

Equivalence Does Not Mean Identical

Even when equivalence is demonstrated, the estimated treatment difference is rarely exactly zero.

Suppose:

$$ \hat{\Delta}=1.1 $$

with:

$$ CI_{90\%}=(-2.0,\;4.0) $$

and equivalence margins of ±5.

The treatment difference is not zero.

Instead, the evidence indicates that the plausible range of treatment differences is entirely within the clinically acceptable interval.

Equivalence means sufficiently similar, not mathematically identical. The equivalence margin defines how much difference can be tolerated while still considering the treatments clinically interchangeable for the specified purpose.

Precision Is Central to Equivalence

Consider two studies with the same observed treatment difference.

Study Estimated Difference 90% CI Margin Conclusion
A 0.5 (−1.5, 2.5) ±5 Equivalent
B 0.5 (−7.0, 8.0) ±5 Not demonstrated

The point estimates are identical. The conclusions differ because the second study is much less precise.

This is one of the defining features of equivalence testing.

Sample Size and Precision

For a simple continuous endpoint:

$$ SE(\hat{\Delta}) \propto \frac{1}{\sqrt{n}} $$

As sample size increases, the confidence interval becomes narrower.

The purpose of increasing sample size in an equivalence trial is therefore not merely to increase the chance of detecting a difference. It is to obtain enough precision to rule out differences larger than the equivalence margin.

Power in an Equivalence Trial

Power is the probability of correctly demonstrating equivalence for a specified true treatment difference.

It is common to assume during planning that:

$$ \Delta_{\text{true}}=0 $$

but power can also be evaluated for values away from zero.

This distinction is useful because the probability of demonstrating equivalence generally decreases as the true treatment difference approaches an equivalence boundary.

Power Depends on the True Difference

Suppose the equivalence interval is:

$$ -5<\Delta<5 $$

A study designed for 90% power when:

$$ \Delta=0 $$

will not necessarily have 90% power when:

$$ \Delta=4 $$

because the true effect is much closer to the upper equivalence boundary.

True Difference Distance from Upper Margin Expected Ability to Demonstrate Equivalence
0 5 units Highest
1 4 units Lower
3 2 units Lower still
4.5 0.5 units Very low
5 0 units Not expected to demonstrate equivalence

The actual power curve should therefore be examined when the design is important or the expected treatment difference is uncertain.

Equivalence Margin vs. Expected Treatment Difference

These two quantities should not be confused.

The equivalence margin is a clinical decision threshold. The expected treatment difference is a planning assumption.

For example:

$$ \Delta_E=5 $$

could be the clinically acceptable difference while the design assumes:

$$ \Delta_{\text{true}}=0 $$

for sample-size calculations.

The study is not trying to prove that the true difference equals zero. It is trying to demonstrate that the difference is sufficiently close to zero to remain within the acceptable region.

Equivalence and Multiple Endpoints

Some trials have multiple efficacy endpoints or multiple components that must all satisfy equivalence requirements.

This introduces an additional multiplicity problem.

For example, a study might require equivalence for:

  • Primary efficacy endpoint
  • Key secondary endpoint
  • Another clinically important outcome

The protocol should clearly identify which endpoint establishes the primary equivalence conclusion and how multiplicity is handled.

Do not automatically treat multiple successful equivalence tests as independent evidence. The overall decision rule should be specified before the trial begins.

Equivalence and Safety

Equivalence in efficacy does not imply equivalence in safety.

A treatment may demonstrate equivalent efficacy while having a different adverse-event profile.

Safety should therefore be analyzed separately using appropriate estimands, summaries, and inferential procedures.

Likewise, an equivalence conclusion for one clinical endpoint does not imply equivalence for every other outcome.

Protocol Considerations

An equivalence protocol should clearly define:

  • Primary endpoint
  • Treatment effect measure
  • Equivalence margins
  • Scientific justification for the margins
  • Type I error
  • Target power
  • Sample-size assumptions
  • Allocation ratio
  • Primary analysis method
  • Analysis populations
  • Missing-data strategy
  • Handling of protocol deviations
  • Intercurrent events
  • Sensitivity analyses
  • Criteria for declaring equivalence

A Complete Equivalence Decision Algorithm

1
Define the clinical endpoint and treatment effect measure.
2
Define the clinically acceptable lower and upper equivalence margins.
3
Specify the Type I error and desired power.
4
Estimate the expected treatment difference and variability.
5
Calculate the required sample size.
6
Randomize and collect the prespecified endpoint.
7
Estimate the treatment difference using the prespecified analysis.
8
Calculate the appropriate confidence interval.
9
Confirm that the entire confidence interval lies inside the equivalence margins.
10
Perform prespecified sensitivity and supportive analyses.

R Implementation: TOST for a Continuous Endpoint

For a simple two-group analysis, the TOST procedure can be implemented directly from the estimated difference and its standard error.

estimate <- 1.2
se <- 1.5

lower_margin <- -5
upper_margin <- 5

alpha <- 0.05

z <- qnorm(1 - alpha)

lower_ci <- estimate - z * se
upper_ci <- estimate + z * se

lower_ci
upper_ci

The resulting 90% confidence interval is approximately:

lower_ci
# approximately -1.27

upper_ci
# approximately 3.67

The equivalence criterion can then be evaluated directly:

equivalent <-
  lower_ci > lower_margin &&
  upper_ci < upper_margin

equivalent

The result is:

TRUE

because the complete interval lies inside:

$$ (-5,+5) $$

Explicit TOST Calculations in R

The two one-sided tests can also be calculated explicitly.

estimate <- 1.2
se <- 1.5

lower_margin <- -5
upper_margin <- 5

alpha <- 0.05

# Test H01: Delta <= lower_margin
t_lower <-
  (estimate - lower_margin) / se

p_lower <-
  1 - pnorm(t_lower)

# Test H02: Delta >= upper_margin
t_upper <-
  (estimate - upper_margin) / se

p_upper <-
  pnorm(t_upper)

p_lower
p_upper

Equivalence is demonstrated when both one-sided p-values are below the prespecified significance level:

equivalent <-
  p_lower < alpha &&
  p_upper < alpha

equivalent

This gives the same inferential conclusion as the confidence-interval approach.

A Compact TOST Function

tost_equivalence <- function(
  estimate,
  se,
  lower_margin,
  upper_margin,
  alpha = 0.05
) {

  z <- qnorm(1 - alpha)

  lower_ci <-
    estimate - z * se

  upper_ci <-
    estimate + z * se

  equivalent <-
    lower_ci > lower_margin &&
    upper_ci < upper_margin

  list(
    estimate = estimate,
    lower_ci = lower_ci,
    upper_ci = upper_ci,
    equivalent = equivalent
  )
}

tost_equivalence(
  estimate = 1.2,
  se = 1.5,
  lower_margin = -5,
  upper_margin = 5
)

Power Calculation by Simulation

For more complex equivalence designs, simulation is often useful.

A basic simulation repeatedly generates treatment-group data, estimates the treatment difference, constructs the confidence interval, and records whether equivalence was demonstrated.

set.seed(123)

nsim <- 10000

n_T <- 99
n_R <- 99

mu_T <- 0
mu_R <- 0

sd <- 12

lower_margin <- -5
upper_margin <- 5

success <- logical(nsim)

for (i in seq_len(nsim)) {

  y_T <- rnorm(
    n_T,
    mean = mu_T,
    sd = sd
  )

  y_R <- rnorm(
    n_R,
    mean = mu_R,
    sd = sd
  )

  estimate <-
    mean(y_T) - mean(y_R)

  se <-
    sqrt(
      sd^2 / n_T +
      sd^2 / n_R
    )

  lower_ci <-
    estimate - qnorm(0.95) * se

  upper_ci <-
    estimate + qnorm(0.95) * se

  success[i] <-
    lower_ci > lower_margin &&
    upper_ci < upper_margin
}

mean(success)

The resulting proportion is an empirical estimate of the probability of demonstrating equivalence under the assumed true difference and variability.

Why Simulation Can Be Valuable

Closed-form sample-size equations are useful, but real equivalence studies can be substantially more complicated.

Simulation can incorporate:

  • Unequal allocation
  • Dropout
  • Non-normal endpoints
  • Repeated measurements
  • Mixed-effects models
  • Binary outcomes
  • Time-to-event endpoints
  • Covariate adjustment
  • Complex missing-data mechanisms

The simulation should reproduce the planned analysis as closely as possible.

Common Equivalence Trial Design Mistakes

  1. Defining equivalence as a nonsignificant superiority test. Failure to reject a difference from zero does not establish equivalence.
  2. Choosing the equivalence margin after seeing the data. The margin must be justified and prespecified.
  3. Using a margin that is too wide. A statistically successful trial may then demonstrate a level of similarity that is not clinically meaningful.
  4. Using a margin that is too narrow without considering feasibility. The resulting sample size may become impractically large.
  5. Looking only at the point estimate. Equivalence depends on the entire confidence interval.
  6. Using the wrong confidence interval. For a 5% TOST analysis, the corresponding two-sided confidence interval is 90%, not 95%.
  7. Ignoring variability uncertainty. Underestimating the SD can result in an underpowered study.
  8. Ignoring missing data. Loss of precision and departures from the intended estimand can materially affect the equivalence conclusion.
  9. Ignoring protocol deviations. In equivalence studies, deviations can bias toward apparent similarity.
  10. Assuming equivalence in efficacy means equivalence in safety. Each important endpoint requires its own appropriate evaluation.
  11. Confusing clinical equivalence with bioequivalence. The underlying concepts are related but the estimands, endpoints, and regulatory frameworks can differ.

What Happens When the Confidence Interval Crosses a Margin?

Suppose the equivalence margin is ±5 and the 90% confidence interval is:

$$ (-3,\;6) $$

The lower bound is acceptable, but the upper bound exceeds +5.

Therefore:

$$ CI_{90\%}\not\subset(-5,5) $$

and equivalence is not demonstrated.

This remains true even if the point estimate is only +1.

What Happens When the Confidence Interval Is Entirely Outside?

Suppose:

$$ CI_{90\%}=(6,\;9) $$

The entire confidence interval is above the upper equivalence margin.

This provides evidence that the treatment difference exceeds the prespecified equivalence region in the positive direction.

The study does not demonstrate equivalence.

The interpretation of superiority depends on the endpoint direction, estimand, and prespecified statistical testing strategy.

What Happens When the Interval Is Very Wide?

Suppose:

$$ CI_{90\%}=(-12,\;14) $$

with equivalence margins of ±5.

The interval crosses both margins.

The study has not demonstrated equivalence because the data are too imprecise to rule out clinically meaningful differences.

Failure to demonstrate equivalence is not proof of non-equivalence. A wide confidence interval may simply indicate that the study lacks sufficient precision.

Equivalence Margin and Clinical Interpretation

The margin translates statistical uncertainty into a clinical decision.

For example, if a difference of five points on a validated clinical scale is considered the largest clinically acceptable difference, then:

$$ \Delta_E=5 $$

defines the statistical boundary.

The confidence interval then answers: Given the observed data, what range of treatment differences remains plausible?

Equivalence is demonstrated when that entire plausible range remains clinically acceptable.

Equivalence and Regulatory Interpretation

Regulatory expectations for equivalence studies can vary substantially depending on the product, endpoint, indication, and development context.

A statistical equivalence calculation should therefore not be treated as a complete regulatory strategy.

The protocol and statistical analysis plan should align with the applicable regulatory guidance and the specific development program.

Sample Size Inflation for Dropout

Suppose the required evaluable sample size is:

$$ N=198 $$

and the anticipated dropout rate is 10%.

A simple inflation calculation is:

$$ N_{\text{enroll}} = \frac{N}{1-0.10} $$
$$ N_{\text{enroll}} = \frac{198}{0.90} = 220 $$

Thus approximately 220 participants would be enrolled to obtain approximately 198 evaluable participants under the assumed dropout rate.

The appropriate inflation method depends on the estimand, analysis method, missing-data assumptions, and whether all randomized participants contribute to the primary analysis.

Allocation Ratio

Equal allocation is often statistically efficient when treatment-group costs and variances are similar.

For unequal allocation, the variance of the treatment difference becomes:

$$ \operatorname{Var}(\bar X_T-\bar X_R) = \frac{\sigma_T^2}{n_T} + \frac{\sigma_R^2}{n_R} $$

Sample-size calculations should therefore account for the actual planned allocation ratio.

Covariate Adjustment

If the primary analysis uses a prespecified covariate-adjusted model, the sample-size calculation should reflect the expected precision of that model.

For example, baseline adjustment may reduce residual variability and improve precision.

However, investigators should not assume an arbitrary variance reduction without justification.

Equivalence With Repeated Measurements

Some clinical trials measure an outcome repeatedly over time.

In such settings, the treatment effect might be defined using:

  • Change from baseline at a specified time point
  • Area under the response curve
  • Average treatment effect over time
  • Model-based contrast

The equivalence margin must correspond to the chosen estimand.

The confidence interval should then be constructed for that specific treatment contrast.

Multiplicity and Hierarchical Testing

If multiple equivalence hypotheses must be satisfied, the statistical testing strategy should be specified in advance.

Possible structures include:

  • Co-primary endpoints
  • Hierarchical testing
  • Multiplicity-adjusted procedures
  • Gatekeeping strategies

The choice depends on the scientific objective and the interpretation of the multiple endpoints.

Interim Analyses

An equivalence trial can include interim analyses, but doing so changes the statistical design.

Repeatedly evaluating equivalence using ordinary unadjusted confidence intervals can inflate the Type I error.

If interim analyses are planned, the monitoring strategy and inferential procedure should be incorporated into the original design.

Bayesian Equivalence Approaches

Although TOST is a standard frequentist framework, Bayesian approaches can also be used to evaluate equivalence.

For example, investigators might calculate:

$$ P(-\Delta_E<\Delta<+\Delta_E\mid\text{data}) $$

which is the posterior probability that the treatment difference lies inside the equivalence region.

A Bayesian decision rule might then require that this probability exceed a prespecified threshold.

Such approaches require prespecified prior distributions and decision rules.

Frequentist vs. Bayesian Interpretation

Framework Typical Equivalence Statement
TOST The confidence interval lies entirely within the equivalence margins.
Bayesian The posterior probability that the effect lies within the equivalence region exceeds a prespecified threshold.

The two frameworks answer related but conceptually different statistical questions.

A Practical Equivalence Design Workflow

1
Define the primary clinical endpoint.
2
Define the treatment effect and estimand.
3
Justify the lower and upper equivalence margins.
4
Specify the Type I error and target power.
5
Specify the expected true treatment difference.
6
Estimate the endpoint variability.
7
Calculate the required sample size.
8
Inflate for anticipated dropout and other prespecified design considerations.
9
Prespecify the primary analysis and analysis populations.
10
Conduct the trial according to the prespecified treatment and assessment schedule.
11
Estimate the treatment difference and construct the appropriate confidence interval.
12
Declare equivalence only if the entire confidence interval lies within the margins.

What Should Be Reported in the Statistical Analysis Plan?

The statistical analysis plan should make the equivalence conclusion completely reproducible.

At minimum, document:

  • Primary endpoint
  • Estimand
  • Treatment effect measure
  • Lower equivalence margin
  • Upper equivalence margin
  • Clinical justification for the margins
  • Type I error
  • Target power
  • Expected treatment difference
  • Variance or SD assumption
  • Sample-size calculation method
  • Allocation ratio
  • Primary statistical model
  • Confidence interval method
  • TOST decision rule
  • Analysis populations
  • Missing-data strategy
  • Protocol-deviation handling
  • Sensitivity analyses
  • Safety analysis

Worked Example Summary

The continuous-endpoint example can be summarized as follows.

Component Value
Design Parallel-group equivalence trial
Endpoint Continuous
Treatment effect \(\mu_T-\mu_R\)
Lower equivalence margin −5
Upper equivalence margin +5
One-sided α 0.05
Target power 90%
Assumed true difference 0
Common SD 12
Approximate N per group 99
Example estimate 1.2
Example SE 1.5
90% CI (−1.27, 3.67)
Equivalence conclusion Demonstrated

The Most Important Concept

The most important concept in equivalence trial design is simple: equivalence must be demonstrated by ruling out clinically meaningful differences, not merely by failing to find a statistically significant difference.

The investigator first defines a clinically meaningful equivalence region:

$$ -\Delta_E<\Delta<+\Delta_E $$

The trial is then designed to provide enough precision to determine whether the observed treatment difference is compatible with that region.

The final statistical conclusion is based on the confidence interval:

$$ CI_{1-2\alpha} \subset (-\Delta_E,+\Delta_E) $$

When the entire interval lies within the margins, equivalence is demonstrated.

When the interval crosses a margin, equivalence is not demonstrated.

Bottom line: An equivalence trial is designed to demonstrate that two treatments are sufficiently similar within a clinically justified range. The equivalence margin must be established before the study and should have a defensible clinical basis. The standard TOST procedure evaluates two one-sided hypotheses, and equivalence at a 5% significance level is commonly expressed as the entire 90% confidence interval for the treatment difference falling inside the prespecified equivalence margins. Sample size is driven heavily by the margin, endpoint variability, Type I error, and desired power. Careful attention to estimands, missing data, protocol deviations, analysis populations, assay sensitivity, and the scientific validity of the reference treatment is essential for a credible equivalence conclusion.

References

Schuirmann, D.J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15, 657–680.
Westlake, W.J. (1976). Symmetrical confidence intervals for bioequivalence trials. Biometrics, 32, 741–744.
Westlake, W.J. (1981). Bioavailability and bioequivalence. In Topics in Pharmaceutical Sciences.
Chow, S.-C. & Liu, J.-P. (2008). Design and Analysis of Bioavailability and Bioequivalence Studies. Chapman & Hall/CRC.
Piaggio, G., Elbourne, D.R., Altman, D.G., Pocock, S.J. & Evans, S.J.W. (2006). Reporting of noninferiority and equivalence randomized trials: extension of the CONSORT statement. JAMA, 295(10), 1152–1160.
FDA. Statistical Approaches to Establishing Bioequivalence. Guidance for Industry.
FDA. Non-Inferiority Clinical Trials to Establish Effectiveness. Guidance for Industry.
EMA. Guideline on the Investigation of Bioequivalence. European Medicines Agency.

Clinical Trials

See these methods in real clinical trials

See the method applied to published trial results, with the estimates, confidence intervals and interpretation explained.