Introduction
Many clinical trials are designed to determine whether a new treatment is better than a control treatment.
Equivalence trials ask a different question: Are the two treatments sufficiently similar that any difference is clinically unimportant?
This distinction is fundamental. A failure to demonstrate superiority does not demonstrate equivalence. Likewise, observing a small difference between treatments does not by itself establish equivalence.
An equivalence trial must be designed specifically to demonstrate that the difference between treatments lies within a prespecified range of clinically acceptable values.
What Is an Equivalence Trial?
An equivalence trial is designed to determine whether the difference between two treatments is sufficiently small to be considered clinically acceptable.
Let the treatment effect be defined as:
where:
- \(\mu_T\) is the mean outcome under the test treatment.
- \(\mu_R\) is the mean outcome under the reference treatment.
- \(\Delta\) is the treatment difference.
Suppose differences between \(-\Delta_E\) and \(+\Delta_E\) are considered clinically acceptable. Then equivalence is defined by:
The quantity \(\Delta_E\) is the equivalence margin.
Equivalence Is Not "No Difference"
This is one of the most important concepts in equivalence trial design.
A conventional statistical test of:
asks whether there is evidence of a difference from exactly zero.
That is not the same question as whether the treatments are sufficiently similar.
A study with a very large sample size could find a statistically significant difference of only 0.2 units even when that difference has no meaningful clinical consequence.
Conversely, a small study might fail to detect a statistically significant difference of 5 units even though such a difference would be clinically important.
Superiority, Non-Inferiority, and Equivalence
These three designs have different objectives.
| Design | Primary Question | Typical Alternative |
|---|---|---|
| Superiority | Is one treatment better? | Difference in one favorable direction |
| Non-inferiority | Is the test treatment not unacceptably worse? | Difference is above a one-sided margin |
| Equivalence | Are the treatments sufficiently similar? | Difference lies within two-sided margins |
The statistical structure follows directly from these questions.
The Equivalence Region
Suppose the clinically acceptable difference is ±5 units. Then:
The equivalence region is therefore:
A treatment difference of:
- −3 units is within the equivalence region.
- 0 units is within the equivalence region.
- +4 units is within the equivalence region.
- +7 units is outside the equivalence region.
- −6 units is outside the equivalence region.
The boundaries are determined by clinical judgment rather than by the observed data.
Choosing the Equivalence Margin
The equivalence margin is arguably the most important design parameter in an equivalence trial.
It should represent the largest difference between treatments that would still be considered clinically acceptable.
A margin should not be selected simply because it produces a convenient sample size.
For an efficacy endpoint, investigators may consider:
- The established effect of the reference treatment
- The variability of the endpoint
- Historical randomized trial evidence
- Clinical relevance of different effect sizes
- Patient perspective
- Regulatory considerations
- The intended use of the new treatment
Symmetric vs. Asymmetric Margins
Not every equivalence margin must be symmetric.
A symmetric margin has:
But some clinical settings may justify different lower and upper limits:
For example, suppose a difference between −2 and +4 units is considered clinically acceptable. Then:
The statistical framework can accommodate asymmetric equivalence margins.
The Equivalence Hypotheses
This is where equivalence trials differ most clearly from conventional hypothesis tests.
Suppose the equivalence limits are:
The null hypothesis is:
The alternative hypothesis is:
In words: the null says the treatments are not equivalent, while the alternative says they are equivalent.
The Two One-Sided Tests Procedure
The standard method for demonstrating equivalence is the Two One-Sided Tests (TOST) procedure.
Instead of testing one two-sided null hypothesis directly, equivalence is demonstrated by conducting two one-sided tests.
The first test evaluates the lower equivalence boundary:
The second test evaluates the upper equivalence boundary:
Both null hypotheses must be rejected for equivalence to be demonstrated.
Why Two Tests?
The two tests address two different ways the treatments could fail to be equivalent.
The Confidence Interval Interpretation
The TOST procedure has an equivalent and often more intuitive confidence interval interpretation.
For a two-sided equivalence test conducted at significance level \(\alpha\), calculate a:
confidence interval for the treatment difference.
If:
then equivalence is demonstrated.
For the common choice:
the corresponding confidence interval is:
Thus, equivalence at the 5% significance level is commonly expressed as the entire 90% confidence interval falling inside the equivalence margins.
A Visual Way to Think About Equivalence
| Confidence Interval | Conclusion |
|---|---|
| Entirely inside the margins | Equivalence demonstrated |
| Crosses either margin | Equivalence not demonstrated |
| Entirely outside one margin | Clearly not equivalent under the specified margins |
| Very wide interval crossing both margins | Data are too imprecise to establish equivalence |
This makes clear why simply observing a small point estimate is insufficient.
A point estimate of zero with a confidence interval from −10 to +10 would not establish equivalence if the margin were ±5.
Worked Example: Continuous Endpoint
Suppose a randomized, parallel-group clinical trial compares a test treatment with a reference treatment using a continuous efficacy endpoint.
Assume that larger values represent better outcomes and that the clinically acceptable difference has been established as ±5 units.
| Design Parameter | Value |
|---|---|
| Endpoint | Continuous |
| Equivalence margin | ±5 units |
| Type I error | 5% |
| Target power | 90% |
| Expected true difference | 0 units |
| Common SD | 12 units |
The treatment effect is defined as:
The equivalence objective is:
Step 1: Define the Hypotheses
The first one-sided test is:
The second is:
Both tests must reject their respective null hypotheses.
Step 2: Suppose the Trial Produces This Result
Assume the estimated treatment difference is:
and the standard error is:
The 90% confidence interval uses:
where:
Therefore:
The margin of error is:
Therefore:
The upper confidence limit exceeds +5. Therefore, the entire confidence interval does not fall inside the equivalence region.
The trial therefore does not demonstrate equivalence.
A Second Example: A Narrower Confidence Interval
Suppose the estimated difference remains:
but the standard error is only:
Then:
giving:
The entire interval lies between −5 and +5.
Therefore, equivalence is demonstrated.
What the Example Demonstrates
The two examples have exactly the same point estimate.
| Scenario | Estimate | 90% CI | Equivalence? |
|---|---|---|---|
| Less precise | 1.2 | (−2.91, 5.31) | No |
| More precise | 1.2 | (−1.27, 3.67) | Yes |
The difference is precision.
This is why sample size is especially important in equivalence trials: the study must be sufficiently precise to exclude clinically meaningful differences.
Sample Size for Two-Group Equivalence Trials
For a simple two-group parallel trial with a continuous endpoint and equal allocation, a useful planning approximation is:
where:
- \(\sigma\) = assumed common standard deviation
- \(\Delta_E\) = equivalence margin
- \(\alpha\) = one-sided TOST significance level
- \(\beta\) = Type II error
- \(n\) = sample size per group
The factor of 2 arises because the variance of the difference between two independent equal-sized group means is:
Worked Sample Size Calculation
Return to the example:
- Equivalence margin = 5
- Common SD = 12
- One-sided α = 0.05
- Power = 90%
The relevant normal quantiles are approximately:
and:
Therefore:
First calculate:
Then:
and:
Thus the approximate required sample size is:
In practice, the final sample size would generally be rounded upward and adjusted for the chosen analysis method, dropout, allocation ratio, and potential use of a t-distribution rather than a normal approximation.
Why the Equivalence Margin Has Such a Large Effect on Sample Size
The sample size is approximately inversely proportional to the square of the equivalence margin:
Therefore, making the margin half as large can require approximately four times the sample size, all else equal.
| Margin | Relative Sample Size |
|---|---|
| \(\Delta_E\) | 1× |
| \(0.75\Delta_E\) | Approximately 1.78× |
| \(0.50\Delta_E\) | Approximately 4× |
| \(0.25\Delta_E\) | Approximately 16× |
This mathematical relationship illustrates why margin selection and trial feasibility must be considered together.
Effect of Variability on Sample Size
Sample size also increases with the square of the assumed standard deviation:
If the true variability is greater than expected, the study may have less precision than anticipated.
For example, increasing the SD from 10 to 12 increases the variance by:
or 44%.
This can have a substantial impact on the required sample size.
Effect of Power on Sample Size
Increasing power increases sample size.
| Power | \(z_{1-\beta}\) | Relative Planning Burden |
|---|---|---|
| 80% | 0.842 | Lower |
| 90% | 1.282 | Higher |
| 95% | 1.645 | Higher still |
The appropriate power depends on the clinical and regulatory context.
Equivalence and Type I Error
In an equivalence study, Type I error concerns incorrectly concluding that the treatments are equivalent when the true difference lies outside the acceptable margins.
The TOST framework uses one-sided tests.
For an overall equivalence conclusion at the conventional 5% level, each component test is typically performed at:
The confidence interval equivalent is therefore a 90% two-sided interval.
Equivalence vs. Failure to Demonstrate Superiority
Consider a superiority study comparing two treatments.
Suppose the study finds:
for a test of whether the treatment difference is statistically significant, and the null hypothesis is not rejected.
That result does not demonstrate equivalence.
The study may simply be too small to rule out clinically meaningful differences.
For example, if the estimated difference is 1 unit but the confidence interval is:
a difference of ±5 or more remains entirely plausible.
The correct conclusion is: equivalence was not established.
It is not: the treatments are equivalent.
Equivalence vs. Non-Inferiority
Non-inferiority typically focuses on one direction.
Suppose negative values indicate worse efficacy for the test treatment. A non-inferiority margin might be:
The question is whether the treatment effect is sufficiently above that boundary.
Equivalence requires both:
and:
Therefore, equivalence is a two-sided requirement while non-inferiority is normally one-sided.
| Characteristic | Non-Inferiority | Equivalence |
|---|---|---|
| Number of margins | One | Two |
| Direction of concern | Usually one direction | Both directions |
| Goal | Rule out unacceptable loss | Rule out unacceptable differences in either direction |
| Typical CI criterion | CI excludes the non-inferiority margin | Entire CI lies within both equivalence margins |
Parallel-Group Equivalence Designs
The simplest equivalence design is a randomized parallel-group trial.
Crossover Equivalence Trials
Equivalence studies are also common in crossover designs, particularly when within-subject comparisons are scientifically appropriate.
For example, each participant might receive both:
- The test treatment
- The reference treatment
The treatment difference is then estimated using within-subject information.
Because the same participant contributes information under both treatments, within-subject variability can be substantially smaller than between-subject variability.
This can make crossover designs efficient when the endpoint and treatment effects are sufficiently stable over time.
Equivalence for Binary Endpoints
The same conceptual framework applies to binary endpoints, although the treatment effect and statistical model differ.
Suppose:
and:
An absolute-risk-difference equivalence criterion might be:
Other effect measures may be appropriate depending on the clinical question, including risk ratios or odds ratios.
Ratio-Based Equivalence
Some equivalence problems are more naturally expressed as ratios rather than differences.
Suppose the relevant parameter is:
An equivalence interval might be:
For example, an equivalence criterion could be expressed as:
This type of ratio-based framework is particularly familiar in bioequivalence studies.
For multiplicative effects, logarithmic transformation is often useful:
which converts a ratio problem into a difference on the log scale.
Bioequivalence vs. Clinical Equivalence
Bioequivalence and clinical equivalence are related but distinct concepts.
Bioequivalence typically concerns pharmacokinetic measures such as:
- AUC
- Cmax
- Other prespecified exposure parameters
Clinical equivalence concerns whether treatments produce sufficiently similar clinical outcomes according to a prespecified margin.
The statistical logic of confidence intervals and equivalence margins is related, but the estimands, endpoints, transformations, and regulatory requirements can differ substantially.
Estimand Considerations
An equivalence trial should define the treatment effect being estimated before the analysis begins.
Important estimand attributes include:
- Population
- Treatment conditions
- Endpoint
- Summary measure
- Handling of intercurrent events
This matters because different analysis strategies can answer different questions about equivalence.
Analysis Populations in Equivalence Trials
Equivalence trials require particularly careful consideration of analysis populations.
Common analysis sets include:
- Intention-to-treat population
- Per-protocol population
- Safety population
In superiority trials, deviations and treatment nonadherence can sometimes make treatment groups appear more similar and therefore bias against finding a difference.
In equivalence trials, that property can be problematic because reduced separation between groups can make equivalence easier to conclude.
Why Both ITT and Per-Protocol Analyses Can Matter
Because different analysis populations can have different biases in equivalence settings, investigators may prespecify complementary analyses.
For example:
| Analysis | Purpose |
|---|---|
| Intent-to-treat | Reflects outcomes according to randomized treatment assignment |
| Per-protocol | Assesses treatment differences among participants sufficiently adherent to the protocol |
| Safety | Characterizes treatment-emergent safety outcomes |
The exact strategy should be prespecified and justified for the trial rather than selected after seeing which analysis supports equivalence.
Missing Data
Missing observations can have a particularly important effect in equivalence trials.
A missing-data strategy should therefore be established before database lock.
Potential approaches depend on the endpoint and estimand and may include:
- Model-based methods
- Multiple imputation
- Mixed-effects models
- Prespecified sensitivity analyses
- Pattern-mixture or tipping-point analyses
The appropriate approach depends on the mechanism and amount of missingness and the scientific estimand.
Assay Sensitivity
An important concern in equivalence and non-inferiority studies is assay sensitivity.
Assay sensitivity refers to the ability of the trial to distinguish effective treatments from ineffective treatments under the conditions of the study.
If the trial is poorly designed or the endpoint is insufficiently sensitive, two ineffective treatments could appear similar.
Therefore, demonstrating equivalence requires more than obtaining a narrow confidence interval. The study must also be scientifically capable of detecting meaningful differences if they exist.
Constancy and Historical Evidence
Historical evidence may be important when selecting equivalence margins and interpreting the reference treatment.
Investigators may need to establish that the reference treatment's historical effect remains reasonably applicable to the current trial setting.
Changes in:
- Patient population
- Background therapy
- Endpoint definitions
- Standard of care
- Comparator implementation
- Trial conduct
can affect the interpretation of historical treatment effects.
Equivalence Does Not Mean Identical
Even when equivalence is demonstrated, the estimated treatment difference is rarely exactly zero.
Suppose:
with:
and equivalence margins of ±5.
The treatment difference is not zero.
Instead, the evidence indicates that the plausible range of treatment differences is entirely within the clinically acceptable interval.
Precision Is Central to Equivalence
Consider two studies with the same observed treatment difference.
| Study | Estimated Difference | 90% CI | Margin | Conclusion |
|---|---|---|---|---|
| A | 0.5 | (−1.5, 2.5) | ±5 | Equivalent |
| B | 0.5 | (−7.0, 8.0) | ±5 | Not demonstrated |
The point estimates are identical. The conclusions differ because the second study is much less precise.
This is one of the defining features of equivalence testing.
Sample Size and Precision
For a simple continuous endpoint:
As sample size increases, the confidence interval becomes narrower.
The purpose of increasing sample size in an equivalence trial is therefore not merely to increase the chance of detecting a difference. It is to obtain enough precision to rule out differences larger than the equivalence margin.
Power in an Equivalence Trial
Power is the probability of correctly demonstrating equivalence for a specified true treatment difference.
It is common to assume during planning that:
but power can also be evaluated for values away from zero.
This distinction is useful because the probability of demonstrating equivalence generally decreases as the true treatment difference approaches an equivalence boundary.
Power Depends on the True Difference
Suppose the equivalence interval is:
A study designed for 90% power when:
will not necessarily have 90% power when:
because the true effect is much closer to the upper equivalence boundary.
| True Difference | Distance from Upper Margin | Expected Ability to Demonstrate Equivalence |
|---|---|---|
| 0 | 5 units | Highest |
| 1 | 4 units | Lower |
| 3 | 2 units | Lower still |
| 4.5 | 0.5 units | Very low |
| 5 | 0 units | Not expected to demonstrate equivalence |
The actual power curve should therefore be examined when the design is important or the expected treatment difference is uncertain.
Equivalence Margin vs. Expected Treatment Difference
These two quantities should not be confused.
The equivalence margin is a clinical decision threshold. The expected treatment difference is a planning assumption.
For example:
could be the clinically acceptable difference while the design assumes:
for sample-size calculations.
The study is not trying to prove that the true difference equals zero. It is trying to demonstrate that the difference is sufficiently close to zero to remain within the acceptable region.
Equivalence and Multiple Endpoints
Some trials have multiple efficacy endpoints or multiple components that must all satisfy equivalence requirements.
This introduces an additional multiplicity problem.
For example, a study might require equivalence for:
- Primary efficacy endpoint
- Key secondary endpoint
- Another clinically important outcome
The protocol should clearly identify which endpoint establishes the primary equivalence conclusion and how multiplicity is handled.
Equivalence and Safety
Equivalence in efficacy does not imply equivalence in safety.
A treatment may demonstrate equivalent efficacy while having a different adverse-event profile.
Safety should therefore be analyzed separately using appropriate estimands, summaries, and inferential procedures.
Likewise, an equivalence conclusion for one clinical endpoint does not imply equivalence for every other outcome.
Protocol Considerations
An equivalence protocol should clearly define:
- Primary endpoint
- Treatment effect measure
- Equivalence margins
- Scientific justification for the margins
- Type I error
- Target power
- Sample-size assumptions
- Allocation ratio
- Primary analysis method
- Analysis populations
- Missing-data strategy
- Handling of protocol deviations
- Intercurrent events
- Sensitivity analyses
- Criteria for declaring equivalence
A Complete Equivalence Decision Algorithm
R Implementation: TOST for a Continuous Endpoint
For a simple two-group analysis, the TOST procedure can be implemented directly from the estimated difference and its standard error.
estimate <- 1.2 se <- 1.5 lower_margin <- -5 upper_margin <- 5 alpha <- 0.05 z <- qnorm(1 - alpha) lower_ci <- estimate - z * se upper_ci <- estimate + z * se lower_ci upper_ci
The resulting 90% confidence interval is approximately:
lower_ci # approximately -1.27 upper_ci # approximately 3.67
The equivalence criterion can then be evaluated directly:
equivalent <- lower_ci > lower_margin && upper_ci < upper_margin equivalent
The result is:
TRUE
because the complete interval lies inside:
Explicit TOST Calculations in R
The two one-sided tests can also be calculated explicitly.
estimate <- 1.2 se <- 1.5 lower_margin <- -5 upper_margin <- 5 alpha <- 0.05 # Test H01: Delta <= lower_margin t_lower <- (estimate - lower_margin) / se p_lower <- 1 - pnorm(t_lower) # Test H02: Delta >= upper_margin t_upper <- (estimate - upper_margin) / se p_upper <- pnorm(t_upper) p_lower p_upper
Equivalence is demonstrated when both one-sided p-values are below the prespecified significance level:
equivalent <- p_lower < alpha && p_upper < alpha equivalent
This gives the same inferential conclusion as the confidence-interval approach.
A Compact TOST Function
tost_equivalence <- function(
estimate,
se,
lower_margin,
upper_margin,
alpha = 0.05
) {
z <- qnorm(1 - alpha)
lower_ci <-
estimate - z * se
upper_ci <-
estimate + z * se
equivalent <-
lower_ci > lower_margin &&
upper_ci < upper_margin
list(
estimate = estimate,
lower_ci = lower_ci,
upper_ci = upper_ci,
equivalent = equivalent
)
}
tost_equivalence(
estimate = 1.2,
se = 1.5,
lower_margin = -5,
upper_margin = 5
)
Power Calculation by Simulation
For more complex equivalence designs, simulation is often useful.
A basic simulation repeatedly generates treatment-group data, estimates the treatment difference, constructs the confidence interval, and records whether equivalence was demonstrated.
set.seed(123)
nsim <- 10000
n_T <- 99
n_R <- 99
mu_T <- 0
mu_R <- 0
sd <- 12
lower_margin <- -5
upper_margin <- 5
success <- logical(nsim)
for (i in seq_len(nsim)) {
y_T <- rnorm(
n_T,
mean = mu_T,
sd = sd
)
y_R <- rnorm(
n_R,
mean = mu_R,
sd = sd
)
estimate <-
mean(y_T) - mean(y_R)
se <-
sqrt(
sd^2 / n_T +
sd^2 / n_R
)
lower_ci <-
estimate - qnorm(0.95) * se
upper_ci <-
estimate + qnorm(0.95) * se
success[i] <-
lower_ci > lower_margin &&
upper_ci < upper_margin
}
mean(success)
The resulting proportion is an empirical estimate of the probability of demonstrating equivalence under the assumed true difference and variability.
Why Simulation Can Be Valuable
Closed-form sample-size equations are useful, but real equivalence studies can be substantially more complicated.
Simulation can incorporate:
- Unequal allocation
- Dropout
- Non-normal endpoints
- Repeated measurements
- Mixed-effects models
- Binary outcomes
- Time-to-event endpoints
- Covariate adjustment
- Complex missing-data mechanisms
The simulation should reproduce the planned analysis as closely as possible.
Common Equivalence Trial Design Mistakes
- Defining equivalence as a nonsignificant superiority test. Failure to reject a difference from zero does not establish equivalence.
- Choosing the equivalence margin after seeing the data. The margin must be justified and prespecified.
- Using a margin that is too wide. A statistically successful trial may then demonstrate a level of similarity that is not clinically meaningful.
- Using a margin that is too narrow without considering feasibility. The resulting sample size may become impractically large.
- Looking only at the point estimate. Equivalence depends on the entire confidence interval.
- Using the wrong confidence interval. For a 5% TOST analysis, the corresponding two-sided confidence interval is 90%, not 95%.
- Ignoring variability uncertainty. Underestimating the SD can result in an underpowered study.
- Ignoring missing data. Loss of precision and departures from the intended estimand can materially affect the equivalence conclusion.
- Ignoring protocol deviations. In equivalence studies, deviations can bias toward apparent similarity.
- Assuming equivalence in efficacy means equivalence in safety. Each important endpoint requires its own appropriate evaluation.
- Confusing clinical equivalence with bioequivalence. The underlying concepts are related but the estimands, endpoints, and regulatory frameworks can differ.
What Happens When the Confidence Interval Crosses a Margin?
Suppose the equivalence margin is ±5 and the 90% confidence interval is:
The lower bound is acceptable, but the upper bound exceeds +5.
Therefore:
and equivalence is not demonstrated.
This remains true even if the point estimate is only +1.
What Happens When the Confidence Interval Is Entirely Outside?
Suppose:
The entire confidence interval is above the upper equivalence margin.
This provides evidence that the treatment difference exceeds the prespecified equivalence region in the positive direction.
The study does not demonstrate equivalence.
The interpretation of superiority depends on the endpoint direction, estimand, and prespecified statistical testing strategy.
What Happens When the Interval Is Very Wide?
Suppose:
with equivalence margins of ±5.
The interval crosses both margins.
The study has not demonstrated equivalence because the data are too imprecise to rule out clinically meaningful differences.
Equivalence Margin and Clinical Interpretation
The margin translates statistical uncertainty into a clinical decision.
For example, if a difference of five points on a validated clinical scale is considered the largest clinically acceptable difference, then:
defines the statistical boundary.
The confidence interval then answers: Given the observed data, what range of treatment differences remains plausible?
Equivalence is demonstrated when that entire plausible range remains clinically acceptable.
Equivalence and Regulatory Interpretation
Regulatory expectations for equivalence studies can vary substantially depending on the product, endpoint, indication, and development context.
A statistical equivalence calculation should therefore not be treated as a complete regulatory strategy.
The protocol and statistical analysis plan should align with the applicable regulatory guidance and the specific development program.
Sample Size Inflation for Dropout
Suppose the required evaluable sample size is:
and the anticipated dropout rate is 10%.
A simple inflation calculation is:
Thus approximately 220 participants would be enrolled to obtain approximately 198 evaluable participants under the assumed dropout rate.
The appropriate inflation method depends on the estimand, analysis method, missing-data assumptions, and whether all randomized participants contribute to the primary analysis.
Allocation Ratio
Equal allocation is often statistically efficient when treatment-group costs and variances are similar.
For unequal allocation, the variance of the treatment difference becomes:
Sample-size calculations should therefore account for the actual planned allocation ratio.
Covariate Adjustment
If the primary analysis uses a prespecified covariate-adjusted model, the sample-size calculation should reflect the expected precision of that model.
For example, baseline adjustment may reduce residual variability and improve precision.
However, investigators should not assume an arbitrary variance reduction without justification.
Equivalence With Repeated Measurements
Some clinical trials measure an outcome repeatedly over time.
In such settings, the treatment effect might be defined using:
- Change from baseline at a specified time point
- Area under the response curve
- Average treatment effect over time
- Model-based contrast
The equivalence margin must correspond to the chosen estimand.
The confidence interval should then be constructed for that specific treatment contrast.
Multiplicity and Hierarchical Testing
If multiple equivalence hypotheses must be satisfied, the statistical testing strategy should be specified in advance.
Possible structures include:
- Co-primary endpoints
- Hierarchical testing
- Multiplicity-adjusted procedures
- Gatekeeping strategies
The choice depends on the scientific objective and the interpretation of the multiple endpoints.
Interim Analyses
An equivalence trial can include interim analyses, but doing so changes the statistical design.
Repeatedly evaluating equivalence using ordinary unadjusted confidence intervals can inflate the Type I error.
If interim analyses are planned, the monitoring strategy and inferential procedure should be incorporated into the original design.
Bayesian Equivalence Approaches
Although TOST is a standard frequentist framework, Bayesian approaches can also be used to evaluate equivalence.
For example, investigators might calculate:
which is the posterior probability that the treatment difference lies inside the equivalence region.
A Bayesian decision rule might then require that this probability exceed a prespecified threshold.
Such approaches require prespecified prior distributions and decision rules.
Frequentist vs. Bayesian Interpretation
| Framework | Typical Equivalence Statement |
|---|---|
| TOST | The confidence interval lies entirely within the equivalence margins. |
| Bayesian | The posterior probability that the effect lies within the equivalence region exceeds a prespecified threshold. |
The two frameworks answer related but conceptually different statistical questions.
A Practical Equivalence Design Workflow
What Should Be Reported in the Statistical Analysis Plan?
The statistical analysis plan should make the equivalence conclusion completely reproducible.
At minimum, document:
- Primary endpoint
- Estimand
- Treatment effect measure
- Lower equivalence margin
- Upper equivalence margin
- Clinical justification for the margins
- Type I error
- Target power
- Expected treatment difference
- Variance or SD assumption
- Sample-size calculation method
- Allocation ratio
- Primary statistical model
- Confidence interval method
- TOST decision rule
- Analysis populations
- Missing-data strategy
- Protocol-deviation handling
- Sensitivity analyses
- Safety analysis
Worked Example Summary
The continuous-endpoint example can be summarized as follows.
| Component | Value |
|---|---|
| Design | Parallel-group equivalence trial |
| Endpoint | Continuous |
| Treatment effect | \(\mu_T-\mu_R\) |
| Lower equivalence margin | −5 |
| Upper equivalence margin | +5 |
| One-sided α | 0.05 |
| Target power | 90% |
| Assumed true difference | 0 |
| Common SD | 12 |
| Approximate N per group | 99 |
| Example estimate | 1.2 |
| Example SE | 1.5 |
| 90% CI | (−1.27, 3.67) |
| Equivalence conclusion | Demonstrated |
The Most Important Concept
The most important concept in equivalence trial design is simple: equivalence must be demonstrated by ruling out clinically meaningful differences, not merely by failing to find a statistically significant difference.
The investigator first defines a clinically meaningful equivalence region:
The trial is then designed to provide enough precision to determine whether the observed treatment difference is compatible with that region.
The final statistical conclusion is based on the confidence interval:
When the entire interval lies within the margins, equivalence is demonstrated.
When the interval crosses a margin, equivalence is not demonstrated.
References
Schuirmann, D.J. (1987).
A comparison of the two one-sided tests procedure and the power
approach for assessing the equivalence of average bioavailability.
Journal of Pharmacokinetics and Biopharmaceutics, 15, 657–680.
Westlake, W.J. (1976).
Symmetrical confidence intervals for bioequivalence trials.
Biometrics, 32, 741–744.
Westlake, W.J. (1981).
Bioavailability and bioequivalence.
In Topics in Pharmaceutical Sciences.
Chow, S.-C. & Liu, J.-P. (2008).
Design and Analysis of Bioavailability and Bioequivalence Studies.
Chapman & Hall/CRC.
Piaggio, G., Elbourne, D.R., Altman, D.G., Pocock, S.J. & Evans, S.J.W.
(2006).
Reporting of noninferiority and equivalence randomized trials:
extension of the CONSORT statement.
JAMA, 295(10), 1152–1160.
FDA.
Statistical Approaches to Establishing Bioequivalence.
Guidance for Industry.
FDA.
Non-Inferiority Clinical Trials to Establish Effectiveness.
Guidance for Industry.
EMA.
Guideline on the Investigation of Bioequivalence.
European Medicines Agency.
See these methods in real clinical trials
See the method applied to published trial results, with the estimates, confidence intervals and interpretation explained.