Tutorials › Biostatistics › Non-Inferiority Trial Design Principles

Confirmatory Trial Design

Non-Inferiority Trial Design Principles

A practical guide to designing and interpreting non-inferiority trials, including the non-inferiority margin, hypotheses, confidence intervals, assay sensitivity, constancy assumptions, sample-size determination, analysis populations, missing data, and a complete worked example.

Advanced 18 min read

What You'll Learn

  • What a non-inferiority trial is designed to establish
  • How the non-inferiority margin defines the acceptable loss of efficacy
  • How non-inferiority hypotheses differ from superiority hypotheses
  • How confidence intervals are used to make the non-inferiority decision
  • Why assay sensitivity and the constancy assumption are fundamental
  • How to calculate sample size and analyze a complete non-inferiority example

Introduction

Many clinical trials are designed to demonstrate that a new treatment is better than an existing treatment.

But superiority is not always the objective.

A new treatment may offer important advantages such as improved tolerability, simpler administration, lower cost, improved adherence, reduced monitoring, shorter treatment duration, or easier storage while providing efficacy that is acceptably close to that of an established active treatment.

In these situations, a non-inferiority trial may be appropriate.

The central question is not: "Is the experimental treatment statistically different from the control?"

Instead, the question is: "Can we rule out a clinically unacceptable loss of efficacy?"

Key idea: A non-inferiority trial does not attempt to prove that two treatments are identical. It attempts to show that the experimental treatment's efficacy is not worse than the active control by more than a prespecified clinically acceptable amount.

Superiority Versus Non-Inferiority

The distinction becomes clearer by comparing the questions asked by the two designs.

Design Primary Question
Superiority Is the experimental treatment better than the comparator?
Non-inferiority Is the experimental treatment not unacceptably worse than the comparator?

A superiority trial is generally concerned with detecting a positive treatment difference.

A non-inferiority trial is concerned with excluding sufficiently large negative differences.

The Basic Setup

Suppose a new treatment, \(T\), is compared with an established active control, \(C\).

Let the treatment effect be defined as:

$$ \Delta=\theta_T-\theta_C $$

where \(\theta\) represents the appropriate measure of efficacy.

For a continuous endpoint, \(\theta\) might be a mean. For a binary endpoint, it might be a response probability. For a time-to-event endpoint, it might be a log hazard ratio.

The direction of the effect must be defined carefully because the interpretation of the non-inferiority margin depends on whether larger or smaller values are better.

First rule of non-inferiority design: Define the treatment effect and its direction before defining the margin. A positive margin cannot be interpreted correctly until the estimand's direction has been established.

What Is the Non-Inferiority Margin?

The most important quantity in a non-inferiority trial is the non-inferiority margin, usually denoted by \(M\) or \(\Delta_{NI}\).

The margin represents the largest loss of treatment effect that can be accepted while still considering the experimental treatment clinically acceptable for the trial's stated purpose.

Suppose larger values of \(\theta\) are better.

Then the experimental treatment is considered non-inferior if its effect is not too far below that of the control.

$$ \Delta=\theta_T-\theta_C $$

and non-inferiority is demonstrated when:

$$ \Delta>-M $$

where \(M>0\) is the prespecified non-inferiority margin.

The margin therefore creates a boundary between an acceptable loss and an unacceptable loss.

Why the Margin Is a Clinical Quantity

The non-inferiority margin should not simply be chosen because it produces a convenient sample size.

Suppose an established treatment has a large and clinically important benefit over placebo.

A new treatment should not be allowed to lose nearly all of that benefit merely because a statistical calculation produces a manageable enrollment target.

The margin must therefore be justified using clinical evidence about the established treatment's effect and the amount of benefit that must be preserved.

Critical principle: The non-inferiority margin is not merely a statistical tuning parameter. It represents a clinically justified amount of potentially tolerable efficacy loss.

Non-Inferiority Hypotheses

For the convention where larger values of \(\theta\) are better, the hypotheses can be written as:

$$ H_0:\Delta\le -M \qquad\text{vs.}\qquad H_A:\Delta>-M $$

The null hypothesis says that the experimental treatment is inferior to the control by at least the non-inferiority margin.

The alternative says that the treatment difference is sufficiently close to zero that unacceptable inferiority can be ruled out.

This is fundamentally different from a superiority hypothesis:

$$ H_0:\Delta\le0 \qquad\text{vs.}\qquad H_A:\Delta>0 $$

In a non-inferiority trial, the null boundary is shifted from zero to the non-inferiority margin.

A Visual Interpretation

Treatment Difference \(\Delta\) Interpretation
\(\Delta\le -M\) Unacceptably inferior
\(-M<\Delta<0\) Some efficacy loss, but within the non-inferiority margin
\(\Delta=0\) No estimated treatment difference
\(\Delta>0\) Experimental treatment favors efficacy

The critical distinction is between statistical equality and acceptable inferiority.

A non-inferiority trial does not need the estimated treatment difference to be zero.

The estimate may favor the control and still support non-inferiority, provided the confidence interval excludes losses larger than the prespecified margin.

The Confidence-Interval Decision Rule

One of the cleanest ways to interpret a non-inferiority trial is through a confidence interval.

Suppose the estimated treatment difference is:

$$ \hat{\Delta}=\hat{\theta}_T-\hat{\theta}_C $$

Construct a two-sided \(100(1-2\alpha)\%\) confidence interval for \(\Delta\). Equivalently, this corresponds to a one-sided test at level \(\alpha\).

For a conventional one-sided \(\alpha=0.025\) non-inferiority test, this is typically represented by a 95% two-sided confidence interval.

Non-inferiority is concluded when the lower confidence bound exceeds the negative margin:

$$ L_{CI}>-M $$

where \(L_{CI}\) is the lower confidence limit.

Practical rule: For a treatment effect defined as experimental minus control, non-inferiority is established when the lower bound of the confidence interval is above \(-M\).

Three Examples of Confidence-Interval Interpretation

Example A: Non-Inferiority Clearly Demonstrated

Suppose:

$$ \hat{\Delta}=-0.02 $$ and: $$ 95\%\,CI=(-0.06,\;0.02) $$

If:

$$ M=0.10 $$

then:

$$ -0.06>-0.10 $$

Therefore, non-inferiority is demonstrated.

Example B: Non-Inferiority Not Demonstrated

Suppose:

$$ \hat{\Delta}=-0.08 $$ and: $$ 95\%\,CI=(-0.14,\;-0.02) $$

With \(M=0.10\):

$$ -0.14<-0.10 $$

The confidence interval includes treatment differences worse than the non-inferiority margin.

Therefore, non-inferiority is not established.

Example C: The Estimate Favors the Control but Non-Inferiority Holds

Suppose:

$$ \hat{\Delta}=-0.04 $$ with: $$ 95\%\,CI=(-0.07,\;-0.01) $$

The estimate favors the control, but if \(M=0.10\), the lower bound remains above \(-0.10\).

Therefore, the experimental treatment can still be declared non-inferior.

Important: Non-inferiority does not mean "the confidence interval crosses zero." That is a superiority concept. Non-inferiority is determined by whether the confidence interval crosses the prespecified non-inferiority margin.

Non-Inferiority Does Not Mean Equivalence

Non-inferiority and equivalence are related but distinct concepts.

A non-inferiority trial generally asks whether the experimental treatment is not too much worse than the comparator.

An equivalence trial asks whether the treatment difference lies within both a lower and upper equivalence margin.

For an equivalence margin \(M\), the hypotheses are conceptually:

$$ -M<\Delta

Thus, equivalence requires ruling out both sufficiently large inferiority and sufficiently large superiority.

Design Question Confidence-Interval Criterion
Superiority Is \(T\) better than \(C\)? Lower bound above 0
Non-inferiority Is \(T\) not unacceptably worse? Lower bound above \(-M\)
Equivalence Are \(T\) and \(C\) sufficiently similar? Entire CI within \((-M,+M)\)

Why an Active Control Is Usually Central

A non-inferiority trial commonly compares the experimental treatment with an established active treatment rather than placebo.

The purpose is to determine whether the experimental treatment preserves an acceptable amount of the established treatment's effect.

This creates a fundamental requirement: the active control must have demonstrated efficacy under circumstances relevant to the current trial.

If the active control performs poorly for reasons unrelated to the experimental treatment, an apparently non-inferior experimental treatment may simply be non-effective.

The Problem of Assay Sensitivity

Assay sensitivity refers to the ability of a trial to distinguish an effective treatment from an ineffective treatment under the trial's conditions.

This concept is especially important in non-inferiority trials.

Consider a trial in which both treatments perform poorly.

If the experimental treatment and active control have similar poor outcomes, the trial may conclude non-inferiority because their difference is small.

That does not necessarily mean the experimental treatment is effective.

The non-inferiority paradox: A trial can observe little difference between two ineffective treatments. Therefore, simply demonstrating "no large difference" is not sufficient unless the trial has preserved the ability to detect treatment effects.

The Constancy Assumption

The constancy assumption is another foundational concept.

Suppose historical trials established that the active control had a substantial benefit over placebo.

The current non-inferiority trial assumes that an appropriate amount of that benefit would still be present under the current trial conditions.

If the historical control effect was:

$$ \theta_C-\theta_P=\delta $$

then the non-inferiority design relies, directly or indirectly, on the assumption that the control would still retain a clinically meaningful effect in the current setting.

This is one reason historical evidence, trial population, endpoint definitions, background therapy, dosing, adherence, and study conduct matter so much in non-inferiority design.

Preserving a Fraction of the Control Effect

Suppose historical evidence suggests that the active control has a benefit of \(E\) over placebo.

A common conceptual approach is to require that the experimental treatment preserve a specified fraction of this effect.

For example, if the control effect is:

$$ E=0.20 $$

and the design requires preservation of 50% of the control effect, then the largest allowable loss would be:

$$ M=0.20(1-0.50)=0.10 $$

This illustrates the logic, but the actual derivation of a non-inferiority margin is more nuanced and depends on the endpoint, historical evidence, clinical judgment, and regulatory/statistical considerations.

Do not mechanically set the margin to an arbitrary fraction. The margin should be justified using the historical evidence and the clinical importance of preserving treatment effect.

Choosing the Endpoint

Non-inferiority design is highly sensitive to the choice and definition of the primary endpoint.

The endpoint should be:

  • Clinically meaningful
  • Measured consistently across treatment groups
  • Supported by reliable historical evidence
  • Appropriate for the active control
  • Defined identically across the trial
  • Available with sufficient completeness

A change in endpoint definition can undermine the historical evidence used to justify the non-inferiority margin.

Endpoint Direction Matters

Suppose higher values indicate better outcomes.

Then:

$$ \Delta=\theta_T-\theta_C $$

and the non-inferiority condition is:

$$ \Delta>-M $$

But suppose lower values indicate better outcomes, as can occur with some biomarkers, symptom scores, or event rates.

The effect definition may instead be:

$$ \Delta=\theta_C-\theta_T $$

so that a positive value favors the experimental treatment.

The corresponding non-inferiority decision must then be written consistently.

Common error: Copying a non-inferiority formula from another trial without checking which direction represents benefit can reverse the interpretation of the entire analysis.

Non-Inferiority With Binary Endpoints

Suppose the primary endpoint is response.

Let:

$$ p_T=P(\text{response}\mid T) $$ and: $$ p_C=P(\text{response}\mid C) $$

Define the risk-difference treatment effect as:

$$ \Delta=p_T-p_C $$

If higher response rates are better, then non-inferiority is:

$$ p_T-p_C>-M $$

The corresponding confidence-interval criterion is:

$$ L_{CI}>-M $$

Non-Inferiority With Continuous Endpoints

For a continuous endpoint, suppose the treatment effect is defined as:

$$ \Delta=\mu_T-\mu_C $$

where \(\mu_T\) and \(\mu_C\) are the treatment-group means.

If higher values indicate better outcomes, the non-inferiority criterion remains:

$$ L_{CI}>-M $$

The major difference is the statistical model used to estimate the treatment difference and its standard error.

Non-Inferiority With Time-to-Event Endpoints

For a time-to-event endpoint, the treatment effect is often represented by a hazard ratio.

Suppose the hazard ratio is defined as:

$$ HR=\frac{\lambda_T}{\lambda_C} $$

where lower hazard is better.

In this setting, a hazard ratio of 1 indicates no treatment difference, while values above 1 indicate a higher hazard for the experimental treatment.

The non-inferiority margin is therefore naturally expressed on the hazard-ratio scale.

For example:

$$ HR_{NI}=1.30 $$

would represent the largest acceptable relative increase in hazard under the specified design.

Non-inferiority would then be demonstrated if the upper confidence bound for the hazard ratio is below the margin:

$$ U_{CI}<1.30 $$
Scale matters: For risk differences, the non-inferiority boundary may be a negative additive difference. For hazard ratios or risk ratios, the margin is typically expressed on a ratio scale. Always define the estimand and margin on the same scale.

One-Sided Versus Two-Sided Testing

Non-inferiority is fundamentally a one-sided testing problem.

The relevant question is whether the experimental treatment could be unacceptably worse than the control.

For a one-sided significance level \(\alpha\), the corresponding confidence interval is commonly a two-sided \(100(1-2\alpha)\%\) interval.

For example:

One-Sided \(\alpha\) Equivalent Two-Sided CI
0.05 90%
0.025 95%
0.01 98%

Thus, a one-sided 2.5% non-inferiority test corresponds to evaluating whether the lower bound of a 95% two-sided confidence interval is above the non-inferiority margin.

Sample Size for a Non-Inferiority Trial

The sample-size calculation depends on:

  • The non-inferiority margin \(M\)
  • The expected treatment effect
  • The variability of the endpoint
  • The type I error
  • The desired power
  • The allocation ratio
  • The statistical test or model
  • The expected dropout or missing-data rate

The margin is particularly influential.

A smaller non-inferiority margin means that the trial must distinguish smaller differences between treatments.

That generally requires a larger sample size.

Design trade-off: Tightening the non-inferiority margin improves the clinical stringency of the trial but can substantially increase the required sample size.

Sample Size for a Continuous Endpoint

Consider a parallel-group trial with equal allocation and a continuous endpoint.

Suppose the expected treatment difference is:

$$ \delta=\mu_T-\mu_C $$

and the non-inferiority margin is \(M\).

The quantity that determines how far the assumed treatment effect lies from the non-inferiority boundary is:

$$ d=\delta+M $$

when larger values are better and the null boundary is \(-M\).

For equal allocation, a commonly used normal-approximation formula for total sample size is:

$$ N \approx \frac{ 2\sigma^2 \left( z_{1-\alpha}+z_{1-\beta} \right)^2 }{ (\delta+M)^2 } $$

where \(\sigma\) is the common standard deviation.

The exact formula depends on the planned analysis and variance assumptions.

Why the Margin Has Such a Large Effect

Notice that the margin appears in the denominator:

$$ (\delta+M)^2 $$

If the expected treatment difference is near zero, then a smaller \(M\) directly reduces the distance from the non-inferiority boundary.

For example, suppose:

$$ \delta=0 $$

and compare:

$$ M=0.20 $$ versus: $$ M=0.10 $$

Ignoring all other changes, halving the margin approximately quadruples the required sample size because the denominator is squared.

Practical implication: The choice of non-inferiority margin is often one of the most consequential design decisions in the entire study.

A Complete Worked Example

Consider a randomized, parallel-group non-inferiority trial comparing a new oral treatment with an established active treatment.

The primary endpoint is a continuous efficacy measure, with larger values representing better outcomes.

Suppose the study assumptions are:

Parameter Planning Value
Expected treatment difference \(\delta\) 0
Common standard deviation \(\sigma\) 1.0
Non-inferiority margin \(M\) 0.30
One-sided type I error 2.5%
Target power 90%
Allocation 1:1

The hypotheses are:

$$ H_0:\mu_T-\mu_C\le-0.30 $$ $$ H_A:\mu_T-\mu_C>-0.30 $$

Step 1: Determine the Distance From the Margin

The assumed treatment difference is:

$$ \delta=0 $$

The non-inferiority margin is:

$$ M=0.30 $$

Therefore:

$$ \delta+M=0.30 $$

This means the assumed treatment effect is 0.30 units above the non-inferiority boundary.

Step 2: Determine the Critical Values

For a one-sided \(\alpha=0.025\) test:

$$ z_{1-\alpha}\approx1.96 $$

For 90% power:

$$ z_{1-\beta}=z_{0.90}\approx1.282 $$

Therefore:

$$ z_{1-\alpha}+z_{1-\beta} = 1.96+1.282 = 3.242 $$

Step 3: Calculate the Initial Sample Size

Using:

$$ N \approx \frac{ 2(1.0)^2(3.242)^2 }{ (0.30)^2 } $$

This gives approximately:

$$ N\approx233.7 $$

Therefore, an initial planning value would be approximately:

$$ \boxed{N\approx234} $$

or approximately 117 patients per treatment group under equal allocation.

Important: This is an illustrative normal-approximation calculation. A production sample-size calculation should use the exact planned analysis, including any covariate adjustment, variance assumptions, repeated measures, dropout assumptions, stratification, or other design features.

Step 4: Inflate for Dropout

Suppose the study anticipates a 10% dropout rate.

A simple inflation is:

$$ N_{inflated} = \frac{234}{1-0.10} $$

which gives:

$$ N_{inflated}\approx260 $$

Thus, approximately 260 randomized patients would be required under this illustrative assumption.

With equal allocation, this corresponds to approximately:

$$ 130\text{ patients per group} $$

Step 5: Define the Analysis

At the end of the trial, suppose the estimated treatment difference is:

$$ \hat{\Delta}=-0.08 $$

and the 95% confidence interval is:

$$ 95\%\,CI=(-0.21,\;0.05) $$

The lower confidence bound is:

$$ L_{CI}=-0.21 $$

The non-inferiority margin is:

$$ -M=-0.30 $$

Because:

$$ -0.21>-0.30 $$

the lower confidence bound is above the non-inferiority margin.

Therefore:

$$ \boxed{\text{Non-inferiority is demonstrated}} $$

What If the Confidence Interval Were Wider?

Suppose instead that the estimated difference were identical:

$$ \hat{\Delta}=-0.08 $$

but the confidence interval were:

$$ 95\%\,CI=(-0.34,\;0.18) $$

The lower bound would be:

$$ -0.34 $$

which is below the margin:

$$ -0.30 $$

Therefore, non-inferiority would not be demonstrated.

The point estimate did not change. The conclusion changed because the uncertainty around the estimate changed. This illustrates why non-inferiority trials require adequate precision, not merely a favorable point estimate.

Non-Inferiority Is Not Established by a Non-Significant Superiority Test

This is one of the most important mistakes to avoid.

Suppose a conventional superiority test produces:

$$ p>0.05 $$

That does not establish non-inferiority.

A non-significant superiority test simply means that the trial did not establish a statistically significant difference in the superiority direction.

The confidence interval could still include clinically important inferiority.

For example:

$$ 95\%\,CI=(-0.40,\;0.10) $$

This interval includes zero, so superiority is not established. But if:

$$ M=0.30 $$

the interval also extends below \(-0.30\), so non-inferiority is not established either.

The Four Possible Interpretations

A useful way to understand the results is to consider the relationship between the confidence interval and the two important boundaries: the non-inferiority margin and zero.

95% CI Interpretation
Entirely above 0 Superiority demonstrated; non-inferiority necessarily also supported
Crosses 0 but remains above \(-M\) Non-inferiority demonstrated; superiority not demonstrated
Extends below \(-M\) and crosses 0 Neither superiority nor non-inferiority demonstrated
Entirely below \(-M\) Inferiority demonstrated relative to the prespecified margin

This framework is particularly useful when explaining results to clinical teams.

Non-Inferiority Followed by Superiority

Sometimes a study is designed first to establish non-inferiority and then, if non-inferiority is demonstrated, to investigate superiority.

For example:

1
Test whether the experimental treatment is non-inferior.
2
If non-inferiority is demonstrated, assess whether the treatment is superior.
3
Interpret the superiority analysis according to the prespecified multiplicity and testing strategy.

This hierarchical approach can be particularly useful when the scientific question is whether the experimental treatment retains acceptable efficacy and may additionally provide superior efficacy.

However, the testing sequence and multiplicity strategy must be prespecified.

Analysis Populations Matter More in Non-Inferiority Trials

In superiority trials, departures from the protocol often tend to dilute differences between treatment groups.

In non-inferiority trials, that dilution can work in an undesirable direction.

If protocol deviations make the treatments look more similar, an analysis that includes those deviations can potentially make it easier to conclude non-inferiority.

For this reason, non-inferiority trials commonly place substantial emphasis on both:

  • The intention-to-treat or full-analysis population
  • The per-protocol population
Important principle: In a non-inferiority trial, agreement between appropriately specified intention-to-treat and per-protocol analyses can provide important reassurance that the non-inferiority conclusion is not being driven by protocol deviations or treatment dilution.

Intention-to-Treat Analysis

The intention-to-treat principle generally analyzes participants according to their randomized treatment assignment.

This preserves the benefits of randomization.

However, in a non-inferiority trial, treatment discontinuation, crossover, nonadherence, and other deviations may reduce the contrast between treatments.

Consequently, the ITT analysis should not automatically be assumed to be the only informative analysis for demonstrating non-inferiority.

Per-Protocol Analysis

A per-protocol analysis attempts to estimate the treatment effect among participants who sufficiently adhered to the protocol's important requirements.

The exact definition must be prespecified.

For example, the protocol may specify criteria involving:

  • Minimum treatment exposure
  • Availability of the primary endpoint assessment
  • Major eligibility violations
  • Prohibited concomitant medications
  • Important treatment deviations
  • Timing of endpoint assessment

The definition should not be created after reviewing treatment-group outcomes.

Missing Data Can Be Particularly Important

Missing data are an important concern in non-inferiority trials because some methods of handling missing observations can make the treatment groups appear artificially similar.

Suppose participants with poor outcomes are disproportionately missing.

The observed difference could become smaller even though the underlying treatment difference is unfavorable.

Therefore, missing-data assumptions should be carefully considered during design rather than only after the database is locked.

Missing Data Sensitivity Analyses

A robust non-inferiority analysis may include sensitivity analyses under different assumptions about missing outcomes.

Examples include:

  • Multiple imputation under a prespecified missing-at-random model
  • Pattern-mixture approaches
  • Tipping-point analyses
  • Worst-case or conservative scenarios when clinically appropriate
  • Other prespecified sensitivity analyses appropriate to the endpoint
Design principle: If the non-inferiority conclusion changes dramatically under plausible missing data assumptions, the strength of the evidence should be interpreted accordingly.

Adherence and Treatment Switching

Adherence can be especially important in a non-inferiority trial.

If participants assigned to different treatments frequently switch treatments or receive inadequate exposure, the observed treatment contrast may be reduced.

That can make the treatments appear more similar.

Therefore, the protocol should define how treatment switching and major nonadherence will be handled in the primary and supportive analyses.

Why Non-Inferiority Trials Can Require More Patients

Non-inferiority trials can require substantial sample sizes.

This may seem counterintuitive because the trial is not trying to demonstrate a positive treatment difference.

The reason is that the trial must obtain enough precision to exclude a clinically unacceptable loss.

Suppose the estimated difference is:

$$ \hat{\Delta}=0 $$

That estimate is favorable.

But if the confidence interval is:

$$ 95\%\,CI=(-0.50,\;0.50) $$

the study cannot rule out substantial inferiority.

A larger sample reduces uncertainty and narrows the confidence interval.

Effect of the Non-Inferiority Margin on Sample Size

Consider the same continuous endpoint with:

$$ \sigma=1 \qquad \delta=0 $$

Suppose the design uses either:

$$ M=0.30 $$ or: $$ M=0.20 $$

The approximate sample-size ratio is:

$$ \frac{N_{0.20}}{N_{0.30}} \approx \left(\frac{0.30}{0.20}\right)^2 = 2.25 $$

Thus, reducing the margin from 0.30 to 0.20 can require roughly 2.25 times as many patients under these simplified assumptions.

Clinical implication: A seemingly small change in the margin can have a major operational impact on enrollment, study duration, cost, and feasibility.

Choice of Analysis Model

The statistical model should be aligned with the endpoint and estimand.

Examples include:

Endpoint Possible Analysis Framework
Continuous ANCOVA, linear model, mixed model, or related methods
Binary Risk difference, risk ratio, odds ratio, or regression-based methods
Time-to-event Cox model or other survival-analysis framework

The non-inferiority margin must be defined on the same effect scale used for the primary decision.

Covariate Adjustment

Covariate adjustment can improve precision when important baseline variables are strongly associated with the endpoint.

For example, a continuous primary endpoint may be analyzed using an ANCOVA model:

$$ Y_i = \beta_0+ \beta_1T_i+ \beta_2X_i+ \epsilon_i $$

where \(T_i\) indicates treatment assignment and \(X_i\) represents a prespecified baseline covariate.

The treatment effect can then be estimated from \(\beta_1\).

The non-inferiority decision is still based on the confidence interval relative to the prespecified margin.

Randomization Is Essential

Non-inferiority trials generally rely on randomization to ensure that the experimental treatment and active control groups are comparable.

Randomization helps prevent systematic differences in:

  • Baseline disease severity
  • Prognostic factors
  • Patient characteristics
  • Clinical management
  • Other sources of confounding

Without adequate randomization, differences in outcomes can be attributed to factors other than treatment.

Blinding Can Be Particularly Valuable

Blinding can help reduce differences in treatment administration, patient behavior, outcome assessment, and decisions about discontinuation.

The importance of blinding depends on the endpoint and intervention.

For subjective endpoints, maintaining blinding may be especially important.

Why Historical Evidence Must Be Credible

A non-inferiority trial frequently relies on historical evidence about the active control.

Suppose historical trials showed:

$$ \text{Control}>\text{Placebo} $$

But the historical studies used a very different population, endpoint, background therapy, dosing regimen, or standard of care.

The historical estimate may no longer accurately represent the treatment effect that would be expected in the current trial.

This creates uncertainty about the amount of efficacy the active control is actually expected to retain.

Historical evidence is part of the statistical design. The credibility of the non-inferiority conclusion depends not only on the current trial's statistical analysis but also on whether the assumptions supporting the active control's effect are credible.

Common Design Mistakes

  1. Choosing the margin for convenience. A margin should be clinically justified rather than selected primarily to reduce sample size.
  2. Confusing non-inferiority with equivalence. Non-inferiority excludes unacceptable inferiority; equivalence requires the entire confidence interval to fall within two bounds.
  3. Using a non-significant superiority test as evidence of non-inferiority. A failure to demonstrate superiority does not establish non-inferiority.
  4. Ignoring the direction of the endpoint. The treatment-effect definition must identify which direction favors the experimental treatment.
  5. Ignoring assay sensitivity. Similar outcomes between two treatments do not prove that both treatments are effective.
  6. Ignoring the constancy assumption. The historical effect of the active control must remain relevant to the current trial.
  7. Using only an ITT analysis. Non-inferiority trials commonly require careful consideration of both intention-to-treat and per-protocol populations.
  8. Defining the per-protocol population after seeing the results. Eligibility criteria should be prespecified.
  9. Underestimating missing-data effects. Missing outcomes can materially influence the treatment contrast.
  10. Changing the non-inferiority margin after seeing the data. The margin is a design parameter and should be fixed before unblinding treatment-group results.
  11. Assuming a favorable point estimate is sufficient. The confidence interval must exclude clinically unacceptable inferiority.

A Practical Non-Inferiority Design Workflow

1
Define the clinical objective of the new treatment.
2
Select the active control and establish why it is an appropriate comparator.
3
Define the primary endpoint and treatment-effect scale.
4
Define which direction of the effect favors the experimental treatment.
5
Review historical evidence supporting the active control's effect.
6
Establish the non-inferiority margin using clinical and historical evidence.
7
Specify the one-sided type I error and desired power.
8
Select the primary statistical model and estimand.
9
Calculate the required sample size.
10
Inflate for dropout, missing data, or other prespecified factors.
11
Prespecify ITT, per-protocol, and sensitivity analyses.
12
Define the final non-inferiority decision using the appropriate confidence interval.

Protocol Elements That Should Be Prespecified

A non-inferiority protocol should provide enough detail that the trial's statistical decision can be reproduced independently.

At minimum, specify:

  • Primary endpoint
  • Estimand
  • Treatment-effect scale
  • Direction of benefit
  • Active control
  • Non-inferiority margin
  • Historical evidence supporting the margin
  • Assumptions concerning the control effect
  • Type I error
  • Target power
  • Sample-size methodology
  • Primary analysis model
  • Confidence-interval method
  • Analysis populations
  • Per-protocol criteria
  • Handling of missing data
  • Sensitivity analyses
  • Treatment adherence rules
  • Treatment-switching rules

Interpreting a Non-Inferiority Result

Suppose the final analysis produces:

$$ \hat{\Delta}=-0.03 $$ and: $$ 95\%\,CI=(-0.11,\;0.05) $$

Suppose:

$$ M=0.15 $$

The lower confidence bound is:

$$ -0.11 $$

and the non-inferiority boundary is:

$$ -0.15 $$

Because:

$$ -0.11>-0.15 $$

non-inferiority is demonstrated.

However, the confidence interval crosses zero, so superiority is not demonstrated by this interval.

A More Difficult Example

Suppose:

$$ \hat{\Delta}=-0.03 $$ and: $$ 95\%\,CI=(-0.19,\;0.13) $$

The point estimate is still only -0.03.

But with:

$$ M=0.15 $$

the lower confidence bound is below the non-inferiority margin:

$$ -0.19<-0.15 $$

Therefore, non-inferiority is not demonstrated.

Lesson: The failure to demonstrate non-inferiority does not necessarily prove that the new treatment is inferior. It may instead indicate that the study was too imprecise to rule out an unacceptable loss.

Failure to Demonstrate Non-Inferiority Is Not Proof of Inferiority

This distinction is important.

Consider:

$$ 95\%\,CI=(-0.19,\;0.13) $$

with \(M=0.15\).

The confidence interval includes treatment differences worse than \(-0.15\), so non-inferiority cannot be concluded.

But the interval also includes zero and positive treatment differences.

The appropriate conclusion is therefore: non-inferiority was not demonstrated, rather than: the experimental treatment was proven inferior.

Superiority After Non-Inferiority

Suppose the confidence interval is:

$$ 95\%\,CI=(0.02,\;0.18) $$

The entire interval is above zero.

Therefore, superiority is demonstrated.

Because the entire interval is also above the non-inferiority boundary, non-inferiority is necessarily supported as well.

This illustrates the relationship:

$$ \text{Superiority} \quad\Rightarrow\quad \text{Non-inferiority} $$

when the superiority and non-inferiority analyses use compatible estimands, scales, and decision rules.

Margin Selection and Clinical Meaning

Suppose a new treatment has substantial advantages in administration.

For example, it might:

  • Require fewer doses
  • Reduce infusion time
  • Improve adherence
  • Reduce monitoring burden
  • Improve tolerability
  • Reduce treatment complexity

The clinical argument for accepting some efficacy loss may therefore be stronger than it would be for a treatment offering no meaningful benefit beyond the active control.

However, this does not mean that the margin should simply be widened until the trial becomes feasible.

The margin still needs to preserve an appropriate amount of established clinical benefit.

Non-Inferiority Margin Versus Clinically Relevant Difference

The non-inferiority margin should not automatically be confused with the smallest clinically important difference used in a superiority study.

These quantities answer different questions.

Quantity Purpose
MCID / clinically important difference Magnitude of benefit considered clinically meaningful
Non-inferiority margin Maximum acceptable loss of established treatment effect
Superiority effect assumption Expected treatment difference used for power calculations

Ratio-Scale Non-Inferiority Margins

For endpoints such as hazard ratios or risk ratios, the margin is often expressed on a ratio scale.

Suppose:

$$ HR_{NI}=1.25 $$

and the estimated hazard ratio is:

$$ \widehat{HR}=1.08 $$

with:

$$ 95\%\,CI=(0.94,\;1.19) $$

Because:

$$ 1.19<1.25 $$

non-inferiority is demonstrated.

The fact that the confidence interval crosses 1 does not prevent a non-inferiority conclusion.

Remember: On ratio scales, 1 is the no-effect value, while the non-inferiority margin may be above 1 or below 1 depending on the endpoint and direction of benefit.

Log-Scale Analysis

Ratio-scale treatment effects are often analyzed on the logarithmic scale.

For a hazard ratio:

$$ \log(HR) $$

the no-effect value becomes:

$$ \log(1)=0 $$

and a hazard-ratio margin \(HR_{NI}\) becomes:

$$ \log(HR_{NI}) $$

The confidence-interval decision can therefore be performed naturally on the log scale and transformed back to the hazard-ratio scale for reporting.

Sample Size and Power Are Not the Whole Design

A common mistake is to treat non-inferiority design as primarily a sample-size problem.

In reality, the sample-size calculation is downstream of several much more fundamental decisions:

1
What clinical benefit must the new treatment preserve?
2
How much loss of efficacy can reasonably be accepted?
3
What historical evidence supports the active control's effect?
4
Will the current trial retain assay sensitivity?
5
What estimand and analysis model will provide the treatment contrast?
6
How much precision is required to rule out unacceptable inferiority?

Regulatory and Statistical Perspective

Non-inferiority trials require particularly strong justification because the interpretation depends on assumptions extending beyond the randomized comparison itself.

The current trial provides a comparison between experimental treatment and active control.

But the conclusion that the experimental treatment is effective can depend on the assumption that the active control would have retained its established effect under the current trial conditions.

Consequently, trial design should carefully document:

  • The historical evidence for the control
  • Why the current population is comparable
  • Why the endpoint is comparable
  • Why treatment duration is comparable
  • Why dosing and adherence are appropriate
  • Why the non-inferiority margin is clinically justified

R Implementation: Continuous Endpoint Sample Size

The illustrative sample-size calculation can be reproduced in R.

alpha <- 0.025
power <- 0.90

sigma <- 1.0
delta <- 0.0
margin <- 0.30

z_alpha <- qnorm(1 - alpha)
z_beta  <- qnorm(power)

N <- 2 * sigma^2 *
     (z_alpha + z_beta)^2 /
     (delta + margin)^2

ceiling(N)

This produces approximately:

# approximately 234

before dropout inflation.

R Implementation: Dropout Inflation

dropout <- 0.10

N_inflated <- ceiling(
  N / (1 - dropout)
)

N_inflated

# approximately 260

This is a simple inflation and should be replaced by a method appropriate to the actual endpoint and planned analysis when more complex missing-data behavior is expected.

R Implementation: Confidence-Interval Decision

The non-inferiority decision can be represented directly in R.

estimate <- -0.08
lower_ci <- -0.21
upper_ci <- 0.05

margin <- 0.30

noninferior <- lower_ci > -margin

noninferior
# TRUE

For this example, the lower confidence bound is above \(-0.30\), so non-inferiority is demonstrated.

R Implementation: Exploring Different Margins

It is useful to examine how the required sample size changes as the non-inferiority margin changes.

sample_size_ni <- function(
  margin,
  sigma = 1,
  delta = 0,
  alpha = 0.025,
  power = 0.90
) {

  z_alpha <- qnorm(1 - alpha)
  z_beta  <- qnorm(power)

  N <- 2 * sigma^2 *
       (z_alpha + z_beta)^2 /
       (delta + margin)^2

  ceiling(N)
}

sample_size_ni(0.40)
sample_size_ni(0.30)
sample_size_ni(0.25)
sample_size_ni(0.20)
sample_size_ni(0.15)

This provides a simple sensitivity analysis demonstrating the operational consequences of different margins.

R Implementation: A Decision Function

ni_decision <- function(
  lower_ci,
  margin
) {

  if (lower_ci > -margin) {
    "Non-inferiority demonstrated"
  } else {
    "Non-inferiority not demonstrated"
  }
}

ni_decision(
  lower_ci = -0.21,
  margin = 0.30
)

Reporting the Result

A clear non-inferiority result should report:

  • The treatment-effect estimate
  • The confidence interval
  • The non-inferiority margin
  • The analysis population
  • The primary analysis method
  • The prespecified decision criterion
  • Supportive analyses

For example:

Example reporting language: The adjusted treatment difference was -0.08 units, with a 95% confidence interval of -0.21 to 0.05. Because the lower confidence limit was above the prespecified non-inferiority margin of -0.30, the criterion for non-inferiority was met.

What Should Be Reported Alongside Non-Inferiority?

A non-inferiority conclusion should not be presented without sufficient context.

Useful supporting information includes:

  • Observed treatment-group outcomes
  • Estimated treatment difference
  • Confidence interval
  • Non-inferiority margin
  • Adherence
  • Protocol deviations
  • Missing data
  • ITT analysis
  • Per-protocol analysis
  • Sensitivity analyses
  • Relevant safety outcomes

Safety Still Matters

Demonstrating non-inferior efficacy does not establish that a treatment has a favorable overall benefit-risk profile.

For example, suppose the experimental treatment is non-inferior in efficacy but causes substantially more serious adverse events.

The clinical interpretation cannot be based solely on the non-inferiority endpoint.

A complete benefit-risk assessment should consider:

  • Efficacy
  • Safety
  • Tolerability
  • Administration burden
  • Adherence
  • Patient preferences
  • Other clinically relevant outcomes

Common Misinterpretations

Misinterpretation Correct Interpretation
"The p-value was greater than 0.05, so the treatments are equivalent." A non-significant superiority test does not establish non-inferiority or equivalence.
"The point estimate favored the control, so the new treatment failed." The estimate can favor control while still demonstrating non-inferiority.
"The confidence interval crossed zero, so non-inferiority failed." Non-inferiority depends on the margin, not zero.
"Non-inferiority means the treatments are identical." It means sufficiently large inferiority has been excluded according to the margin.
"A wider margin is always better because it reduces sample size." The margin must preserve an appropriate amount of established benefit.
"Failure to demonstrate non-inferiority proves inferiority." Failure may reflect insufficient precision or other uncertainty.
"ITT alone is automatically the most conservative analysis." In non-inferiority trials, treatment dilution can make ITT analyses favor similarity.

Worked Example Summary

Component Value
Study type Randomized parallel-group non-inferiority trial
Endpoint Continuous efficacy endpoint
Direction Higher values are better
Expected treatment difference 0
Standard deviation 1.0
Non-inferiority margin 0.30
One-sided \(\alpha\) 0.025
Target power 90%
Initial total N Approximately 234
Dropout assumption 10%
Inflated N Approximately 260
Observed difference -0.08
95% CI (-0.21, 0.05)
Non-inferiority boundary -0.30
Conclusion Non-inferiority demonstrated

The Most Important Concept

The most important idea in a non-inferiority trial is simple: the goal is to rule out an unacceptable loss of efficacy, not to prove that the two treatments are identical.

The entire design follows from that principle.

First, investigators define the clinical effect that matters.

Then they establish how much of the active control's established benefit must be preserved.

That leads to the non-inferiority margin.

The sample size is then selected so that the study has sufficient precision to exclude inferiority beyond that margin.

Finally, the trial is interpreted using a confidence interval.

$$ \boxed{ L_{CI}>-M \quad\Rightarrow\quad \text{Non-inferiority demonstrated} } $$

For ratio-scale endpoints, the corresponding criterion is expressed on the appropriate ratio scale, such as:

$$ \boxed{ U_{CI}
Bottom line: A rigorous non-inferiority trial requires much more than a sample-size calculation. The investigator must define a clinically defensible non-inferiority margin, establish a credible active control effect, preserve assay sensitivity, consider the constancy assumption, prespecify the estimand and analysis model, account carefully for adherence and missing data, and use the appropriate confidence-interval boundary for the final decision. A result that fails to demonstrate non-inferiority is not automatically evidence of inferiority, while a non-significant superiority test is not evidence of non-inferiority. The central statistical question is whether clinically unacceptable inferiority can be ruled out with adequate confidence.

References

ICH E9. Statistical Principles for Clinical Trials. International Council for Harmonisation.
ICH E10. Choice of Control Group and Related Issues in Clinical Trials. International Council for Harmonisation.
ICH E9(R1). Addendum on Estimands and Sensitivity Analysis in Clinical Trials. International Council for Harmonisation.
Piaggio, G., Elbourne, D.R., Pocock, S.J., Evans, S.J.W. & Altman, D.G. (2012). Reporting of noninferiority and equivalence randomized trials: extension of the CONSORT 2010 statement. JAMA, 308(24), 2594–2604.
Schuirmann, D.J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15, 657–680.
Blackwelder, W.C. (1982). Proving the null hypothesis in clinical trials. Controlled Clinical Trials, 3(4), 345–353.
Committee for Proprietary Medicinal Products. Points to Consider on Switching Between Superiority and Non-Inferiority. European Medicines Agency.

Clinical Trials

See these methods in real clinical trials

See the method applied to published trial results, with the estimates, confidence intervals and interpretation explained.

ENGAGE AF-TIMI 48
Independent statistical analysis of ENGAGE AF-TIMI 48, evaluating edoxaban versus warfarin for stroke and systemic embolic events in patients with atrial fibrillation.
Phase 3 · n = 21,105
PRoFESS
An independent statistical analysis of PRoFESS (NCT00153062), including its randomized double-blind design, recurrent-stroke endpoints, Cox proportional-hazards analyses, non-inferiority framework, hazard ratios, confidence…
Phase 4 · n = 20,332
DISCOVER
Independent statistical analysis of the DISCOVER phase 3 trial evaluating F/TAF versus F/TDF for HIV-1 pre-exposure prophylaxis, including non-inferiority testing, rate-ratio analysis,…
Phase 3 · n = 5,399
FLAME
Independent statistical analysis of the FLAME phase 3 trial of QVA149 versus LABA/ICS in COPD, including exacerbation rates, time-to-event analyses, non-inferiority methodology,…
Phase 3 · n = 3,362
SWOG 9346
Independent statistical analysis of SWOG 9346, a randomized phase 3 trial of intermittent versus continuous hormonal therapy in men with stage IV…
Phase 3 · n = 3,040
ACTG A5279
Independent statistical analysis of ACTG A5279 (NCT01404312), a randomized phase 3 trial comparing a rifapentine-plus-isoniazid regimen with an isoniazid regimen for tuberculosis…
Phase 3 · n = 3,000
See all 60 trials using non-inferiority trial design →