Introduction
Many clinical trials are designed to demonstrate that a new treatment is better than an existing treatment.
But superiority is not always the objective.
A new treatment may offer important advantages such as improved tolerability, simpler administration, lower cost, improved adherence, reduced monitoring, shorter treatment duration, or easier storage while providing efficacy that is acceptably close to that of an established active treatment.
In these situations, a non-inferiority trial may be appropriate.
The central question is not: "Is the experimental treatment statistically different from the control?"
Instead, the question is: "Can we rule out a clinically unacceptable loss of efficacy?"
Superiority Versus Non-Inferiority
The distinction becomes clearer by comparing the questions asked by the two designs.
| Design | Primary Question |
|---|---|
| Superiority | Is the experimental treatment better than the comparator? |
| Non-inferiority | Is the experimental treatment not unacceptably worse than the comparator? |
A superiority trial is generally concerned with detecting a positive treatment difference.
A non-inferiority trial is concerned with excluding sufficiently large negative differences.
The Basic Setup
Suppose a new treatment, \(T\), is compared with an established active control, \(C\).
Let the treatment effect be defined as:
where \(\theta\) represents the appropriate measure of efficacy.
For a continuous endpoint, \(\theta\) might be a mean. For a binary endpoint, it might be a response probability. For a time-to-event endpoint, it might be a log hazard ratio.
The direction of the effect must be defined carefully because the interpretation of the non-inferiority margin depends on whether larger or smaller values are better.
What Is the Non-Inferiority Margin?
The most important quantity in a non-inferiority trial is the non-inferiority margin, usually denoted by \(M\) or \(\Delta_{NI}\).
The margin represents the largest loss of treatment effect that can be accepted while still considering the experimental treatment clinically acceptable for the trial's stated purpose.
Suppose larger values of \(\theta\) are better.
Then the experimental treatment is considered non-inferior if its effect is not too far below that of the control.
and non-inferiority is demonstrated when:
where \(M>0\) is the prespecified non-inferiority margin.
The margin therefore creates a boundary between an acceptable loss and an unacceptable loss.
Why the Margin Is a Clinical Quantity
The non-inferiority margin should not simply be chosen because it produces a convenient sample size.
Suppose an established treatment has a large and clinically important benefit over placebo.
A new treatment should not be allowed to lose nearly all of that benefit merely because a statistical calculation produces a manageable enrollment target.
The margin must therefore be justified using clinical evidence about the established treatment's effect and the amount of benefit that must be preserved.
Non-Inferiority Hypotheses
For the convention where larger values of \(\theta\) are better, the hypotheses can be written as:
The null hypothesis says that the experimental treatment is inferior to the control by at least the non-inferiority margin.
The alternative says that the treatment difference is sufficiently close to zero that unacceptable inferiority can be ruled out.
This is fundamentally different from a superiority hypothesis:
In a non-inferiority trial, the null boundary is shifted from zero to the non-inferiority margin.
A Visual Interpretation
| Treatment Difference \(\Delta\) | Interpretation |
|---|---|
| \(\Delta\le -M\) | Unacceptably inferior |
| \(-M<\Delta<0\) | Some efficacy loss, but within the non-inferiority margin |
| \(\Delta=0\) | No estimated treatment difference |
| \(\Delta>0\) | Experimental treatment favors efficacy |
The critical distinction is between statistical equality and acceptable inferiority.
A non-inferiority trial does not need the estimated treatment difference to be zero.
The estimate may favor the control and still support non-inferiority, provided the confidence interval excludes losses larger than the prespecified margin.
The Confidence-Interval Decision Rule
One of the cleanest ways to interpret a non-inferiority trial is through a confidence interval.
Suppose the estimated treatment difference is:
Construct a two-sided \(100(1-2\alpha)\%\) confidence interval for \(\Delta\). Equivalently, this corresponds to a one-sided test at level \(\alpha\).
For a conventional one-sided \(\alpha=0.025\) non-inferiority test, this is typically represented by a 95% two-sided confidence interval.
Non-inferiority is concluded when the lower confidence bound exceeds the negative margin:
where \(L_{CI}\) is the lower confidence limit.
Three Examples of Confidence-Interval Interpretation
Example A: Non-Inferiority Clearly Demonstrated
Suppose:
If:
then:
Therefore, non-inferiority is demonstrated.
Example B: Non-Inferiority Not Demonstrated
Suppose:
With \(M=0.10\):
The confidence interval includes treatment differences worse than the non-inferiority margin.
Therefore, non-inferiority is not established.
Example C: The Estimate Favors the Control but Non-Inferiority Holds
Suppose:
The estimate favors the control, but if \(M=0.10\), the lower bound remains above \(-0.10\).
Therefore, the experimental treatment can still be declared non-inferior.
Non-Inferiority Does Not Mean Equivalence
Non-inferiority and equivalence are related but distinct concepts.
A non-inferiority trial generally asks whether the experimental treatment is not too much worse than the comparator.
An equivalence trial asks whether the treatment difference lies within both a lower and upper equivalence margin.
For an equivalence margin \(M\), the hypotheses are conceptually:
Thus, equivalence requires ruling out both sufficiently large inferiority and sufficiently large superiority.
| Design | Question | Confidence-Interval Criterion |
|---|---|---|
| Superiority | Is \(T\) better than \(C\)? | Lower bound above 0 |
| Non-inferiority | Is \(T\) not unacceptably worse? | Lower bound above \(-M\) |
| Equivalence | Are \(T\) and \(C\) sufficiently similar? | Entire CI within \((-M,+M)\) |
Why an Active Control Is Usually Central
A non-inferiority trial commonly compares the experimental treatment with an established active treatment rather than placebo.
The purpose is to determine whether the experimental treatment preserves an acceptable amount of the established treatment's effect.
This creates a fundamental requirement: the active control must have demonstrated efficacy under circumstances relevant to the current trial.
If the active control performs poorly for reasons unrelated to the experimental treatment, an apparently non-inferior experimental treatment may simply be non-effective.
The Problem of Assay Sensitivity
Assay sensitivity refers to the ability of a trial to distinguish an effective treatment from an ineffective treatment under the trial's conditions.
This concept is especially important in non-inferiority trials.
Consider a trial in which both treatments perform poorly.
If the experimental treatment and active control have similar poor outcomes, the trial may conclude non-inferiority because their difference is small.
That does not necessarily mean the experimental treatment is effective.
The Constancy Assumption
The constancy assumption is another foundational concept.
Suppose historical trials established that the active control had a substantial benefit over placebo.
The current non-inferiority trial assumes that an appropriate amount of that benefit would still be present under the current trial conditions.
If the historical control effect was:
then the non-inferiority design relies, directly or indirectly, on the assumption that the control would still retain a clinically meaningful effect in the current setting.
This is one reason historical evidence, trial population, endpoint definitions, background therapy, dosing, adherence, and study conduct matter so much in non-inferiority design.
Preserving a Fraction of the Control Effect
Suppose historical evidence suggests that the active control has a benefit of \(E\) over placebo.
A common conceptual approach is to require that the experimental treatment preserve a specified fraction of this effect.
For example, if the control effect is:
and the design requires preservation of 50% of the control effect, then the largest allowable loss would be:
This illustrates the logic, but the actual derivation of a non-inferiority margin is more nuanced and depends on the endpoint, historical evidence, clinical judgment, and regulatory/statistical considerations.
Choosing the Endpoint
Non-inferiority design is highly sensitive to the choice and definition of the primary endpoint.
The endpoint should be:
- Clinically meaningful
- Measured consistently across treatment groups
- Supported by reliable historical evidence
- Appropriate for the active control
- Defined identically across the trial
- Available with sufficient completeness
A change in endpoint definition can undermine the historical evidence used to justify the non-inferiority margin.
Endpoint Direction Matters
Suppose higher values indicate better outcomes.
Then:
and the non-inferiority condition is:
But suppose lower values indicate better outcomes, as can occur with some biomarkers, symptom scores, or event rates.
The effect definition may instead be:
so that a positive value favors the experimental treatment.
The corresponding non-inferiority decision must then be written consistently.
Non-Inferiority With Binary Endpoints
Suppose the primary endpoint is response.
Let:
Define the risk-difference treatment effect as:
If higher response rates are better, then non-inferiority is:
The corresponding confidence-interval criterion is:
Non-Inferiority With Continuous Endpoints
For a continuous endpoint, suppose the treatment effect is defined as:
where \(\mu_T\) and \(\mu_C\) are the treatment-group means.
If higher values indicate better outcomes, the non-inferiority criterion remains:
The major difference is the statistical model used to estimate the treatment difference and its standard error.
Non-Inferiority With Time-to-Event Endpoints
For a time-to-event endpoint, the treatment effect is often represented by a hazard ratio.
Suppose the hazard ratio is defined as:
where lower hazard is better.
In this setting, a hazard ratio of 1 indicates no treatment difference, while values above 1 indicate a higher hazard for the experimental treatment.
The non-inferiority margin is therefore naturally expressed on the hazard-ratio scale.
For example:
would represent the largest acceptable relative increase in hazard under the specified design.
Non-inferiority would then be demonstrated if the upper confidence bound for the hazard ratio is below the margin:
One-Sided Versus Two-Sided Testing
Non-inferiority is fundamentally a one-sided testing problem.
The relevant question is whether the experimental treatment could be unacceptably worse than the control.
For a one-sided significance level \(\alpha\), the corresponding confidence interval is commonly a two-sided \(100(1-2\alpha)\%\) interval.
For example:
| One-Sided \(\alpha\) | Equivalent Two-Sided CI |
|---|---|
| 0.05 | 90% |
| 0.025 | 95% |
| 0.01 | 98% |
Thus, a one-sided 2.5% non-inferiority test corresponds to evaluating whether the lower bound of a 95% two-sided confidence interval is above the non-inferiority margin.
Sample Size for a Non-Inferiority Trial
The sample-size calculation depends on:
- The non-inferiority margin \(M\)
- The expected treatment effect
- The variability of the endpoint
- The type I error
- The desired power
- The allocation ratio
- The statistical test or model
- The expected dropout or missing-data rate
The margin is particularly influential.
A smaller non-inferiority margin means that the trial must distinguish smaller differences between treatments.
That generally requires a larger sample size.
Sample Size for a Continuous Endpoint
Consider a parallel-group trial with equal allocation and a continuous endpoint.
Suppose the expected treatment difference is:
and the non-inferiority margin is \(M\).
The quantity that determines how far the assumed treatment effect lies from the non-inferiority boundary is:
when larger values are better and the null boundary is \(-M\).
For equal allocation, a commonly used normal-approximation formula for total sample size is:
where \(\sigma\) is the common standard deviation.
The exact formula depends on the planned analysis and variance assumptions.
Why the Margin Has Such a Large Effect
Notice that the margin appears in the denominator:
If the expected treatment difference is near zero, then a smaller \(M\) directly reduces the distance from the non-inferiority boundary.
For example, suppose:
and compare:
Ignoring all other changes, halving the margin approximately quadruples the required sample size because the denominator is squared.
A Complete Worked Example
Consider a randomized, parallel-group non-inferiority trial comparing a new oral treatment with an established active treatment.
The primary endpoint is a continuous efficacy measure, with larger values representing better outcomes.
Suppose the study assumptions are:
| Parameter | Planning Value |
|---|---|
| Expected treatment difference \(\delta\) | 0 |
| Common standard deviation \(\sigma\) | 1.0 |
| Non-inferiority margin \(M\) | 0.30 |
| One-sided type I error | 2.5% |
| Target power | 90% |
| Allocation | 1:1 |
The hypotheses are:
Step 1: Determine the Distance From the Margin
The assumed treatment difference is:
The non-inferiority margin is:
Therefore:
This means the assumed treatment effect is 0.30 units above the non-inferiority boundary.
Step 2: Determine the Critical Values
For a one-sided \(\alpha=0.025\) test:
For 90% power:
Therefore:
Step 3: Calculate the Initial Sample Size
Using:
This gives approximately:
Therefore, an initial planning value would be approximately:
or approximately 117 patients per treatment group under equal allocation.
Step 4: Inflate for Dropout
Suppose the study anticipates a 10% dropout rate.
A simple inflation is:
which gives:
Thus, approximately 260 randomized patients would be required under this illustrative assumption.
With equal allocation, this corresponds to approximately:
Step 5: Define the Analysis
At the end of the trial, suppose the estimated treatment difference is:
and the 95% confidence interval is:
The lower confidence bound is:
The non-inferiority margin is:
Because:
the lower confidence bound is above the non-inferiority margin.
Therefore:
What If the Confidence Interval Were Wider?
Suppose instead that the estimated difference were identical:
but the confidence interval were:
The lower bound would be:
which is below the margin:
Therefore, non-inferiority would not be demonstrated.
Non-Inferiority Is Not Established by a Non-Significant Superiority Test
This is one of the most important mistakes to avoid.
Suppose a conventional superiority test produces:
That does not establish non-inferiority.
A non-significant superiority test simply means that the trial did not establish a statistically significant difference in the superiority direction.
The confidence interval could still include clinically important inferiority.
For example:
This interval includes zero, so superiority is not established. But if:
the interval also extends below \(-0.30\), so non-inferiority is not established either.
The Four Possible Interpretations
A useful way to understand the results is to consider the relationship between the confidence interval and the two important boundaries: the non-inferiority margin and zero.
| 95% CI | Interpretation |
|---|---|
| Entirely above 0 | Superiority demonstrated; non-inferiority necessarily also supported |
| Crosses 0 but remains above \(-M\) | Non-inferiority demonstrated; superiority not demonstrated |
| Extends below \(-M\) and crosses 0 | Neither superiority nor non-inferiority demonstrated |
| Entirely below \(-M\) | Inferiority demonstrated relative to the prespecified margin |
This framework is particularly useful when explaining results to clinical teams.
Non-Inferiority Followed by Superiority
Sometimes a study is designed first to establish non-inferiority and then, if non-inferiority is demonstrated, to investigate superiority.
For example:
This hierarchical approach can be particularly useful when the scientific question is whether the experimental treatment retains acceptable efficacy and may additionally provide superior efficacy.
However, the testing sequence and multiplicity strategy must be prespecified.
Analysis Populations Matter More in Non-Inferiority Trials
In superiority trials, departures from the protocol often tend to dilute differences between treatment groups.
In non-inferiority trials, that dilution can work in an undesirable direction.
If protocol deviations make the treatments look more similar, an analysis that includes those deviations can potentially make it easier to conclude non-inferiority.
For this reason, non-inferiority trials commonly place substantial emphasis on both:
- The intention-to-treat or full-analysis population
- The per-protocol population
Intention-to-Treat Analysis
The intention-to-treat principle generally analyzes participants according to their randomized treatment assignment.
This preserves the benefits of randomization.
However, in a non-inferiority trial, treatment discontinuation, crossover, nonadherence, and other deviations may reduce the contrast between treatments.
Consequently, the ITT analysis should not automatically be assumed to be the only informative analysis for demonstrating non-inferiority.
Per-Protocol Analysis
A per-protocol analysis attempts to estimate the treatment effect among participants who sufficiently adhered to the protocol's important requirements.
The exact definition must be prespecified.
For example, the protocol may specify criteria involving:
- Minimum treatment exposure
- Availability of the primary endpoint assessment
- Major eligibility violations
- Prohibited concomitant medications
- Important treatment deviations
- Timing of endpoint assessment
The definition should not be created after reviewing treatment-group outcomes.
Missing Data Can Be Particularly Important
Missing data are an important concern in non-inferiority trials because some methods of handling missing observations can make the treatment groups appear artificially similar.
Suppose participants with poor outcomes are disproportionately missing.
The observed difference could become smaller even though the underlying treatment difference is unfavorable.
Therefore, missing-data assumptions should be carefully considered during design rather than only after the database is locked.
Missing Data Sensitivity Analyses
A robust non-inferiority analysis may include sensitivity analyses under different assumptions about missing outcomes.
Examples include:
- Multiple imputation under a prespecified missing-at-random model
- Pattern-mixture approaches
- Tipping-point analyses
- Worst-case or conservative scenarios when clinically appropriate
- Other prespecified sensitivity analyses appropriate to the endpoint
Adherence and Treatment Switching
Adherence can be especially important in a non-inferiority trial.
If participants assigned to different treatments frequently switch treatments or receive inadequate exposure, the observed treatment contrast may be reduced.
That can make the treatments appear more similar.
Therefore, the protocol should define how treatment switching and major nonadherence will be handled in the primary and supportive analyses.
Why Non-Inferiority Trials Can Require More Patients
Non-inferiority trials can require substantial sample sizes.
This may seem counterintuitive because the trial is not trying to demonstrate a positive treatment difference.
The reason is that the trial must obtain enough precision to exclude a clinically unacceptable loss.
Suppose the estimated difference is:
That estimate is favorable.
But if the confidence interval is:
the study cannot rule out substantial inferiority.
A larger sample reduces uncertainty and narrows the confidence interval.
Effect of the Non-Inferiority Margin on Sample Size
Consider the same continuous endpoint with:
Suppose the design uses either:
The approximate sample-size ratio is:
Thus, reducing the margin from 0.30 to 0.20 can require roughly 2.25 times as many patients under these simplified assumptions.
Choice of Analysis Model
The statistical model should be aligned with the endpoint and estimand.
Examples include:
| Endpoint | Possible Analysis Framework |
|---|---|
| Continuous | ANCOVA, linear model, mixed model, or related methods |
| Binary | Risk difference, risk ratio, odds ratio, or regression-based methods |
| Time-to-event | Cox model or other survival-analysis framework |
The non-inferiority margin must be defined on the same effect scale used for the primary decision.
Covariate Adjustment
Covariate adjustment can improve precision when important baseline variables are strongly associated with the endpoint.
For example, a continuous primary endpoint may be analyzed using an ANCOVA model:
where \(T_i\) indicates treatment assignment and \(X_i\) represents a prespecified baseline covariate.
The treatment effect can then be estimated from \(\beta_1\).
The non-inferiority decision is still based on the confidence interval relative to the prespecified margin.
Randomization Is Essential
Non-inferiority trials generally rely on randomization to ensure that the experimental treatment and active control groups are comparable.
Randomization helps prevent systematic differences in:
- Baseline disease severity
- Prognostic factors
- Patient characteristics
- Clinical management
- Other sources of confounding
Without adequate randomization, differences in outcomes can be attributed to factors other than treatment.
Blinding Can Be Particularly Valuable
Blinding can help reduce differences in treatment administration, patient behavior, outcome assessment, and decisions about discontinuation.
The importance of blinding depends on the endpoint and intervention.
For subjective endpoints, maintaining blinding may be especially important.
Why Historical Evidence Must Be Credible
A non-inferiority trial frequently relies on historical evidence about the active control.
Suppose historical trials showed:
But the historical studies used a very different population, endpoint, background therapy, dosing regimen, or standard of care.
The historical estimate may no longer accurately represent the treatment effect that would be expected in the current trial.
This creates uncertainty about the amount of efficacy the active control is actually expected to retain.
Common Design Mistakes
- Choosing the margin for convenience. A margin should be clinically justified rather than selected primarily to reduce sample size.
- Confusing non-inferiority with equivalence. Non-inferiority excludes unacceptable inferiority; equivalence requires the entire confidence interval to fall within two bounds.
- Using a non-significant superiority test as evidence of non-inferiority. A failure to demonstrate superiority does not establish non-inferiority.
- Ignoring the direction of the endpoint. The treatment-effect definition must identify which direction favors the experimental treatment.
- Ignoring assay sensitivity. Similar outcomes between two treatments do not prove that both treatments are effective.
- Ignoring the constancy assumption. The historical effect of the active control must remain relevant to the current trial.
- Using only an ITT analysis. Non-inferiority trials commonly require careful consideration of both intention-to-treat and per-protocol populations.
- Defining the per-protocol population after seeing the results. Eligibility criteria should be prespecified.
- Underestimating missing-data effects. Missing outcomes can materially influence the treatment contrast.
- Changing the non-inferiority margin after seeing the data. The margin is a design parameter and should be fixed before unblinding treatment-group results.
- Assuming a favorable point estimate is sufficient. The confidence interval must exclude clinically unacceptable inferiority.
A Practical Non-Inferiority Design Workflow
Protocol Elements That Should Be Prespecified
A non-inferiority protocol should provide enough detail that the trial's statistical decision can be reproduced independently.
At minimum, specify:
- Primary endpoint
- Estimand
- Treatment-effect scale
- Direction of benefit
- Active control
- Non-inferiority margin
- Historical evidence supporting the margin
- Assumptions concerning the control effect
- Type I error
- Target power
- Sample-size methodology
- Primary analysis model
- Confidence-interval method
- Analysis populations
- Per-protocol criteria
- Handling of missing data
- Sensitivity analyses
- Treatment adherence rules
- Treatment-switching rules
Interpreting a Non-Inferiority Result
Suppose the final analysis produces:
Suppose:
The lower confidence bound is:
and the non-inferiority boundary is:
Because:
non-inferiority is demonstrated.
However, the confidence interval crosses zero, so superiority is not demonstrated by this interval.
A More Difficult Example
Suppose:
The point estimate is still only -0.03.
But with:
the lower confidence bound is below the non-inferiority margin:
Therefore, non-inferiority is not demonstrated.
Failure to Demonstrate Non-Inferiority Is Not Proof of Inferiority
This distinction is important.
Consider:
with \(M=0.15\).
The confidence interval includes treatment differences worse than \(-0.15\), so non-inferiority cannot be concluded.
But the interval also includes zero and positive treatment differences.
The appropriate conclusion is therefore: non-inferiority was not demonstrated, rather than: the experimental treatment was proven inferior.
Superiority After Non-Inferiority
Suppose the confidence interval is:
The entire interval is above zero.
Therefore, superiority is demonstrated.
Because the entire interval is also above the non-inferiority boundary, non-inferiority is necessarily supported as well.
This illustrates the relationship:
when the superiority and non-inferiority analyses use compatible estimands, scales, and decision rules.
Margin Selection and Clinical Meaning
Suppose a new treatment has substantial advantages in administration.
For example, it might:
- Require fewer doses
- Reduce infusion time
- Improve adherence
- Reduce monitoring burden
- Improve tolerability
- Reduce treatment complexity
The clinical argument for accepting some efficacy loss may therefore be stronger than it would be for a treatment offering no meaningful benefit beyond the active control.
However, this does not mean that the margin should simply be widened until the trial becomes feasible.
The margin still needs to preserve an appropriate amount of established clinical benefit.
Non-Inferiority Margin Versus Clinically Relevant Difference
The non-inferiority margin should not automatically be confused with the smallest clinically important difference used in a superiority study.
These quantities answer different questions.
| Quantity | Purpose |
|---|---|
| MCID / clinically important difference | Magnitude of benefit considered clinically meaningful |
| Non-inferiority margin | Maximum acceptable loss of established treatment effect |
| Superiority effect assumption | Expected treatment difference used for power calculations |
Ratio-Scale Non-Inferiority Margins
For endpoints such as hazard ratios or risk ratios, the margin is often expressed on a ratio scale.
Suppose:
and the estimated hazard ratio is:
with:
Because:
non-inferiority is demonstrated.
The fact that the confidence interval crosses 1 does not prevent a non-inferiority conclusion.
Log-Scale Analysis
Ratio-scale treatment effects are often analyzed on the logarithmic scale.
For a hazard ratio:
the no-effect value becomes:
and a hazard-ratio margin \(HR_{NI}\) becomes:
The confidence-interval decision can therefore be performed naturally on the log scale and transformed back to the hazard-ratio scale for reporting.
Sample Size and Power Are Not the Whole Design
A common mistake is to treat non-inferiority design as primarily a sample-size problem.
In reality, the sample-size calculation is downstream of several much more fundamental decisions:
Regulatory and Statistical Perspective
Non-inferiority trials require particularly strong justification because the interpretation depends on assumptions extending beyond the randomized comparison itself.
The current trial provides a comparison between experimental treatment and active control.
But the conclusion that the experimental treatment is effective can depend on the assumption that the active control would have retained its established effect under the current trial conditions.
Consequently, trial design should carefully document:
- The historical evidence for the control
- Why the current population is comparable
- Why the endpoint is comparable
- Why treatment duration is comparable
- Why dosing and adherence are appropriate
- Why the non-inferiority margin is clinically justified
R Implementation: Continuous Endpoint Sample Size
The illustrative sample-size calculation can be reproduced in R.
alpha <- 0.025
power <- 0.90
sigma <- 1.0
delta <- 0.0
margin <- 0.30
z_alpha <- qnorm(1 - alpha)
z_beta <- qnorm(power)
N <- 2 * sigma^2 *
(z_alpha + z_beta)^2 /
(delta + margin)^2
ceiling(N)
This produces approximately:
# approximately 234
before dropout inflation.
R Implementation: Dropout Inflation
dropout <- 0.10 N_inflated <- ceiling( N / (1 - dropout) ) N_inflated # approximately 260
This is a simple inflation and should be replaced by a method appropriate to the actual endpoint and planned analysis when more complex missing-data behavior is expected.
R Implementation: Confidence-Interval Decision
The non-inferiority decision can be represented directly in R.
estimate <- -0.08 lower_ci <- -0.21 upper_ci <- 0.05 margin <- 0.30 noninferior <- lower_ci > -margin noninferior # TRUE
For this example, the lower confidence bound is above \(-0.30\), so non-inferiority is demonstrated.
R Implementation: Exploring Different Margins
It is useful to examine how the required sample size changes as the non-inferiority margin changes.
sample_size_ni <- function(
margin,
sigma = 1,
delta = 0,
alpha = 0.025,
power = 0.90
) {
z_alpha <- qnorm(1 - alpha)
z_beta <- qnorm(power)
N <- 2 * sigma^2 *
(z_alpha + z_beta)^2 /
(delta + margin)^2
ceiling(N)
}
sample_size_ni(0.40)
sample_size_ni(0.30)
sample_size_ni(0.25)
sample_size_ni(0.20)
sample_size_ni(0.15)
This provides a simple sensitivity analysis demonstrating the operational consequences of different margins.
R Implementation: A Decision Function
ni_decision <- function(
lower_ci,
margin
) {
if (lower_ci > -margin) {
"Non-inferiority demonstrated"
} else {
"Non-inferiority not demonstrated"
}
}
ni_decision(
lower_ci = -0.21,
margin = 0.30
)
Reporting the Result
A clear non-inferiority result should report:
- The treatment-effect estimate
- The confidence interval
- The non-inferiority margin
- The analysis population
- The primary analysis method
- The prespecified decision criterion
- Supportive analyses
For example:
What Should Be Reported Alongside Non-Inferiority?
A non-inferiority conclusion should not be presented without sufficient context.
Useful supporting information includes:
- Observed treatment-group outcomes
- Estimated treatment difference
- Confidence interval
- Non-inferiority margin
- Adherence
- Protocol deviations
- Missing data
- ITT analysis
- Per-protocol analysis
- Sensitivity analyses
- Relevant safety outcomes
Safety Still Matters
Demonstrating non-inferior efficacy does not establish that a treatment has a favorable overall benefit-risk profile.
For example, suppose the experimental treatment is non-inferior in efficacy but causes substantially more serious adverse events.
The clinical interpretation cannot be based solely on the non-inferiority endpoint.
A complete benefit-risk assessment should consider:
- Efficacy
- Safety
- Tolerability
- Administration burden
- Adherence
- Patient preferences
- Other clinically relevant outcomes
Common Misinterpretations
| Misinterpretation | Correct Interpretation |
|---|---|
| "The p-value was greater than 0.05, so the treatments are equivalent." | A non-significant superiority test does not establish non-inferiority or equivalence. |
| "The point estimate favored the control, so the new treatment failed." | The estimate can favor control while still demonstrating non-inferiority. |
| "The confidence interval crossed zero, so non-inferiority failed." | Non-inferiority depends on the margin, not zero. |
| "Non-inferiority means the treatments are identical." | It means sufficiently large inferiority has been excluded according to the margin. |
| "A wider margin is always better because it reduces sample size." | The margin must preserve an appropriate amount of established benefit. |
| "Failure to demonstrate non-inferiority proves inferiority." | Failure may reflect insufficient precision or other uncertainty. |
| "ITT alone is automatically the most conservative analysis." | In non-inferiority trials, treatment dilution can make ITT analyses favor similarity. |
Worked Example Summary
| Component | Value |
|---|---|
| Study type | Randomized parallel-group non-inferiority trial |
| Endpoint | Continuous efficacy endpoint |
| Direction | Higher values are better |
| Expected treatment difference | 0 |
| Standard deviation | 1.0 |
| Non-inferiority margin | 0.30 |
| One-sided \(\alpha\) | 0.025 |
| Target power | 90% |
| Initial total N | Approximately 234 |
| Dropout assumption | 10% |
| Inflated N | Approximately 260 |
| Observed difference | -0.08 |
| 95% CI | (-0.21, 0.05) |
| Non-inferiority boundary | -0.30 |
| Conclusion | Non-inferiority demonstrated |
The Most Important Concept
The most important idea in a non-inferiority trial is simple: the goal is to rule out an unacceptable loss of efficacy, not to prove that the two treatments are identical.
The entire design follows from that principle.
First, investigators define the clinical effect that matters.
Then they establish how much of the active control's established benefit must be preserved.
That leads to the non-inferiority margin.
The sample size is then selected so that the study has sufficient precision to exclude inferiority beyond that margin.
Finally, the trial is interpreted using a confidence interval.
For ratio-scale endpoints, the corresponding criterion is expressed on the appropriate ratio scale, such as:
References
ICH E9.
Statistical Principles for Clinical Trials.
International Council for Harmonisation.
ICH E10.
Choice of Control Group and Related Issues in Clinical Trials.
International Council for Harmonisation.
ICH E9(R1).
Addendum on Estimands and Sensitivity Analysis in Clinical Trials.
International Council for Harmonisation.
Piaggio, G., Elbourne, D.R., Pocock, S.J., Evans, S.J.W. & Altman, D.G.
(2012).
Reporting of noninferiority and equivalence randomized trials:
extension of the CONSORT 2010 statement.
JAMA, 308(24), 2594–2604.
Schuirmann, D.J. (1987).
A comparison of the two one-sided tests procedure and the power
approach for assessing the equivalence of average bioavailability.
Journal of Pharmacokinetics and Biopharmaceutics, 15, 657–680.
Blackwelder, W.C. (1982).
Proving the null hypothesis in clinical trials.
Controlled Clinical Trials, 3(4), 345–353.
Committee for Proprietary Medicinal Products.
Points to Consider on Switching Between Superiority and
Non-Inferiority.
European Medicines Agency.
See these methods in real clinical trials
See the method applied to published trial results, with the estimates, confidence intervals and interpretation explained.