Introduction
Biosimilar development is fundamentally a comparative exercise. The objective is not to demonstrate that the proposed product is molecularly identical to the reference product.
Instead, the developer must build a structured body of evidence showing that the proposed product is highly similar to the reference product and that any observed differences do not translate into clinically meaningful differences.
This creates a distinctive statistical problem.
Traditional clinical development often asks whether a treatment is better than placebo or whether one treatment is non-inferior to another.
Biosimilar development instead asks whether the remaining uncertainty between two biological products is sufficiently small, and sufficiently understood, to support a conclusion of biosimilarity.
FDA's current framework evaluates evidence from comparisons of the proposed biosimilar and reference product in the context of the reference product's previously established safety and effectiveness.
EMA similarly describes biosimilar development as a comprehensive, stepwise comparability exercise intended to demonstrate high similarity and exclude clinically meaningful differences.
What Is Biosimilarity?
A biosimilar is a biological medicinal product that contains a version of the active substance of an already authorized reference biological product.
Because biological products are produced using living systems and are structurally complex, some degree of natural or manufacturing-related variability is expected.
Consequently, the statistical question is not:
Instead, the scientific objective is to determine whether any observed differences are inconsistent with the level of similarity expected to be clinically acceptable.
The Biosimilar Evidence Hierarchy
The development program is generally organized as a sequence of increasingly clinically integrated comparisons.
This is not simply a sequence of independent hypothesis tests.
Evidence generated at an earlier stage influences the amount and nature of evidence needed at later stages.
Why Statistical Assessment Is Different From a Generic Equivalence Trial
A conventional equivalence trial may focus primarily on one clinical endpoint.
A biosimilar program can contain hundreds of analytical measurements across multiple quality attributes, several functional assays, PK endpoints, immunogenicity variables, and potentially clinical endpoints.
The statistical framework therefore needs to address:
- Multiple quality attributes
- Different sources of variability
- Reference-product lot-to-lot variation
- Assay precision
- Acceptance ranges
- Equivalence or similarity margins
- Multiplicity and aggregation
- Missing observations
- Intercurrent events
- Between-subject and within-subject variability
- Uncertainty in the estimated treatment difference or ratio
Analytical Similarity
Analytical similarity is often the foundation of a biosimilar development program.
The objective is to characterize the proposed biosimilar and reference product using appropriately sensitive analytical methods.
Examples of potentially relevant characteristics include:
- Primary structure
- Higher-order structure
- Post-translational modifications
- Glycosylation
- Charge variants
- Size variants
- Purity
- Aggregation
- Binding characteristics
- Biological activity
- Potency
The exact set of quality attributes is product-specific.
FDA's September 2025 final guidance provides recommendations for designing and evaluating comparative analytical studies for therapeutic protein biosimilars and emphasizes the scientific assessment of comparative quality attributes.
Critical Quality Attributes
A critical quality attribute is a physical, chemical, biological, or microbiological property or characteristic that should be appropriately controlled to ensure product quality.
From a statistical perspective, each attribute may have:
- A distribution in the reference product
- Within-batch variability
- Between-batch variability
- Analytical measurement error
- A biological interpretation
- A potential clinical relevance
This means that a simple comparison of two means is often insufficient.
Reference-Product Variability
Suppose a quality attribute has reference-product batch measurements:
A biosimilar batch with a result of 100.3 may be entirely consistent with the reference distribution.
A result of 103.0 may be numerically different from the reference mean but could still require interpretation in the context of analytical variability, reference-product variability, and clinical relevance.
Acceptance Ranges
An acceptance range defines a region within which an estimated difference or ratio is considered sufficiently compatible with similarity.
For a ratio-based comparison, a generic acceptance region can be written as:
where:
- \(L\) = lower similarity limit
- \(U\) = upper similarity limit
- \(\theta\) = estimated relative difference or ratio
The limits should be scientifically justified.
They should not simply be selected because they are convenient statistically.
Confidence Intervals and Similarity
A common way of expressing uncertainty is through a confidence interval.
Suppose the estimated ratio is:
and the 90% confidence interval is:
If the prespecified acceptance limits are:
then the complete confidence interval lies inside the acceptance region.
This provides evidence consistent with the prespecified similarity criterion.
Interactive Example: Confidence Interval Within the Similarity Region
A confidence interval that extends outside a prespecified acceptance range does not automatically mean the product is clinically different; it means the specified statistical criterion has not been met and the evidence requires further assessment.
Why 80–125% Is Not a Universal Biosimilar Rule
The interval:
is familiar from conventional bioequivalence assessment.
However, it should not be treated as a universal acceptance interval for every biosimilar statistical comparison.
The appropriate margin depends on:
- The endpoint
- The variability of the reference product
- The clinical relevance of potential differences
- The assay characteristics
- The mechanism of action
- The pharmacological context
- The regulatory framework
EMA's clinical pharmacology Q&A notes that 80–125% may be used for PK comparisons in the absence of rigorous justification, but emphasizes that acceptance criteria should be prespecified and justified.
Ratio Versus Difference
Statistical comparisons can be expressed as either differences or ratios.
For an additive comparison:
For a multiplicative comparison:
The appropriate scale depends on the endpoint and scientific model.
PK exposure measures such as AUC and Cmax are commonly analyzed on the logarithmic scale because ratios are naturally represented by differences on that scale.
Log Transformation for PK
Suppose the AUC values are positively skewed.
The analysis can use:
The treatment difference on the log scale is:
Exponentiating gives the geometric-mean ratio:
PK Similarity
Pharmacokinetics are a central component of biosimilar development.
Common PK endpoints include:
- AUC0-t
- AUC0-inf
- Cmax
- Other product-specific exposure measures
The appropriate endpoints depend on the product and route of administration.
EMA emphasizes that biosimilar PK assessment should reflect distribution and elimination as well as absorption, because biologically derived products may have complex disposition characteristics.
Typical PK Statistical Model
For a randomized crossover study, a log-transformed analysis can be represented conceptually as:
The exact model depends on the study design and regulatory agreement.
Parallel-Group PK Studies
Parallel-group studies may require adjustment for important baseline characteristics.
For example, body weight can influence PK for some biological products.
An ANCOVA model may therefore be considered:
EMA specifically notes that ANCOVA can be acceptable in parallel-group biosimilar PK comparisons when covariates are prespecified and justified.
PK Confidence Intervals
Suppose the geometric mean ratio for AUC is:
with a 90% confidence interval:
If the prespecified acceptance interval is 0.80 to 1.25, the PK criterion is satisfied for this endpoint.
Pharmacodynamic Similarity
Pharmacodynamic endpoints can be highly informative when there is a sensitive, clinically relevant biomarker linked to the mechanism of action.
Examples can include:
- Biomarkers of target engagement
- Receptor occupancy
- Pharmacological response markers
- Cell-count responses
- Other product-specific PD biomarkers
A PD endpoint may be analyzed using:
- Mean difference
- Ratio of means
- Area under the effect curve
- Maximum effect
- Time to maximum effect
- Model-based longitudinal analysis
Clinical Efficacy Studies
Historically, biosimilar programs frequently included comparative efficacy studies designed to demonstrate that the biosimilar was neither meaningfully better nor worse than the reference product.
The regulatory landscape is evolving.
FDA issued draft guidance in October 2025 describing considerations for determining whether comparative efficacy studies are needed for therapeutic protein biosimilars. The document is explicitly draft and is not for implementation.
EMA went further in its March 2026 adopted reflection paper on a tailored clinical approach. The paper explains that, for well-characterized biological products, robust analytical comparability combined with PK and appropriate safety data may in some circumstances provide sufficient evidence without a dedicated comparative efficacy study.
Equivalence Versus Non-Inferiority
A biosimilar comparative efficacy study may use an equivalence framework when the objective is to exclude clinically important differences in either direction.
For an efficacy difference:
an equivalence criterion can be expressed as:
where \(M\) is the prespecified equivalence margin.
The null hypothesis for an equivalence test is therefore conceptually:
against:
Why "No Significant Difference" Is Not Enough
Suppose a comparative efficacy study produces:
That does not establish equivalence.
A non-significant difference may simply reflect inadequate precision.
For example, an estimate could be:
with a very wide confidence interval:
A conclusion of equivalence depends on whether the entire confidence interval lies within the prespecified equivalence margin.
Choosing the Equivalence Margin
The equivalence margin is one of the most important statistical assumptions in a comparative efficacy study.
A defensible margin should be based on clinical considerations and the established efficacy of the reference product.
It should not simply be selected to make the sample size convenient.
The margin should be justified before unblinding the comparative analysis.
Sample-Size Planning
For a simple two-arm equivalence comparison of a continuous endpoint, a conceptual sample-size expression is:
where:
- \(\sigma\) = standard deviation
- \(M\) = equivalence margin
- \(\alpha\) = type I error component
- \(1-\beta\) = desired power
This is only a simplified planning equation.
Actual biosimilar study calculations may need to account for:
- Ratio-based endpoints
- Log transformation
- Unequal allocation
- Dropout
- Repeated measurements
- Covariate adjustment
- Stratification
- Binary response endpoints
- Non-normal endpoints
- Multiple primary endpoints
Example Sample-Size Calculation
Suppose an equivalence study uses:
- Equivalence margin = 10 units
- Standard deviation = 20 units
- Power = 90%
- Two-sided alpha = 5%
The simplified calculation is:
which produces approximately:
before accounting for dropout and other design considerations.
Multiplicity Across Endpoints
Biosimilar studies can involve multiple endpoints.
For example:
| Domain | Potential Endpoint |
|---|---|
| PK | AUC |
| PK | Cmax |
| PD | Maximum effect |
| Clinical | Response rate |
| Safety | Adverse-event incidence |
| Immunogenicity | ADA incidence |
Not every endpoint necessarily belongs to one conventional multiplicity family.
The statistical analysis plan should clearly distinguish:
- Primary similarity criteria
- Supportive endpoints
- Descriptive analyses
- Safety summaries
- Exploratory analyses
Analytical Similarity and Multiplicity
Analytical similarity presents a particularly interesting multiplicity problem.
A development program may compare dozens or hundreds of attributes.
It is usually inappropriate to treat every individual comparison as though it were a standalone confirmatory hypothesis test.
Instead, the statistical framework should reflect the role of each quality attribute and the overall comparability strategy.
EMA's statistical reflection paper specifically addresses statistical methodology for comparative quality attributes, including the comparison objective, sampling strategy, variability, acceptance ranges, and inferential methods.
Distribution-Based Comparisons
For continuous analytical attributes, useful summaries may include:
- Means
- Standard deviations
- Variances
- Quantiles
- Reference-product ranges
- Probability distributions
- Confidence intervals
- Prediction intervals
The choice depends on the nature of the attribute and the scientific question.
Difference in Means
A simple comparison can be represented as:
with confidence interval:
This approach may be useful for some normally distributed quality attributes, but it should not be applied automatically to every attribute.
Ratio of Means
For multiplicative attributes:
the logarithmic scale may provide a more appropriate analysis:
Reference Variability
Suppose the reference product has substantial lot-to-lot variability.
The statistical assessment should distinguish:
- True product variability
- Sampling variability
- Analytical assay variability
- Batch effects
- Measurement error
Failure to separate these sources can produce either an overly permissive or overly restrictive similarity assessment.
Batch Selection
The batches selected for analytical comparison should be scientifically representative.
Potential considerations include:
- Manufacturing history
- Commercial relevance
- Lot age
- Storage conditions
- Manufacturing sites
- Reference-product variability
Statistical analysis cannot rescue a poorly designed sampling strategy.
Immunogenicity
Immunogenicity is an important component of biosimilar development.
Relevant outcomes can include:
- Anti-drug antibody incidence
- Neutralizing antibody incidence
- Time to antibody development
- Antibody persistence
- Antibody titers
- Effects on PK
- Effects on efficacy
- Effects on safety
The analysis may include categorical comparisons, time-to-event methods, longitudinal analyses, and exposure-response evaluations.
Binary Immunogenicity Endpoints
Suppose ADA incidence is:
| Group | ADA Positive | Total | Rate |
|---|---|---|---|
| Biosimilar | 22 | 200 | 11.0% |
| Reference | 24 | 198 | 12.1% |
The crude risk difference is:
or −1.1 percentage points.
The interpretation should consider the confidence interval and the prespecified analysis rather than the point estimate alone.
Time-to-Immunogenicity
If timing of ADA development is clinically important, survival-analysis methods can be considered.
For example:
- Kaplan-Meier curves
- Cox proportional hazards models
- Restricted mean survival time
The choice should be driven by the estimand and scientific question.
Missing Data
Missing data are especially important in clinical comparative studies.
Potential causes include:
- Study withdrawal
- Adverse events
- Loss to follow-up
- Assay failure
- Protocol deviations
- Sample handling problems
Missingness should not automatically be handled using simple last-observation carried forward.
Instead, the statistical analysis plan should define appropriate estimands, analysis populations, and sensitivity analyses.
Estimands in Biosimilar Studies
The estimand framework provides a structured way to specify exactly what treatment comparison is being estimated.
Important dimensions include:
- Population
- Treatment condition
- Variable
- Intercurrent events
- Summary measure
For example, a clinical efficacy estimand might specify whether treatment discontinuation because of adverse events is handled using a treatment-policy, hypothetical, composite, or other strategy.
Analysis Populations
A biosimilar clinical study may define several analysis populations.
| Population | Typical Purpose |
|---|---|
| Full analysis set | Primary effectiveness or estimand analysis where appropriate |
| Per-protocol set | Assessment among participants with sufficiently compliant conduct |
| Safety set | Subjects receiving study treatment |
| PK population | Subjects with evaluable PK profiles |
| Immunogenicity population | Subjects with relevant post-baseline samples |
For equivalence studies, interpretation of both full-analysis and per-protocol analyses can be particularly important.
Concordance Between Analysis Populations
Suppose the primary analysis demonstrates equivalence in the per-protocol set but not in the full analysis set.
This discrepancy requires investigation.
Conversely, agreement between appropriately defined populations strengthens confidence in the robustness of the conclusion.
Sensitivity Analyses
Sensitivity analyses can assess whether the similarity conclusion depends on specific assumptions.
Examples include:
- Alternative missing-data assumptions
- Alternative covariance structures
- Alternative baseline adjustments
- Alternative analysis populations
- Exclusion of influential observations
- Alternative handling of protocol deviations
Protocol Deviations
Protocol deviations can be particularly important in equivalence studies.
A deviation that affects treatment exposure, assessment timing, or endpoint measurement can influence the ability of the study to distinguish products.
The analysis plan should prospectively define which deviations are considered important and how they affect analysis populations.
Outliers
An outlier should not be removed simply because it makes a similarity criterion fail.
The analyst should determine whether the observation represents:
- A genuine biological value
- Measurement error
- Sample contamination
- Data-entry error
- Protocol deviation
- Assay failure
Any exclusion should be supported by predefined criteria and documented.
Reference Product Drift
Reference products can evolve over time because of manufacturing changes.
This creates a statistical and scientific question about which reference batches should represent the relevant reference-product distribution.
The solution is not simply to average every historical batch available.
The selection should reflect the scientific and regulatory context of the development program.
FDA Versus EMA: High-Level Comparison
| Topic | FDA | EMA |
|---|---|---|
| Overall principle | Totality of evidence supporting biosimilarity | Stepwise comprehensive comparability exercise |
| Analytical similarity | Central component of biosimilar assessment | Central component of biosimilar assessment |
| PK | Important component of clinical pharmacology assessment | Important component of biosimilarity assessment |
| PD | Used when scientifically informative | Used when appropriate and sensitive |
| Comparative efficacy | Need is product- and evidence-dependent; FDA issued 2025 draft recommendations | Tailored clinical approach may reduce or eliminate need in suitable cases |
| Immunogenicity | Important clinical component | Important clinical component |
| Statistical acceptance | Endpoint- and study-specific | Endpoint- and product-specific |
This table is intentionally high-level. Regulatory requirements should always be assessed against the current product-specific guidance and agency advice.
FDA: Current Statistical Direction
FDA's September 2025 final guidance on comparative analytical assessment provides the current finalized framework for analytical similarity assessment of therapeutic protein biosimilars.
FDA also maintains a clinical pharmacology guidance describing the clinical pharmacology data used to support biosimilarity determinations.
The FDA's October 2025 draft guidance on comparative efficacy studies signals a more selective approach to when comparative efficacy studies may be needed. Because it remains draft, sponsors should not treat its recommendations as binding requirements.
EMA: Current Statistical Direction
EMA's overarching biosimilar guideline remains an important foundation for the European framework. The current effective overarching guideline is Revision 1, while a new revision was opened for consultation in July 2026.
More importantly for the current clinical-development strategy, EMA adopted the Reflection Paper on a Tailored Clinical Approach in Biosimilar Development in March 2026.
The reflection paper recognizes that dedicated comparative efficacy studies may not be required for well-characterized biological products when analytical and PK evidence is sufficiently informative, together with appropriate safety and other relevant evidence.
Why the Analytical Assessment Can Reduce Clinical Uncertainty
Consider two hypothetical development programs.
| Evidence | Program A | Program B |
|---|---|---|
| Structural similarity | Very strong | Strong |
| Functional similarity | Very strong | Moderate |
| PK similarity | Very strong | Strong |
| PD similarity | Strong | Not available |
| Residual uncertainty | Low | Higher |
| Potential clinical efficacy requirement | May be reduced depending on regulatory assessment | May require additional evidence |
The statistical strategy should therefore be driven by residual uncertainty, not by a fixed checklist applied identically to every molecule.
Totality of Evidence
A biosimilarity conclusion can be represented conceptually as:
This is not a literal arithmetic formula.
It represents the conceptual integration of different evidence streams.
Residual Uncertainty
A useful way to think about the development program is:
The purpose of each study is to reduce a specific type of uncertainty.
| Study | Primary Uncertainty Addressed |
|---|---|
| Structural characterization | Does the molecular structure differ? |
| Functional assay | Does biological activity differ? |
| PK study | Does systemic exposure differ? |
| PD study | Does pharmacological response differ? |
| Immunogenicity study | Does immune response differ? |
| Clinical efficacy study | Does remaining clinical uncertainty matter? |
Statistical Power in Biosimilar Development
Power should be considered carefully because a biosimilar study is often designed to establish similarity rather than superiority.
For an equivalence study:
Conceptually, power is the probability that the confidence interval will fall entirely inside the prespecified similarity region when the products are sufficiently similar.
Precision Is Critical
Suppose two studies have the same estimated treatment ratio:
Study A:
Study B:
Both point estimates are identical.
But Study A provides substantially more precise evidence.
This illustrates why sample size and variance control are central to biosimilar study design.
Confidence Intervals Are More Informative Than P-Values
A p-value can answer whether a null hypothesis is rejected under a particular test.
A confidence interval additionally shows:
- The estimated magnitude of the difference
- The direction of the difference
- The uncertainty surrounding the estimate
- Whether the complete interval fits within the similarity margin
For biosimilarity, these features are particularly important.
Common Statistical Mistake: Testing for Equality
Suppose an analyst performs:
and obtains:
It is incorrect to conclude:
"The products are statistically equivalent."
The appropriate interpretation is that the analysis did not provide evidence against the null hypothesis of equality.
That is not the same as demonstrating that any possible difference is sufficiently small.
Common Statistical Mistake: Using an Arbitrary Margin
Suppose the observed difference is only 2%.
An analyst might choose an equivalence margin of 20% and conclude equivalence.
But if a 20% difference would be clinically meaningful, the analysis is not scientifically persuasive.
Common Statistical Mistake: Ignoring Assay Variability
Suppose an analytical assay has a coefficient of variation of 12%.
A measured difference of 5% may be indistinguishable from assay variability.
A measured difference of 30% may be much more difficult to explain as random measurement variation.
The statistical model should therefore account for the properties of the measurement process.
Common Statistical Mistake: Overinterpreting One Quality Attribute
A biosimilar may show a modest difference in one quality attribute while remaining highly similar across many other characteristics.
That single difference should be evaluated in context.
Questions include:
- Is the attribute clinically relevant?
- Is the difference reproducible?
- Does it affect function?
- Is it within reference-product variability?
- Does it affect PK?
- Does it affect immunogenicity?
- Does it alter the mechanism of action?
R Implementation: PK Equivalence
A simple log-transformed PK analysis in R can be implemented using a linear model.
fit <- lm( log(AUC) ~ TRT + SEQUENCE + PERIOD, data = pk ) summary(fit) coef_ci <- confint( fit, "TRTTest", level = 0.90 ) gmr <- exp(coef(fit)["TRTTest"]) ci_gmr <- exp(coef_ci) gmr ci_gmr
The exact model depends on the study design.
For a crossover study, subject effects and the appropriate period and sequence structure should be incorporated according to the prespecified analysis.
R Implementation: Equivalence Confidence Interval
For a continuous endpoint, the estimated difference and confidence interval can be compared with the equivalence limits.
estimate <- -0.8 lower <- -5.2 upper <- 3.6 margin <- 10 equivalent <- lower > -margin && upper < margin equivalent
The result is TRUE only when the entire confidence
interval lies within the prespecified equivalence region.
R Implementation: Ratio-Based Similarity
estimate <- 1.03 lower <- 0.98 upper <- 1.08 lower_margin <- 0.80 upper_margin <- 1.25 similar <- lower > lower_margin && upper < upper_margin similar
Again, the limits shown here are illustrative. Actual biosimilar acceptance criteria must be scientifically justified and prespecified.
R Implementation: Analytical Quality Attributes
Suppose a dataset contains:
product batch attribute value
A simple descriptive comparison might begin with:
library(dplyr)
summary_stats <- data %>%
group_by(product, attribute) %>%
summarise(
n = n(),
mean = mean(value, na.rm = TRUE),
sd = sd(value, na.rm = TRUE),
median = median(value, na.rm = TRUE),
.groups = "drop"
)
This is only a descriptive starting point.
The regulatory similarity assessment may require more sophisticated analyses depending on the attribute and agreed analytical strategy.
R Visualization of Analytical Similarity
library(ggplot2)
ggplot(
data,
aes(
x = product,
y = value
)
) +
geom_boxplot() +
facet_wrap(~attribute, scales = "free_y") +
labs(
x = NULL,
y = "Measured Attribute"
) +
theme_classic()
Distributional graphics can help identify:
- Differences in central tendency
- Differences in spread
- Outliers
- Batch effects
- Potential shifts in the reference distribution
Forest Plot for Similarity
A forest plot can summarize confidence intervals across multiple endpoints.
library(ggplot2)
ggplot(
results,
aes(
x = estimate,
y = endpoint
)
) +
geom_errorbarh(
aes(
xmin = lower,
xmax = upper
),
height = 0.15
) +
geom_point() +
geom_vline(
xintercept = 1,
linetype = "dashed"
) +
theme_classic()
For ratio-based endpoints, the reference value is:
For difference-based endpoints, the reference value is:
Recommended Statistical Workflow
Statistical Analysis Plan
The SAP should clearly define:
- Analysis populations
- Primary and secondary endpoints
- Baseline definitions
- Transformation rules
- Statistical models
- Covariates
- Random effects where applicable
- Confidence interval level
- Acceptance criteria
- Missing-data handling
- Protocol-deviation rules
- Outlier handling
- Sensitivity analyses
- Multiplicity strategy
- Software and validation procedures
Pre-Specification Is Especially Important
A similarity margin selected after reviewing the results is not credible as a confirmatory criterion.
Likewise, selectively changing the analysis population after seeing which population produces the desired conclusion undermines the interpretation.
The critical statistical decisions should therefore be established before unblinding or database lock, as applicable to the analysis.
Quality-Control Checklist
Example Statistical Assessment
| Endpoint | Estimate | 90% CI | Illustrative Criterion | Interpretation |
|---|---|---|---|---|
| AUC ratio | 1.03 | 0.98–1.08 | 0.80–1.25 | Within range |
| Cmax ratio | 0.97 | 0.91–1.04 | 0.80–1.25 | Within range |
| PD difference | 1.2 units | −2.1 to 4.5 | −10 to +10 | Within range |
| Clinical response difference | −1.5% | −7.8 to 4.8% | −10 to +10% | Within range |
This table is illustrative only.
In an actual submission, each criterion must be scientifically justified and aligned with the applicable regulatory strategy.
What Happens When a Criterion Fails?
Suppose a PK endpoint produces:
with:
If the upper similarity limit is 1.25, the confidence interval extends just beyond the limit.
The appropriate response is not simply to discard the study.
Questions should include:
- Was the study appropriately designed?
- Was the endpoint measured reliably?
- Was there unexpected variability?
- Were important covariates imbalanced?
- Were there assay problems?
- Were there anti-drug antibodies?
- Was the observed difference reproducible?
- What do the analytical and functional data show?
EMA's clinical pharmacology Q&A specifically discusses investigation of root causes when PK confidence intervals fail the prespecified acceptance limits and emphasizes interpretation of the complete evidence rather than simply ignoring an unsuccessful study.
Why Failed Studies Cannot Simply Be Hidden
Suppose two PK studies are performed.
| Study | Result |
|---|---|
| Study 1 | Meets similarity criterion |
| Study 2 | Fails similarity criterion |
The existence of Study 1 does not automatically make Study 2 irrelevant.
The discrepancy should be investigated and incorporated into the overall scientific assessment.
This is an important principle for regulatory statistics:
Interchangeability Is Different From Biosimilarity
In the United States, a biosimilar designation and an interchangeable designation are related but distinct regulatory concepts.
A product can satisfy biosimilarity requirements without automatically satisfying the additional requirements for interchangeability.
Therefore, statistical planning for an interchangeability program should not simply copy the statistical strategy used for an ordinary biosimilarity demonstration.
FDA's guidance framework separately addresses considerations for demonstrating interchangeability.
Switching and Multiple-Period Designs
Where interchangeability is being investigated, a study may involve switching between reference and biosimilar products.
Such designs introduce additional statistical considerations, including:
- Sequence effects
- Period effects
- Carryover considerations
- Within-subject variability
- Repeated exposure
- Immunogenicity
- Safety following switching
The analysis must be designed around the specific regulatory question rather than simply around generic crossover-study methods.
EMA's Tailored Clinical Approach
The 2026 EMA reflection paper is particularly important for statistical programmers because it reinforces the idea that the clinical program should be proportionate to residual uncertainty.
For sufficiently well-characterized biological substances, analytical and functional similarity may provide highly sensitive evidence of similarity, while PK and appropriate safety information can address remaining clinical uncertainty.
The implication is that statistical planning should begin with the question:
rather than:
EMA's adopted reflection paper explicitly states that comparative efficacy studies may not be required for well-characterized biological substances when analytical and PK evidence is sufficiently informative, together with appropriate safety evidence.
Regulatory Strategy and Statistical Strategy
The statistical plan should therefore be integrated with the scientific development strategy.
| Scientific Question | Statistical Question |
|---|---|
| Are structures similar? | How should structural attributes be compared? |
| Are functional activities similar? | What scale and acceptance criteria are appropriate? |
| Is exposure similar? | What PK endpoints and confidence intervals are needed? |
| Is PD response similar? | What biomarker and similarity criterion are justified? |
| Is immune response similar? | What incidence, timing, and titer analyses are appropriate? |
| Is clinical uncertainty resolved? | Is a comparative efficacy study necessary? |
Common Mistakes
- Treating biosimilarity as ordinary superiority testing. The statistical objective is different.
- Assuming 80–125% applies to every biosimilar comparison. Acceptance criteria are endpoint- and scientifically dependent.
- Using "p > 0.05" as proof of similarity. Non-significance is not evidence of equivalence.
- Choosing the margin after reviewing results. Similarity margins must be scientifically justified and prospectively defined.
- Ignoring reference-product variability. The reference distribution is central to analytical similarity.
- Ignoring assay variability. Observed differences must be interpreted in light of measurement precision.
- Automatically requiring a large efficacy trial. Current FDA and EMA approaches increasingly emphasize scientifically tailored clinical development.
- Automatically eliminating all clinical efficacy data. A tailored approach does not mean that clinical evidence is universally unnecessary.
- Ignoring failed or discordant studies. All meaningful evidence should be integrated into the totality-of-evidence assessment.
- Confusing biosimilarity with interchangeability. The regulatory questions are related but distinct.
- Using unvalidated analytical datasets. The statistical result is only as reliable as the underlying measurements.
- Overinterpreting individual quality attributes. An attribute should be interpreted in the context of function and clinical relevance.
What Should a Statistical Report Contain?
A statistical report supporting biosimilarity should allow a reviewer to understand the complete analytical pathway.
At minimum, consider including:
- Study objectives
- Statistical estimands
- Analysis populations
- Endpoint definitions
- Data transformations
- Statistical models
- Variance assumptions
- Similarity margins
- Confidence intervals
- Missing-data handling
- Protocol-deviation handling
- Sensitivity analyses
- Individual-study results
- Integrated interpretation
Suggested Analytical Similarity Table
| Attribute | Reference Mean | Biosimilar Mean | Difference | 95% CI | Assessment |
|---|---|---|---|---|---|
| Attribute A | 100.2 | 100.7 | 0.5 | −1.2 to 2.2 | Similar |
| Attribute B | 42.1 | 41.8 | −0.3 | −1.1 to 0.5 | Similar |
| Attribute C | 7.8 | 8.0 | 0.2 | −0.1 to 0.5 | Review in context |
The exact presentation should follow the statistical methodology agreed for the development program.
Statistical Interpretation of Quality Attributes
An attribute can be statistically different without necessarily being clinically meaningful.
Conversely, a statistically non-significant difference may still require scientific attention if the estimate is large or the study is imprecise.
Therefore, the reviewer should consider:
and:
The complete evidence package is what supports the regulatory conclusion.
Validation of Statistical Programming
Biosimilar analyses may be particularly sensitive to programming errors because small changes in confidence intervals, transformations, or reference populations can affect conclusions.
Programming QC should include:
- Independent programming
- Double-programming of key endpoints
- Dataset-level reconciliation
- Table-to-source verification
- Confidence-interval validation
- Margin validation
- Graphical review
- Extreme-value review
- Traceability to source data
Example of a Statistical QC Check in R
results %>%
mutate(
criterion_met =
lower_ci > lower_margin &
upper_ci < upper_margin
) %>%
select(
endpoint,
estimate,
lower_ci,
upper_ci,
lower_margin,
upper_margin,
criterion_met
)
The output should be independently checked against the prespecified statistical criteria.
Reporting the Overall Conclusion
A biosimilarity conclusion should not be reduced to a single p-value.
A stronger summary describes:
- The analytical evidence
- The functional evidence
- The PK evidence
- The PD evidence where applicable
- The immunogenicity evidence
- The clinical evidence where applicable
- Any residual uncertainties
- The rationale for the overall conclusion
Example Overall Interpretation
This wording is intentionally generic and should not be copied into a regulatory submission without adapting it to the actual product, data, and regulatory strategy.
A Practical Decision Framework
The Most Important Statistical Concepts
| Concept | Why It Matters |
|---|---|
| Similarity margin | Defines what difference is scientifically acceptable. |
| Confidence interval | Quantifies uncertainty around the estimated difference or ratio. |
| Reference variability | Defines the natural range against which similarity is assessed. |
| Assay variability | Determines how precisely a quality attribute can be measured. |
| Precision | Determines whether the study can meaningfully exclude important differences. |
| Estimand | Defines exactly what treatment comparison is being estimated. |
| Sensitivity analysis | Tests robustness of the similarity conclusion. |
| Totality of evidence | Integrates analytical, clinical pharmacology, safety, immunogenicity, and clinical information. |
Regulatory References
FDA.
Development of Therapeutic Protein Biosimilars: Comparative Analytical
Assessment and Other Quality-Related Considerations. Final Guidance,
September 2025.
FDA.
Clinical Pharmacology Data to Support a Demonstration of Biosimilarity to a
Reference Product. Final Guidance, December 2016.
FDA.
Scientific Considerations in Demonstrating Biosimilarity to a Reference
Product: Updated Recommendations for Assessing the Need for Comparative
Efficacy Studies. Draft Guidance, October 2025.
EMA.
Guideline on Similar Biological Medicinal Products. CHMP/437/04 Rev.1,
effective April 2015; revision under consultation in 2026.
EMA.
Similar Biological Medicinal Products Containing Biotechnology-Derived
Proteins as Active Substance: Non-Clinical and Clinical Issues.
EMEA/CHMP/BMWP/42832/2005 Rev.1.
EMA.
Reflection Paper on a Tailored Clinical Approach in Biosimilar Development.
EMA/CHMP/BMWP/60916/2025, adopted March 2026.
EMA.
Reflection Paper on Statistical Methodology for the Comparative Assessment of
Quality Attributes in Drug Development. EMA/CHMP/138502/2017.