Tutorials › Biostatistics › EMA Statistical Guidance Overview

Regulatory Statistics

EMA Statistical Guidance Overview

A practical guide to the European Medicines Agency statistical framework for clinical trials, including ICH E9, estimands, confirmatory analysis, multiplicity, missing data, covariate adjustment, subgroup analysis, time-to-event endpoints, adaptive designs, and statistical reporting.

Intermediate 22 min read

What You'll Learn

  • How EMA statistical guidance fits together
  • The role of ICH E9 and ICH E9(R1)
  • How estimands connect clinical questions to analyses
  • How EMA approaches missing data and intercurrent events
  • How multiplicity, subgroup analysis, and covariates affect inference
  • How to translate regulatory guidance into a statistical analysis plan

Introduction

European clinical-trial statistical practice is not defined by one single document called "EMA statistical guidance." Instead, statistical expectations are distributed across a framework of EMA guidelines, ICH guidelines, scientific guidelines, methodological guidance, product-area guidance, and study-specific regulatory interaction.

For statisticians working on pharmaceutical development programs, this can initially appear complicated.

The practical challenge is therefore not simply learning individual guidelines. It is understanding which statistical principle answers which development question and how those principles fit together.

Key idea: EMA statistical expectations are best viewed as a connected framework. ICH E9 provides foundational statistical principles for clinical trials, ICH E9(R1) extends that framework through estimands and sensitivity analysis, and additional guidance addresses specific issues such as multiplicity, missing data, subgroup analysis, covariate adjustment, adaptive designs, and disease-specific endpoints.

The Regulatory Statistical Framework

A useful conceptual model is to think of the statistical framework as a series of questions.

1
What clinical question is being answered?
Define the treatment effect and target population.
2
What is the estimand?
Specify the population, treatment conditions, variable, intercurrent-event strategy, and summary measure.
3
What design can answer the question?
Specify randomization, comparator, allocation, timing, sample size, and follow-up.
4
What analysis estimates the treatment effect?
Choose the statistical model, covariates, contrasts, and inference procedure.
5
How will uncertainty and departures from assumptions be evaluated?
Prespecify sensitivity analyses and appropriate robustness checks.

A Practical Map of EMA Statistical Guidance

The following interactive map summarizes the major statistical topics encountered during clinical development.

Figure 1. Statistical Guidance Framework for Clinical Trials
Hover over a topic to see its primary statistical purpose. The relationships are conceptual rather than a formal hierarchy of regulatory documents.
ICH E9 Statistical principles ICH E9(R1) Estimands & sensitivity Methodological Topics Design & analysis Confirmatory Inference Hypotheses & CI Missing Data Assumptions & sensitivity Subgroups & Covariates Heterogeneity & precision Special Designs Adaptive & specialized

The framework emphasizes how foundational statistical principles connect to specific design and analysis questions.

ICH E9: The Statistical Foundation

One of the most important documents for understanding EMA statistical expectations is ICH E9: Statistical Principles for Clinical Trials.

ICH E9 provides a framework for designing and analyzing clinical trials using sound statistical principles.

Its scope includes concepts such as:

  • Trial objectives
  • Trial design
  • Primary and secondary variables
  • Analysis sets
  • Sample-size determination
  • Randomization
  • Interim analyses
  • Statistical methods
  • Missing observations
  • Multiplicity
  • Reporting of results
Practical interpretation: ICH E9 should not be treated as a collection of programming rules. It is a framework for ensuring that the statistical design and analysis answer the scientific question posed by the clinical trial.

The Statistical Question Comes First

A common mistake is to begin with the statistical model.

For example:

"We will analyze the endpoint using ANCOVA."

This statement is incomplete.

Before choosing ANCOVA, the statistician should understand:

  • What treatment effect is of interest?
  • Which patients are part of the target population?
  • What outcome is being summarized?
  • What happens if treatment is discontinued?
  • What happens if rescue medication is administered?
  • What happens if the endpoint cannot be observed?

These questions lead naturally to the concept of the estimand.

ICH E9(R1) and the Estimand Framework

ICH E9(R1) introduced an important refinement to the statistical framework: the distinction between the clinical question and the statistical method used to estimate it.

An estimand describes the treatment effect of interest in a way that specifies key attributes of the scientific question.

The framework considers:

  • Population: Who is being described?
  • Treatment condition: What treatment strategies are being compared?
  • Variable: What outcome is being measured?
  • Intercurrent events: How are events occurring after treatment initiation handled?
  • Population-level summary: How is the individual-level outcome summarized?

The general idea can be expressed as:

$$ \text{Clinical Question} \rightarrow \text{Estimand} \rightarrow \text{Estimator} \rightarrow \text{Estimate} $$

This sequence is one of the most important ideas in modern regulatory statistics.

Why Estimands Matter

Consider a randomized trial evaluating a new treatment.

Some patients discontinue treatment before the planned endpoint.

There are several possible scientific questions.

For example:

  • What would the treatment effect be if patients remained on treatment?
  • What is the effect of assigning patients to the treatment strategy, regardless of later discontinuation?
  • What would happen if treatment were discontinued because of toxicity?
  • What would happen if rescue medication became available?

These are different questions.

Consequently, they can lead to different estimands and potentially different statistical analyses.

Intercurrent Events

An intercurrent event is an event occurring after treatment initiation that affects either the interpretation or existence of the outcome measurement.

Examples include:

  • Treatment discontinuation
  • Rescue medication
  • Switching treatment
  • Use of prohibited medication
  • Death
  • Organ transplantation
  • Initiation of another therapy

The important question is not simply:

$$ \text{What should we do with the missing value?} $$

Instead, the first question is:

$$ \text{What treatment effect do we want to estimate after the event occurs?} $$

Estimand Strategies

ICH E9(R1) describes several strategies for handling intercurrent events.

Strategy Concept
Treatment policy The outcome is considered regardless of whether the intercurrent event occurs.
Composite The intercurrent event is incorporated into the definition of the outcome.
While on treatment The treatment effect is defined over the period before a specified intercurrent event.
Principal stratum The effect is evaluated within a population defined by post-randomization event behavior.
Hypothetical The treatment effect is defined under a specified hypothetical scenario in which the event did not occur.

The appropriate strategy depends on the clinical question rather than on a generic preference for one missing-data method.

Estimand and Statistical Model Are Not the Same Thing

This distinction is fundamental.

The estimand defines what effect is wanted.

The estimator is the statistical procedure used to estimate that effect.

For example, a treatment-policy estimand for a continuous endpoint might be estimated using:

  • MMRM
  • ANCOVA
  • Repeated-measures models
  • Multiple imputation
  • Other prespecified estimators

The same broad clinical question can potentially be estimated using different methods under different assumptions.

Do not write an SAP by starting with the model. Start by defining the estimand. Then select an estimator whose assumptions and implementation are appropriate for that estimand.

Primary Endpoints

Regulatory statistical guidance places particular emphasis on the prespecification of the primary endpoint.

A primary endpoint should be clearly defined before unblinding and should be closely connected to the primary clinical objective.

Examples include:

  • Change from baseline in a continuous clinical score
  • Binary response
  • Time to progression
  • Overall survival
  • Event-free survival
  • Rate of recurrent events

Continuous Endpoints

For a continuous endpoint, a treatment effect may be expressed as a difference between adjusted means.

For two treatment groups:

$$ \Delta = \mu_T-\mu_C $$

where:

  • \(\mu_T\) = population mean under treatment
  • \(\mu_C\) = population mean under control

A confidence interval should generally accompany the treatment-effect estimate.

$$ \widehat{\Delta} \pm z_{1-\alpha/2} SE(\widehat{\Delta}) $$

The exact inferential procedure depends on the endpoint, design, estimator, and statistical assumptions.

Binary Endpoints

Binary outcomes require careful definition of both the response criterion and the estimand.

For a response probability:

$$ p_T=P(Y=1\mid T) $$

and:

$$ p_C=P(Y=1\mid C) $$

Possible measures of treatment effect include:

  • Risk difference
  • Risk ratio
  • Odds ratio

For example, the risk difference is:

$$ RD=p_T-p_C $$

The choice of effect measure should be scientifically justified and aligned with the estimand and endpoint interpretation.

Time-to-Event Endpoints

Time-to-event endpoints are common in regulatory clinical trials.

Examples include:

  • Overall survival
  • Progression-free survival
  • Time to treatment failure
  • Time to recurrence
  • Duration of response

The Kaplan-Meier estimator is frequently used to describe the survival function.

$$ \widehat{S}(t) = \prod_{t_i\leq t} \left( 1-\frac{d_i}{n_i} \right) $$

where \(d_i\) represents events at time \(t_i\), and \(n_i\) represents the number at risk immediately before that time.

Hazard Ratios

The Cox proportional-hazards model is frequently used to estimate relative hazard.

$$ h(t\mid X) = h_0(t)\exp(\beta X) $$

For a binary treatment indicator, the hazard ratio is:

$$ HR=e^\beta $$

However, the hazard ratio should not automatically be interpreted as a ratio of median survival times.

Important: A hazard ratio is a model-based relative measure of instantaneous event rates. Its interpretation depends on the model and assumptions, including the proportional-hazards assumption when that model is used.

Multiplicity

Multiplicity is one of the most important regulatory statistical issues in confirmatory clinical trials.

Suppose a study tests \(m\) hypotheses, each at significance level \(\alpha\).

If all hypotheses are tested independently without adjustment, the probability of at least one false-positive result can exceed \(\alpha\).

For independent tests, a simplified family-wise error expression is:

$$ FWER = 1-(1-\alpha)^m $$

For example, with five independent tests at \(\alpha=0.05\):

$$ 1-(1-0.05)^5 \approx0.226 $$

Thus, the nominal 5% level does not automatically control the overall family-wise error rate when multiple confirmatory hypotheses are tested.

Sources of Multiplicity

Multiplicity can arise from:

  • Multiple primary endpoints
  • Multiple treatment comparisons
  • Multiple doses
  • Multiple time points
  • Multiple populations
  • Multiple hypotheses
  • Repeated testing
  • Interim analyses

The important question is not simply whether there are many p-values.

The important question is which hypotheses belong to the confirmatory family and what claims the study intends to make.

Multiplicity Strategies

Approach Concept
Bonferroni Allocate the overall type I error across hypotheses.
Holm Sequentially reject ordered hypotheses while controlling family-wise error.
Hierarchical testing Test hypotheses according to a prespecified sequence.
Gatekeeping Testing of one family depends on results from another family.
Graph-based procedures Represent alpha allocation and recycling through a graphical structure.

Multiplicity and the SAP

A statement such as:

"Multiplicity will be adjusted as appropriate."

is generally insufficient for a confirmatory statistical analysis plan.

The SAP should identify:

  • The hypothesis family
  • The individual hypotheses
  • The type I error level
  • The testing sequence
  • The alpha allocation
  • The decision rules
  • The treatment comparisons
  • The relationship between primary and secondary endpoints

Sample Size

Sample-size determination should be linked to the primary objective and statistical assumptions.

For a simplified two-arm comparison of means with equal allocation, a generic sample-size relationship is:

$$ n \approx \frac{ 2\sigma^2 \left( z_{1-\alpha/2}+z_{1-\beta} \right)^2 }{ \Delta^2 } $$

where:

  • \(\sigma^2\) = variance
  • \(\Delta\) = target treatment difference
  • \(\alpha\) = type I error probability
  • \(1-\beta\) = target power

Real trial calculations may require adjustments for:

  • Dropout
  • Unequal allocation
  • Stratification
  • Repeated measurements
  • Multiplicity
  • Interim analyses
  • Event-driven recruitment
  • Non-inferiority margins

Clinically Meaningful Differences

A statistically significant difference is not automatically a clinically important difference.

For a continuous endpoint, a statistically significant result may be:

$$ \widehat{\Delta}=1.2 $$

while the clinically meaningful difference might be:

$$ \Delta_{\text{clinical}}=5 $$

Therefore, study design should consider the magnitude of treatment effect that would be clinically relevant, not merely an effect that can be detected statistically.

Superiority Trials

A superiority trial is designed to demonstrate that one treatment provides a beneficial effect compared with another treatment.

The statistical hypothesis may be written as:

$$ H_0:\Delta=0 $$

against:

$$ H_1:\Delta\neq0 $$

The exact hypothesis depends on the endpoint and direction of the treatment effect.

Non-Inferiority Trials

Non-inferiority trials have a fundamentally different objective.

The purpose is generally to demonstrate that the experimental treatment is not unacceptably worse than the active comparator by more than a prespecified margin.

Let \(\Delta\) represent the treatment difference and \(M\) the non-inferiority margin.

A simplified formulation is:

$$ H_0:\Delta\leq-M $$

against:

$$ H_1:\Delta>-M $$

The choice of the non-inferiority margin requires substantial clinical and historical justification.

Non-inferiority is not simply "showing a non-significant difference." Failure to demonstrate superiority is fundamentally different from demonstrating that a treatment satisfies a prespecified non-inferiority criterion.

Covariate Adjustment

Covariate adjustment can improve precision when baseline variables are prognostic for the outcome.

For a continuous endpoint, a simplified ANCOVA model is:

$$ Y_i = \beta_0 + \beta_1T_i + \beta_2X_i + \varepsilon_i $$

where:

  • \(Y_i\) = outcome
  • \(T_i\) = treatment indicator
  • \(X_i\) = baseline covariate
  • \(\beta_1\) = adjusted treatment effect

Potential covariates may include:

  • Baseline disease severity
  • Age
  • Baseline biomarker
  • Geographic region
  • Stratification variables

Covariate selection should be prespecified and clinically justified.

Baseline Adjustment Is Not the Same as Post-Randomization Adjustment

Baseline covariates are measured before treatment and therefore have a different statistical role from variables affected by treatment.

Post-randomization variables can be affected by treatment and may introduce bias if they are handled incorrectly.

Practical rule: Baseline prognostic covariates can often improve precision. Post-randomization variables require substantially more careful causal and estimand-based consideration.

Stratification Factors

Many randomized trials use stratification factors during randomization.

Examples include:

  • Disease stage
  • Prior treatment
  • Geographic region
  • Biomarker status

If stratification was used, the SAP should explain how the stratification variables enter the primary analysis.

The statistical analysis should be consistent with the design and randomization scheme.

Subgroup Analyses

Subgroup analyses are often clinically important but statistically challenging.

Common subgroups include:

  • Sex
  • Age group
  • Race or ethnicity where appropriate
  • Disease severity
  • Biomarker status
  • Geographic region
  • Prior therapy

A common mistake is to perform many subgroup-specific hypothesis tests and interpret each one independently.

Instead, the central question is whether there is evidence that the treatment effect differs across subgroups.

Interaction Testing

A simplified treatment-by-subgroup model is:

$$ Y = \beta_0 + \beta_1T + \beta_2S + \beta_3(T\times S) + \varepsilon $$

The interaction coefficient \(\beta_3\) addresses whether the treatment effect differs according to subgroup status.

This is conceptually different from asking whether the treatment effect is statistically significant within one subgroup but not another.

Classic statistical error: "Significant in subgroup A but not significant in subgroup B" does not by itself establish that the treatment effects differ between A and B. A formal comparison of effects, typically through an interaction or equivalent heterogeneity assessment, is required.

Missing Data

Missing data are unavoidable in many clinical trials.

Common causes include:

  • Withdrawal of consent
  • Loss to follow-up
  • Adverse events
  • Treatment discontinuation
  • Administrative reasons
  • COVID-related or other external disruptions
  • Death

A simplistic approach is to classify missingness as one of three categories:

  • MCAR — Missing Completely At Random
  • MAR — Missing At Random
  • MNAR — Missing Not At Random

But the estimand framework emphasizes that the handling of missing observations should be connected to the clinical question and assumptions.

MAR and Conditional Missingness

Under a simplified missing-at-random assumption:

$$ P(M=1\mid Y_{\text{obs}},Y_{\text{mis}}) = P(M=1\mid Y_{\text{obs}}) $$

where \(M\) represents missingness.

This does not mean that missingness is unrelated to observed patient characteristics.

It means that, conditional on appropriate observed information, the probability of missingness does not additionally depend on the unobserved outcome.

Why "LOCF" Is Not a Universal Solution

Last observation carried forward, or LOCF, has historically been used as a simple method for filling in missing outcomes.

However, carrying the last observation forward imposes strong assumptions about the patient's subsequent outcome.

For example:

$$ Y_{24}=Y_{12} $$

implicitly assumes that the Week 12 value adequately represents the Week 24 outcome.

That assumption may be implausible in many disease settings.

Modern principle: Missing-data handling should be justified in terms of the estimand and plausible assumptions rather than selected simply because it is easy to program.

Sensitivity Analysis

A primary analysis depends on assumptions.

A strong statistical analysis evaluates whether conclusions are robust to plausible departures from those assumptions.

For example:

1
Define the primary estimand.
2
Specify the primary estimator and assumptions.
3
Identify plausible departures from those assumptions.
4
Construct sensitivity analyses targeting those departures.
5
Assess whether the clinical conclusion changes.

Primary Analysis and Sensitivity Analysis

Component Primary Analysis Sensitivity Analysis
Purpose Estimate the prespecified primary treatment effect Evaluate robustness
Assumptions Primary assumptions Plausible alternative assumptions
Prespecification Required Should generally be prespecified
Interpretation Primary evidence Support for or challenge to robustness

Analysis Sets

Clinical trials commonly define several analysis populations.

Population General Purpose
Full Analysis Set Approximates the intention-to-treat principle and preserves randomized treatment comparisons.
Per-Protocol Set Focuses on patients sufficiently compliant with prespecified protocol requirements.
Safety Set Usually includes patients according to treatment actually received or exposure definition.
Other specialized sets May be defined for specific endpoints, pharmacokinetics, biomarkers, or protocol populations.

The exact definition of each set should be prespecified.

Intention-to-Treat Principle

The intention-to-treat principle is closely associated with randomized comparisons.

The central concept is that patients should generally be analyzed according to the treatment to which they were randomized for the primary comparison, consistent with the trial objective and prespecified analysis.

This preserves the benefits of randomization.

Important distinction: The intention-to-treat principle concerns treatment assignment and analysis population. It does not by itself specify how missing outcomes or intercurrent events should be handled. Those questions are addressed through the estimand and associated analysis strategy.

Interim Analyses

Interim analyses can provide information before the planned final analysis.

Potential purposes include:

  • Stopping for overwhelming efficacy
  • Stopping for futility
  • Sample-size reassessment
  • Safety monitoring
  • Adaptive design decisions

However, repeated access to accumulating data can affect type I error.

Therefore, confirmatory interim analyses generally require prespecified statistical procedures.

Group Sequential Designs

Suppose a trial has \(K\) planned looks at the accumulating data.

If each analysis were conducted independently at the ordinary significance level, the overall false-positive probability could exceed the intended level.

Group-sequential methods control the overall type I error through appropriately chosen critical boundaries.

The basic principle is:

$$ P(\text{at least one false rejection}) \leq \alpha $$

under the prespecified design.

Adaptive Designs

Adaptive designs allow prespecified changes to aspects of a trial based on accumulating information.

Potential adaptations include:

  • Sample-size modification
  • Dropping treatment arms
  • Changing allocation ratios
  • Population enrichment
  • Selection of doses
  • Stopping for efficacy or futility

The statistical validity of an adaptive design depends on how the adaptation is defined, when it occurs, who has access to the information, and how inference is protected.

Key principle: "Adaptive" does not mean "unplanned." A scientifically credible adaptive design requires a prespecified decision framework and appropriate control of operating characteristics.

Bayesian Methods

Bayesian methods can be used in clinical development when appropriately justified.

Bayes' theorem can be expressed as:

$$ p(\theta\mid y) \propto p(y\mid\theta)p(\theta) $$

where:

  • \(p(\theta)\) = prior distribution
  • \(p(y\mid\theta)\) = likelihood
  • \(p(\theta\mid y)\) = posterior distribution

Regulatory acceptability depends on the specific application, assumptions, prior specification, design characteristics, and operating properties.

Simulation

Simulation can be particularly valuable for complex trial designs.

It can evaluate:

  • Type I error
  • Power
  • Bias
  • Coverage probability
  • Decision probabilities
  • Operating characteristics
  • Impact of missing data
  • Adaptive decision rules

For a proposed design, simulation can estimate:

$$ \widehat{Power} = \frac{ \text{Number of simulated trials rejecting }H_0 }{ \text{Total simulated trials} } $$

Clinical Trial Design

Statistical guidance cannot be separated from trial design.

Important design decisions include:

  • Parallel-group versus crossover design
  • Superiority versus non-inferiority
  • Randomization ratio
  • Blinding
  • Control selection
  • Endpoint timing
  • Follow-up duration
  • Stratification
  • Interim analyses
  • Sample size

Comparator Selection

The statistical analysis is only as meaningful as the clinical comparison it represents.

Potential comparators include:

  • Placebo
  • Standard of care
  • Active comparator
  • Historical control
  • External control

The choice affects interpretation of the estimated treatment effect and can affect the assumptions required for inference.

Historical Controls

Historical controls may be useful in selected settings, particularly when randomized controlled trials are difficult or unethical.

However, historical comparisons are vulnerable to:

  • Changes in patient population
  • Changes in standard of care
  • Changes in endpoint definitions
  • Changes in diagnostic methods
  • Changes in supportive care
  • Selection bias
  • Temporal trends

Consequently, an apparently large treatment effect may partly reflect differences between populations rather than treatment efficacy.

External Controls and Real-World Data

External control data can provide useful context in selected development settings.

However, the fundamental statistical problem remains:

$$ \text{Observed Difference} = \text{Treatment Effect} + \text{Population Differences} + \text{Other Bias} $$

Careful design and adjustment are therefore required when external data are used for comparative inference.

Repeated Measures

Many clinical endpoints are measured repeatedly over time.

For example:

Patient   Baseline   Week 4   Week 8   Week 12   Week 24
001       62         55       51       48        45
002       58         57       55       54        53
003       64         60       59       61        63

Repeated-measures models can account for within-patient correlation.

A simplified model may be written:

$$ Y_{ij} = \mu + T_i + V_j + (TV)_{ij} + \varepsilon_{ij} $$

where \(i\) indexes patients and \(j\) indexes visits.

MMRM

A mixed model for repeated measures can be useful for longitudinal continuous endpoints.

A simplified formulation is:

$$ Y = X\beta + \varepsilon $$

with:

$$ \varepsilon \sim N(0,\Sigma) $$

where \(\Sigma\) represents the within-patient covariance structure.

The covariance structure should be appropriate for the endpoint and study design.

Data Transformation

Some endpoints require transformation before analysis.

For example, a logarithmic transformation is:

$$ Y^*=\log(Y) $$

A log-scale treatment difference can sometimes be transformed into a ratio scale.

For example:

$$ Ratio=e^{\Delta} $$

Interpretability should be considered carefully when selecting transformed scales.

Non-Normal Data

Not every clinical endpoint follows a normal distribution.

Examples include:

  • Counts
  • Proportions
  • Ordinal outcomes
  • Highly skewed biomarkers
  • Time-to-event outcomes
  • Recurrent events

The statistical model should reflect the endpoint's measurement scale and distributional properties.

Ordinal Endpoints

Ordinal outcomes contain ordered categories without necessarily having equal distances between categories.

Examples include:

  • Disease severity grades
  • Clinical response categories
  • Functional status categories

An ordinal model may use cumulative logits:

$$ \log \left( \frac{P(Y\leq k)} {P(Y>k)} \right) = \alpha_k-\beta X $$

The appropriateness of the model depends on the endpoint and its assumptions.

Safety Analysis

Statistical principles apply to safety as well as efficacy.

Safety analyses commonly examine:

  • Adverse events
  • Serious adverse events
  • Adverse events leading to discontinuation
  • Deaths
  • Laboratory abnormalities
  • Vital signs
  • Exposure-adjusted event rates

Safety interpretation often emphasizes descriptive statistics rather than formal hypothesis testing for every possible comparison.

Exposure-Adjusted Incidence Rates

When treatment exposure differs between groups, event rates can sometimes be summarized relative to exposure time.

A simplified incidence-rate calculation is:

$$ IR = \frac{\text{Number of events}} {\text{Total person-time at risk}} $$

The exact definition of exposure time and handling of recurrent events should be prespecified.

Multiplicity in Safety

Safety monitoring differs from confirmatory efficacy testing.

A trial may evaluate many adverse events without treating every descriptive comparison as a confirmatory hypothesis test.

The statistical objective is often to identify clinically important signals rather than to establish a single formal superiority claim.

Do not automatically apply an efficacy-style multiplicity adjustment to every safety table. The appropriate approach depends on the purpose of the analysis and the clinical question.

Biomarker and Subpopulation Analyses

Biomarker-defined populations can introduce additional statistical complexity.

Examples include:

  • Biomarker-positive patients
  • Biomarker-negative patients
  • High versus low biomarker expression
  • Genomic subgroups
  • Enriched populations

If treatment effect depends on biomarker status, the analysis may involve an interaction:

$$ \beta_3 = \text{Treatment}\times\text{Biomarker interaction} $$

The biomarker strategy should ideally be integrated into the overall clinical development strategy rather than introduced only after observing the data.

Multiplicity From Biomarker Exploration

Testing many biomarkers can produce apparently impressive findings by chance.

If \(m\) independent hypotheses are tested, the chance of at least one false positive increases with \(m\).

Therefore, exploratory biomarker findings often require confirmation in subsequent studies.

Protocol, SAP, and TFL Consistency

A regulatory statistical program should maintain traceability from protocol to analysis plan to outputs.

A
Protocol: defines the clinical objectives and key endpoints.
B
Estimand: defines the treatment effect of interest.
C
SAP: specifies the estimator, analysis population, assumptions, and inference.
D
ADaM: implements the analysis-ready derivations.
E
TFLs: present the resulting evidence.

Statistical Programming Implications

EMA statistical expectations ultimately have practical consequences for statistical programming.

The programmer may need to implement:

  • Analysis populations
  • Baseline derivations
  • Endpoint derivations
  • Visit windows
  • Intercurrent-event flags
  • Imputation datasets
  • Multiplicity procedures
  • Covariate-adjusted models
  • Subgroup analyses
  • Sensitivity analyses
  • Confidence intervals
  • Graphical diagnostics

CDISC and Regulatory Statistics

Statistical methodology and data standards solve different problems.

Layer Primary Purpose
Protocol Defines clinical objectives and study design
Estimand Defines the treatment effect of interest
SAP Defines statistical implementation
SDTM Structures collected clinical-trial data
ADaM Structures analysis-ready data
TFLs Communicate statistical results

The statistical analysis must remain traceable across these layers.

Example: From Clinical Question to Analysis

Consider a hypothetical randomized trial evaluating Treatment A versus Treatment B for a continuous symptom score at Week 24.

Step 1: Clinical Question

Does Treatment A improve symptoms compared with Treatment B at Week 24?

Step 2: Estimand

Suppose the treatment-policy strategy is selected for treatment discontinuation.

Step 3: Endpoint

Change from baseline in symptom score at Week 24.

Step 4: Treatment Effect

$$ \Delta = \mu_A-\mu_B $$

Step 5: Estimator

A prespecified covariate-adjusted longitudinal model could be used, depending on the design and estimand.

Step 6: Inference

Report:

  • Adjusted treatment difference
  • Standard error
  • Confidence interval
  • P-value where appropriate

Step 7: Sensitivity

Evaluate plausible alternative assumptions about missing outcomes and intercurrent events.

Example Statistical Analysis Specification

Primary endpoint:
Change from baseline in symptom score at Week 24.

Population:
Full Analysis Set.

Primary estimand:
Treatment-policy strategy for treatment discontinuation.

Estimator:
Prespecified repeated-measures model with treatment,
visit, treatment-by-visit interaction, baseline score,
and randomized stratification factors as appropriate.

Treatment effect:
Adjusted between-group difference at Week 24.

Inference:
Two-sided 95% confidence interval and prespecified
hypothesis test.

Sensitivity analyses:
Alternative assumptions addressing missing data and
intercurrent events.

Clinical Interpretation of P-Values

A p-value measures the compatibility of the observed data with a specified null hypothesis under the assumptions of the statistical procedure.

It does not directly measure:

  • The probability that the treatment works
  • The magnitude of clinical benefit
  • The probability that the null hypothesis is true
  • The importance of the result

For example:

$$ p=0.003 $$

does not mean there is a 99.7% probability that the treatment is effective.

Confidence Intervals

Confidence intervals provide information about the estimated treatment effect and its uncertainty.

Suppose the treatment difference is:

$$ \widehat{\Delta}=4.2 $$

with a 95% confidence interval:

$$ [1.1,\;7.3] $$

This communicates considerably more information than a p-value alone.

The interval suggests both the direction and the plausible precision of the estimated effect under the model and assumptions used.

Statistical Significance vs Clinical Relevance

Result Potential Interpretation
Small effect, narrow CI, p < 0.05 Statistically convincing but potentially clinically modest
Large effect, wide CI, p > 0.05 Potentially important but imprecisely estimated
Large effect, narrow CI excluding no effect Strong statistical and potentially clinical evidence
CI crosses clinically important thresholds Interpretation may remain uncertain

Multiplicity and Confidence Intervals

When multiple confirmatory hypotheses are tested, the interpretation of confidence intervals should be consistent with the multiplicity strategy.

A nominal 95% confidence interval attached to one of many unadjusted comparisons does not necessarily provide the desired family-wise coverage for the entire hypothesis family.

Therefore, the reporting strategy should align with the inferential procedure.

Missing Data and Regulatory Interpretation

A missing-data strategy should be clinically interpretable.

Consider a patient who discontinues treatment because of worsening disease.

Simply treating the subsequent observation as missing may obscure an important clinical event.

The estimand framework asks whether the event should:

  • Be incorporated into the outcome
  • Be handled under a treatment-policy strategy
  • Define a hypothetical scenario
  • Define a while-on-treatment effect
  • Lead to a different estimand altogether

Regulatory Review Perspective

A useful way to anticipate statistical review is to ask:

  • What exactly is the treatment effect being claimed?
  • Was that effect prespecified?
  • Does the analysis estimate the stated estimand?
  • Were intercurrent events handled consistently with the estimand?
  • Are missing-data assumptions plausible?
  • Were sensitivity analyses adequate?
  • Was multiplicity controlled where required?
  • Were subgroup claims supported by appropriate interaction assessments?
  • Are the analysis populations clearly defined?
  • Can every important result be reproduced from the analysis datasets?
Regulatory principle: A statistically sophisticated analysis is not necessarily a good analysis if it answers a different question from the one posed by the clinical trial.

Common Statistical Mistakes

  1. Starting with the statistical model. Define the clinical question and estimand first.
  2. Treating ICH E9 as a programming manual. E9 provides statistical principles rather than code-level implementation instructions.
  3. Ignoring intercurrent events. Treatment discontinuation and rescue medication can change the meaning of an endpoint.
  4. Using a generic missing-data method without justification. The method should follow from the estimand and assumptions.
  5. Calling every subgroup p-value evidence of heterogeneity. Differences between subgroup-specific p-values do not establish treatment interaction.
  6. Ignoring multiplicity. Multiple confirmatory claims require an appropriate inferential strategy.
  7. Confusing non-significance with non-inferiority. Non-inferiority requires a prespecified margin and appropriate analysis.
  8. Using post-randomization covariates casually. Variables affected by treatment require careful causal consideration.
  9. Reporting only p-values. Treatment-effect estimates and confidence intervals are important for interpretation.
  10. Failing to prespecify important analysis decisions. Post hoc analytical flexibility can undermine confirmatory inference.

EMA Statistical Guidance vs ICH Guidance

It is useful to distinguish between EMA and ICH roles.

Source Role
ICH International harmonization of technical requirements among participating regulatory regions.
EMA European regulatory framework and implementation of applicable scientific and methodological guidance.
ICH E9 Foundational statistical principles for clinical trials.
ICH E9(R1) Estimands and sensitivity analysis framework.
EMA therapeutic-area guidance Additional endpoint and disease-specific expectations.
EMA methodological guidance Specific statistical or methodological issues.

Consequently, a statistician preparing an EU regulatory submission should not look only for documents whose title contains the word "statistics."

Relevant statistical expectations can also appear in disease-specific and endpoint-specific guidance.

Therapeutic-Area Guidance

Statistical requirements can differ according to therapeutic area.

Examples include:

  • Oncology
  • Neurology
  • Cardiovascular disease
  • Rare diseases
  • Vaccines
  • Anti-infective therapy
  • Psychiatry
  • Metabolic disease

A statistical analysis plan should therefore consider both the general statistical framework and any applicable disease-specific guidance.

Rare Diseases and Small Populations

Small-population trials create statistical challenges because conventional large-sample methods may provide limited information.

Challenges include:

  • Limited sample size
  • Rare endpoints
  • Heterogeneous disease
  • Difficulty recruiting controls
  • Limited prior information
  • Large uncertainty

The key principle is not that conventional statistics are abandoned, but that the statistical design must recognize the information constraints.

Precision Becomes Especially Important

With small samples, an estimated treatment effect may have substantial uncertainty.

For example:

$$ \widehat{\Delta}=12 \qquad 95\%\,CI=[-2,\;26] $$

The point estimate suggests a potentially important effect, but the confidence interval demonstrates substantial uncertainty.

Data Cutoffs

Regulatory analyses generally require clearly defined data cutoffs.

A cutoff specifies the date through which relevant information is included.

For example:

Data cutoff:
31 December 2026

Primary efficacy follow-up:
24 weeks

Safety follow-up:
30 days after last dose

Different endpoints may use different analysis rules, but the conventions should be prespecified and traceable.

Protocol Deviations

Protocol deviations can affect analysis populations and interpretation.

Examples include:

  • Incorrect treatment administration
  • Eligibility violations
  • Prohibited medication
  • Missed assessments
  • Major timing deviations

The impact of protocol deviations should be evaluated systematically rather than simply removing inconvenient observations.

Blinding and Statistical Integrity

Statistical decisions should be protected from unnecessary access to unblinded treatment information.

Important processes may include:

  • Blinded data review
  • Independent statistical programming
  • Database lock procedures
  • Predefined analysis specifications
  • Independent validation
  • Controlled unblinding

Statistical Programming Quality Control

Regulatory statistical analyses require reproducibility.

A robust programming workflow should include:

1
Define the analysis specification.
2
Develop analysis datasets.
3
Program primary analyses.
4
Program sensitivity analyses.
5
Generate tables, listings, and figures.
6
Perform independent QC.
7
Reconcile outputs against source data and specifications.

Traceability

A regulatory statistical result should be traceable from the reported output back to the underlying patient-level information.

A simplified traceability chain is:

$$ \text{Source Data} \rightarrow \text{SDTM} \rightarrow \text{ADaM} \rightarrow \text{Analysis} \rightarrow \text{TFL} $$

For example, a treatment difference in a primary efficacy table should be traceable to:

  • The endpoint definition
  • The analysis population
  • The analysis dataset
  • The statistical model
  • The treatment contrast
  • The underlying patient observations

Example: Complete Analysis Chain

Stage Example
Clinical objective Compare symptom improvement at Week 24
Estimand Treatment-policy effect
Endpoint Change from baseline at Week 24
Population Full Analysis Set
Estimator Prespecified repeated-measures model
Contrast Treatment A − Treatment B
Inference Two-sided confidence interval and hypothesis test
Sensitivity Alternative missing-data assumptions
Output Primary efficacy table

How to Read an EMA Statistical Guideline

When reading a regulatory statistical document, do not read every section with the same objective.

Instead, identify:

1
Scope: What types of trials or analyses does the guidance cover?
2
Definitions: Which statistical concepts are formally defined?
3
Recommendations: Which practices are encouraged or discouraged?
4
Exceptions: Under what circumstances can an alternative approach be justified?
5
Implementation: What does the guidance imply for the protocol and SAP?

Guideline Language Matters

Regulatory documents may use different levels of normative language.

A statistician should distinguish among:

  • Requirements arising from applicable regulation
  • Guideline recommendations
  • Preferred methodological practices
  • Examples
  • Areas where alternative approaches may be scientifically justified

An alternative statistical approach should generally be accompanied by a clear scientific rationale.

Pre-Specification Is Central

A recurring principle across regulatory statistics is prespecification.

Important decisions should be established before the relevant treatment effects are examined.

This includes:

  • Primary endpoint
  • Primary estimand
  • Analysis population
  • Statistical model
  • Covariates
  • Multiplicity strategy
  • Missing-data strategy
  • Sensitivity analyses
  • Subgroup analyses
  • Interim analyses
Why this matters: Prespecification reduces the opportunity to choose analytical methods based on which approach produces the most favorable observed result.

Post Hoc Analyses

Post hoc analyses are not automatically invalid.

They can be useful for:

  • Exploring unexpected findings
  • Generating hypotheses
  • Understanding treatment heterogeneity
  • Investigating safety signals
  • Supporting future studies

However, post hoc findings generally require more cautious interpretation than prespecified confirmatory analyses.

Confirmatory vs Exploratory Analysis

Characteristic Confirmatory Exploratory
Objective Test prespecified clinical hypotheses Generate or investigate hypotheses
Prespecification Strongly expected May be less extensive
Multiplicity Important for confirmatory claims Interpret with caution
Interpretation Supports regulatory claims Hypothesis-generating
Post hoc flexibility Limited Greater, but must be disclosed

Practical SAP Checklist

Before finalizing a statistical analysis plan for an EMA-facing clinical trial, ask:

1
Is the primary clinical question clearly stated?
2
Is the primary estimand explicitly defined?
3
Are intercurrent events addressed?
4
Is the analysis population clearly defined?
5
Is the primary estimator consistent with the estimand?
6
Are multiplicity procedures prespecified?
7
Are missing-data assumptions explicit?
8
Are sensitivity analyses clinically meaningful?
9
Are subgroup analyses prespecified and appropriately interpreted?
10
Can the analysis be reproduced from the analysis datasets?

EMA Statistical Guidance: A Practical Hierarchy

A useful mental model is:

$$ \text{Clinical Objective} \rightarrow \text{Estimand} \rightarrow \text{Design} \rightarrow \text{Analysis} \rightarrow \text{Sensitivity} \rightarrow \text{Interpretation} $$

Each stage constrains the next.

For example, if the clinical objective changes, the estimand may change. If the estimand changes, the appropriate estimator may change. If the estimator changes, the required sensitivity analyses may also change.

What Statistical Programmers Should Take Away

Statistical programming is not simply the final implementation step after the statistical methodology has been decided.

A programmer who understands the regulatory statistical framework can identify important inconsistencies earlier.

For example:

  • An ADaM variable may not implement the intended estimand.
  • A treatment-discontinuation flag may be missing.
  • A subgroup variable may not match the SAP definition.
  • A multiplicity adjustment may not match the planned hypothesis hierarchy.
  • A sensitivity analysis may use a different analysis population unintentionally.
  • A table may report an unadjusted estimate when the SAP requires an adjusted estimate.

What Statistical Programmers Should Validate

Area Programming Check
Population Analysis flags match SAP definitions
Baseline Baseline selection is reproducible
Endpoint Derived endpoint matches protocol/SAP
Intercurrent events Event flags and post-event handling are correct
Model Covariates, visits, interactions, and contrasts match specification
Multiplicity Adjusted inference follows the prespecified procedure
Missing data Imputation or modeling assumptions are implemented correctly
Subgroups Definitions and interaction terms are correct
Outputs TFLs reconcile to analysis datasets

EMA Statistical Guidance and R

The statistical principles described by EMA and ICH are software-independent.

They can be implemented using:

  • SAS
  • R
  • Python
  • Validated statistical software
  • Specialized clinical-trial analysis systems

For example, a simple adjusted linear model in R might be:

fit <- lm(
  CHG ~ TRT01P + BASE + STRATUM,
  data = adsl
)

summary(fit)

The important regulatory question is not whether R or SAS was used.

The important question is whether the implementation correctly represents the prespecified statistical method and can be appropriately validated.

Example: Multiplicity in R

For exploratory illustration, adjusted p-values can be obtained using standard procedures.

p_values <- c(
  endpoint1 = 0.012,
  endpoint2 = 0.031,
  endpoint3 = 0.087,
  endpoint4 = 0.004
)

p.adjust(
  p_values,
  method = "holm"
)

In a regulatory submission, the exact multiplicity procedure should follow the prespecified testing strategy rather than simply selecting a convenient software option after seeing the results.

Example: Confidence Intervals in R

fit <- lm(
  CHG ~ TRT01P + BASE,
  data = adrs
)

confint(
  fit,
  "TRT01PActive"
)

The resulting estimate and interval should be interpreted in the context of the specified model and treatment contrast.

Regulatory Statistics Is About More Than P-Values

A mature statistical analysis should answer several questions simultaneously:

  • What is the treatment effect?
  • How large is it?
  • How precise is it?
  • What assumptions were required?
  • Are the conclusions robust?
  • How do results vary across clinically relevant populations?
  • What happened after treatment discontinuation?
  • How were missing observations handled?

This is why modern regulatory statistics emphasizes estimands, effect estimates, confidence intervals, and sensitivity analysis rather than relying on a p-value alone.

A Complete Regulatory Statistical Workflow

1
Define the clinical objective.
2
Define the target population.
3
Define the treatment comparison.
4
Define the endpoint.
5
Define intercurrent-event strategies.
6
Specify the estimand.
7
Select the estimator.
8
Determine sample size and operating characteristics.
9
Define multiplicity control.
10
Define sensitivity analyses.
11
Implement analysis datasets and programs.
12
Validate and interpret the results.

Final Statistical Checklist

Question Check
Clinical question Is the treatment question unambiguous?
Estimand Are all key attributes specified?
Intercurrent events Is each important event addressed?
Endpoint Is the endpoint operationally defined?
Population Is the analysis set prespecified?
Estimator Does the method estimate the intended estimand?
Multiplicity Are confirmatory claims appropriately protected?
Missing data Are assumptions and sensitivity analyses appropriate?
Subgroups Are subgroup conclusions based on appropriate comparisons?
Precision Are confidence intervals reported?
Reproducibility Can results be traced from source data to TFLs?
Interpretation Are statistical and clinical relevance distinguished?

The Most Important Concept

The central lesson of EMA statistical guidance is not a particular statistical test.

It is the principle of maintaining a coherent connection between the clinical question, trial design, estimand, statistical analysis, assumptions, and interpretation.

The progression can be summarized as:

$$ \boxed{ \text{Clinical Question} \rightarrow \text{Estimand} \rightarrow \text{Estimator} \rightarrow \text{Estimate} \rightarrow \text{Interpretation} } $$

If those components are aligned, the statistical analysis has a clear scientific purpose.

If they are not aligned, even a technically sophisticated model can answer the wrong question.

Bottom line: EMA statistical guidance should be viewed as a connected methodological framework rather than a single statistical rulebook. ICH E9 provides the foundational principles, while ICH E9(R1) places the clinical question and estimand at the center of modern trial analysis. Other methodological and therapeutic-area guidance addresses issues such as multiplicity, missing data, covariate adjustment, subgroup analysis, survival endpoints, adaptive designs, and specialized populations. For statisticians and programmers, the practical objective is to translate these principles into a prespecified, reproducible, traceable, and clinically meaningful analysis.

References

International Council for Harmonisation (ICH). E9: Statistical Principles for Clinical Trials. ICH Harmonised Guideline.

International Council for Harmonisation (ICH). E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials. ICH Harmonised Guideline.

European Medicines Agency. ICH E9 Statistical Principles for Clinical Trials. European regulatory implementation of the ICH statistical principles.

European Medicines Agency. ICH E9(R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials. European regulatory implementation of the estimand framework.

European Medicines Agency. Guideline on Adjustment for Baseline Covariates in Clinical Trials. Scientific and methodological guidance for covariate adjustment.

European Medicines Agency. Guideline on Multiplicity Issues in Clinical Trials. Methodological considerations for multiple testing and confirmatory inference.

European Medicines Agency. Guideline on Clinical Trials in Small Populations. Statistical and methodological considerations for rare diseases and small populations.

European Medicines Agency. Guideline on the Investigation of Subgroups in Confirmatory Clinical Trials. Considerations for treatment-effect heterogeneity and subgroup analysis.