Introduction
A randomized clinical trial is usually designed to estimate a treatment effect in a prespecified target population. Regulatory review, however, rarely stops with the overall treatment effect.
Regulators also need to understand whether the observed benefit and risk are reasonably applicable across important subsets of the intended population.
These subsets may be defined by demographic characteristics such as age or sex, disease characteristics such as disease severity or stage, prior treatment, geographic region, baseline risk, organ function, or other clinically relevant characteristics.
This is the purpose of subgroup analysis.
This distinction is fundamental.
A subgroup can have a nonsignificant treatment effect simply because it is small. Conversely, a statistically significant subgroup finding can arise by chance when many subgroups are examined.
Regulatory interpretation therefore considers the magnitude, precision, consistency, biological plausibility, prespecification, and clinical context of subgroup findings rather than relying on a single p-value.
What Is a Subgroup?
A subgroup is a subset of trial participants defined by one or more baseline characteristics.
Examples include:
- Age category
- Sex
- Race or ethnicity
- Geographic region
- Disease severity
- Prior therapy
- Baseline biomarker status
- Renal function
- Hepatic function
- Baseline risk category
- Performance status
- Body weight
- Comorbidity status
The defining characteristic is that the subgroup variable is generally known at or before treatment assignment and can therefore be used to characterize heterogeneity of the randomized treatment effect.
Subgroup Analysis Is Not the Same as Subsetting After Treatment
One of the most important statistical distinctions is between a baseline subgroup and a subgroup defined by something that occurs after randomization.
For example, consider:
- Age at baseline — appropriate subgroup characteristic
- Baseline biomarker — appropriate subgroup characteristic
- Prior treatment — appropriate subgroup characteristic
- Patients who discontinued treatment — potentially post-randomization subset
- Patients who developed an adverse event — post-randomization subset
- Patients who achieved response — outcome-defined subset
The latter categories can create serious interpretational problems because treatment itself may influence membership in the subset.
The Regulatory Question
The overall treatment effect answers:
where \(\theta\) represents the treatment effect in the primary analysis population.
A subgroup analysis asks a different question:
For example:
The regulatory question is usually not simply whether each subgroup has a statistically significant treatment effect.
Instead, the key question is whether the treatment effects differ materially across subgroups.
The Interaction Is the Central Statistical Concept
Suppose a binary subgroup variable \(G\) defines two groups. A simple treatment-by-subgroup model can be written as:
where:
- \(T\) = treatment indicator
- \(G\) = subgroup indicator
- \(T\times G\) = treatment-by-subgroup interaction
- \(\beta_3\) = parameter describing differential treatment effect between subgroups
The interaction term is therefore fundamentally different from the treatment effect within one subgroup.
Why "Significant in One Subgroup, Not Significant in Another" Is Wrong
Consider the following hypothetical results:
| Subgroup | Treatment Effect | 95% CI | p-value |
|---|---|---|---|
| Age <65 | −2.8 | −4.1 to −1.5 | <0.001 |
| Age ≥65 | −2.1 | −4.5 to 0.3 | 0.087 |
It is tempting to conclude:
That conclusion does not follow.
The older subgroup has a wider confidence interval because it may contain fewer participants.
The point estimates are actually fairly similar:
The correct question is whether there is convincing evidence that the treatment effects differ.
An Actual Regulatory-Style Forest Plot
Forest plots are one of the most common ways to present subgroup treatment effects.
The values are simulated for educational purposes. A forest plot should be interpreted using treatment-effect estimates, confidence intervals, subgroup sizes, clinical plausibility, and interaction evidence rather than by visually checking whether individual confidence intervals cross the null value.
How to Read a Forest Plot
A subgroup forest plot usually contains:
- The subgroup definition
- The number of patients in each subgroup
- The treatment-effect estimate
- A confidence interval
- A vertical line representing the null effect
- Sometimes an interaction p-value
For a hazard ratio:
is the usual null value.
For an odds ratio or risk ratio, the null value is also:
For a mean difference:
is the null value.
Confidence Intervals Are Essential
A subgroup point estimate without a confidence interval is difficult to interpret.
Suppose two subgroups have:
| Subgroup | HR | 95% CI |
|---|---|---|
| Subgroup A | 0.72 | 0.58–0.89 |
| Subgroup B | 0.72 | 0.39–1.34 |
The point estimates are identical.
The difference is precision.
The second subgroup may simply have substantially fewer events.
A regulatory reviewer should therefore distinguish:
- Observed effect magnitude
- Statistical precision
- Event count
- Subgroup sample size
- Evidence for treatment-effect heterogeneity
What Does a Treatment-by-Subgroup Interaction Test?
Suppose the treatment effects are:
The interaction hypothesis is approximately:
against:
The interaction test therefore directly addresses whether the treatment effect differs between the subgroups.
Interaction P-Values Need Careful Interpretation
An interaction p-value below a conventional threshold can provide evidence of heterogeneity.
But an interaction p-value above that threshold does not prove that the treatment effects are identical.
For example:
means that the analysis did not provide strong statistical evidence of a difference in treatment effects.
It does not establish:
with certainty.
Statistical Significance Is Not Clinical Significance
A statistically significant interaction may be clinically trivial.
Conversely, a clinically important difference may fail to reach statistical significance when subgroup sample sizes are small.
For example:
could represent a clinically important difference.
But if subgroup 2 contains very few events, the confidence interval may be extremely wide.
Regulatory interpretation therefore requires both statistical and clinical judgment.
Prespecification Matters
A credible subgroup strategy should be developed before the primary results are known.
The statistical analysis plan should identify important subgroup variables and explain how they will be analyzed.
Examples include:
- Age categories
- Sex
- Race and ethnicity
- Geographic region
- Baseline disease severity
- Prior treatment
- Biomarker status
- Relevant baseline prognostic factors
Prespecification helps distinguish a scientifically motivated subgroup investigation from a search through the data for a favorable result.
Why Regulators Care About Prespecified Subgroups
Suppose a trial examines 20 subgroup variables.
Even if treatment truly has the same effect everywhere, random variation can produce apparently unusual results.
If the sponsor searches enough subgroups, one or more apparently favorable or unfavorable findings can occur by chance.
Prespecification therefore improves interpretability.
How Many Subgroups Should Be Planned?
There is no universal magic number.
The appropriate set depends on:
- The disease
- The mechanism of action
- The intended treatment population
- Known prognostic factors
- Known effect modifiers
- Safety considerations
- Demographic representation
- Regulatory expectations
A subgroup should have a scientific reason for inclusion.
Common Regulatory Subgroup Domains
| Domain | Examples | Typical Rationale |
|---|---|---|
| Demographic | Age, sex, race, ethnicity | Generalizability and potential effect modification |
| Disease | Stage, severity, disease duration | Baseline prognosis and disease biology |
| Prior therapy | Prior treatment, treatment line | Potential treatment-effect modification |
| Biomarker | Positive/negative status | Mechanistic or predictive hypothesis |
| Geographic | Region, country grouping | Population and practice differences |
| Organ function | Renal or hepatic function | Exposure, safety, and efficacy considerations |
| Baseline risk | Low, intermediate, high risk | Prognostic heterogeneity |
Demographic Subgroups Have a Special Role
Regulators expect sponsors to consider whether the study population adequately represents the intended population.
Important demographic variables can include:
- Sex
- Age
- Race
- Ethnicity
- Geographic region
However, subgroup analysis cannot compensate for inadequate enrollment.
If only a very small number of participants from a demographic group are enrolled, a precise subgroup treatment-effect estimate may not be possible.
Subgroup Representation and Modern Trial Planning
Modern regulatory expectations increasingly emphasize inclusion of a clinical trial population that reflects the patients likely to receive the product.
FDA's current diversity guidance discusses both demographic characteristics and non-demographic characteristics of populations, including characteristics such as organ dysfunction, comorbid conditions, disabilities, extremes of body weight, and populations with low disease prevalence.
This means subgroup considerations should begin during trial design, not only after database lock.
Small Subgroups
Small subgroup sizes are one of the most common problems in regulatory subgroup analysis.
Suppose:
| Subgroup | Treatment N | Control N |
|---|---|---|
| Majority population | 450 | 450 |
| Small subgroup | 28 | 25 |
The small subgroup may produce a highly unstable estimate.
For a time-to-event endpoint, the number of events may be even more important than the number of participants.
For example:
may provide very little information about treatment-effect heterogeneity.
Wide Confidence Intervals Are Informative
A wide confidence interval should not automatically be interpreted as evidence against treatment benefit.
It usually means the estimate is imprecise.
For example:
The point estimate suggests benefit, but the confidence interval indicates substantial uncertainty.
The appropriate conclusion is not:
A more appropriate statement is:
Forest Plot With a Small Subgroup
The small subgroup is not evidence of treatment failure. Its confidence interval demonstrates limited information and uncertainty.
Do Confidence Intervals Need to Overlap?
No.
A common visual rule is:
This is not a reliable statistical rule.
The appropriate comparison is based on the statistical model and treatment-by- subgroup interaction or another appropriate heterogeneity analysis.
What Does "Consistency" Mean?
Consistency does not necessarily mean identical numerical estimates.
Suppose the treatment effects are:
These are not identical.
But they may be entirely compatible with a common underlying treatment effect, especially when confidence intervals are considered.
Regulatory assessment is therefore concerned with whether there is credible and clinically meaningful heterogeneity, not whether every estimate is numerically identical.
Heterogeneity Can Be Quantitative or Qualitative
Consider two types of heterogeneity.
Quantitative Heterogeneity
The treatment effect points in the same direction but differs in magnitude.
Qualitative Heterogeneity
The apparent direction of effect differs.
Qualitative heterogeneity is generally more concerning because the apparent benefit in one subgroup is accompanied by an apparent lack of benefit or potential harm in another.
Biological Plausibility
A subgroup finding is more credible when there is a plausible scientific explanation for it.
For example, a treatment designed to inhibit a specific molecular pathway may reasonably be expected to have different effects according to a validated biomarker.
In contrast, a dramatic difference discovered only after examining dozens of unrelated baseline characteristics may be much less convincing.
Internal Consistency Across Studies
Subgroup interpretation becomes stronger when similar patterns appear across independent randomized trials.
For example, suppose three trials produce:
| Trial | Subgroup A | Subgroup B |
|---|---|---|
| Trial 1 | HR 0.70 | HR 0.78 |
| Trial 2 | HR 0.74 | HR 0.82 |
| Trial 3 | HR 0.69 | HR 0.76 |
The consistency of the pattern provides substantially more context than one isolated subgroup analysis.
Pooling Across Studies
Regulatory submissions may include subgroup analyses across multiple controlled trials when appropriate.
Pooling can increase the amount of information available for important subgroups.
However, pooled analyses require careful consideration of:
- Trial design
- Patient population
- Endpoint definition
- Treatment regimen
- Follow-up
- Control group
- Study heterogeneity
- Statistical model
A pooled estimate should not be created simply because the individual subgroups are small.
Multiplicity
Subgroup analyses create a multiplicity problem when many hypotheses are examined.
Suppose a trial evaluates:
- Age
- Sex
- Race
- Ethnicity
- Region
- Prior treatment
- Biomarker status
- Disease severity
- Baseline risk
- Performance status
If each produces multiple subgroup comparisons, many statistical tests may be performed.
The more opportunities there are to observe an apparently unusual result, the greater the chance of observing one by chance.
Exploratory Versus Confirmatory Subgroup Claims
This distinction is essential.
| Feature | Exploratory Subgroup | Confirmatory Subgroup |
|---|---|---|
| Purpose | Assess consistency or generate hypotheses | Test a prespecified treatment-effect claim |
| Prespecified | Preferably | Essential |
| Multiplicity | Interpret findings cautiously | Must be addressed within the testing strategy |
| Evidence standard | Supportive/descriptive | Confirmatory |
| Post hoc findings | Can be useful | Generally weak as confirmatory evidence |
A subgroup that was not part of the confirmatory testing strategy should not usually be presented as though it were a prespecified confirmatory claim.
When a Subgroup Is Part of the Confirmatory Strategy
Some trials are explicitly designed to establish treatment benefit in a specific subgroup.
For example, a trial may target:
In that setting, the subgroup may be the primary analysis population rather than a secondary exploratory subgroup.
The statistical testing strategy should then be aligned with the scientific claim.
Hierarchical Testing and Subgroups
If multiple hypotheses are part of a confirmatory strategy, the sponsor should define how statistical error will be controlled.
Possible approaches include:
- Hierarchical testing
- Gatekeeping procedures
- Alpha allocation
- Closed testing procedures
- Other prespecified multiplicity-control strategies
The appropriate method depends on the claims being tested and the structure of the trial.
Do You Need to Adjust Every Exploratory Subgroup P-Value?
Not necessarily.
An exploratory subgroup analysis can be descriptive.
The key is to clearly identify it as exploratory and avoid presenting a large collection of nominal p-values as if each represented an independent confirmatory conclusion.
Continuous Subgroup Variables
Age, body weight, renal function, and other variables are naturally continuous.
Artificial categorization can lose information.
For example:
may be convenient for presentation, but the biological relationship between age and treatment effect may not have a discontinuity at exactly 65 years.
When scientifically appropriate, continuous interaction models can be considered.
where \(A\) is age as a continuous variable.
Why Arbitrary Cut Points Can Be Dangerous
Suppose investigators try:
- Age 60
- Age 65
- Age 70
- Age 75
- Age 80
and report whichever cut point produces the strongest apparent difference.
That is a form of data-driven subgroup selection.
The resulting finding may be difficult to interpret because the cut point was chosen after examining the data.
Subgroup Definitions Must Be Reproducible
A regulatory submission should make it possible to determine exactly who belongs to each subgroup.
For example, "older patients" is inadequate.
A reproducible definition might be:
The precise definition should match the protocol, SAP, analysis dataset, and tables.
Stratified Randomization Does Not Automatically Solve Subgroup Analysis
A trial may stratify randomization by region or disease severity.
That improves balance in the randomized allocation.
It does not automatically establish that the treatment effect is consistent within each stratum.
The subgroup analysis still requires appropriate estimation and interpretation.
Covariate Adjustment Versus Subgroup Analysis
A baseline covariate can improve precision of the overall treatment-effect estimate.
A subgroup analysis asks whether treatment effects differ according to that covariate.
These are related but different questions.
| Analysis | Main Question |
|---|---|
| Covariate adjustment | Can baseline information improve estimation of the overall treatment effect? |
| Subgroup analysis | Does the treatment effect differ across levels of a baseline characteristic? |
| Interaction analysis | Is there statistical evidence that treatment effects differ? |
Subgroup Analysis and the Estimand Framework
ICH E9(R1) emphasizes defining the treatment effect of interest through an estimand framework.
This matters for subgroup analysis because the subgroup must be defined in a way that corresponds to the clinical question.
For example, the question might be:
That is conceptually different from:
The second definition could involve information that occurs after treatment and therefore require a different causal framework.
Subgroup Analysis of Time-to-Event Endpoints
For survival endpoints, the most common treatment-effect measure is a hazard ratio.
A subgroup forest plot might display:
where \(g\) denotes the subgroup.
A Cox model with an interaction term can be written as:
coxph(
Surv(time, status) ~
TRT * SUBGROUP,
data = analysis_data
)
The interaction term assesses whether the treatment effect differs across subgroup levels.
Subgroup Analysis of Binary Endpoints
For a binary endpoint, treatment effects may be expressed as:
- Risk difference
- Risk ratio
- Odds ratio
For example:
The treatment-by-subgroup interaction can then evaluate whether the risk difference varies across subgroup levels.
Subgroup Analysis of Continuous Endpoints
For a continuous endpoint, treatment effects may be represented by a mean difference:
An ANCOVA-style model can incorporate treatment, subgroup, baseline covariates, and treatment-by-subgroup interaction.
Baseline Adjustment
Baseline adjustment may improve precision for continuous outcomes.
For example:
lm(
change_from_baseline ~
TRT * SUBGROUP +
baseline_value,
data = analysis_data
)
The interaction term remains the key component for evaluating whether treatment effects differ by subgroup.
Forest Plot Versus Subgroup Table
A forest plot is visually efficient.
A table provides additional detail.
A regulatory submission will often need both.
| Feature | Forest Plot | Detailed Table |
|---|---|---|
| Effect estimate | Yes | Yes |
| Confidence interval | Yes | Yes |
| Sample size | Usually | Yes |
| Events | Optional | Often important |
| Raw responder counts | Usually no | Yes |
| Interaction p-value | Often | Often |
| Detailed descriptive statistics | Limited | Excellent |
Regulatory-Style Forest Plot Design
A useful forest plot should allow the reviewer to determine quickly:
- What the treatment effect is
- How precise it is
- How many patients contribute
- Whether effects are directionally consistent
- Whether there is evidence of interaction
Avoid excessive decoration.
The objective is rapid scientific interpretation.
Forest Plot Example With Demographic Subgroups
Demographic subgroup analyses should be interpreted together with the number of participants, precision of estimates, overall evidence, and adequacy of representation.
What Regulators Look for in a Subgroup Table
A strong subgroup table usually provides enough information to distinguish effect magnitude from uncertainty.
Depending on endpoint and analysis, useful fields include:
- Subgroup definition
- Treatment sample size
- Control sample size
- Number of events
- Effect estimate
- Confidence interval
- Interaction p-value
- Relevant descriptive statistics
Safety Subgroups
Subgroup analysis is not limited to efficacy.
Regulators also need to understand whether safety risks differ across important patient groups.
Examples include:
- Age
- Sex
- Race and ethnicity
- Renal function
- Hepatic function
- Body weight
- Comorbidities
- Concomitant medications
Safety subgroup analysis can be especially challenging because adverse events may be uncommon.
A very small subgroup may produce no observed events even though the true risk is not zero.
Exposure Matters in Safety Subgroups
If treatment exposure differs substantially between subgroups, comparing raw adverse-event counts can be misleading.
Depending on the question, regulators may consider:
- Number of exposed participants
- Patient-years of exposure
- Treatment duration
- Exposure-adjusted incidence rates
- Severity and seriousness
- Temporal relationship
Race and Ethnicity Analyses
Race and ethnicity analyses require particular care.
They may be relevant to:
- Population representation
- Generalizability
- Potential differences in treatment effect
- Safety
- Pharmacokinetics
- Pharmacogenetic considerations
However, race and ethnicity are not interchangeable concepts and should not be treated as purely biological proxies.
The categories used in the analysis should be clearly defined and consistent with the data-collection strategy.
Geographic Subgroups
Regional differences may arise from:
- Patient characteristics
- Standard of care
- Diagnostic practices
- Endpoint ascertainment
- Clinical management
- Enrollment differences
- Random variation
A geographic interaction should therefore not automatically be interpreted as a biological treatment difference.
Subgroup Analysis in Multiregional Trials
For large global trials, a common structure is:
- Overall population
- Major geographic region
- Key countries where sufficiently represented
Country-level analyses can become extremely unstable when individual countries contain few participants.
Biomarker Subgroups
Biomarker-defined subgroups can be particularly important when the mechanism of action suggests predictive enrichment.
For example:
The treatment-by-biomarker interaction can evaluate whether the treatment effect differs according to biomarker status.
A biomarker subgroup can become especially important when the treatment claim is intended to apply only to the biomarker-defined population.
Prognostic Versus Predictive Subgroups
This distinction is extremely useful.
Prognostic Factor
A prognostic factor is associated with outcome regardless of treatment.
For example, patients with advanced disease may have worse outcomes under both treatment and control.
Predictive Factor
A predictive factor modifies the relative treatment effect.
The key statistical feature is the interaction:
A variable can be strongly prognostic without being predictive of treatment benefit.
Subgroup Analysis and Baseline Imbalance
Randomization balances treatment groups in expectation, not necessarily in every small subgroup.
A small subgroup may show apparent baseline imbalances simply because of random variation.
This is another reason to avoid interpreting subgroup estimates without confidence intervals and clinical context.
Subgroup Sample Size Planning
If a subgroup is important enough to support a potential regulatory claim, its representation should be considered during trial planning.
Important questions include:
- How many participants are expected?
- How many events are expected?
- Is the subgroup prevalence known?
- Can enrollment targets be monitored?
- Will the study have useful precision within the subgroup?
- Is an interaction test realistically powered?
Why Interaction Tests Are Often Underpowered
A trial can have adequate power to detect the overall treatment effect while having poor power to detect treatment-effect heterogeneity.
This occurs because the interaction comparison effectively asks whether two treatment effects differ.
For example:
may be estimated much less precisely than the overall treatment effect.
Therefore, failure to detect an interaction does not necessarily mean the trial has established homogeneity.
Do You Need to Power Every Subgroup?
Usually not.
It is generally unrealistic to design a typical confirmatory trial to have adequate power for every demographic and clinical subgroup.
Instead, sponsors should identify subgroups that are especially important for the scientific and regulatory question.
For those subgroups, enrollment and precision should be considered explicitly.
Subgroup Analysis Workflow
Data Structure for Subgroup Analysis
A typical analysis dataset might contain:
| Variable | Description |
|---|---|
| USUBJID | Unique participant identifier |
| TRT01P | Planned treatment |
| AGE | Age at baseline |
| SEX | Sex subgroup |
| RACE | Race category |
| ETHNIC | Ethnicity category |
| REGION | Geographic region |
| BIOMARK | Baseline biomarker subgroup |
| BASE | Baseline value |
| AVAL | Analysis outcome |
| EVENT | Event indicator where applicable |
Creating Subgroup Variables
Suppose age is analyzed using a prespecified 65-year cutoff.
analysis_data$AGEGR1 <-
ifelse(
analysis_data$AGE < 65,
"<65",
"≥65"
)
The important point is not the programming syntax.
The important point is that the cutoff should be defined independently of the observed treatment results whenever it is intended to support a prespecified analysis.
R Example: Continuous Endpoint
fit <- lm( CHANGE ~ TRT * AGEGR1 + BASE, data = analysis_data ) summary(fit)
The treatment-by-age-group interaction can be examined through the model coefficient corresponding to:
TRT:AGEGR1
The subgroup-specific treatment effects should then be estimated directly, rather than inferred solely from the significance of individual treatment coefficients.
R Example: Binary Endpoint
fit <- glm( RESPONSE ~ TRT * AGEGR1, family = binomial(), data = analysis_data ) summary(fit)
Depending on the estimand and analysis specification, adjusted models may include additional prespecified baseline covariates.
R Example: Time-to-Event Endpoint
library(survival)
fit <- coxph(
Surv(TIME, EVENT) ~
TRT * AGEGR1,
data = analysis_data
)
summary(fit)
The model estimates the treatment effect and the treatment-by-subgroup interaction on the log-hazard scale.
Building a Forest Plot in R
A simple forest plot can be constructed after a subgroup-analysis dataset has been prepared.
library(ggplot2)
ggplot(
subgroup_results,
aes(
y = SUBGROUP,
x = ESTIMATE
)
) +
geom_vline(
xintercept = 1,
linetype = "dashed"
) +
geom_errorbarh(
aes(
xmin = LCL,
xmax = UCL
),
height = 0.15
) +
geom_point(
size = 2.5
) +
scale_x_log10() +
labs(
x = "Hazard Ratio",
y = NULL
) +
theme_classic()
For modern versions of ggplot2, the confidence
interval can instead be drawn with horizontal geom_errorbar() after mapping the appropriate
coordinates.
Example Long-Format Subgroup Results
subgroup_results <- data.frame(
SUBGROUP = c(
"Overall",
"Age <65",
"Age ≥65",
"Male",
"Female",
"Biomarker positive",
"Biomarker negative"
),
ESTIMATE = c(
0.74,
0.68,
0.82,
0.71,
0.77,
0.62,
0.91
),
LCL = c(
0.65,
0.56,
0.65,
0.59,
0.63,
0.51,
0.70
),
UCL = c(
0.84,
0.83,
1.04,
0.86,
0.94,
0.76,
1.19
)
)
Adding Interaction P-Values
A regulatory forest plot may include interaction p-values alongside groups of subgroups.
For example:
| Subgroup Variable | Interaction p-value |
|---|---|
| Age | 0.48 |
| Sex | 0.72 |
| Race | 0.31 |
| Biomarker | 0.018 |
The biomarker interaction deserves further investigation because it provides more evidence of differential treatment effect than the other variables.
It does not automatically prove that biomarker status is a predictive biomarker.
What to Do When an Interaction Is Interesting
Suppose:
A reasonable regulatory investigation would include:
- Reviewing the subgroup estimates and confidence intervals
- Checking subgroup sizes and event counts
- Reviewing the biological rationale
- Checking whether the subgroup was prespecified
- Examining related endpoints
- Examining consistency across studies
- Considering multiplicity
- Assessing whether the magnitude is clinically meaningful
- Determining whether the finding supports a regulatory claim
Do Not Stop at the Interaction P-Value
A p-value is evidence about statistical compatibility with the null interaction hypothesis.
It does not tell you:
- Whether the difference is clinically important
- Whether the subgroup is biologically plausible
- Whether the analysis was prespecified
- Whether the result replicates
- Whether the subgroup is sufficiently represented
- Whether the result is robust to alternative definitions
Sensitivity Analyses
Important subgroup findings may benefit from sensitivity analyses.
For example:
- Alternative but clinically defensible subgroup definitions
- Adjusted versus unadjusted models
- Different analysis populations
- Alternative endpoint definitions
- Alternative handling of missing data
- Pooling across relevant trials
The goal is not to search for a favorable result.
The goal is to determine whether the finding is robust.
Subgroup Definitions Should Not Be Reverse-Engineered
Suppose investigators observe that treatment effects appear different at age 72.
They then define:
and report this as though it were a prespecified subgroup.
That undermines credibility.
If the threshold was selected after reviewing treatment results, it should be identified as exploratory and interpreted accordingly.
Subgroup Analysis and Missing Data
Missing subgroup information can create another problem.
For example, suppose race is missing for a meaningful fraction of participants.
The analysis should define how missing subgroup classifications are handled.
Possible approaches include:
- Explicit "Missing" category
- Prespecified imputation
- Exclusion from the subgroup comparison
- Sensitivity analyses
The correct approach depends on the variable and analysis objective.
Post Hoc Subgroup Analysis
Post hoc analyses are common during regulatory review.
They can be useful for:
- Explaining unexpected findings
- Investigating safety signals
- Understanding heterogeneity
- Generating future hypotheses
- Assessing generalizability
But the evidentiary status should remain clear.
A post hoc finding generally requires more caution than a prespecified hypothesis.
Three Regulatory Scenarios
A useful way to think about subgroup evidence is through three broad situations.
Scenario 1: Overall Efficacy Is Convincing
The overall treatment effect is statistically persuasive and clinically relevant.
The subgroup question becomes:
This is generally the most straightforward situation.
Scenario 2: Overall Efficacy Is Borderline
The overall result may be statistically persuasive but clinically borderline, or the overall benefit-risk assessment may be uncertain.
A subgroup result may receive increased attention.
However, identifying a favorable subgroup after observing the data does not automatically establish a credible subgroup-specific indication.
Scenario 3: Overall Efficacy Is Not Persuasive
The temptation to search for a subgroup with a positive result is particularly strong.
This is also where multiplicity and post hoc selection become especially important.
Subgroup Analysis in Labeling
A subgroup finding can potentially affect product labeling when the evidence supports a meaningful difference in efficacy or safety.
However, labeling implications require more than an isolated nominally significant subgroup p-value.
Regulatory assessment can consider:
- Magnitude of effect
- Consistency across trials
- Statistical evidence
- Biological rationale
- Population representation
- Safety
- Benefit-risk balance
- Clinical relevance
Subgroup Analysis in the Integrated Summary of Effectiveness
For a regulatory submission, subgroup analyses can extend beyond one pivotal study.
Integrated analyses can help determine whether apparent subgroup differences are reproduced across controlled studies.
The goal is to distinguish:
from:
Forest Plot Across Multiple Trials
This figure is illustrative. Cross-study comparisons require careful assessment of endpoint definitions, populations, treatment regimens, follow-up, and statistical models.
Regulatory Expectations: FDA
FDA review considers whether clinical-trial data support the safety and effectiveness conclusions for the intended population.
Demographic subgroup information has long been part of FDA review, including assessment of sex, age, race, and ethnicity.
FDA's more recent guidance on enhancing participation in clinical trials also emphasizes enrolling populations that represent patients likely to use the medical product.
The implication for statisticians is important:
Regulatory Expectations: EMA
EMA has a dedicated guideline on the investigation of subgroups in confirmatory clinical trials.
The guideline emphasizes that subgroup investigation is an important component of clinical-trial planning and inference and focuses on the credibility of findings when assessing heterogeneity and applicability of treatment effects.
EMA distinguishes exploratory subgroup investigations from situations in which subgroup hypotheses are part of a confirmatory testing strategy.
FDA and EMA: Similarities
| Principle | Regulatory Expectation |
|---|---|
| Prespecification | Important for credible confirmatory interpretation |
| Demographic assessment | Important for safety, efficacy, and generalizability |
| Interaction | More informative than comparing separate subgroup p-values |
| Precision | Confidence intervals and sample sizes matter |
| Multiplicity | Important when subgroup hypotheses are confirmatory |
| Clinical relevance | Statistical findings require clinical interpretation |
| Consistency | Findings across trials can strengthen interpretation |
FDA and EMA: Important Practical Distinction
The agencies are broadly aligned on the principle that subgroup analyses should not be interpreted mechanically.
However, their guidance documents differ in structure and emphasis.
Sponsors should therefore consult the applicable current agency-specific guidance for the development program rather than assuming that one generic subgroup-analysis template satisfies every regulatory context.
Subgroup Analysis and Clinical Study Reports
A CSR should make the subgroup analysis traceable to the statistical analysis plan.
The report should identify:
- Which subgroups were prespecified
- Which analyses were exploratory
- How subgroup variables were defined
- Which analysis population was used
- Which treatment-effect measure was used
- How missing subgroup information was handled
- How interactions were evaluated
- How multiplicity was addressed when applicable
- How small subgroup sizes were interpreted
Suggested CSR Table Structure
| Subgroup | Treatment N | Control N | Estimate | 95% CI | Interaction p |
|---|---|---|---|---|---|
| Overall | 500 | 498 | 0.74 | 0.65–0.84 | — |
| Age <65 | 270 | 262 | 0.70 | 0.58–0.85 | 0.48 |
| Age ≥65 | 230 | 236 | 0.79 | 0.65–0.97 | 0.48 |
| Male | 290 | 280 | 0.71 | 0.59–0.86 | 0.72 |
| Female | 210 | 218 | 0.78 | 0.64–0.96 | 0.72 |
How to Write a Regulatory Subgroup Conclusion
A strong conclusion should describe the evidence rather than overstate it.
What Not to Write
Avoid: "The treatment did not work in older patients because their p-value was greater than 0.05."
Avoid: "There was no subgroup heterogeneity because the interaction p-value was greater than 0.05."
A Better Interpretation Framework
Common Mistakes in Regulatory Subgroup Analysis
- Comparing subgroup p-values. A significant p-value in one subgroup and a nonsignificant p-value in another does not establish treatment-effect heterogeneity.
- Ignoring interaction tests. The treatment-by-subgroup interaction is usually the direct statistical assessment of differential treatment effect.
- Overinterpreting a nonsignificant interaction. Failure to detect heterogeneity is not proof of identical treatment effects.
- Ignoring confidence intervals. A point estimate without its precision can be misleading.
- Ignoring subgroup size. Small subgroups may produce highly unstable estimates.
- Searching for favorable post hoc cut points. Data-driven thresholds reduce credibility.
- Calling exploratory findings confirmatory. Post hoc findings should be identified clearly.
- Ignoring multiplicity. Many subgroup investigations increase the opportunity for chance findings.
- Using post-randomization variables as ordinary subgroups. This can introduce selection bias and causal interpretation problems.
- Ignoring representation. A subgroup cannot be evaluated reliably if it was barely represented.
- Overinterpreting country-level results. Very small country populations may produce unstable estimates.
- Reporting only favorable subgroups. Regulatory review requires a balanced assessment.
- Assuming prognostic means predictive. A subgroup can have worse outcomes without modifying treatment effect.
- Using a forest plot as the analysis. The forest plot is a display; the statistical model produces the estimates.
Regulatory Subgroup Checklist
Practical Example: Interpreting a Subgroup Finding
Suppose the overall trial demonstrates:
Now suppose age-specific analyses produce:
and:
The estimates differ numerically.
But the difference is modest:
Suppose the interaction p-value is:
A reasonable conclusion would be that the treatment effect appears broadly consistent across age groups, while acknowledging uncertainty within each subgroup.
Another Example: Potential Heterogeneity
Now consider:
with:
This pattern deserves much more attention.
The next questions should be:
- Was the subgroup prespecified?
- How many patients and events were in each subgroup?
- Is the interaction clinically meaningful?
- Is there a biological rationale?
- Does the same pattern appear for secondary endpoints?
- Does it appear in other trials?
- Could the result be due to multiplicity?
- Could differences in baseline risk explain the pattern?
Only after those questions are addressed should the finding influence the overall regulatory interpretation.
Subgroup Analysis Quality Control
Subgroup analyses require rigorous programming validation.
At minimum, verify:
- Every participant has the correct treatment assignment.
- Every subgroup variable is derived according to the SAP.
- Subgroup definitions are mutually exclusive where intended.
- Subgroup categories are exhaustive where intended.
- Missing values are handled consistently.
- Participants are not duplicated.
- Treatment totals reconcile with the primary analysis.
- Subgroup totals reconcile with the overall population.
- Effect estimates reproduce independently calculated results.
- Confidence intervals are correct.
- Interaction p-values match the specified model.
- Forest-plot values match the analysis tables.
Traceability
Every regulatory subgroup estimate should be traceable:
Recommended Statistical Analysis Plan Language
A generic SAP might describe the approach along the following lines:
The exact language should be adapted to the study's endpoint, estimand, analysis model, and regulatory strategy.
How to Present Subgroup Findings to Regulators
A strong presentation follows a logical sequence.
- Present the overall treatment effect.
- Present key subgroup estimates.
- Show confidence intervals.
- Show subgroup sample sizes or event counts.
- Present interaction assessments.
- Discuss clinically important differences.
- Discuss uncertainty in small subgroups.
- Discuss biological plausibility.
- Compare findings across studies where appropriate.
- State clearly whether findings are confirmatory or exploratory.
What a Reviewer Should Be Able to Determine Quickly
After reviewing a subgroup table or forest plot, a reviewer should be able to answer:
- What is the overall treatment effect?
- What is the treatment effect in each important subgroup?
- How precise are those estimates?
- Are subgroup effects directionally consistent?
- Is there evidence of meaningful heterogeneity?
- Are important populations adequately represented?
- Are any concerning safety differences apparent?
- Were the analyses prespecified?
- Are unusual findings biologically plausible?
- Do findings replicate elsewhere?
The Most Important Statistical Distinction
The most important concept in subgroup analysis can be summarized as:
A statistically significant treatment effect in one subgroup and a nonsignificant treatment effect in another does not establish that the treatment effects differ.
The relevant comparison is:
and its uncertainty.
The Most Important Regulatory Distinction
The most important regulatory distinction is:
The objective is not to identify whichever subgroup gives the smallest p-value.
The objective is to determine whether the totality of evidence supports generalization of the treatment effect and whether any clinically important heterogeneity requires additional explanation.
FDA/EMA-Oriented Practical Summary
| Question | Good Practice |
|---|---|
| Was the subgroup prespecified? | Identify this explicitly. |
| Is the subgroup clinically relevant? | Provide scientific rationale. |
| Is the subgroup sufficiently represented? | Report sample size and events. |
| Are treatment effects shown? | Provide estimates and confidence intervals. |
| Was heterogeneity assessed? | Use an appropriate treatment-by-subgroup interaction or related method. |
| Was multiplicity considered? | Address it for confirmatory claims and interpret exploratory findings cautiously. |
| Is the result clinically important? | Consider magnitude, not only p-value. |
| Is the finding biologically plausible? | Evaluate mechanism and prior evidence. |
| Does the finding replicate? | Examine other trials or relevant datasets where appropriate. |
| Could the finding be chance? | Consider multiplicity, sample size, and precision. |
Final Checklist for Statistical Programmers
Bottom Line
References
European Medicines Agency.
Guideline on the investigation of subgroups in confirmatory clinical trials.
EMA/CHMP/539146/2013. Effective 1 August 2019.
EMA scientific guideline .
U.S. Food and Drug Administration.
Enhancing Participation in Clinical Trials — Eligibility Criteria, Enrollment
Practices, and Trial Designs. Guidance for Industry. December 2025.
FDA guidance .
U.S. Food and Drug Administration.
Integrated Summary of Effectiveness Guidance for Industry.
FDA guidance .
International Council for Harmonisation.
ICH E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials.
ICH E9(R1) .
European Medicines Agency.
Multiplicity issues in clinical trials.
EMA scientific guideline .