Tutorials › Biostatistics › Subgroup Analysis: Regulatory Expectations

Regulatory Statistics

Subgroup Analysis: Regulatory Expectations

A practical guide to designing, analyzing, interpreting, and reporting subgroup analyses in clinical trials, including regulatory expectations, prespecification, treatment-by-subgroup interactions, heterogeneity, multiplicity, demographic representation, forest plots, small subgroup sizes, and implementation in R.

Intermediate 20 min read

What You'll Learn

  • What regulators mean by a clinically credible subgroup analysis
  • Why prespecification and biological rationale matter
  • How to evaluate treatment-by-subgroup interactions
  • Why lack of statistical significance does not prove consistency
  • How demographic representation affects interpretation and generalizability
  • How to create regulatory-style subgroup forest plots in R

Introduction

A randomized clinical trial is usually designed to estimate a treatment effect in a prespecified target population. Regulatory review, however, rarely stops with the overall treatment effect.

Regulators also need to understand whether the observed benefit and risk are reasonably applicable across important subsets of the intended population.

These subsets may be defined by demographic characteristics such as age or sex, disease characteristics such as disease severity or stage, prior treatment, geographic region, baseline risk, organ function, or other clinically relevant characteristics.

This is the purpose of subgroup analysis.

Key idea: A regulatory subgroup analysis is not primarily an exercise in finding a subgroup with a statistically significant result. Its central purpose is to evaluate whether the treatment effect is reasonably consistent across relevant patient subsets, whether important heterogeneity exists, and whether the overall trial result can be generalized to the intended population.

This distinction is fundamental.

A subgroup can have a nonsignificant treatment effect simply because it is small. Conversely, a statistically significant subgroup finding can arise by chance when many subgroups are examined.

Regulatory interpretation therefore considers the magnitude, precision, consistency, biological plausibility, prespecification, and clinical context of subgroup findings rather than relying on a single p-value.

What Is a Subgroup?

A subgroup is a subset of trial participants defined by one or more baseline characteristics.

Examples include:

  • Age category
  • Sex
  • Race or ethnicity
  • Geographic region
  • Disease severity
  • Prior therapy
  • Baseline biomarker status
  • Renal function
  • Hepatic function
  • Baseline risk category
  • Performance status
  • Body weight
  • Comorbidity status

The defining characteristic is that the subgroup variable is generally known at or before treatment assignment and can therefore be used to characterize heterogeneity of the randomized treatment effect.

Subgroup Analysis Is Not the Same as Subsetting After Treatment

One of the most important statistical distinctions is between a baseline subgroup and a subgroup defined by something that occurs after randomization.

For example, consider:

  • Age at baseline — appropriate subgroup characteristic
  • Baseline biomarker — appropriate subgroup characteristic
  • Prior treatment — appropriate subgroup characteristic
  • Patients who discontinued treatment — potentially post-randomization subset
  • Patients who developed an adverse event — post-randomization subset
  • Patients who achieved response — outcome-defined subset

The latter categories can create serious interpretational problems because treatment itself may influence membership in the subset.

Regulatory principle: Do not casually interpret an analysis restricted to patients who experienced a post-randomization event as though it were an ordinary baseline subgroup analysis. Treatment assignment can affect whether patients enter that subset, creating selection and causal interpretation problems.

The Regulatory Question

The overall treatment effect answers:

$$ H_0:\theta=0 $$

where \(\theta\) represents the treatment effect in the primary analysis population.

A subgroup analysis asks a different question:

$$ \theta_g = \text{treatment effect within subgroup }g $$

For example:

$$ \theta_{\text{age}<65} \qquad\text{vs.}\qquad \theta_{\text{age}\geq65} $$

The regulatory question is usually not simply whether each subgroup has a statistically significant treatment effect.

Instead, the key question is whether the treatment effects differ materially across subgroups.

The Interaction Is the Central Statistical Concept

Suppose a binary subgroup variable \(G\) defines two groups. A simple treatment-by-subgroup model can be written as:

$$ Y = \beta_0 + \beta_1T + \beta_2G + \beta_3(T\times G) + \epsilon $$

where:

  • \(T\) = treatment indicator
  • \(G\) = subgroup indicator
  • \(T\times G\) = treatment-by-subgroup interaction
  • \(\beta_3\) = parameter describing differential treatment effect between subgroups

The interaction term is therefore fundamentally different from the treatment effect within one subgroup.

Why "Significant in One Subgroup, Not Significant in Another" Is Wrong

Consider the following hypothetical results:

Subgroup Treatment Effect 95% CI p-value
Age <65 −2.8 −4.1 to −1.5 <0.001
Age ≥65 −2.1 −4.5 to 0.3 0.087

It is tempting to conclude:

Incorrect: "The treatment works in younger patients but not in older patients because only the younger subgroup is statistically significant."

That conclusion does not follow.

The older subgroup has a wider confidence interval because it may contain fewer participants.

The point estimates are actually fairly similar:

$$ -2.8 \quad\text{vs.}\quad -2.1 $$

The correct question is whether there is convincing evidence that the treatment effects differ.

Remember: "Significant versus nonsignificant" is not the same as "different versus not different." A formal comparison of subgroup treatment effects requires an appropriate interaction or heterogeneity assessment.

An Actual Regulatory-Style Forest Plot

Forest plots are one of the most common ways to present subgroup treatment effects.

Figure 1. Example Subgroup Forest Plot
Simulated hazard ratios for the primary endpoint. Values below 1 favor treatment. The interaction p-value assesses evidence that treatment effects differ across levels of each subgroup variable.

The values are simulated for educational purposes. A forest plot should be interpreted using treatment-effect estimates, confidence intervals, subgroup sizes, clinical plausibility, and interaction evidence rather than by visually checking whether individual confidence intervals cross the null value.

How to Read a Forest Plot

A subgroup forest plot usually contains:

  • The subgroup definition
  • The number of patients in each subgroup
  • The treatment-effect estimate
  • A confidence interval
  • A vertical line representing the null effect
  • Sometimes an interaction p-value

For a hazard ratio:

$$ HR=1 $$

is the usual null value.

For an odds ratio or risk ratio, the null value is also:

$$ 1 $$

For a mean difference:

$$ MD=0 $$

is the null value.

Confidence Intervals Are Essential

A subgroup point estimate without a confidence interval is difficult to interpret.

Suppose two subgroups have:

Subgroup HR 95% CI
Subgroup A 0.72 0.58–0.89
Subgroup B 0.72 0.39–1.34

The point estimates are identical.

The difference is precision.

The second subgroup may simply have substantially fewer events.

A regulatory reviewer should therefore distinguish:

  • Observed effect magnitude
  • Statistical precision
  • Event count
  • Subgroup sample size
  • Evidence for treatment-effect heterogeneity

What Does a Treatment-by-Subgroup Interaction Test?

Suppose the treatment effects are:

$$ \theta_1 \quad\text{and}\quad \theta_2 $$

The interaction hypothesis is approximately:

$$ H_0:\theta_1-\theta_2=0 $$

against:

$$ H_A:\theta_1-\theta_2\neq0 $$

The interaction test therefore directly addresses whether the treatment effect differs between the subgroups.

Interaction P-Values Need Careful Interpretation

An interaction p-value below a conventional threshold can provide evidence of heterogeneity.

But an interaction p-value above that threshold does not prove that the treatment effects are identical.

For example:

$$ p_{\text{interaction}}=0.42 $$

means that the analysis did not provide strong statistical evidence of a difference in treatment effects.

It does not establish:

$$ \theta_1=\theta_2 $$

with certainty.

Better regulatory language: "Treatment effects appeared generally consistent across the prespecified subgroups, with no compelling evidence of treatment-effect heterogeneity." Avoid saying "the treatment effect was proven identical across all subgroups."

Statistical Significance Is Not Clinical Significance

A statistically significant interaction may be clinically trivial.

Conversely, a clinically important difference may fail to reach statistical significance when subgroup sample sizes are small.

For example:

$$ HR_1=0.70 \qquad HR_2=1.05 $$

could represent a clinically important difference.

But if subgroup 2 contains very few events, the confidence interval may be extremely wide.

Regulatory interpretation therefore requires both statistical and clinical judgment.

Prespecification Matters

A credible subgroup strategy should be developed before the primary results are known.

The statistical analysis plan should identify important subgroup variables and explain how they will be analyzed.

Examples include:

  • Age categories
  • Sex
  • Race and ethnicity
  • Geographic region
  • Baseline disease severity
  • Prior treatment
  • Biomarker status
  • Relevant baseline prognostic factors

Prespecification helps distinguish a scientifically motivated subgroup investigation from a search through the data for a favorable result.

Post hoc does not mean unusable. Exploratory post hoc analyses can be informative. The key requirement is that their exploratory status is made explicit and that the findings are not presented with the same confirmatory credibility as prespecified hypotheses.

Why Regulators Care About Prespecified Subgroups

Suppose a trial examines 20 subgroup variables.

Even if treatment truly has the same effect everywhere, random variation can produce apparently unusual results.

If the sponsor searches enough subgroups, one or more apparently favorable or unfavorable findings can occur by chance.

Prespecification therefore improves interpretability.

How Many Subgroups Should Be Planned?

There is no universal magic number.

The appropriate set depends on:

  • The disease
  • The mechanism of action
  • The intended treatment population
  • Known prognostic factors
  • Known effect modifiers
  • Safety considerations
  • Demographic representation
  • Regulatory expectations

A subgroup should have a scientific reason for inclusion.

Good subgroup planning: Start with the clinical question, not with the variables that happen to be available in the database.

Common Regulatory Subgroup Domains

Domain Examples Typical Rationale
Demographic Age, sex, race, ethnicity Generalizability and potential effect modification
Disease Stage, severity, disease duration Baseline prognosis and disease biology
Prior therapy Prior treatment, treatment line Potential treatment-effect modification
Biomarker Positive/negative status Mechanistic or predictive hypothesis
Geographic Region, country grouping Population and practice differences
Organ function Renal or hepatic function Exposure, safety, and efficacy considerations
Baseline risk Low, intermediate, high risk Prognostic heterogeneity

Demographic Subgroups Have a Special Role

Regulators expect sponsors to consider whether the study population adequately represents the intended population.

Important demographic variables can include:

  • Sex
  • Age
  • Race
  • Ethnicity
  • Geographic region

However, subgroup analysis cannot compensate for inadequate enrollment.

If only a very small number of participants from a demographic group are enrolled, a precise subgroup treatment-effect estimate may not be possible.

Representation comes before analysis. A large, sophisticated subgroup-analysis program cannot recover information that was never collected because the relevant population was poorly represented in the trial.

Subgroup Representation and Modern Trial Planning

Modern regulatory expectations increasingly emphasize inclusion of a clinical trial population that reflects the patients likely to receive the product.

FDA's current diversity guidance discusses both demographic characteristics and non-demographic characteristics of populations, including characteristics such as organ dysfunction, comorbid conditions, disabilities, extremes of body weight, and populations with low disease prevalence.

This means subgroup considerations should begin during trial design, not only after database lock.

Small Subgroups

Small subgroup sizes are one of the most common problems in regulatory subgroup analysis.

Suppose:

Subgroup Treatment N Control N
Majority population 450 450
Small subgroup 28 25

The small subgroup may produce a highly unstable estimate.

For a time-to-event endpoint, the number of events may be even more important than the number of participants.

For example:

$$ n=53 \qquad \text{but only} \qquad E=8 $$

may provide very little information about treatment-effect heterogeneity.

Wide Confidence Intervals Are Informative

A wide confidence interval should not automatically be interpreted as evidence against treatment benefit.

It usually means the estimate is imprecise.

For example:

$$ HR=0.68 \qquad 95\%\,CI=(0.31,\;1.49) $$

The point estimate suggests benefit, but the confidence interval indicates substantial uncertainty.

The appropriate conclusion is not:

"The treatment does not work in this subgroup."

A more appropriate statement is:

"The subgroup estimate was imprecise because of the limited amount of information available; the data do not provide a reliable basis for concluding that the treatment effect differs from that observed in the overall population."

Forest Plot With a Small Subgroup

Figure 2. Precision Depends on Subgroup Information
Simulated treatment effects illustrating how a small subgroup can produce a wide confidence interval even when its point estimate is similar to the overall estimate.

The small subgroup is not evidence of treatment failure. Its confidence interval demonstrates limited information and uncertainty.

Do Confidence Intervals Need to Overlap?

No.

A common visual rule is:

"If the confidence intervals overlap, the subgroup effects are not different."

This is not a reliable statistical rule.

The appropriate comparison is based on the statistical model and treatment-by- subgroup interaction or another appropriate heterogeneity analysis.

What Does "Consistency" Mean?

Consistency does not necessarily mean identical numerical estimates.

Suppose the treatment effects are:

$$ HR_1=0.70 \qquad HR_2=0.78 \qquad HR_3=0.74 $$

These are not identical.

But they may be entirely compatible with a common underlying treatment effect, especially when confidence intervals are considered.

Regulatory assessment is therefore concerned with whether there is credible and clinically meaningful heterogeneity, not whether every estimate is numerically identical.

Heterogeneity Can Be Quantitative or Qualitative

Consider two types of heterogeneity.

Quantitative Heterogeneity

The treatment effect points in the same direction but differs in magnitude.

$$ HR_1=0.65 \qquad HR_2=0.82 $$

Qualitative Heterogeneity

The apparent direction of effect differs.

$$ HR_1=0.65 \qquad HR_2=1.15 $$

Qualitative heterogeneity is generally more concerning because the apparent benefit in one subgroup is accompanied by an apparent lack of benefit or potential harm in another.

Regulatory perspective: A numerical difference between subgroup estimates does not automatically indicate clinically important heterogeneity. Direction, magnitude, precision, clinical plausibility, and consistency across studies should all be considered.

Biological Plausibility

A subgroup finding is more credible when there is a plausible scientific explanation for it.

For example, a treatment designed to inhibit a specific molecular pathway may reasonably be expected to have different effects according to a validated biomarker.

In contrast, a dramatic difference discovered only after examining dozens of unrelated baseline characteristics may be much less convincing.

Internal Consistency Across Studies

Subgroup interpretation becomes stronger when similar patterns appear across independent randomized trials.

For example, suppose three trials produce:

Trial Subgroup A Subgroup B
Trial 1 HR 0.70 HR 0.78
Trial 2 HR 0.74 HR 0.82
Trial 3 HR 0.69 HR 0.76

The consistency of the pattern provides substantially more context than one isolated subgroup analysis.

Pooling Across Studies

Regulatory submissions may include subgroup analyses across multiple controlled trials when appropriate.

Pooling can increase the amount of information available for important subgroups.

However, pooled analyses require careful consideration of:

  • Trial design
  • Patient population
  • Endpoint definition
  • Treatment regimen
  • Follow-up
  • Control group
  • Study heterogeneity
  • Statistical model

A pooled estimate should not be created simply because the individual subgroups are small.

Multiplicity

Subgroup analyses create a multiplicity problem when many hypotheses are examined.

Suppose a trial evaluates:

  • Age
  • Sex
  • Race
  • Ethnicity
  • Region
  • Prior treatment
  • Biomarker status
  • Disease severity
  • Baseline risk
  • Performance status

If each produces multiple subgroup comparisons, many statistical tests may be performed.

The more opportunities there are to observe an apparently unusual result, the greater the chance of observing one by chance.

Exploratory Versus Confirmatory Subgroup Claims

This distinction is essential.

Feature Exploratory Subgroup Confirmatory Subgroup
Purpose Assess consistency or generate hypotheses Test a prespecified treatment-effect claim
Prespecified Preferably Essential
Multiplicity Interpret findings cautiously Must be addressed within the testing strategy
Evidence standard Supportive/descriptive Confirmatory
Post hoc findings Can be useful Generally weak as confirmatory evidence

A subgroup that was not part of the confirmatory testing strategy should not usually be presented as though it were a prespecified confirmatory claim.

When a Subgroup Is Part of the Confirmatory Strategy

Some trials are explicitly designed to establish treatment benefit in a specific subgroup.

For example, a trial may target:

$$ \text{Biomarker-positive patients} $$

In that setting, the subgroup may be the primary analysis population rather than a secondary exploratory subgroup.

The statistical testing strategy should then be aligned with the scientific claim.

Hierarchical Testing and Subgroups

If multiple hypotheses are part of a confirmatory strategy, the sponsor should define how statistical error will be controlled.

Possible approaches include:

  • Hierarchical testing
  • Gatekeeping procedures
  • Alpha allocation
  • Closed testing procedures
  • Other prespecified multiplicity-control strategies

The appropriate method depends on the claims being tested and the structure of the trial.

Do You Need to Adjust Every Exploratory Subgroup P-Value?

Not necessarily.

An exploratory subgroup analysis can be descriptive.

The key is to clearly identify it as exploratory and avoid presenting a large collection of nominal p-values as if each represented an independent confirmatory conclusion.

Common reporting error: Listing dozens of subgroup p-values and highlighting only those below 0.05 can create the appearance of a systematic hypothesis-testing program when the analysis was actually exploratory.

Continuous Subgroup Variables

Age, body weight, renal function, and other variables are naturally continuous.

Artificial categorization can lose information.

For example:

$$ \text{Age}<65 \qquad\text{vs.}\qquad \text{Age}\geq65 $$

may be convenient for presentation, but the biological relationship between age and treatment effect may not have a discontinuity at exactly 65 years.

When scientifically appropriate, continuous interaction models can be considered.

$$ Y = \beta_0+\beta_1T+\beta_2A+ \beta_3(T\times A)+\epsilon $$

where \(A\) is age as a continuous variable.

Why Arbitrary Cut Points Can Be Dangerous

Suppose investigators try:

  • Age 60
  • Age 65
  • Age 70
  • Age 75
  • Age 80

and report whichever cut point produces the strongest apparent difference.

That is a form of data-driven subgroup selection.

The resulting finding may be difficult to interpret because the cut point was chosen after examining the data.

Subgroup Definitions Must Be Reproducible

A regulatory submission should make it possible to determine exactly who belongs to each subgroup.

For example, "older patients" is inadequate.

A reproducible definition might be:

Example: Age subgroup defined using age at informed consent, categorized as <65 years and ≥65 years.

The precise definition should match the protocol, SAP, analysis dataset, and tables.

Stratified Randomization Does Not Automatically Solve Subgroup Analysis

A trial may stratify randomization by region or disease severity.

That improves balance in the randomized allocation.

It does not automatically establish that the treatment effect is consistent within each stratum.

The subgroup analysis still requires appropriate estimation and interpretation.

Covariate Adjustment Versus Subgroup Analysis

A baseline covariate can improve precision of the overall treatment-effect estimate.

A subgroup analysis asks whether treatment effects differ according to that covariate.

These are related but different questions.

Analysis Main Question
Covariate adjustment Can baseline information improve estimation of the overall treatment effect?
Subgroup analysis Does the treatment effect differ across levels of a baseline characteristic?
Interaction analysis Is there statistical evidence that treatment effects differ?

Subgroup Analysis and the Estimand Framework

ICH E9(R1) emphasizes defining the treatment effect of interest through an estimand framework.

This matters for subgroup analysis because the subgroup must be defined in a way that corresponds to the clinical question.

For example, the question might be:

$$ \text{What is the treatment effect in biomarker-positive patients?} $$

That is conceptually different from:

$$ \text{What is the treatment effect among patients who eventually became biomarker-positive?} $$

The second definition could involve information that occurs after treatment and therefore require a different causal framework.

Practical rule: Define the subgroup as part of the clinical question and estimand strategy before deciding how to analyze it.

Subgroup Analysis of Time-to-Event Endpoints

For survival endpoints, the most common treatment-effect measure is a hazard ratio.

A subgroup forest plot might display:

$$ HR_g = \frac{h_T(t\mid g)} {h_C(t\mid g)} $$

where \(g\) denotes the subgroup.

A Cox model with an interaction term can be written as:

coxph(
  Surv(time, status) ~
    TRT * SUBGROUP,
  data = analysis_data
)

The interaction term assesses whether the treatment effect differs across subgroup levels.

Subgroup Analysis of Binary Endpoints

For a binary endpoint, treatment effects may be expressed as:

  • Risk difference
  • Risk ratio
  • Odds ratio

For example:

$$ RD_g = P(Y=1\mid T=1,g) - P(Y=1\mid T=0,g) $$

The treatment-by-subgroup interaction can then evaluate whether the risk difference varies across subgroup levels.

Subgroup Analysis of Continuous Endpoints

For a continuous endpoint, treatment effects may be represented by a mean difference:

$$ MD_g = \bar{Y}_{T,g} - \bar{Y}_{C,g} $$

An ANCOVA-style model can incorporate treatment, subgroup, baseline covariates, and treatment-by-subgroup interaction.

Baseline Adjustment

Baseline adjustment may improve precision for continuous outcomes.

For example:

lm(
  change_from_baseline ~
    TRT * SUBGROUP +
    baseline_value,
  data = analysis_data
)

The interaction term remains the key component for evaluating whether treatment effects differ by subgroup.

Forest Plot Versus Subgroup Table

A forest plot is visually efficient.

A table provides additional detail.

A regulatory submission will often need both.

Feature Forest Plot Detailed Table
Effect estimate Yes Yes
Confidence interval Yes Yes
Sample size Usually Yes
Events Optional Often important
Raw responder counts Usually no Yes
Interaction p-value Often Often
Detailed descriptive statistics Limited Excellent

Regulatory-Style Forest Plot Design

A useful forest plot should allow the reviewer to determine quickly:

  • What the treatment effect is
  • How precise it is
  • How many patients contribute
  • Whether effects are directionally consistent
  • Whether there is evidence of interaction

Avoid excessive decoration.

The objective is rapid scientific interpretation.

Forest Plot Example With Demographic Subgroups

Figure 3. Demographic Subgroup Analysis
Simulated risk ratios for an efficacy endpoint. Values below 1 favor the investigational treatment.

Demographic subgroup analyses should be interpreted together with the number of participants, precision of estimates, overall evidence, and adequacy of representation.

What Regulators Look for in a Subgroup Table

A strong subgroup table usually provides enough information to distinguish effect magnitude from uncertainty.

Depending on endpoint and analysis, useful fields include:

  • Subgroup definition
  • Treatment sample size
  • Control sample size
  • Number of events
  • Effect estimate
  • Confidence interval
  • Interaction p-value
  • Relevant descriptive statistics

Safety Subgroups

Subgroup analysis is not limited to efficacy.

Regulators also need to understand whether safety risks differ across important patient groups.

Examples include:

  • Age
  • Sex
  • Race and ethnicity
  • Renal function
  • Hepatic function
  • Body weight
  • Comorbidities
  • Concomitant medications

Safety subgroup analysis can be especially challenging because adverse events may be uncommon.

A very small subgroup may produce no observed events even though the true risk is not zero.

Zero observed events do not prove zero risk. When subgroup safety counts are small, the absence of observed events should be interpreted in the context of exposure and statistical uncertainty.

Exposure Matters in Safety Subgroups

If treatment exposure differs substantially between subgroups, comparing raw adverse-event counts can be misleading.

Depending on the question, regulators may consider:

  • Number of exposed participants
  • Patient-years of exposure
  • Treatment duration
  • Exposure-adjusted incidence rates
  • Severity and seriousness
  • Temporal relationship

Race and Ethnicity Analyses

Race and ethnicity analyses require particular care.

They may be relevant to:

  • Population representation
  • Generalizability
  • Potential differences in treatment effect
  • Safety
  • Pharmacokinetics
  • Pharmacogenetic considerations

However, race and ethnicity are not interchangeable concepts and should not be treated as purely biological proxies.

The categories used in the analysis should be clearly defined and consistent with the data-collection strategy.

Geographic Subgroups

Regional differences may arise from:

  • Patient characteristics
  • Standard of care
  • Diagnostic practices
  • Endpoint ascertainment
  • Clinical management
  • Enrollment differences
  • Random variation

A geographic interaction should therefore not automatically be interpreted as a biological treatment difference.

Subgroup Analysis in Multiregional Trials

For large global trials, a common structure is:

  • Overall population
  • Major geographic region
  • Key countries where sufficiently represented

Country-level analyses can become extremely unstable when individual countries contain few participants.

Practical principle: Do not create dozens of country-level efficacy claims simply because the database contains participants from dozens of countries.

Biomarker Subgroups

Biomarker-defined subgroups can be particularly important when the mechanism of action suggests predictive enrichment.

For example:

$$ \text{Biomarker Positive} \qquad \text{vs.} \qquad \text{Biomarker Negative} $$

The treatment-by-biomarker interaction can evaluate whether the treatment effect differs according to biomarker status.

A biomarker subgroup can become especially important when the treatment claim is intended to apply only to the biomarker-defined population.

Prognostic Versus Predictive Subgroups

This distinction is extremely useful.

Prognostic Factor

A prognostic factor is associated with outcome regardless of treatment.

For example, patients with advanced disease may have worse outcomes under both treatment and control.

Predictive Factor

A predictive factor modifies the relative treatment effect.

The key statistical feature is the interaction:

$$ T\times G $$

A variable can be strongly prognostic without being predictive of treatment benefit.

Important: A subgroup with poorer prognosis is not automatically a subgroup in which the treatment works differently.

Subgroup Analysis and Baseline Imbalance

Randomization balances treatment groups in expectation, not necessarily in every small subgroup.

A small subgroup may show apparent baseline imbalances simply because of random variation.

This is another reason to avoid interpreting subgroup estimates without confidence intervals and clinical context.

Subgroup Sample Size Planning

If a subgroup is important enough to support a potential regulatory claim, its representation should be considered during trial planning.

Important questions include:

  • How many participants are expected?
  • How many events are expected?
  • Is the subgroup prevalence known?
  • Can enrollment targets be monitored?
  • Will the study have useful precision within the subgroup?
  • Is an interaction test realistically powered?

Why Interaction Tests Are Often Underpowered

A trial can have adequate power to detect the overall treatment effect while having poor power to detect treatment-effect heterogeneity.

This occurs because the interaction comparison effectively asks whether two treatment effects differ.

For example:

$$ \theta_1-\theta_2 $$

may be estimated much less precisely than the overall treatment effect.

Therefore, failure to detect an interaction does not necessarily mean the trial has established homogeneity.

Do You Need to Power Every Subgroup?

Usually not.

It is generally unrealistic to design a typical confirmatory trial to have adequate power for every demographic and clinical subgroup.

Instead, sponsors should identify subgroups that are especially important for the scientific and regulatory question.

For those subgroups, enrollment and precision should be considered explicitly.

Subgroup Analysis Workflow

1
Define the clinical question and intended population.
2
Identify clinically meaningful baseline subgroup variables.
3
Prespecify key subgroup definitions in the protocol and SAP.
4
Assess expected representation and information within important subgroups.
5
Define the treatment-effect measure and statistical model.
6
Estimate treatment effects within subgroups.
7
Estimate treatment-by-subgroup interaction or another appropriate heterogeneity measure.
8
Review confidence intervals, sample sizes, and event counts.
9
Assess clinical plausibility and consistency across related analyses.
10
Clearly distinguish confirmatory findings from exploratory findings.

Data Structure for Subgroup Analysis

A typical analysis dataset might contain:

Variable Description
USUBJID Unique participant identifier
TRT01P Planned treatment
AGE Age at baseline
SEX Sex subgroup
RACE Race category
ETHNIC Ethnicity category
REGION Geographic region
BIOMARK Baseline biomarker subgroup
BASE Baseline value
AVAL Analysis outcome
EVENT Event indicator where applicable

Creating Subgroup Variables

Suppose age is analyzed using a prespecified 65-year cutoff.

analysis_data$AGEGR1 <-
  ifelse(
    analysis_data$AGE < 65,
    "<65",
    "≥65"
  )

The important point is not the programming syntax.

The important point is that the cutoff should be defined independently of the observed treatment results whenever it is intended to support a prespecified analysis.

R Example: Continuous Endpoint

fit <- lm(
  CHANGE ~ TRT * AGEGR1 + BASE,
  data = analysis_data
)

summary(fit)

The treatment-by-age-group interaction can be examined through the model coefficient corresponding to:

TRT:AGEGR1

The subgroup-specific treatment effects should then be estimated directly, rather than inferred solely from the significance of individual treatment coefficients.

R Example: Binary Endpoint

fit <- glm(
  RESPONSE ~ TRT * AGEGR1,
  family = binomial(),
  data = analysis_data
)

summary(fit)

Depending on the estimand and analysis specification, adjusted models may include additional prespecified baseline covariates.

R Example: Time-to-Event Endpoint

library(survival)

fit <- coxph(
  Surv(TIME, EVENT) ~
    TRT * AGEGR1,
  data = analysis_data
)

summary(fit)

The model estimates the treatment effect and the treatment-by-subgroup interaction on the log-hazard scale.

Building a Forest Plot in R

A simple forest plot can be constructed after a subgroup-analysis dataset has been prepared.

library(ggplot2)

ggplot(
  subgroup_results,
  aes(
    y = SUBGROUP,
    x = ESTIMATE
  )
) +
  geom_vline(
    xintercept = 1,
    linetype = "dashed"
  ) +
  geom_errorbarh(
    aes(
      xmin = LCL,
      xmax = UCL
    ),
    height = 0.15
  ) +
  geom_point(
    size = 2.5
  ) +
  scale_x_log10() +
  labs(
    x = "Hazard Ratio",
    y = NULL
  ) +
  theme_classic()

For modern versions of ggplot2, the confidence interval can instead be drawn with horizontal geom_errorbar() after mapping the appropriate coordinates.

Example Long-Format Subgroup Results

subgroup_results <- data.frame(
  SUBGROUP = c(
    "Overall",
    "Age <65",
    "Age ≥65",
    "Male",
    "Female",
    "Biomarker positive",
    "Biomarker negative"
  ),
  ESTIMATE = c(
    0.74,
    0.68,
    0.82,
    0.71,
    0.77,
    0.62,
    0.91
  ),
  LCL = c(
    0.65,
    0.56,
    0.65,
    0.59,
    0.63,
    0.51,
    0.70
  ),
  UCL = c(
    0.84,
    0.83,
    1.04,
    0.86,
    0.94,
    0.76,
    1.19
  )
)

Adding Interaction P-Values

A regulatory forest plot may include interaction p-values alongside groups of subgroups.

For example:

Subgroup Variable Interaction p-value
Age 0.48
Sex 0.72
Race 0.31
Biomarker 0.018

The biomarker interaction deserves further investigation because it provides more evidence of differential treatment effect than the other variables.

It does not automatically prove that biomarker status is a predictive biomarker.

What to Do When an Interaction Is Interesting

Suppose:

$$ p_{\text{interaction}}=0.018 $$

A reasonable regulatory investigation would include:

  • Reviewing the subgroup estimates and confidence intervals
  • Checking subgroup sizes and event counts
  • Reviewing the biological rationale
  • Checking whether the subgroup was prespecified
  • Examining related endpoints
  • Examining consistency across studies
  • Considering multiplicity
  • Assessing whether the magnitude is clinically meaningful
  • Determining whether the finding supports a regulatory claim

Do Not Stop at the Interaction P-Value

A p-value is evidence about statistical compatibility with the null interaction hypothesis.

It does not tell you:

  • Whether the difference is clinically important
  • Whether the subgroup is biologically plausible
  • Whether the analysis was prespecified
  • Whether the result replicates
  • Whether the subgroup is sufficiently represented
  • Whether the result is robust to alternative definitions

Sensitivity Analyses

Important subgroup findings may benefit from sensitivity analyses.

For example:

  • Alternative but clinically defensible subgroup definitions
  • Adjusted versus unadjusted models
  • Different analysis populations
  • Alternative endpoint definitions
  • Alternative handling of missing data
  • Pooling across relevant trials

The goal is not to search for a favorable result.

The goal is to determine whether the finding is robust.

Subgroup Definitions Should Not Be Reverse-Engineered

Suppose investigators observe that treatment effects appear different at age 72.

They then define:

$$ \text{Age}<72 \qquad\text{vs.}\qquad \text{Age}\geq72 $$

and report this as though it were a prespecified subgroup.

That undermines credibility.

If the threshold was selected after reviewing treatment results, it should be identified as exploratory and interpreted accordingly.

Subgroup Analysis and Missing Data

Missing subgroup information can create another problem.

For example, suppose race is missing for a meaningful fraction of participants.

The analysis should define how missing subgroup classifications are handled.

Possible approaches include:

  • Explicit "Missing" category
  • Prespecified imputation
  • Exclusion from the subgroup comparison
  • Sensitivity analyses

The correct approach depends on the variable and analysis objective.

Do not silently exclude participants with missing subgroup data. The resulting population may differ from the intended analysis population and can affect interpretation of representativeness.

Post Hoc Subgroup Analysis

Post hoc analyses are common during regulatory review.

They can be useful for:

  • Explaining unexpected findings
  • Investigating safety signals
  • Understanding heterogeneity
  • Generating future hypotheses
  • Assessing generalizability

But the evidentiary status should remain clear.

A post hoc finding generally requires more caution than a prespecified hypothesis.

Three Regulatory Scenarios

A useful way to think about subgroup evidence is through three broad situations.

Scenario 1: Overall Efficacy Is Convincing

The overall treatment effect is statistically persuasive and clinically relevant.

The subgroup question becomes:

$$ \text{Does the conclusion appear applicable across important subgroups?} $$

This is generally the most straightforward situation.

Scenario 2: Overall Efficacy Is Borderline

The overall result may be statistically persuasive but clinically borderline, or the overall benefit-risk assessment may be uncertain.

A subgroup result may receive increased attention.

However, identifying a favorable subgroup after observing the data does not automatically establish a credible subgroup-specific indication.

Scenario 3: Overall Efficacy Is Not Persuasive

The temptation to search for a subgroup with a positive result is particularly strong.

This is also where multiplicity and post hoc selection become especially important.

Important: A positive subgroup discovered after an unsuccessful overall trial should not automatically rescue the trial. Credibility requires a strong prespecified scientific rationale, adequate information, appropriate statistical evidence, and convincing clinical context.

Subgroup Analysis in Labeling

A subgroup finding can potentially affect product labeling when the evidence supports a meaningful difference in efficacy or safety.

However, labeling implications require more than an isolated nominally significant subgroup p-value.

Regulatory assessment can consider:

  • Magnitude of effect
  • Consistency across trials
  • Statistical evidence
  • Biological rationale
  • Population representation
  • Safety
  • Benefit-risk balance
  • Clinical relevance

Subgroup Analysis in the Integrated Summary of Effectiveness

For a regulatory submission, subgroup analyses can extend beyond one pivotal study.

Integrated analyses can help determine whether apparent subgroup differences are reproduced across controlled studies.

The goal is to distinguish:

$$ \text{real treatment-effect heterogeneity} $$

from:

$$ \text{random variation within individual studies}. $$

Forest Plot Across Multiple Trials

Figure 4. Subgroup Consistency Across Three Simulated Trials
Illustrative hazard ratios for the same subgroup across independent randomized trials. Repetition of a similar direction and magnitude strengthens the descriptive evidence of consistency.

This figure is illustrative. Cross-study comparisons require careful assessment of endpoint definitions, populations, treatment regimens, follow-up, and statistical models.

Regulatory Expectations: FDA

FDA review considers whether clinical-trial data support the safety and effectiveness conclusions for the intended population.

Demographic subgroup information has long been part of FDA review, including assessment of sex, age, race, and ethnicity.

FDA's more recent guidance on enhancing participation in clinical trials also emphasizes enrolling populations that represent patients likely to use the medical product.

The implication for statisticians is important:

FDA-oriented principle: Subgroup analysis begins with appropriate representation and data collection. The analysis should then characterize whether treatment effects and safety findings are reasonably applicable across important patient characteristics.

Regulatory Expectations: EMA

EMA has a dedicated guideline on the investigation of subgroups in confirmatory clinical trials.

The guideline emphasizes that subgroup investigation is an important component of clinical-trial planning and inference and focuses on the credibility of findings when assessing heterogeneity and applicability of treatment effects.

EMA distinguishes exploratory subgroup investigations from situations in which subgroup hypotheses are part of a confirmatory testing strategy.

EMA-oriented principle: Subgroup evidence should be evaluated in the context of the overall trial, with attention to heterogeneity, internal consistency, statistical uncertainty, clinical relevance, and the credibility of the subgroup finding.

FDA and EMA: Similarities

Principle Regulatory Expectation
Prespecification Important for credible confirmatory interpretation
Demographic assessment Important for safety, efficacy, and generalizability
Interaction More informative than comparing separate subgroup p-values
Precision Confidence intervals and sample sizes matter
Multiplicity Important when subgroup hypotheses are confirmatory
Clinical relevance Statistical findings require clinical interpretation
Consistency Findings across trials can strengthen interpretation

FDA and EMA: Important Practical Distinction

The agencies are broadly aligned on the principle that subgroup analyses should not be interpreted mechanically.

However, their guidance documents differ in structure and emphasis.

Sponsors should therefore consult the applicable current agency-specific guidance for the development program rather than assuming that one generic subgroup-analysis template satisfies every regulatory context.

Subgroup Analysis and Clinical Study Reports

A CSR should make the subgroup analysis traceable to the statistical analysis plan.

The report should identify:

  • Which subgroups were prespecified
  • Which analyses were exploratory
  • How subgroup variables were defined
  • Which analysis population was used
  • Which treatment-effect measure was used
  • How missing subgroup information was handled
  • How interactions were evaluated
  • How multiplicity was addressed when applicable
  • How small subgroup sizes were interpreted

Suggested CSR Table Structure

Subgroup Treatment N Control N Estimate 95% CI Interaction p
Overall 500 498 0.74 0.65–0.84 —
Age <65 270 262 0.70 0.58–0.85 0.48
Age ≥65 230 236 0.79 0.65–0.97 0.48
Male 290 280 0.71 0.59–0.86 0.72
Female 210 218 0.78 0.64–0.96 0.72

How to Write a Regulatory Subgroup Conclusion

A strong conclusion should describe the evidence rather than overstate it.

Example: "Treatment effects were generally consistent across the prespecified demographic and disease-characteristic subgroups. Although point estimates varied across subgroups, confidence intervals were generally overlapping and there was no compelling evidence of treatment-by-subgroup interaction. Some subgroups contained relatively few participants or events, resulting in imprecise estimates; therefore, the absence of evidence for heterogeneity in these subgroups should be interpreted cautiously."

What Not to Write

Avoid: "The treatment was effective in all subgroups because all subgroup p-values were significant."

Avoid: "The treatment did not work in older patients because their p-value was greater than 0.05."

Avoid: "There was no subgroup heterogeneity because the interaction p-value was greater than 0.05."

A Better Interpretation Framework

1
Look at the overall treatment effect. Start with the primary analysis.
2
Look at subgroup point estimates. Determine whether effects are broadly similar or materially different.
3
Look at confidence intervals. Assess precision and uncertainty.
4
Look at the interaction. Directly assess evidence for differential treatment effect.
5
Look at subgroup size and events. Determine whether the analysis contains enough information to be informative.
6
Look at biological plausibility. Ask whether the finding has a credible clinical explanation.
7
Look across related endpoints and studies. Determine whether the pattern replicates.
8
Classify the evidence appropriately. Distinguish confirmatory evidence from exploratory observations.

Common Mistakes in Regulatory Subgroup Analysis

  1. Comparing subgroup p-values. A significant p-value in one subgroup and a nonsignificant p-value in another does not establish treatment-effect heterogeneity.
  2. Ignoring interaction tests. The treatment-by-subgroup interaction is usually the direct statistical assessment of differential treatment effect.
  3. Overinterpreting a nonsignificant interaction. Failure to detect heterogeneity is not proof of identical treatment effects.
  4. Ignoring confidence intervals. A point estimate without its precision can be misleading.
  5. Ignoring subgroup size. Small subgroups may produce highly unstable estimates.
  6. Searching for favorable post hoc cut points. Data-driven thresholds reduce credibility.
  7. Calling exploratory findings confirmatory. Post hoc findings should be identified clearly.
  8. Ignoring multiplicity. Many subgroup investigations increase the opportunity for chance findings.
  9. Using post-randomization variables as ordinary subgroups. This can introduce selection bias and causal interpretation problems.
  10. Ignoring representation. A subgroup cannot be evaluated reliably if it was barely represented.
  11. Overinterpreting country-level results. Very small country populations may produce unstable estimates.
  12. Reporting only favorable subgroups. Regulatory review requires a balanced assessment.
  13. Assuming prognostic means predictive. A subgroup can have worse outcomes without modifying treatment effect.
  14. Using a forest plot as the analysis. The forest plot is a display; the statistical model produces the estimates.

Regulatory Subgroup Checklist

1
Is the subgroup clinically meaningful?
2
Was it prespecified?
3
Is the subgroup defined using baseline information?
4
Is the subgroup sufficiently represented?
5
Are treatment effects and confidence intervals reported?
6
Has treatment-by-subgroup interaction been evaluated where appropriate?
7
Have small sample sizes and event counts been considered?
8
Has multiplicity been considered for confirmatory claims?
9
Is there a plausible clinical or biological explanation?
10
Is the finding consistent across related analyses or studies?

Practical Example: Interpreting a Subgroup Finding

Suppose the overall trial demonstrates:

$$ HR=0.74 \qquad 95\%\,CI=(0.65,\;0.84) $$

Now suppose age-specific analyses produce:

$$ HR_{<65}=0.70 $$

and:

$$ HR_{\geq65}=0.79 $$

The estimates differ numerically.

But the difference is modest:

$$ 0.79-0.70=0.09 $$

Suppose the interaction p-value is:

$$ p_{\mathrm{interaction}}=0.48 $$

A reasonable conclusion would be that the treatment effect appears broadly consistent across age groups, while acknowledging uncertainty within each subgroup.

Another Example: Potential Heterogeneity

Now consider:

$$ HR_1=0.55 \qquad HR_2=1.08 $$

with:

$$ p_{\mathrm{interaction}}=0.012 $$

This pattern deserves much more attention.

The next questions should be:

  • Was the subgroup prespecified?
  • How many patients and events were in each subgroup?
  • Is the interaction clinically meaningful?
  • Is there a biological rationale?
  • Does the same pattern appear for secondary endpoints?
  • Does it appear in other trials?
  • Could the result be due to multiplicity?
  • Could differences in baseline risk explain the pattern?

Only after those questions are addressed should the finding influence the overall regulatory interpretation.

Subgroup Analysis Quality Control

Subgroup analyses require rigorous programming validation.

At minimum, verify:

  • Every participant has the correct treatment assignment.
  • Every subgroup variable is derived according to the SAP.
  • Subgroup definitions are mutually exclusive where intended.
  • Subgroup categories are exhaustive where intended.
  • Missing values are handled consistently.
  • Participants are not duplicated.
  • Treatment totals reconcile with the primary analysis.
  • Subgroup totals reconcile with the overall population.
  • Effect estimates reproduce independently calculated results.
  • Confidence intervals are correct.
  • Interaction p-values match the specified model.
  • Forest-plot values match the analysis tables.

Traceability

Every regulatory subgroup estimate should be traceable:

A
Source data
B
Analysis dataset
C
Subgroup derivation
D
Statistical model
E
Treatment-effect estimate
F
Confidence interval and interaction assessment
G
CSR table and forest plot

Recommended Statistical Analysis Plan Language

A generic SAP might describe the approach along the following lines:

Illustrative wording: "Treatment effects will be summarized within prespecified subgroups defined by baseline demographic and disease characteristics. For each subgroup, the treatment effect and corresponding two-sided 95% confidence interval will be presented. Where appropriate, treatment-by-subgroup interaction terms will be used to assess evidence of heterogeneity of treatment effect. Subgroup analyses are primarily intended to evaluate consistency and applicability of the treatment effect and will be interpreted in the context of subgroup sample size, event counts, precision, multiplicity, biological plausibility, and the overall trial results. Unless explicitly included in the confirmatory testing strategy, subgroup analyses will be considered exploratory."

The exact language should be adapted to the study's endpoint, estimand, analysis model, and regulatory strategy.

How to Present Subgroup Findings to Regulators

A strong presentation follows a logical sequence.

  1. Present the overall treatment effect.
  2. Present key subgroup estimates.
  3. Show confidence intervals.
  4. Show subgroup sample sizes or event counts.
  5. Present interaction assessments.
  6. Discuss clinically important differences.
  7. Discuss uncertainty in small subgroups.
  8. Discuss biological plausibility.
  9. Compare findings across studies where appropriate.
  10. State clearly whether findings are confirmatory or exploratory.

What a Reviewer Should Be Able to Determine Quickly

After reviewing a subgroup table or forest plot, a reviewer should be able to answer:

  • What is the overall treatment effect?
  • What is the treatment effect in each important subgroup?
  • How precise are those estimates?
  • Are subgroup effects directionally consistent?
  • Is there evidence of meaningful heterogeneity?
  • Are important populations adequately represented?
  • Are any concerning safety differences apparent?
  • Were the analyses prespecified?
  • Are unusual findings biologically plausible?
  • Do findings replicate elsewhere?

The Most Important Statistical Distinction

The most important concept in subgroup analysis can be summarized as:

$$ \boxed{ \text{Difference between treatment effects} \neq \text{difference between subgroup p-values} } $$

A statistically significant treatment effect in one subgroup and a nonsignificant treatment effect in another does not establish that the treatment effects differ.

The relevant comparison is:

$$ \boxed{ \theta_1-\theta_2 } $$

and its uncertainty.

The Most Important Regulatory Distinction

The most important regulatory distinction is:

$$ \boxed{ \text{Consistency assessment} \neq \text{search for a significant subgroup} } $$

The objective is not to identify whichever subgroup gives the smallest p-value.

The objective is to determine whether the totality of evidence supports generalization of the treatment effect and whether any clinically important heterogeneity requires additional explanation.

FDA/EMA-Oriented Practical Summary

Question Good Practice
Was the subgroup prespecified? Identify this explicitly.
Is the subgroup clinically relevant? Provide scientific rationale.
Is the subgroup sufficiently represented? Report sample size and events.
Are treatment effects shown? Provide estimates and confidence intervals.
Was heterogeneity assessed? Use an appropriate treatment-by-subgroup interaction or related method.
Was multiplicity considered? Address it for confirmatory claims and interpret exploratory findings cautiously.
Is the result clinically important? Consider magnitude, not only p-value.
Is the finding biologically plausible? Evaluate mechanism and prior evidence.
Does the finding replicate? Examine other trials or relevant datasets where appropriate.
Could the finding be chance? Consider multiplicity, sample size, and precision.

Final Checklist for Statistical Programmers

1
Read the protocol and SAP subgroup sections.
2
Confirm all prespecified subgroup variables.
3
Verify subgroup derivation against source data.
4
Check subgroup population counts.
5
Check event counts for time-to-event endpoints.
6
Fit the prespecified treatment-effect model.
7
Estimate subgroup-specific effects and confidence intervals.
8
Estimate treatment-by-subgroup interaction where specified.
9
Reconcile estimates with independent calculations.
10
Validate tables, listings, and forest plots against one another.

Bottom Line

Regulatory subgroup analysis is fundamentally an exercise in credibility. A credible analysis begins with a clinically meaningful subgroup definition and appropriate representation of the intended population. It uses prespecified methods where possible, reports treatment-effect estimates with confidence intervals, evaluates treatment-by-subgroup interaction rather than comparing isolated p-values, considers sample size and event counts, accounts for multiplicity when subgroup findings are confirmatory, and interprets heterogeneity using clinical and biological context. Exploratory findings can be valuable, but their evidentiary status should remain explicit. A forest plot is an effective way to display the evidence, but the statistical and regulatory interpretation must come from the complete body of evidence rather than from the visual appearance of individual confidence intervals.

References

European Medicines Agency. Guideline on the investigation of subgroups in confirmatory clinical trials. EMA/CHMP/539146/2013. Effective 1 August 2019. EMA scientific guideline .

U.S. Food and Drug Administration. Enhancing Participation in Clinical Trials — Eligibility Criteria, Enrollment Practices, and Trial Designs. Guidance for Industry. December 2025. FDA guidance .

U.S. Food and Drug Administration. Integrated Summary of Effectiveness Guidance for Industry. FDA guidance .

International Council for Harmonisation. ICH E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials. ICH E9(R1) .

European Medicines Agency. Multiplicity issues in clinical trials. EMA scientific guideline .