Introduction
A conventional parallel-group clinical trial assigns each participant to one treatment and follows that participant throughout the study.
A crossover trial takes a different approach. The same participant receives more than one treatment during different periods of the study.
For example, a participant might receive the investigational treatment during Period 1 and placebo during Period 2, while another participant receives the opposite sequence.
This can make crossover designs highly efficient when the disease and treatment effect are sufficiently stable over time and when treatment effects do not persist into subsequent periods.
But those assumptions are critical.
A crossover trial is not simply a parallel trial with two measurements per patient. The analysis must account for the repeated measurements, treatment sequence, period, and potential residual effects from treatment administered in an earlier period.
When Is a Crossover Design Appropriate?
Crossover designs are most useful when the clinical setting has characteristics that make repeated treatment exposure scientifically reasonable.
Typical favorable characteristics include:
- The condition is relatively stable over the study period.
- The treatment effect develops and disappears within a reasonable period.
- A suitable washout period can separate treatments.
- The endpoint can be measured repeatedly.
- Receiving the treatments in either order is clinically acceptable.
- The treatment does not permanently alter the disease course.
Common applications include pharmacokinetic studies, pharmacodynamic studies, symptom-control studies, bioequivalence studies, and some chronic disease trials.
Parallel vs. Crossover Design
Consider a simple comparison between Treatment A and Treatment B.
In a parallel-group study, one participant may receive A while another receives B.
In a crossover study, the same participant can receive both:
The second sequence is essential because it prevents treatment from being completely confounded with calendar period.
| Feature | Parallel Group | 2×2 Crossover |
|---|---|---|
| Participant receives | One treatment | Both treatments |
| Primary comparison | Between participants | Within participants |
| Between-subject variability | Directly contributes to error | Partly removed by within-subject comparison |
| Period effects | Usually less central | Must be modeled |
| Carryover concern | No crossover carryover | Potentially important |
| Washout required | Usually no | Often yes |
The Basic 2×2 Crossover Design
The simplest and most widely recognized crossover design is the 2×2 crossover.
There are two treatments and two periods.
Let:
- A = investigational treatment
- B = control or comparator
- Period 1 = first treatment period
- Period 2 = second treatment period
Participants are randomized to one of two treatment sequences:
| Sequence | Period 1 | Washout | Period 2 |
|---|---|---|---|
| AB | A | Yes | B |
| BA | B | Yes | A |
Randomization to sequence is important because otherwise treatment could be confounded with period.
Why Randomize the Sequence?
Suppose every participant received A first and B second.
If the outcome improves during the second period, it would be impossible to know whether the improvement was caused by B or simply by a period effect.
For example, participants might become familiar with a study procedure, symptoms might improve naturally over time, or an outcome might drift because of disease progression.
Randomizing participants to AB and BA distributes those temporal effects across the two treatment sequences.
The Four Fundamental Effects
Four concepts appear repeatedly in crossover trial methodology:
- Treatment effect
- Period effect
- Sequence effect
- Carryover effect
They are related, but they are not interchangeable.
Treatment Effect
The treatment effect is the effect of A relative to B.
This is usually the primary parameter of interest.
Period Effect
A period effect occurs when outcomes differ systematically between Period 1 and Period 2 regardless of treatment.
Period effects can arise from time trends, learning, natural history, seasonality, treatment discontinuation, or other temporal mechanisms.
Sequence Effect
A sequence effect occurs when participants randomized to AB differ systematically from participants randomized to BA.
Because sequence is randomized, a sequence difference can be investigated as part of the design.
However, sequence effects can be difficult to interpret because treatment sequence is also associated with treatment exposure order.
Carryover Effect
A carryover effect occurs when treatment administered in one period continues to influence the outcome in a later period.
For example, suppose A produces a biological effect lasting four weeks but the washout period lasts only one week.
A participant receiving A in Period 1 may still be partially exposed to the effect of A during Period 2.
The Data Structure
A crossover dataset usually contains multiple observations for each participant.
For a 2×2 design, a simplified dataset might look like:
| Subject | Sequence | Period | Treatment | Outcome |
|---|---|---|---|---|
| 001 | AB | 1 | A | 82 |
| 001 | AB | 2 | B | 74 |
| 002 | BA | 1 | B | 69 |
| 002 | BA | 2 | A | 78 |
| 003 | AB | 1 | A | 91 |
| 003 | AB | 2 | B | 84 |
The same participant appears more than once.
This means the observations cannot generally be treated as independent.
That within-subject correlation is one of the defining statistical features of the crossover design.
The Within-Subject Difference
For an AB participant, define:
where \(Y_{iA}\) and \(Y_{iB}\) are the participant's outcomes under A and B.
For a BA participant, the same treatment-specific difference can be calculated:
The important point is that the difference is always defined using treatment labels, not simply Period 1 minus Period 2.
Why Within-Subject Comparisons Can Be Efficient
Suppose the outcome can be represented as:
where \(b_i\) represents a participant-specific effect.
If the same participant receives both treatments, the participant effect largely cancels when comparing treatment outcomes within that participant:
The stable participant-specific component \(b_i\) disappears.
This is the fundamental statistical efficiency advantage of crossover designs.
The Role of Within-Subject Correlation
Let the variance of an outcome under each treatment be \(\sigma^2\), and let the within-subject correlation between the two treatment measurements be \(\rho\).
Then the variance of the treatment difference is:
As \(\rho\) increases, the variance of the within-subject difference decreases.
For example, if:
then:
rather than \(2\sigma^2\) when the observations are independent.
Basic Statistical Model
A conventional model for a continuous endpoint in a 2×2 crossover study can be written as:
where:
- \(\mu\) = overall mean
- \(S_i\) = subject effect
- \(P_j\) = period effect
- \(T_k\) = treatment effect
- \(\epsilon_{ijk}\) = residual error
The subject effect is generally modeled as a random effect in a mixed-effects framework.
and:
Mixed-Effects Model Formulation
A practical linear mixed model is:
with:
and:
The treatment coefficient is the primary treatment contrast.
What About Sequence?
Sequence is typically included as a fixed effect when it is part of the prespecified crossover analysis.
A model may therefore be written as:
where \(S_i\) represents treatment sequence.
This allows the model to distinguish:
- Treatment differences
- Period differences
- Sequence differences
- Subject-level variability
A Complete Worked Example
Consider a crossover study evaluating an investigational antihypertensive treatment against placebo.
The primary endpoint is change in systolic blood pressure after each treatment period.
Lower values represent better blood-pressure control.
Each participant receives both treatments.
| Design Component | Value |
|---|---|
| Treatments | Investigational treatment A and placebo B |
| Periods | 2 |
| Sequences | AB and BA |
| Washout | 2 weeks |
| Primary endpoint | Change in systolic blood pressure |
| Analysis | Mixed-effects model |
| Primary contrast | A minus B |
Suppose 40 participants are randomized:
- 20 to sequence AB
- 20 to sequence BA
The resulting treatment-period data are summarized below.
Worked Example: Individual Treatment Differences
For illustration, consider eight participants from the full dataset.
| Subject | Sequence | A | B | A − B |
|---|---|---|---|---|
| 001 | AB | −8 | −5 | −3 |
| 002 | AB | −10 | −6 | −4 |
| 003 | AB | −7 | −4 | −3 |
| 004 | AB | −11 | −7 | −4 |
| 005 | BA | −9 | −5 | −4 |
| 006 | BA | −8 | −6 | −2 |
| 007 | BA | −12 | −7 | −5 |
| 008 | BA | −9 | −6 | −3 |
Every participant has a negative treatment difference.
For example, Subject 001 has:
The negative difference indicates that Treatment A produced a greater reduction in systolic blood pressure than B for this participant.
Estimating the Treatment Effect
The treatment effect can be estimated from the mean of the within-participant differences in a simple balanced 2×2 setting.
For the eight illustrative participants:
Therefore:
The estimated treatment effect is therefore:
The interpretation is that Treatment A is associated with an average systolic blood-pressure reduction 3.5 mmHg greater than Treatment B in this illustrative dataset.
Testing the Treatment Effect
Suppose the primary hypothesis is:
For a simple within-subject analysis, the test statistic is:
where \(s_D\) is the sample standard deviation of the within-subject differences.
This is mathematically equivalent to a paired \(t\)-test when the design and analysis are appropriately reduced to one difference per participant.
Why the Paired t-Test Can Be Useful
The paired \(t\)-test provides a useful conceptual introduction because it directly analyzes:
and tests whether:
It therefore makes the within-subject nature of the crossover design immediately visible.
However, the paired analysis does not by itself provide the full flexibility of a mixed model.
The Period Effect in the Worked Example
Suppose the mean outcome in Period 1 is:
and the mean outcome in Period 2 is:
The crude period difference is:
This suggests a time-period difference of approximately 0.9 mmHg.
That difference should not automatically be interpreted as evidence that one treatment is superior.
Because each sequence receives both treatments in opposite orders, period and treatment can be separated through the prespecified model.
Sequence and Treatment Are Not the Same Thing
Suppose the AB group has an overall mean outcome different from the BA group.
That is a sequence difference.
It does not automatically mean that Treatment A has a different biological effect from Treatment B.
The treatment comparison is based on treatment exposure within the randomized crossover structure.
Carryover: The Most Important Design Issue
Consider a hypothetical treatment whose biological effect persists for several weeks after dosing stops.
Suppose:
- Period 1 lasts 4 weeks.
- The washout lasts 1 week.
- Period 2 lasts 4 weeks.
If Treatment A's effect persists beyond the one-week washout, then a participant receiving A first may begin Period 2 with residual A activity.
The Period 2 outcome is therefore no longer a clean measurement of the Period 2 treatment.
This can bias or complicate the treatment comparison.
Why Washout Is So Important
The washout period should be long enough, based on pharmacology and clinical behavior, to minimize residual treatment effects.
For a pharmacologic treatment, investigators may consider:
- Elimination half-life
- Active metabolites
- Pharmacodynamic half-life
- Target engagement
- Receptor occupancy
- Biological persistence
- Clinical response duration
A washout based only on plasma concentration can be inadequate if the clinical or pharmacodynamic effect persists substantially longer.
Carryover Is Not Simply a p-Value
A common but problematic strategy is:
"Test for carryover, and if the p-value is greater than 0.05, ignore carryover."
This is not a robust substitute for appropriate design.
Failure to detect carryover statistically does not establish that meaningful carryover is absent.
The most defensible approach is to design the trial so that carryover is scientifically implausible or sufficiently minimized through an appropriate washout.
Baseline Measurements
Many crossover studies collect a baseline measurement before each treatment period.
For example:
| Subject | Period | Baseline | Post-Treatment | Change |
|---|---|---|---|---|
| 001 | 1 | 150 | 142 | −8 |
| 001 | 2 | 147 | 142 | −5 |
The primary endpoint might then be the change from period-specific baseline.
This can improve interpretability and account for differences in pre-period status.
The exact choice of endpoint should be prespecified.
Baseline Adjustment in the Model
A mixed model can include baseline as a covariate:
where \(B_{ij}\) represents the appropriate baseline measure.
Whether baseline adjustment is preferable to analyzing change from baseline depends on the endpoint definition, estimand, design, and statistical analysis plan.
Period 1 Is Special
Period 1 has no previous treatment period.
Therefore, genuine carryover into Period 1 cannot exist.
Carryover becomes relevant primarily when analyzing subsequent periods.
This asymmetry is one reason crossover models require more thought than simply adding a treatment indicator to a repeated-measures dataset.
Treatment-by-Period Interaction
Investigators may also consider whether the treatment effect differs by period.
A treatment-by-period interaction can be represented as:
A large interaction can indicate that the treatment contrast is not stable across periods.
However, an interaction term should not automatically be interpreted as proof of carryover.
It may reflect:
- Period-dependent treatment effects
- Time trends
- Residual treatment effects
- Random variation
- Other treatment-by-time mechanisms
Modeling the Correlation
The repeated observations from the same participant are correlated.
A mixed model handles this by introducing a participant-level random effect.
A more general covariance structure can also be used when justified.
For a simple two-period design, a random-intercept model is often sufficient for the basic analysis, although the final covariance structure should be chosen according to the trial design and analysis plan.
Mixed Model in R
A simple implementation using lme4 is:
library(lme4) fit <- lmer( outcome ~ treatment + period + sequence + (1 | subject), data = crossover ) summary(fit)
Here:
treatmentestimates the treatment contrast.periodadjusts for systematic period differences.sequencerepresents treatment-order assignment.(1 | subject)accounts for repeated observations within participant.
Using Estimated Marginal Means
For a clinical trial, estimated marginal means are often useful for reporting adjusted treatment means and treatment contrasts.
library(emmeans) emm <- emmeans( fit, ~ treatment ) pairs(emm)
The resulting treatment contrast can provide:
- Estimated treatment difference
- Standard error
- Confidence interval
- Adjusted p-value, where applicable
Confidence Intervals
Suppose the model estimates:
with standard error:
For a large-sample 95% confidence interval:
giving:
or approximately:
The interpretation would be that the estimated treatment difference is approximately 3.5 mmHg in favor of Treatment A, with a 95% confidence interval for A minus B from approximately −5.46 to −1.54 mmHg.
The exact interval in an actual analysis would depend on the fitted model and degrees-of-freedom method.
Sample Size for a Crossover Trial
Sample size calculations for crossover studies are based on the variability of the within-subject treatment difference, not simply the overall between-subject standard deviation.
For a paired comparison, the approximate sample size for a two-sided test is:
where:
- \(\sigma_D\) = standard deviation of within-subject differences
- \(\Delta\) = clinically important treatment difference
- \(\alpha\) = two-sided type I error
- \(1-\beta\) = desired power
Equivalently:
The exact formula used in a protocol should account for the planned analysis and any anticipated dropout.
Worked Sample Size Calculation
Suppose investigators expect:
- Clinically meaningful difference: 4 mmHg
- SD of within-subject differences: 8 mmHg
- Two-sided \(\alpha=0.05\)
- Power = 90%
Using:
the approximate sample size is:
Therefore:
which gives approximately:
Thus, approximately 84 evaluable participants would be required under this simplified paired-difference calculation.
Dropout Inflation
Suppose the study requires 84 evaluable participants and investigators expect 10% attrition.
The enrollment target is approximately:
Therefore, approximately:
would be enrolled after rounding upward.
Why Crossover Designs Can Require Fewer Participants
The advantage becomes clearer by considering the relationship between within-subject correlation and variance.
If two treatment measurements have variance \(\sigma^2\) and correlation \(\rho\), then:
If \(\rho=0\), the within-subject difference has variance:
If \(\rho=0.80\), it becomes:
Thus, strong within-person correlation can make the treatment comparison much more precise.
But Efficiency Is Not Free
The statistical efficiency of a crossover design comes with additional assumptions and operational complexity.
Participants must:
- Remain in the study long enough to receive both treatments.
- Be able to safely receive both treatments.
- Have sufficiently stable disease.
- Have treatment effects that can wash out adequately.
- Return for repeated assessments.
A crossover trial can therefore have a smaller statistical sample size while requiring a longer and more complicated study for each participant.
Missing Data Are Particularly Important
Suppose a participant completes Period 1 but withdraws before Period 2.
The participant has contributed useful data, but there is no complete within-subject treatment difference.
A simple paired analysis would have difficulty using that participant fully.
A mixed-effects model can often use the available observations under its missing-at-random assumptions.
Missingness Can Still Bias the Study
Mixed models do not magically solve informative dropout.
If participants discontinue because a treatment is ineffective or because of treatment-related toxicity, missingness may depend on unobserved outcomes.
The analysis should therefore include appropriate sensitivity analyses when missing-data mechanisms could plausibly affect the treatment estimate.
Period 2 Missingness Can Be Especially Important
If treatment A is administered in Period 1 to one sequence and Period 2 to the other, differential dropout between sequences can reduce the information available for one treatment.
For example, if Treatment A is poorly tolerated, participants receiving A in Period 1 may discontinue before receiving B.
This can create both statistical and interpretational complications.
Therefore, dropout patterns should be examined by:
- Sequence
- Period
- Treatment
- Reason for discontinuation
Washout Period Design
Washout duration should be based on the expected persistence of the treatment effect.
For pharmacokinetic elimination, a rough conceptual relationship is:
where:
After \(m\) half-lives, the fraction remaining is:
For example, after five half-lives:
or approximately 3.1% of the initial concentration remains.
But this is only a pharmacokinetic illustration. The appropriate clinical washout cannot be determined from plasma half-life alone when pharmacodynamic or biological effects persist longer.
Sequence Balance
A simple 2×2 crossover is generally balanced when approximately equal numbers of participants are assigned to AB and BA.
For example:
| Sequence | Participants |
|---|---|
| AB | 42 |
| BA | 42 |
| Total | 84 |
Balanced randomization improves the symmetry of treatment exposure across periods and helps protect against loss of efficiency.
Randomization
Randomization should be performed before treatment assignment.
A typical randomization process is:
Blinding in Crossover Trials
Crossover trials can be double-blind, single-blind, or open-label.
When feasible, both treatments should be made indistinguishable to reduce expectation and assessment bias.
Blinding may be particularly important for subjective endpoints such as:
- Pain scores
- Symptom severity
- Patient-reported outcomes
- Quality-of-life measures
For objective endpoints, blinding remains valuable but the magnitude of bias may differ.
Analyzing a Binary Endpoint
Not every crossover study has a continuous endpoint.
Suppose the outcome is binary:
- Responder
- Non-responder
The analysis should account for the paired/repeated nature of the observations.
Potential approaches include:
- Conditional logistic models
- Generalized linear mixed models
- Marginal models using generalized estimating equations
- Exact methods in small samples
The appropriate method depends on the estimand, endpoint definition, missingness, and design.
Analyzing Time-to-Event Outcomes
Standard crossover designs are less natural for time-to-event outcomes because the outcome process can be permanently altered by the first treatment.
For example, if Treatment A prevents an event permanently, a participant cannot simply "wash out" that prevention and then experience an independent risk under Treatment B.
Bioequivalence Crossover Trials
Crossover designs are especially important in pharmacokinetic bioequivalence.
A standard 2×2 design may compare:
- Test formulation T
- Reference formulation R
Participants receive:
| Sequence | Period 1 | Period 2 |
|---|---|---|
| TR | T | R |
| RT | R | T |
Pharmacokinetic endpoints such as AUC and \(C_{\max}\) are commonly analyzed after logarithmic transformation.
The treatment difference on the log scale can then be exponentiated to obtain a geometric mean ratio.
For example, if:
then:
or 108%.
Log Transformation
For positively skewed endpoints such as AUC or \(C_{\max}\), the logarithmic transformation can stabilize variance and make treatment effects more naturally interpretable as ratios.
If:
then a difference on the log scale corresponds to a ratio on the original scale.
Analyzing the Log-Transformed Endpoint in R
crossover$log_auc <- log(crossover$auc) fit_auc <- lmer( log_auc ~ treatment + period + sequence + (1 | subject), data = crossover ) library(emmeans) emm_auc <- emmeans( fit_auc, ~ treatment ) contrast_auc <- contrast( emm_auc, method = "revpairwise" ) summary( contrast_auc, infer = TRUE )
If the estimated treatment difference is \(d\), the corresponding ratio is:
exp(d)
Classical Crossover ANOVA
Historically, crossover trials were often analyzed using analysis-of-variance approaches.
A traditional model might include:
- Sequence
- Subject nested within sequence
- Period
- Treatment
- Residual error
For a simple balanced 2×2 design, this framework is closely related to analysis of within-subject treatment differences.
Modern clinical-trial analyses often use mixed-effects models because they provide greater flexibility for incomplete observations and additional covariates.
A Simple Difference-Based Analysis
Suppose each participant completes both periods.
The analysis dataset can be converted to one row per participant:
| Subject | Sequence | Outcome A | Outcome B | Difference A − B |
|---|---|---|---|---|
| 001 | AB | −8 | −5 | −3 |
| 002 | AB | −10 | −6 | −4 |
| 003 | BA | −7 | −5 | −2 |
A simple model for the difference could then be:
where sequence can be included as an explanatory variable if appropriate.
The intercept estimates the treatment difference under the reference sequence, while the sequence coefficient accounts for a difference between sequences.
Why Not Just Analyze Period 2?
One sometimes encounters the idea of ignoring Period 1 and analyzing only the second period.
This can waste information and does not automatically solve carryover.
More importantly, if Period 2 contains different treatment assignments across sequences, simply analyzing Period 2 becomes a parallel-group comparison and throws away the within-subject information from Period 1.
Such an approach might be relevant in special designs or analyses motivated by specific carryover concerns, but it should not be adopted casually.
Baseline Covariates
A crossover model may also include baseline characteristics such as:
- Age
- Sex
- Baseline disease severity
- Renal function
- Body weight
- Other prespecified prognostic factors
For example:
Covariate adjustment should generally be prespecified and clinically justified.
Interaction With Baseline
If the scientific question concerns whether treatment effects differ according to baseline status, an interaction can be introduced:
The interaction coefficient estimates how the treatment effect changes with baseline.
Such analyses should generally be considered prespecified subgroup or effect-modification analyses rather than automatically included in the primary model.
Carryover and the Classical 2×2 Model
Suppose the outcome in Period 2 contains a residual effect from Period 1.
A simplified conceptual model could be:
where \(C_{i1}\) represents the residual effect of the Period 1 treatment.
If carryover is clinically meaningful, the interpretation of the usual treatment contrast becomes more complicated.
This is why the crossover design should be selected only after considering the expected persistence of treatment effects.
Direct Effects vs. Residual Effects
It is useful to distinguish:
- Direct treatment effect: effect caused by the treatment currently administered.
- Residual/carryover effect: effect in the current period attributable to a previous treatment.
The ideal crossover design attempts to make the residual component negligible.
Period Effects Can Be Real Even Without Carryover
Suppose there is no carryover at all.
The outcome can still change between Period 1 and Period 2 because of time.
For example, patients might improve naturally:
A period effect therefore does not imply carryover.
Carryover effect = residual influence of a treatment received in an earlier period.
They are conceptually different.
Sequence Effect and Carryover
In a simple 2×2 crossover, a strong sequence difference can raise concern about differential treatment-order effects.
However, sequence effect and carryover effect should not be treated as synonyms.
A sequence difference can result from several mechanisms, including:
- Carryover
- Chance imbalance
- Time-dependent disease behavior
- Sequence-specific adherence
- Other interactions with treatment order
Adherence
Crossover trials provide repeated exposure opportunities, so adherence should be evaluated separately by treatment and period.
Potential summaries include:
- Percentage of prescribed doses taken
- Number of missed doses
- Exposure duration
- Protocol deviations
- Premature treatment discontinuation
Adherence summaries should not automatically replace the primary intention-to- treat-style estimand.
Estimand Considerations
Modern clinical trial analysis should distinguish the statistical model from the clinical question being estimated.
For a crossover study, investigators should specify:
- Population
- Treatment conditions
- Endpoint
- Time point
- Summary measure
- Intercurrent-event strategy
For example, the primary estimand might concern the average difference in change from baseline between Treatment A and Treatment B among randomized participants under a specified strategy for treatment discontinuation.
The crossover model is then selected to estimate that target appropriately.
Intercurrent Events
Examples of intercurrent events include:
- Permanent treatment discontinuation
- Rescue medication
- Protocol deviations
- Hospitalization
- Adverse-event-related treatment interruption
The SAP should explain how these events affect the primary analysis.
This is particularly important because a crossover study can be disrupted if a participant cannot complete the second treatment period.
Period Length
The treatment period must be long enough for the endpoint to reflect the treatment effect.
A period that is too short can produce an immature treatment measurement.
A period that is unnecessarily long can increase:
- Dropout
- Nonadherence
- Calendar-time effects
- Study burden
The period length therefore represents a clinical and statistical design decision rather than merely an operational scheduling choice.
Washout vs. Run-In
A run-in period occurs before randomized treatment exposure and may be used to stabilize treatment or assess eligibility.
A washout period separates treatment periods to reduce residual effects.
They serve different purposes.
What Happens If There Is No Washout?
Some crossover designs intentionally omit a washout when the treatment effect is known to be sufficiently short-lived or when a continuous treatment design is otherwise appropriate.
The absence of a washout is not automatically invalid.
But the investigator must have a strong scientific rationale that treatment effects from the prior period will not compromise the subsequent measurement.
Four-Period Crossover Designs
Not all crossover studies are 2×2.
A 4-period design may repeat each treatment:
| Sequence | Period 1 | Period 2 | Period 3 | Period 4 |
|---|---|---|---|---|
| ABAB | A | B | A | B |
| BABA | B | A | B | A |
Replicated crossover designs can be useful when within-subject variability is important or when the study has specialized objectives such as bioequivalence variance estimation.
The analysis becomes more complex because additional periods and repeated treatment exposures must be modeled.
Latin-Square Crossover Designs
When there are more than two treatments, Latin-square designs can balance treatment order across periods.
For three treatments:
| Sequence | Period 1 | Period 2 | Period 3 |
|---|---|---|---|
| ABC | A | B | C |
| BCA | B | C | A |
| CAB | C | A | B |
Additional sequences can be used depending on the desired balance and number of treatments.
Balanced Incomplete Designs
When there are many treatments, asking every participant to receive every treatment can become impractical.
Balanced incomplete crossover designs allow each participant to receive only a subset of treatments while maintaining structured balance across treatment comparisons.
These designs require careful consideration of estimability, treatment order, period effects, and participant burden.
Crossover Designs in Pharmacokinetics
Crossover designs are especially natural in pharmacokinetic studies because within-participant PK variability can be substantially lower than between- participant variability.
For example, each participant may receive:
- Test formulation
- Reference formulation
with PK sampling performed after each treatment.
The primary comparison can then be made within participant.
PK Parameters and Log Transformation
Common PK parameters include:
- AUC
- \(C_{\max}\)
- \(T_{\max}\)
- Half-life
AUC and \(C_{\max}\) are commonly right-skewed and may be analyzed on the logarithmic scale.
\(T_{\max}\), being a time parameter, often requires different methods.
Power and Correlation
Because crossover power depends on within-subject variability, pilot or historical data should provide an estimate of:
or an equivalent within-subject variance parameter.
Using a between-subject SD in a crossover sample-size calculation can produce a substantially incorrect enrollment target.
Sensitivity to the Within-Subject SD
Suppose the clinically meaningful difference is fixed at 4 units.
| Within-Subject SD | Relative Sample Size |
|---|---|
| 4 | Lower |
| 6 | Moderate |
| 8 | Higher |
| 10 | Substantially higher |
Because sample size is proportional to the variance:
underestimating the within-subject standard deviation can substantially underestimate the required sample size.
R Sample Size Illustration
alpha <- 0.05 power <- 0.90 delta <- 4 sd_diff <- 8 n <- 2 * (qnorm(1 - alpha / 2) + qnorm(power))^2 * sd_diff^2 / delta^2 ceiling(n)
This returns an approximate participant count based on the paired-difference normal approximation.
For a final clinical trial calculation, software designed specifically for the planned crossover model may be preferable.
Analyzing the Worked Dataset in R
Suppose the dataset is in long format with variables:
subject sequence period treatment baseline outcome
A change-from-baseline endpoint can be created as:
crossover$change <- crossover$outcome - crossover$baseline
The primary mixed model can then be specified as:
library(lme4)
fit <- lmer(
change ~ treatment + period + sequence +
(1 | subject),
data = crossover
)
summary(fit)
Obtaining the Treatment Contrast
library(emmeans) emm <- emmeans( fit, ~ treatment ) contrast( emm, method = "revpairwise", infer = c(TRUE, TRUE) )
The resulting contrast should be reported with its estimate, confidence interval, and p-value where hypothesis testing is part of the prespecified analysis.
Model Diagnostics
For a continuous endpoint, examine:
- Residual distributions
- Residual-versus-fitted plots
- Potential outliers
- Influential observations
- Random-effect assumptions
- Covariance structure
For example:
plot(fit) qqnorm(resid(fit)) qqline(resid(fit))
The goal is not to demand perfect normality, but to determine whether the model provides a reasonable representation of the data and whether influential observations materially affect conclusions.
Outliers
Crossover studies can be particularly sensitive to unusual within-subject differences.
Suppose most treatment differences are around 3 units, but one participant has a difference of 25 units.
That participant can substantially influence the estimated treatment effect.
Investigators should distinguish:
- Data errors
- Protocol deviations
- Genuine extreme observations
A genuine extreme observation should not simply be removed because it weakens statistical significance.
Protocol Deviations
Potential crossover-specific protocol deviations include:
- Incorrect treatment sequence
- Insufficient washout
- Wrong treatment administered
- Period timing outside predefined windows
- Incomplete treatment exposure
- Premature discontinuation
The SAP should define how such deviations affect analysis populations and sensitivity analyses.
Intent-to-Treat and Per-Protocol Analyses
The primary analysis population should follow the trial's prespecified estimand and analysis strategy.
A per-protocol analysis can be useful as a sensitivity analysis when crossover deviations are scientifically important.
However, selectively excluding participants based on their observed treatment response can introduce bias.
Common Mistakes in Crossover Analysis
- Ignoring period. Treatment and time are not interchangeable.
- Ignoring within-subject correlation. The two observations from one participant are not independent.
- Assuming a washout automatically eliminates carryover. The adequacy of washout depends on pharmacologic and clinical persistence.
- Testing carryover only after the trial. A nonsignificant test does not establish that clinically important carryover is absent.
- Using a parallel-group analysis without considering the crossover structure. This can waste within-subject information.
- Using the wrong treatment contrast. Period 1 minus Period 2 is not necessarily A minus B.
- Using between-subject SD for sample size. The relevant variance is generally the within-subject treatment-difference variance.
- Dropping participants with one missing period automatically. Mixed models can often use available observations.
- Assuming sequence imbalance is harmless. Severe imbalance can reduce efficiency and complicate interpretation.
- Overinterpreting sequence effects as carryover. A sequence difference has multiple possible explanations.
- Ignoring treatment-period interaction. A treatment effect that changes across periods may require additional investigation.
- Choosing the analysis after seeing the results. The primary model and treatment contrast should be prespecified.
A Practical Crossover Trial Design Workflow
What Should Be in the Statistical Analysis Plan?
A crossover SAP should make the analysis reproducible from the protocol and database structure.
At minimum, specify:
- Primary estimand
- Primary endpoint
- Treatment contrast
- Treatment coding
- Period coding
- Sequence coding
- Random-effects structure
- Fixed-effects structure
- Covariance assumptions
- Baseline adjustment
- Confidence interval method
- Degrees-of-freedom method
- Handling of missing observations
- Protocol deviation rules
- Sensitivity analyses
- Subgroup analyses
- Multiplicity strategy where applicable
A Model SAP Statement
For an illustrative continuous endpoint, the SAP might state:
The actual SAP should provide sufficient detail to specify the exact coding, covariance structure, degrees-of-freedom method, estimand, and handling of intercurrent events.
Interpreting the Final Treatment Estimate
Suppose the primary analysis produces:
with:
The interpretation depends on the endpoint direction.
If lower values represent improvement, Treatment A is estimated to produce a 3.2-unit greater improvement than Treatment B.
The confidence interval quantifies uncertainty around that estimated treatment contrast.
Statistical Significance vs. Clinical Importance
Suppose a very large crossover study estimates:
with a highly significant p-value.
That does not automatically mean the treatment effect is clinically meaningful.
The clinically relevant question is whether a difference of 0.8 units matters to patients or clinicians.
Therefore, the protocol should define the clinically meaningful treatment difference before the study whenever possible.
Crossover Designs and Multiplicity
A simple 2×2 crossover has one primary treatment comparison.
Multiplicity becomes more important when the study contains:
- Multiple treatments
- Multiple primary endpoints
- Multiple time points
- Multiple treatment contrasts
- Multiple subgroup claims
The multiplicity strategy should be prespecified rather than developed after examining the results.
When a Crossover Design Should Be Avoided
A crossover design may be inappropriate when:
- The disease is rapidly progressive.
- The treatment produces irreversible changes.
- The first treatment permanently changes the response to later treatment.
- There is no credible washout period.
- The endpoint has substantial secular drift.
- Treatment discontinuation creates lasting effects.
- Patients cannot safely receive both treatments.
- The first treatment may alter eligibility for the second.
Example: Why Disease Progression Can Break the Design
Suppose a treatment is intended to slow an irreversible neurodegenerative process.
A participant receives A for six months and then B after washout.
Even if the pharmacologic effect of A disappears, the participant's disease state may have changed permanently during Period 1.
The Period 2 response under B is therefore not comparable to what would have happened under B six months earlier.
This is a fundamental reason why crossover designs are generally unsuitable for many progressive or irreversible diseases.
Example: A Good Crossover Setting
Now consider a short-acting treatment for a stable condition where:
- Each treatment produces an effect within hours.
- The effect disappears within one day.
- A several-day washout is feasible.
- The endpoint can be measured repeatedly.
- The disease state remains reasonably stable.
This setting is much more naturally suited to crossover methodology.
Reporting Results
A crossover trial report should clearly describe:
- Number randomized to each sequence
- Number completing each period
- Reasons for discontinuation
- Washout duration
- Protocol deviations
- Treatment exposure
- Period-specific descriptive statistics
- Primary treatment contrast
- Confidence interval
- P-value when applicable
- Sensitivity analyses
Illustrative Results Table
| Analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| Adjusted Treatment A − B | −3.2 | −5.1 to −1.3 | 0.002 |
| Period 2 − Period 1 | −0.9 | −1.8 to 0.0 | 0.052 |
| Sequence effect | 1.1 | −0.8 to 3.0 | 0.25 |
The primary result is the treatment contrast.
Period and sequence estimates provide supporting information about the design and temporal structure of the data.
Do Not Turn Secondary Effects Into the Primary Conclusion
The existence of a period effect does not mean the treatment comparison is invalid.
Likewise, a sequence effect does not automatically establish carryover.
The interpretation must consider:
- The prespecified model
- The design assumptions
- Clinical plausibility
- Magnitude of effects
- Confidence intervals
- Sensitivity analyses
Sensitivity Analysis Strategy
A robust crossover analysis may include:
- Primary mixed-effects analysis
- Complete-case analysis
- Within-subject difference analysis
- Alternative covariance structure
- Analysis excluding major protocol deviations
- Analysis using alternative baseline definitions
- Missing-data sensitivity analysis
The purpose is to determine whether the primary treatment conclusion is robust to reasonable alternative assumptions.
A Complete Crossover Analysis Checklist
| Question | What to Check |
|---|---|
| Is crossover clinically appropriate? | Stable disease and reversible treatment effects |
| Is washout adequate? | PK, PD, biological and clinical persistence |
| Are sequences randomized? | AB and BA appropriately balanced |
| Is the endpoint appropriate? | Repeatedly measurable and interpretable |
| Is period included? | Prespecified adjustment |
| Is within-subject correlation handled? | Mixed model or appropriate paired method |
| Is the treatment contrast explicit? | A minus B or B minus A |
| Is carryover addressed? | Primarily through design and washout |
| Is missing data addressed? | Primary and sensitivity strategies |
| Is sample size based on the right variance? | Within-subject variability |
R: Complete Illustrative Workflow
library(lme4)
library(emmeans)
# Long-format crossover dataset
#
# subject
# sequence
# period
# treatment
# baseline
# outcome
# Change from period-specific baseline
crossover$change <-
crossover$outcome -
crossover$baseline
# Primary mixed-effects model
fit <- lmer(
change ~ treatment + period + sequence +
(1 | subject),
data = crossover
)
# Model summary
summary(fit)
# Adjusted treatment means
emm <- emmeans(
fit,
~ treatment
)
# Treatment contrast
trt_contrast <- contrast(
emm,
method = "revpairwise"
)
summary(
trt_contrast,
infer = TRUE
)
# Diagnostics
plot(fit)
qqnorm(resid(fit))
qqline(resid(fit))
R: Creating Treatment Differences
For a simple complete 2×2 dataset, treatment differences can also be calculated directly.
library(dplyr)
library(tidyr)
wide <- crossover %>%
select(subject, treatment, change) %>%
pivot_wider(
names_from = treatment,
values_from = change
)
wide$A_minus_B <-
wide$A - wide$B
summary(wide$A_minus_B)
This is useful for descriptive checks and for understanding the basic within-subject treatment comparison.
R: Paired Analysis as a Sensitivity Check
t.test( wide$A, wide$B, paired = TRUE )
This analysis tests the average within-subject treatment difference.
It should not automatically replace the primary mixed model when the protocol specifies a more comprehensive model.
R: Checking Sequence Balance
table(crossover$sequence)
prop.table(
table(
unique(crossover[c("subject", "sequence")])$sequence
)
)
The objective is to verify that randomization produced the intended allocation across sequences.
R: Summarizing Treatment by Period
crossover %>%
group_by(
sequence,
period,
treatment
) %>%
summarise(
n = n(),
mean = mean(change, na.rm = TRUE),
sd = sd(change, na.rm = TRUE),
.groups = "drop"
)
These descriptive summaries can reveal unusual period-specific patterns before the primary model is interpreted.
How to Think About the Analysis
A useful mental model is:
The Most Important Statistical Concept
The defining advantage of a crossover trial is not simply that every participant receives two treatments.
It is that the treatment comparison can exploit the fact that the same participant provides responses under both treatment conditions.
The treatment contrast can therefore be viewed as:
rather than comparing two completely different groups of people.
If individuals differ substantially from one another but are relatively stable within themselves, this can greatly improve precision.
The Most Important Design Concept
The statistical efficiency of the crossover design is secondary to the validity of the treatment comparison.
Before choosing a crossover design, ask:
- Can every participant safely receive both treatments?
- Will the first treatment leave a clinically meaningful residual effect?
- Can an adequate washout be implemented?
- Will the underlying disease remain sufficiently stable?
- Can the endpoint be measured repeatedly?
If the answer to these questions is unfavorable, a parallel-group design may be scientifically more appropriate even if it requires more participants.
Worked Example Summary
| Component | Illustrative Value |
|---|---|
| Design | 2×2 randomized crossover |
| Treatment sequences | AB and BA |
| Participants per sequence | 20 |
| Total participants | 40 |
| Periods | 2 |
| Washout | 2 weeks |
| Primary endpoint | Change in systolic blood pressure |
| Primary contrast | A − B |
| Illustrative treatment estimate | −3.5 mmHg |
| Illustrative interpretation | A produces greater reduction than B |
| Primary analysis | Linear mixed-effects model |
Final Takeaways
A crossover trial is a powerful design when its underlying scientific assumptions are appropriate.
The central concepts are:
- Each participant receives multiple treatments.
- Randomization determines treatment sequence.
- Within-subject comparisons provide statistical efficiency.
- Period effects must be considered.
- Sequence describes treatment order and should not be confused with treatment effect.
- Carryover is primarily controlled through appropriate study design and washout.
- Repeated observations require appropriate statistical modeling.
- Mixed-effects models provide a flexible framework for continuous crossover endpoints.
- Sample size should be based on within-subject variability.
- Missing observations and intercurrent events should be addressed prospectively.
References
Jones, B. & Kenward, M.G. (2015).
Design and Analysis of Cross-Over Trials.
3rd Edition. Chapman & Hall/CRC.
Senn, S. (2002).
Cross-over Trials in Clinical Research.
2nd Edition. Wiley.
Grizzle, J.E. (1965).
The two-period change-over design and its use in clinical trials.
Biometrics, 21, 467–480.
Wellek, S. & Blettner, M. (2012).
On the proper use of the crossover design in clinical trials.
Deutsches Ärzteblatt International, 109(15), 276–281.
Chow, S.C. & Liu, J.P. (2013).
Design and Analysis of Clinical Trials: Concepts and Methodologies.
3rd Edition. Wiley.
Iannuccelli, M. et al.
Statistical considerations in crossover clinical trials.
Clinical trial methodology literature on treatment, period, sequence, and
carryover effects.