Tutorials › AP Statistics › Comparing Treatments Using Crossover Designs

Choosing a mean-inference procedure · Tutorial 772 of 1000

Comparing Treatments Using Crossover Designs

See why crossover experiments are analyzed as paired data, how to define the differences, and when a paired t procedure gives a useful comparison.

Intermediate 10 min read

What You'll Learn

  • Identify a crossover design in which every subject receives both treatments.
  • Define the order of subtraction for each subject’s paired difference.
  • Explain why pairing can reduce variation from differences among subjects.
  • Check the conditions for paired t inference in a crossover experiment.
  • Use a paired t test or interval to compare the treatments’ mean responses.
  • Recognize how treatment order and carryover can limit a crossover study.

When Every Subject Receives Both Treatments

Some experiments compare treatments by assigning different subjects to separate groups. A crossover design instead has each subject receive both treatments, usually in separate periods. For example, each participant might try two study methods, each device might be tested under two settings, or each plant might receive two treatments at different times. Because the two responses from one subject are linked, the comparison is a paired problem.

As explained in “Experiments Comparing Two Treatments,” the design determines the procedure. A crossover study does not create two independent groups just because the data can be arranged in two columns. Each row links the measurements from the same subject, so the analysis focuses on a difference for each subject. This is the same paired-data logic introduced in “Paired t Versus Two-Sample t.”

Definition: In a crossover design, each subject receives both treatments, often in a randomized order. The two responses from a subject form a pair, and the analysis compares the treatments using one difference per subject.

The order of subtraction must match the research question. If treatment A is a new method and treatment B is a standard method, define \(d=\text{response under B}-\text{response under A}\) when a positive difference would mean the new method reduces the response. Then calculate that difference for every subject. The mean of these differences, \(\bar{x}_d\), estimates the population mean difference, \(\mu_d\), for the stated order.

For example, if the response is time to finish a task, defining \(d=B-A\) means a positive difference indicates that a subject took longer under the standard method. That would be consistent with the new method taking less time. Reversing the subtraction would reverse the sign of every difference, the hypotheses’ direction, and the interpretation—but not the strength of evidence in a two-sided test.

$$ d_i=\text{response for subject }i\text{ under A} -\text{response for subject }i\text{ under B} $$

This formula is only one possible order. State the order in words and use it consistently. The paired t procedure treats the \(n\) differences as one sample. It tests or estimates \(\mu_d\), not two unrelated population means. For a test of no mean difference, the null value is zero.

$$ H_0:\mu_d=0 \qquad H_a:\mu_d\ne 0,\quad \mu_d>0,\quad\text{or}\quad \mu_d<0 $$

The paired t statistic is the sample mean difference minus the null value, divided by the estimated standard error of the differences. Its degrees of freedom are \(n-1\), where \(n\) is the number of complete pairs. A confidence interval for the mean difference uses the same sample of differences.

$$ t=\frac{\bar{x}_d-0}{s_d/\sqrt{n}} \qquad\qquad \bar{x}_d\ \pm\ t^*\frac{s_d}{\sqrt{n}} $$

Here, \(\bar{x}_d\) is the mean of the paired differences, \(s_d\) is their sample standard deviation, and \(t^*\) is the critical value for the chosen confidence level with \(n-1\) degrees of freedom. Pairing can make the differences less variable than the individual responses because each subject serves as their own comparison. It controls for stable differences among subjects, such as one person generally completing a task faster than another.

Order, Carryover, and Conditions

A crossover design has a feature that a simple before-and-after comparison may not: treatments occur in periods. If everyone receives A first and B second, changes over time could be confused with treatment differences. For example, participants may become more practiced, tired, or familiar with a measurement process. Randomizing which treatment comes first helps prevent order from being systematically tied to one treatment.

A second concern is carryover. A treatment may continue affecting a subject after its period ends. If that happens, the response during the second treatment can depend partly on the first treatment. Researchers may use a suitable waiting period, called a washout period, or otherwise design the study to reduce carryover. Randomizing treatment order does not automatically eliminate a lasting treatment effect. If carryover is plausible and not addressed, the paired differences may not cleanly represent the treatments’ direct comparison.

Conditions: The response is quantitative, and each subject has a measurement under both treatments. The subjects or treatment orders should come from an appropriate random sample or randomized experiment. Differences from distinct subjects should be independent; apply the 10% condition if sampling without replacement from a finite population. For a small number of pairs, the distribution of the differences should have no strong skewness or outliers. The design should also make period effects and carryover unlikely to distort the comparison.

For paired t inference, inspect the differences, not the two treatment distributions separately. A paired t test is a one-sample t test applied to the list of differences. With a small sample, a graph of the differences should not show strong skewness or outliers. With a larger sample, t procedures are generally more robust, but extreme outliers can still be problematic. The treatment responses could each look unusual on their own while the differences are reasonably well behaved, or the reverse.

Randomly assigning treatment order can support a cause-and-effect interpretation of the comparison for the experimental subjects, provided the study is carried out appropriately and carryover is controlled. It does not by itself make those subjects representative of a broader population. As discussed in “Generalizing Conclusions to a Population,” generalization depends on how the subjects were selected.

Worked Examples: Paired Comparisons in Crossover Studies

Worked Example: Comparing Two Task Methods

Eight volunteers each complete a task using a new method and a standard method. Treatment order is randomized, and enough time is allowed between trials to make a lasting effect from the first method unlikely. Let \(d=\text{standard-method time}-\text{new-method time}\), in minutes. The eight differences are \(2, 1, 3, 0, 2, -1, 1, 0\). A plot of the differences shows no strong skewness or outliers. Is there evidence that the new method has a lower mean completion time?

State: Let \(\mu_d\) be the true mean difference in completion time, standard method minus new method, for subjects like those in this experiment. A positive \(\mu_d\) means the new method has a lower mean completion time. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\).

Plan: Use a paired t test on the subject-by-subject differences. Completion time is quantitative, and each volunteer used both methods, so the data are paired. Treatment order was randomized, and the waiting time is intended to limit carryover. Assume the volunteers’ differences are independent. Since \(n=8\) is small, check the differences’ shape: the stated plot has no strong skewness or outliers. These conditions support the procedure. The study does not describe a random sample, so a conclusion should be limited to subjects like those tested rather than generalized automatically to all people.

Do: The sum of the differences is \(8\), so \(\bar{x}_d=8/8=1\) minute. Their deviations from 1 are \(1,0,2,-1,1,-2,0,-1\), and the sum of squared deviations is \(12\). Thus \(s_d=\sqrt{12/7}\approx1.309\) minutes, and the standard error is \(1.309/\sqrt{8}\approx0.463\) minutes. The test statistic is:

$$ t=\frac{1-0}{\sqrt{12/7}/\sqrt{8}} =\frac{1}{\sqrt{12/56}} \approx 2.160 $$

There are \(8-1=7\) degrees of freedom. For the greater-than alternative, the p-value is approximately \(0.034\), rounded. Assuming the true mean difference is zero, this is the probability of obtaining a t statistic of \(2.160\) or greater in the direction of a lower mean completion time with the new method.

Conclude: At \(\alpha=0.05\), \(0.034<0.05\), so reject \(H_0\). The experiment provides convincing evidence that the new method produces a lower mean completion time than the standard method for subjects like those tested, assuming the crossover design’s order and carryover safeguards were effective.

Worked Example: Estimating a Difference in Battery Life

Ten randomly selected technicians each test two battery settings on the same device, with the setting order randomized. The response is hours of operation. Define \(d=\text{hours under setting A}-\text{hours under setting B}\). The resulting differences have mean \(\bar{x}_d=2.4\) hours and standard deviation \(s_d=1.6\) hours. The differences have no strong skewness or outliers, and the testing arrangement makes carryover unlikely. Find and interpret a 95% confidence interval for the population mean difference.

State: Let \(\mu_d\) be the true mean difference in battery life, setting A minus setting B, for the population represented by the random sample of technicians. We want to estimate \(\mu_d\).

Plan: Use a paired t interval, treating the ten differences as one sample. Each technician tested both settings, the order was randomized, and the differences are reasonably free of strong skewness and outliers. The technicians were randomly sampled, and if they were sampled without replacement, the population should be at least ten times the sample size for the 10% condition. The design also addresses the concern that one setting might systematically be tested first.

Do: There are \(9\) degrees of freedom. The 95% critical value is \(t^*\approx2.262\). The standard error is \(1.6/\sqrt{10}\approx0.506\) hours, so the margin of error is \(2.262(0.506)\approx1.144\) hours. The interval is:

$$ 2.4\pm2.262\left(\frac{1.6}{\sqrt{10}}\right) =2.4\pm1.144 \approx(1.256,\ 3.544)\text{ hours} $$

Conclude: We are 95% confident that the population mean battery life under setting A is between about 1.256 and 3.544 hours greater than under setting B. This interval estimates the mean of the paired differences, with A minus B as the stated order; it is not an interval for two separate means.

Worked Example: A Small Difference in Sleep Duration

Six participants each try two evening routines in randomized order. Let \(d=\text{hours of sleep after routine A}-\text{hours after routine B}\). The differences are \(1,2,0,1,-1,1\) hours. A plot of these differences shows no strong skewness or outliers, and the study uses a suitable interval between trials to reduce carryover. Is there evidence that the routines produce different mean sleep durations?

State: Let \(\mu_d\) be the true mean difference in sleep duration, routine A minus routine B, for subjects like those in the study. The question asks whether the means differ, so \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\).

Plan: Use a paired t test on the six differences because every participant tried both routines. Treatment order was randomized, carryover is addressed, and the differences show no strong skewness or outliers. We assume differences from distinct participants are independent. With only six participants, the shape check is important. The participants are not described as a random sample, so broad generalization is not warranted.

Do: The differences sum to \(4\), giving \(\bar{x}_d=4/6\approx0.667\) hours. The sum of squared differences is \(8\), and the sample sum of squares is \(8-4^2/6=16/3\). Therefore \(s_d=\sqrt{(16/3)/5}=\sqrt{16/15}\approx1.033\) hours. The standard error is \(1.033/\sqrt{6}\approx0.422\) hours, and:

$$ t=\frac{0.667-0}{1.033/\sqrt{6}} \approx1.581 $$

The degrees of freedom are \(6-1=5\), and the two-sided p-value is approximately \(0.175\), rounded. If the true mean difference were zero, this is the probability of obtaining a t statistic at least as far from zero as \(1.581\) in either direction.

Conclude: At \(\alpha=0.05\), \(0.175>0.05\), so fail to reject \(H_0\). The data do not provide convincing evidence that the two evening routines produce different mean sleep durations for subjects like those tested. This result does not prove the routines have equal mean effects.

Common Mistakes and AP Exam Tips

  • Treating the two measurements as independent groups: The same subject contributes both responses. Calculate one difference per subject and use a paired t procedure, rather than an unpooled two-sample t test.
  • Changing the subtraction order partway through: Define the difference before calculating it. State what a positive difference means, and align the hypotheses and conclusion with that definition.
  • Checking the treatment distributions instead of the differences: For paired t inference, the small-sample shape condition applies to the distribution of the differences.
  • Assuming random treatment order solves carryover: Randomizing order helps with order effects, but a lasting effect of the first treatment may still affect the second response. Describe the design safeguards and be cautious if carryover is plausible.
  • Calling the result proof of equality after a large p-value: If you fail to reject, say the data do not provide convincing evidence for the alternative. Do not conclude that the treatments have identical effects.
  • Claiming generalization based only on random order: Randomizing treatment order supports the experiment’s comparison; it does not make the subjects representative of a wider population.

For full credit, identify the paired design, define the population mean difference in a clear order, and name the paired t procedure. In the Plan step, check the conditions using the subjects and the differences, including the small-sample shape check when needed. In the conclusion, interpret the result in the original units and context.

Key takeaway: In a crossover experiment, each subject receives both treatments, so analyze the subject-level differences with a paired t procedure. Define the subtraction order, check the differences and the treatment-order design, and consider whether carryover could distort the comparison.

Check Your Understanding

Use the crossover-design ideas to answer each question.

  1. Each participant tests two music settings, and the response is reading time. What makes the observations paired rather than independent?
  2. If \(d=\text{time under B}-\text{time under A}\), what does a positive mean difference say about the treatments’ average times?
  3. For paired t inference with 12 complete subjects, what are the degrees of freedom, and which data should be checked for skewness and outliers?
  4. Why might randomizing which treatment comes first help, and why does it not necessarily eliminate carryover?
  5. A paired t test fails to reject \(H_0:\mu_d=0\). Write an appropriate evidence conclusion without claiming the treatments are equal.