Tutorials › AP Statistics › Matching Subjects Into Pairs

Choosing a mean-inference procedure · Tutorial 773 of 1000

Matching Subjects Into Pairs

Recognize a matched-pairs design, analyze one treatment difference per pair, and distinguish what matching can control from what random assignment can establish.

Intermediate 11 min read

What You'll Learn

  • Identify how matching subjects on relevant characteristics creates genuine pairs.
  • Distinguish matched-pairs experiments from crossover designs and independent-group experiments.
  • Define a difference for each pair in an order that answers the research question.
  • Choose a paired t procedure and check conditions using the differences.
  • Explain how random assignment within pairs and subject selection affect conclusions.
  • Avoid treating matching as proof that all differences between subjects have been controlled.

When Similar Subjects Receive Different Treatments

In a crossover design, each subject receives both treatments, so the two responses from that subject form a pair. There is another way to create paired data: match two different subjects with similar characteristics, then assign one subject in each pair to each treatment. For example, researchers might match students with similar prior test scores, athletes with similar performance records, or plants with similar initial heights.

The characteristics used for matching should be relevant to the response. If two matched subjects are alike in an important way that affects the outcome, comparing their responses can reduce the impact of that source of variation. But the two subjects are still different individuals, and matching does not make them identical. The design determines the analysis: use the pairwise differences, rather than treating all responses in the two treatment groups as unrelated.

Definition: In a matched-pairs design, two distinct subjects with similar characteristics are linked into a pair, and one subject in each pair receives each treatment. The response comparison is made within each pair by calculating one difference per pair.

This is different from a crossover design, where the same subject receives both treatments. In a matched-pairs experiment, each subject receives only one treatment, but the two subjects’ responses are linked because they were deliberately matched. As in “Paired t Versus Two-Sample t,” it is the design-based link—not the number of columns in a data table—that tells you whether the data are paired.

From Matched Subjects to a Mean Difference

Choose and state an order for subtracting the responses. For instance, if treatment A is a new tutoring program and treatment B is the usual program, you might define \(d=\text{score under A}-\text{score under B}\). A positive difference then means the subject receiving A in that pair had the higher score. Repeat the same subtraction order for every pair.

The population parameter is \(\mu_d\), the true mean of the differences for the population represented by the pairs, using the stated order. Each pair contributes one value to the sample of differences. The paired t procedure is a one-sample t procedure on those differences, as explained in “Comparing Treatments Using Crossover Designs.” It is not a two-sample t procedure on the original treatment groups.

$$ d_i=\text{response for the subject receiving A in pair }i -\text{response for the subject receiving B in pair }i $$

For a question about whether the treatments have different mean responses, the hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\). A directional alternative is appropriate only when the research question specifies a direction. The paired t statistic uses the mean and standard deviation of the differences, and its degrees of freedom are one less than the number of pairs.

$$ t=\frac{\bar{x}_d-0}{s_d/\sqrt{n}} \qquad\qquad df=n-1 $$

Here, \(n\) is the number of complete pairs, not the total number of subjects. The standard error estimates how much the sample mean difference \(\bar{x}_d\) would vary from sample to sample. Reversing the order of subtraction changes the signs of the differences and the direction of the interpretation, so settle the order before stating hypotheses or drawing a conclusion.

How Matching and Assignment Work Together

In a well-designed matched-pairs experiment, researchers first form pairs using relevant characteristics and then randomly assign one subject in each pair to treatment A and the other to treatment B. That assignment makes the treatment comparison within each pair fair on average. It also helps separate treatment effects from differences between pairs, such as one pair having much greater prior experience than another.

Matching alone does not establish cause and effect. If subjects choose their own treatments, or if researchers assign treatments without randomization, the groups may differ in other ways even after matching. As covered in “Statistical Significance in Observational Studies,” a statistically significant result in an observational study is evidence of an association, not proof that a treatment caused the difference.

Matching also does not guarantee a more precise result. It often helps when the matching characteristic is strongly related to the response, because subjects in a pair tend to have similar outcomes apart from treatment. If the characteristic is not useful, or if matched subjects differ in other important ways, the differences may still be variable. For inference, the relevant spread is the spread of the pairwise differences.

Conditions: The response is quantitative, and subjects are deliberately matched into pairs with one response under each treatment per pair. For a randomized experiment, treatment assignment should be randomized within each pair. The pairs should come from an appropriate random sample or randomized experiment, and differences from distinct pairs should be independent; apply the 10% condition when sampling without replacement. For a small number of pairs, the distribution of the differences should have no strong skewness or outliers.

Check the differences, not the two treatment groups separately, when assessing the shape condition for paired t inference. If subjects were randomly sampled from a finite population without replacement, the 10% condition concerns the number of sampled pairs relative to the number of eligible pairs. If the pairs were not randomly sampled, be careful about generalizing beyond the subjects or population the design can represent.

Worked Examples: Choosing and Using Paired t

Worked Example: Matching Seedlings by Initial Height

A greenhouse team randomly selects ten pairs of seedlings with similar initial heights. Within each pair, one seedling is randomly assigned fertilizer A and the other fertilizer B. The response is growth in centimeters over four weeks. Define \(d=\text{growth with A}-\text{growth with B}\). The ten differences are \(1, 2, -1, 3, 0, 2, 1, -2, 2, 2\) centimeters. A plot of the differences shows no strong skewness or outliers. Is there evidence that fertilizer A produces greater mean growth?

State: Let \(\mu_d\) be the true mean difference in four-week growth, fertilizer A minus fertilizer B, for seedlings represented by the random sample. A positive difference means more growth with A. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\).

Plan: Use a paired t test on the ten within-pair differences. Growth is quantitative, seedlings were matched by initial height, and one seedling in each pair was randomly assigned to each fertilizer. The pairs were randomly selected, and we assume differences from distinct pairs are independent. With only ten pairs, the stated plot’s check for strong skewness and outliers is important. If sampling was without replacement, the population of eligible pairs should be at least ten times the sample of ten pairs to meet the 10% condition. These conditions support the paired t procedure.

Do: The differences sum to \(10\), so \(\bar{x}_d=10/10=1\) centimeter. The sum of squared deviations from 1 is \(22\), giving \(s_d=\sqrt{22/9}\approx1.563\) centimeters. The standard error is \(1.563/\sqrt{10}\approx0.494\) centimeter. Therefore:

$$ t=\frac{1-0}{\sqrt{22/9}/\sqrt{10}} =\frac{1}{\sqrt{22/90}} \approx2.023 $$

There are \(10-1=9\) degrees of freedom. For the greater-than alternative, the p-value is approximately \(0.037\), rounded. Assuming the true mean difference is zero, this is the probability of obtaining a t statistic of \(2.023\) or greater in the direction of greater growth with A.

Conclude: At \(\alpha=0.05\), \(0.037<0.05\), so reject \(H_0\). The data provide convincing evidence that fertilizer A produces greater mean four-week growth than fertilizer B for seedlings represented by this experiment. Because treatment was randomly assigned within pairs, the experiment supports a cause-and-effect conclusion for its experimental setting.

Worked Example: Estimating a Matched Difference in Quiz Scores

A school randomly selects 12 pairs of students with similar scores on a prior quiz. Within each pair, one student is randomly assigned to a new review activity and the other to the usual review. Let \(d=\text{new-activity score}-\text{usual-review score}\), measured in points. The differences have \(\bar{x}_d=4.5\) points and \(s_d=3.0\) points, with no strong skewness or outliers. Find and interpret a 95% confidence interval for the population mean difference.

State: Let \(\mu_d\) be the true mean difference in quiz scores, new activity minus usual review, for the population represented by the randomly selected matched pairs. We want to estimate \(\mu_d\).

Plan: Use a paired t interval for the 12 pairwise differences. The response is quantitative, students are matched on prior quiz scores, and treatment was randomly assigned within each pair. The pairs were randomly sampled, and the differences have no strong skewness or outliers. Assume the differences from distinct pairs are independent and, if sampled without replacement, the population has at least ten times as many eligible pairs as the sample. These conditions support the interval.

Do: The degrees of freedom are \(12-1=11\), so the 95% critical value is \(t^*\approx2.201\). The standard error is \(3.0/\sqrt{12}\approx0.866\) points, and the margin of error is \(2.201(0.866)\approx1.905\) points. Thus:

$$ 4.5\pm2.201\left(\frac{3.0}{\sqrt{12}}\right) =4.5\pm1.905 \approx(2.595,\ 6.405)\text{ points} $$

Conclude: We are 95% confident that the population mean score with the new review activity is between about 2.595 and 6.405 points higher than with the usual review. The interval estimates the mean of matched-pair differences in the stated order, not the difference between two unrelated sample means.

Worked Example: Matching Patients Without Random Treatment Assignment

A health program evaluates recovery time in days. Researchers randomly sample seven pairs of patients who are similar in age and initial symptom severity. In each pair, one patient received a new program and the other received usual care; treatment was not randomly assigned. Define \(d=\text{recovery time with usual care}-\text{recovery time with the program}\). The differences are \(2, 1, 4, -1, 3, 0, 2\) days. A plot shows no strong skewness or outliers. Is there evidence of an association between treatment received and mean recovery time?

State: Let \(\mu_d\) be the true mean difference in recovery time, usual care minus the program, for the population represented by the sampled matched pairs. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\). A positive difference means longer recovery under usual care.

Plan: Use a paired t test because each pair was deliberately formed from similar patients and contributes one difference. Recovery time is quantitative, the patients were randomly sampled, and the differences from distinct pairs are assumed independent. With seven pairs, the stated shape check is important. If sampling was without replacement, the eligible population should contain at least 70 pairs for the 10% condition. Treatment was not randomly assigned, so a result can support an association but cannot by itself show that the program caused a change in recovery time.

Do: The differences sum to \(11\), so \(\bar{x}_d=11/7\approx1.571\) days. The sum of squared differences is \(35\); the sum of squared deviations is \(35-11^2/7=124/7\). Therefore \(s_d=\sqrt{(124/7)/6}=\sqrt{62/21}\approx1.718\) days. The standard error is \(1.718/\sqrt{7}\approx0.650\) days, and:

$$ t=\frac{1.571-0}{1.718/\sqrt{7}} \approx2.419 $$

There are \(7-1=6\) degrees of freedom. For the greater-than alternative, the p-value is approximately \(0.026\), rounded. Assuming the population mean difference is zero, this is the probability of obtaining a t statistic of \(2.419\) or greater in the direction of longer recovery with usual care.

Conclude: At \(\alpha=0.05\), \(0.026<0.05\), so reject \(H_0\). The data provide convincing evidence of an association between treatment received and mean recovery time for patients represented by the sample, with longer recovery associated with usual care. Because treatment was not randomly assigned, this finding does not establish that the program caused shorter recovery.

Common Mistakes and AP Exam Tips

  • Calling every similar pair a crossover: In a matched-pairs design, two different subjects receive one treatment each. In a crossover, the same subject receives both. Both designs use paired differences, but describe the design accurately.
  • Using an unpooled two-sample t procedure because there are two treatment groups: Deliberate matching creates a link between responses. Analyze one difference per pair with a paired t procedure.
  • Matching on a characteristic but not randomizing treatment: Matching can improve comparability, but it is not a substitute for random assignment. Do not claim a causal effect from a matched observational study.
  • Checking the wrong distributions: For paired t inference, assess the distribution of the pairwise differences, especially when the number of pairs is small.
  • Counting subjects instead of pairs: If there are \(n\) pairs, the paired t procedure has \(n\) differences and \(n-1\) degrees of freedom, even though there are \(2n\) subjects.
  • Leaving the subtraction order unclear: State the order and explain what a positive difference means. Match the hypotheses and conclusion to that order.

For full credit, identify the deliberate matching, define the population mean difference and subtraction order, and name the paired t procedure. In the Plan step, check random selection or assignment as appropriate, independence of pairs, the 10% condition when relevant, and the shape of the differences. In the conclusion, distinguish evidence of an association from evidence of a cause-and-effect relationship.

Key takeaway: When distinct subjects are matched by relevant characteristics and one subject in each pair receives each treatment, analyze one treatment difference per pair with a paired t procedure. Matching creates the pairing; random assignment within pairs supports a cause-and-effect conclusion.

Check Your Understanding

Use the matched-pairs ideas to answer each question.

  1. Researchers match cyclists on recent race times and assign one cyclist in each pair to each training plan. What makes this a paired design even though no cyclist uses both plans?
  2. If \(d=\text{response under plan A}-\text{response under plan B}\), what does a positive \(\mu_d\) mean?
  3. A study has 15 complete matched pairs. How many differences are analyzed, and what are the paired t degrees of freedom?
  4. Why should the shape check for a small-sample paired t procedure focus on the differences rather than each treatment group separately?
  5. Researchers match patients but allow patients to choose their treatments. If a paired t test finds convincing evidence of a mean difference, what can the study conclude—and what can it not establish?