Tutorials › AP Statistics › Two-Sample Versus Paired Test on the Same Numbers

Two-sample t hypothesis tests · Tutorial 739 of 1000

Two-Sample Versus Paired Test on the Same Numbers

Compare paired and two-sample t tests on the same measurements, and see why the data’s design—not the calculator output—determines which result is valid.

Intermediate 10 min read

What You'll Learn

  • Identify whether observations are genuinely paired or come from independent groups.
  • Compare the standard errors and conclusions from paired and unpaired analyses of repeated measurements.
  • Explain why pairing unrelated observations does not create a valid paired t test.
  • See how changing the matches while keeping the same measurements can change a paired result.
  • Use participant or unit identifiers to check that each calculated difference reflects a real pair.

The Design Determines the Test

A set of measurements does not tell you by itself whether to use a paired t test or a two-sample t test. The study design does. If the same people are measured twice, or if units are deliberately matched, the analysis uses one difference per genuine pair. If the groups consist of separate, unlinked units, the analysis compares two independent sample means.

In “Common Errors in Two-Sample t Tests,” you saw why a paired procedure is not appropriate without genuine pairs. Here we make the consequences visible: we will analyze the same measurements using the correct design and then see what changes when the data are treated as if they came from a different design. A calculator can produce an answer for an inappropriate procedure; that does not make the answer valid.

Key check: Before calculating, identify the observational units and the study’s links between them. For repeated measurements, preserve each unit’s actual match and calculate its difference. For independent groups, do not invent matches between units.

What Changes When Paired Data Are Treated as Independent?

For paired data, define a difference consistently for every pair—for example, after minus before—and analyze the sample of differences using a one-sample t procedure. The standard error is based on the standard deviation of those differences, \(s_d/\sqrt{n}\), where \(n\) is the number of pairs. The earlier tutorial “Writing a Complete Paired t Test Solution” explains the conditions and interpretation for that procedure.

A two-sample t test instead treats the two sets of measurements as independent samples. Its standard error is based on the two separate sample standard deviations: \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). This calculation does not use which first measurement belongs with which second measurement. When people are measured twice, ignoring the matches discards information about how their measurements move together.

Definition: A genuine pair is a design-based link between two observations, such as two measurements on the same person or measurements on deliberately matched units. Pairing means analyzing the within-pair differences, not simply placing observations in rows.

Worked Example: Repeated Measurements Analyzed Both Ways

In a fictional study, six adult cyclists are randomly sampled. Each cyclist’s recovery time after exercise is measured before and after following a hydration routine. The measurements, in minutes, are shown below. Assume the sample is less than 10% of the target population of cyclists.

CyclistBeforeAfterAfter − Before
12018−2
22422−2
32219−3
42623−3
52523−2
62320−3

These are paired data because each after measurement belongs to the same cyclist as its before measurement. The research question is whether the population mean recovery time differs before and after the routine. Use a two-sided test at \(\alpha=0.05\).

1
State.
Let \(\mu_d\) be the true mean change in recovery time, defined as after minus before, for the target population of adult cyclists under these measurement conditions. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\).
2
Plan and check conditions.
Use a paired t test because each cyclist has both measurements, giving six genuine pairs. The cyclists were randomly sampled, and the 10% condition is met by assumption. With only six pairs, the differences should be checked for strong skewness or outliers; these six differences are clustered between −2 and −3 minutes, with no apparent outlier. These conditions support using a t procedure for the mean difference.
3
Do.
The mean difference is \(\bar{x}_d=(-2-2-3-3-2-3)/6=-2.5\) minutes. The deviations from this mean are \(0.5,0.5,-0.5,-0.5,0.5,-0.5\), whose squared values sum to \(1.5\). Thus \(s_d^2=1.5/5=0.3\), and \(s_d=\sqrt{0.3}\approx0.5477\) minutes.
$$ SE_{\bar{x}_d}=\frac{s_d}{\sqrt{n}} =\frac{0.5477}{\sqrt{6}} \approx0.2236\text{ minutes}, \qquad t=\frac{\bar{x}_d-0}{SE_{\bar{x}_d}} =\frac{-2.5}{0.2236} \approx-11.180, \qquad df=6-1=5. $$

The two-sided p-value is the probability, assuming \(\mu_d=0\), of a t statistic at least as far from zero as \(-11.180\), in either direction, with 5 degrees of freedom. A calculator gives \(p\approx0.0001\), rounded.

4
Conclude in context.
Because \(0.0001<0.05\), reject \(H_0\). These data provide convincing evidence that the true mean change in recovery time for the target population of adult cyclists is not zero.

Now treat the before and after measurements incorrectly as two independent samples. The before mean is \(140/6\approx23.333\) minutes, and the after mean is \(125/6\approx20.833\) minutes, so their difference is still \(-2.5\) minutes. The before sum of squared deviations is \(23.333\), giving \(s_1^2=23.333/5\approx4.667\). The after sum is \(22.833\), giving \(s_2^2=22.833/5\approx4.567\).

$$ SE_{\bar{x}_1-\bar{x}_2} =\sqrt{\frac{4.667}{6}+\frac{4.567}{6}} \approx1.2405\text{ minutes}, \qquad t=\frac{-2.5}{1.2405}\approx-2.015. $$

The Welch degrees of freedom are approximately 10, and the two-sided p-value is about \(0.0715\), rounded. This unpaired analysis would fail to reject at the 0.05 level, unlike the paired analysis. But the measurements were not independent: the same cyclists provided both values. The two-sample result answers a procedure’s hypothetical independent-samples question, not the question supported by this study design. The paired test is the valid analysis.

The study describes a randomly sampled group measured before and after a routine; it does not say that the routine was randomly assigned or that measurement order was randomized. Therefore, the test result is evidence of a mean change, but it does not by itself establish that the routine caused the change.

What Changes When Independent Data Are Paired Artificially?

The reverse mistake is to make pairs from independent groups just because the sample sizes match or the values can be arranged in rows. As “Identifying Two Independent Samples” explains, a real design link—not a convenient layout—is what makes observations paired. Without genuine matches, a difference calculated between arbitrary rows has no meaningful interpretation.

Worked Example: Matching Unrelated Observations by Row

A fictional researcher randomly samples five independent units from each of two production lines and records a quantitative quality score. The scores are Line A: \(10,12,14,16,18\), and Line B: \(11,13,15,17,19\). No unit appears in both groups, and the observations were not matched.

The appropriate comparison is an unpooled two-sample t test. The group means are \(\bar{x}_A=14\) and \(\bar{x}_B=15\). For each group, the sum of squared deviations from its mean is 40, so each sample variance is \(40/(5-1)=10\). The standard error, test statistic for equal means, Welch degrees of freedom, and two-sided p-value are:

$$ SE=\sqrt{\frac{10}{5}+\frac{10}{5}}=2, \qquad t=\frac{14-15}{2}=-0.500, \qquad df=8, \qquad p\approx0.631. $$

The two-sample procedure is appropriate because the groups contain different, independent units; assume the samples are each less than 10% of their populations and the plots show no strong skewness or outliers. For illustration, suppose someone pairs the values by their displayed row. The differences \(A-B\) are then \(-1,-1,-1,-1,-1\), with a standard deviation of zero. A paired t statistic cannot be calculated because its standard error would be zero. More importantly, these row-by-row differences have no design-based meaning.

Even the arbitrary pairing can change if the same Line B scores are put in a different order. For example, pairing Line A in its displayed order with Line B ordered as \(19,11,17,13,15\) produces differences \(-9,1,-3,3,3\). Their mean is still \(-1\), but their squared deviations from \(-1\) sum to 104, so \(s_d^2=104/4=26\) and \(SE=\sqrt{26/5}\approx2.2804\). The resulting statistic is \(t=-1/2.2804\approx-0.439\), with 4 degrees of freedom and a two-sided p-value of about \(0.684\).

The two arbitrary pairings give different paired-test calculations for the same group measurements. Neither paired result is valid: the study did not define those pairs. The independent two-sample analysis uses the actual design and does not depend on how the scores are ordered within either list.

Pairing Information Cannot Be Recovered From the Group Summaries

A further lesson is that the two groups’ means and standard deviations do not tell you how repeated measurements are paired. If participant identifiers are lost or the matches are scrambled, the marginal measurements may remain unchanged while the paired differences—and therefore the paired standard error and test statistic—change. Never choose matches based on which arrangement gives a preferred result.

Worked Example: Same Measurements, Different Matches

Return to the cyclist measurements. Keep all six before values and all six after values exactly as recorded, but imagine that the after values have been incorrectly reassigned to different cyclists in the order \(23,23,20,18,22,19\). This is the same collection of after measurements, but it is no longer the actual participant matching.

The artificial differences are \(3,-1,-2,-8,-3,-4\), with mean \(-15/6=-2.5\) minutes. Their deviations from \(-2.5\) are \(5.5,1.5,0.5,-5.5,-0.5,-1.5\); the squared deviations sum to 65.5. Thus \(s_d^2=65.5/5=13.1\), and \(s_d=\sqrt{13.1}\approx3.6194\) minutes.

$$ SE_{\bar{x}_d}=\sqrt{\frac{13.1}{6}}\approx1.4776\text{ minutes}, \qquad t=\frac{-2.5}{1.4776}\approx-1.692, \qquad df=5. $$

The two-sided p-value is about \(0.151\), rounded, rather than the approximately \(0.0001\) from the true matches. The independent two-sample calculation would also be unchanged from the earlier example, because the two groups’ measurements and summary statistics have not changed. The paired calculation changes because it depends on which observations belong together.

This reassignment is not a valid analysis of the cyclists: the study records each person’s before and after measurements, so those actual matches must be retained. The example shows why the pairing identifiers are part of the data, not a formatting choice.

Common Mistakes and AP Exam Tips

  • Choosing by appearance rather than design: Equal sample sizes or a table with aligned rows does not establish pairing. Name the units and explain the genuine link.
  • Discarding the matched differences: For repeated measurements, state which value is subtracted from which, then use one difference per pair. Reversing the subtraction changes the sign and interpretation.
  • Pairing independent units after data collection: Sorting values, matching similar scores, or pairing the first observations in two lists does not create a paired design. Use the unpooled two-sample t procedure when the groups are independent.
  • Assuming matching always makes the p-value smaller: The paired standard error depends on the variation in the actual differences. It can be smaller or larger than the independent-samples standard error; the design, not the desired result, determines the method.
  • Reporting both tests as equally acceptable: When asked to choose a procedure, identify the one supported by the design. You may compare an incorrect calculation to explain the consequence, but label it clearly as invalid for that study.
  • Overstating the conclusion: A small p-value is evidence about a population mean difference under the stated conditions. Whether it supports a cause-and-effect conclusion depends on the study design, including random assignment.

For full credit, an AP response should make the design link explicit. Say that the same subjects were measured twice and define the direction of the differences, or state that the two groups contain independent units with no pairing. Then use the corresponding test and interpret its result in context. If two analyses of the same measurements disagree, explain which design assumption is actually justified rather than choosing the more impressive p-value.

Key takeaway: The data’s genuine links determine the procedure. Analyze actual repeated or matched observations with their pairwise differences; analyze unlinked groups with an unpooled two-sample t test. Arbitrary matching changes calculations but cannot change the study design.

Check Your Understanding

For each situation, identify the design and explain which comparison is appropriate.

  1. The same 18 students take a skills assessment before and after a workshop. What makes these observations paired, and what quantity should be analyzed?
  2. Two independently sampled groups have 12 observations each. Does the matching sample size justify a paired t test? Explain.
  3. Why can the paired standard error change when measurements are reassigned to different people, even if both groups’ means and standard deviations stay the same?
  4. In the cyclist example, why is the two-sample result not a valid alternative to the paired result?
  5. A researcher randomly assigns separate machines to either of two settings and measures each machine once. Which design feature tells you whether a two-sample or paired procedure is appropriate?