The Link Between Observations Decides the Procedure
The previous tutorial, “One-Sample t Versus Two-Sample t,” showed that procedure choice depends on the target and how the data were collected. Now there is another important question: are the two sets of measurements linked? If the same people are measured twice, the observations are paired. Treating them as two independent groups would ignore information built into the design.
A paired t procedure analyzes the differences within genuine pairs. A two-sample t procedure compares means from two separate, unlinked groups. The two procedures can both involve two sets of numbers, but their parameters and calculations differ.
The link must come from the design, not from the fact that two lists have the same number of values. Each before measurement must belong to the same person as its after measurement, or each unit must be deliberately matched to its partner. Two unrelated groups of equal size are still independent groups.
Turn Each Pair Into One Difference
For paired data, state the subtraction order before calculating. For example, if \(d=\text{before}-\text{after}\), a positive difference means the measurement was lower after the event or intervention. Reversing the order would reverse the signs and change the interpretation, so keep the chosen order consistent.
Let \(\mu_d\) represent the true mean of the pairwise differences for the population of interest. The paired procedure uses the sample mean \(\bar{d}\) and sample standard deviation \(s_d\) of the differences. It does not treat the before and after measurements as two independent samples.
Here, \(n\) is the number of pairs, and \(\mu_{d,0}\) is the value specified by the null hypothesis, often 0. This is the one-sample t statistic applied to the differences. In contrast, the two-sample t statistic from “The Two-Sample t Test Statistic” uses the separate sample means and estimates variability across two independent groups.
Notice the kind of independence required: the two observations within one pair are linked, but the differences from different pairs should be independent. Do not claim that every individual measurement is independent when the design deliberately connects measurements within pairs.
Worked Example: Resting Pulse for the Same 18 Adults
Worked Example: Resting Pulse Before and After a Routine
A researcher randomly selects 18 adults enrolled in a community wellness program and records each person’s resting pulse, in beats per minute, before and after a four-week breathing routine. The research question is whether the population mean pulse is lower after the routine. The following invented data are the 18 differences, defined as before minus after:
State: Let \(\mu_d\) be the true mean difference, in beats per minute, for adults in the population represented by the random sample, where \(d=\text{before}-\text{after}\). The question asks whether the mean difference is positive, so the hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\).
Plan: Use a paired t test because each before measurement is linked to an after measurement from the same adult. Equivalently, use a one-sample t test on the 18 differences. The random sample supports inference to the population it represents. Assuming that population contains at least 180 adults, the 10% condition is satisfied because \(18\leq 0.10(180)\). The differences from different adults are treated as independent. The listed differences are roughly symmetric around 3.0, with no apparent outlier, so using a t procedure is reasonable for this sample of 18 pairs.
Do: The sum of the differences is 54.0, so
The deviations from 3.0 are \(-1.8,-1.6,\ldots,-0.2,0.2,\ldots,1.6,1.8\). Their squared deviations sum to \(22.8\), so the sample standard deviation is
Under the null hypothesis, \(\mu_{d,0}=0\). The test statistic is therefore
The one-sided p-value from a t distribution with 17 degrees of freedom is less than 0.0001. This is the probability, assuming the true mean difference is 0, of obtaining a t statistic of at least 10.99 in the direction of a positive mean difference.
Conclude: Because the p-value is less than 0.05, reject \(H_0\). The data provide convincing evidence that the population mean before-minus-after resting pulse difference is positive for adults represented by this random sample. This conclusion describes a mean change; because the study does not describe random assignment to a comparison condition, it does not by itself establish that the routine caused the change.
Worked Example: Matching Makes a Difference
Worked Example: Two Mulches on Matched Seedlings
A student compares two mulches using six pairs of seedlings. Within each pair, the seedlings are planted close together in similar conditions; one receives Mulch A and the other receives Mulch B. After four weeks, the student records each seedling’s height. The differences, defined as A minus B, are \(1.2,-0.4,0.8,1.0,-0.2,\) and \(0.6\) centimeters.
Identify the link: The seedlings are deliberately matched in pairs based on planting location. The measurements are not from two unlinked groups, even though different seedlings receive the two mulches.
Identify the target and summarize the differences: Let \(\mu_d\) be the true mean difference in four-week height for the population of matched seedlings, with \(d=\text{height under A}-\text{height under B}\). The sample differences sum to \(3.0\) centimeters, giving
Choose the procedure: A paired t procedure is the relevant method because the design links seedlings into six pairs. It analyzes the six differences, not the two collections of six heights as independent samples. Before carrying out inference, the student would need to check the randomization or sampling design, independence across pairs, and the distribution of the differences. With only six pairs, the shape and any outliers in those differences deserve particular attention.
Worked Example: Equal Group Sizes Do Not Create Pairs
Worked Example: Two Unmatched Study Groups
A school compares weekly study hours for 18 students who attend an optional study workshop and 18 different students who do not attend. Students were not matched, and no student appears in both groups. The workshop group’s sample mean is 7.4 hours with sample standard deviation 1.8 hours. The other group’s sample mean is 6.6 hours with sample standard deviation 1.5 hours.
Identify the target and design: Let \(\mu_W\) and \(\mu_N\) be the true mean weekly study hours for the populations represented by the workshop and non-workshop groups. The target is \(\mu_W-\mu_N\). The 18 students in one group do not correspond to the 18 students in the other group; equal sample sizes do not make observations paired.
Choose the procedure: This is a two-sample t setting because the question compares the means of two separate, unlinked groups. The observed difference is \(7.4-6.6=0.8\) hour. For illustration, the unpooled estimated standard error is
The corresponding statistic for a null difference of 0 is \(0.8/0.552\approx1.45\). This uses the two-sample structure. Subtracting a workshop student’s hours from an arbitrarily selected non-workshop student’s hours would invent pairs that the study did not create.
A Reliable Choice Routine
Before doing calculations, apply this routine to the question and the data collection:
State what is measured and its units.
Ask whether the same unit contributes both observations or whether units were deliberately matched.
Choose an order such as before minus after or A minus B, and explain what positive differences mean in context.
Paired data target \(\mu_d\) and call for paired t. Two unlinked groups target \(\mu_1-\mu_2\) and call for two-sample t.
Assess the sampling or assignment design, independence at the relevant level, and the distribution needed for t inference.
This routine prevents a common shortcut: “There are two columns, so use two-sample t.” Two columns may contain repeated measurements on the same people, matched observations, or genuinely independent groups. The design distinguishes these cases.
Common Mistakes and AP Exam Tips
- Using two-sample t for before-and-after measurements: The same person’s measurements are linked. Define one difference per person and analyze the differences with paired t.
- Calling the paired observations independent: Measurements within a pair are linked. A full-credit explanation says that the differences from distinct pairs should be independent.
- Assuming equal sample sizes mean paired data: Two groups can have equal sizes without any deliberate matching. State what links each observation before choosing paired t.
- Leaving the subtraction order unstated: Define \(d\) in words, including which measurement is subtracted from which. Interpret the sign using that order.
- Using the wrong parameter: Paired inference concerns \(\mu_d\), the population mean of the pairwise differences. Two-sample inference concerns \(\mu_1-\mu_2\), the difference between two population means.
- Claiming causation from a before-and-after comparison alone: A paired design accounts for differences among individuals, but it does not automatically rule out other changes over time. Match causal language to the study design.
A concise AP response might say: “Because the same 18 adults were measured before and after, the observations are paired. Define \(d=\text{before}-\text{after}\); the target is \(\mu_d\), so a paired t procedure is appropriate if its conditions are met.” For independent groups, identify the two populations and explain that no observations are deliberately linked.
Check Your Understanding
For each situation, identify whether the data are paired or independent and name the appropriate mean-inference procedure.
- A clinic records each of 25 patients’ cholesterol levels before and after a program. What is the same-unit-twice clue, and what difference could be defined?
- A researcher compares reaction times for 20 people who drank tea with reaction times for 20 different people who drank water. No one is matched. Are the groups paired just because both samples contain 20 people?
- A farm pairs plots with similar soil conditions and applies Fertilizer A to one plot and Fertilizer B to the other in each pair. What data should a paired t procedure analyze?
- For a before-minus-after difference, a positive \(d\) means the after measurement is lower. If the research question asks whether the mean after measurement is higher, which direction of alternative hypothesis would match this definition of \(d\)?
- In paired data, which observations should be independent for the paired t procedure: the two measurements within each pair, or the differences from separate pairs?