From Twelve Score Pairs to One Test
A paired t test turns each student’s before-and-after scores into a single difference. The test then uses those 12 differences to assess a claim about the true mean change. In “Running a Paired t Test on the Calculator,” you learned how to enter the difference list; here, we work through a complete test with numerical scores and interpret each step.
We will define the difference as after score minus before score. With that order, a positive difference means a higher score after tutoring, and a negative difference means a lower score. As in “Defining the Difference Variable” and “The Parameter \(\mu_d\) in Paired Inference,” the parameter \(\mu_d\) describes the true mean of these differences in the population of interest.
The arithmetic is only one part of the response. As in “Conditions for One-Sample Versus Paired Data,” check the design and the distribution of the differences. With 12 pairs, the shape assessment matters: use the graph-based approach from “What to Check With a Small Sample of 12 Observations,” rather than relying on a large-sample rule.
Worked Examples
Worked Example: Testing for a Mean Change in Tutoring Scores
A fictional school has 150 students who completed a tutoring program. A random sample of 12 students took the same skills test before and after tutoring. Their scores, out of 100, are shown below. Define \(d_i=\text{after}_i-\text{before}_i\), in score points. Test whether the true mean difference is nonzero at \(\alpha=0.05\). For these differences, a Normal probability plot is reasonably linear and the data display shows no strong skewness or pronounced outlier.
| Student | Before | After | Difference \(d_i\) |
|---|---|---|---|
| 1 | 62 | 58 | -4 |
| 2 | 70 | 67 | -3 |
| 3 | 68 | 65 | -3 |
| 4 | 75 | 73 | -2 |
| 5 | 81 | 80 | -1 |
| 6 | 73 | 72 | -1 |
| 7 | 66 | 67 | 1 |
| 8 | 78 | 79 | 1 |
| 9 | 72 | 74 | 2 |
| 10 | 84 | 87 | 3 |
| 11 | 69 | 72 | 3 |
| 12 | 77 | 81 | 4 |
State. Let \(\mu_d\) be the true mean after-minus-before test-score difference, in points, for all students in this school’s tutoring-program population. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\), with \(\alpha=0.05\). A two-sided alternative is appropriate because the question asks whether the mean changed in either direction.
Plan. Use a paired t test because the same student took both tests, and the analysis uses one difference per student. The 12 students were randomly sampled from the 150 students in the population, supporting generalization to that population. Since \(12\) is less than 10% of \(150\), the 10% condition is met. Each student contributes one difference, and different students are separate individuals, so independence between pairs is reasonable. Because \(n=12\) is small, assess the shape of the differences; the stated Normal probability plot and data display provide reasonable support for the shape condition.
Do. The differences sum to zero, so \(\bar d=0/12=0\). Their squared values sum to \(80\). Using the sample standard deviation formula,
The standard error can also be checked by squaring it: \((80/11)/12=80/132\), whose square root is approximately \(0.7785\). Thus the test statistic and degrees of freedom are
For the two-sided alternative, the p-value is the probability of a t statistic at least as far from zero as \(0\), assuming \(H_0\) is true. That probability is \(p=1.000\). A TI-84 T-Test in Data mode, using the difference list and selecting \(\ne\mu_0\), gives the same result.
Conclude. Because \(1.000>0.05\), fail to reject \(H_0\). These data do not provide convincing evidence that the true mean after-minus-before score difference for students in this tutoring-program population is nonzero. This does not prove that the mean change is exactly zero. Also, a before-and-after comparison without random assignment does not by itself show that tutoring caused any change.
Worked Example: Testing for an Increase in Scores
In another fictional school, 12 students are randomly selected from 180 students who attended a tutoring course. Let each difference be after minus before, in points. The differences are \(1,2,3,4,0,1,2,3,4,-1,2,3\). The question is whether the population mean score difference is greater than zero. A graph-based check finds no pronounced outlier or strong departure from an approximately Normal pattern.
State and plan. Let \(\mu_d\) be the true mean after-minus-before score difference, in points, for students represented by this population. Test \(H_0:\mu_d=0\) against \(H_a:\mu_d>0\) at \(\alpha=0.05\). Use a paired t test because the scores are linked within students. The students were randomly selected, \(12\) is less than 10% of \(180\), different students contribute separate differences, and the stated graph-based check supports the shape condition for this small sample.
Do. The differences sum to \(24\), so \(\bar d=24/12=2\). Their squared values sum to \(74\). The sum of squared deviations is \(74-24^2/12=74-48=26\). Therefore,
For \(H_a:\mu_d>0\), use the upper-tail probability for \(t=4.506\) with 11 degrees of freedom. A calculator gives \(p\approx0.0004\) (rounded). Since \(0.0004<0.05\), reject \(H_0\).
Conclude. The sample provides convincing evidence that the true mean after-minus-before score difference for students represented by this population is greater than zero. This supports a positive average difference, but the before-and-after design alone does not establish that the tutoring course caused the increase.
Worked Example: A Positive Sample Mean Is Not Enough
A third fictional school randomly samples 12 students from 160 students who participated in a study-skills program. The after-minus-before score differences, in points, are \(-3,-2,-1,0,0,1,1,2,2,3,4,5\). Test whether the true mean difference is nonzero at \(\alpha=0.05\). A Normal probability plot is reasonably linear, with no pronounced outlier.
State and plan. Let \(\mu_d\) be the true mean after-minus-before score difference, in points, for students represented by the population. Use \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\). A paired t test is appropriate because each difference comes from one student measured twice. The students were randomly sampled, \(12\) is less than 10% of \(160\), separate students contribute independent differences, and the graph-based evidence is reasonable for the shape condition.
Do. The differences sum to \(12\), so \(\bar d=12/12=1\). Their squared values sum to \(74\), giving a sum of squared deviations of \(74-12^2/12=62\). Thus,
The two-sided p-value for \(t=1.459\) with 11 degrees of freedom is approximately \(0.1725\) (rounded). Since \(0.1725>0.05\), fail to reject \(H_0\).
Conclude. Although the sample mean difference is positive, the sample does not provide convincing evidence that the true mean after-minus-before score difference for students represented by this population is nonzero. The direction of the sample mean alone does not settle the test.
What a Complete AP Response Communicates
A complete paired t test connects the study question to the data and then to the conclusion. The four-step organization used in “Writing the Full Four-Step One-Sample t Test” also works here, with the paired differences as the data. State the parameter and hypotheses, explain why the paired t procedure and its conditions are appropriate, show the calculation and p-value, and conclude in context.
- Keep the subtraction order visible. If \(d_i=\text{after}_i-\text{before}_i\), use that definition in the table, hypotheses, and conclusion. Reversing it reverses the signs and changes the meaning of a one-sided alternative.
- Count pairs, not measurements. Twelve students provide 12 differences, so \(n=12\) and \(df=11\), not \(n=24\) or \(df=23\).
- Choose the alternative from the research question. Use a two-sided alternative for change in either direction and a one-sided alternative only when the question specifies a direction. Do not choose the tail after looking at the sign of \(\bar d\) or \(t\).
- State what the p-value assumes. For example: “Assuming the true mean difference is zero, the probability of a t statistic at least as extreme as the observed value in either direction is \(0.1725\).”
- Do not say that failing to reject proves no change. Say that the data do not provide convincing evidence for the alternative claim. This distinction is emphasized in “Fail to Reject H0 Does Not Mean H0 Is True.”
- Separate a mean-change conclusion from a causal claim. Pairing handles the relationship between each student’s two scores. It does not, by itself, rule out other explanations for a before-and-after difference.
Check Your Understanding
Use the paired t test ideas from the examples to answer each question.
- If a student’s before score is 74 and after score is 81, what is the after-minus-before difference, and what does its sign mean?
- A study has 12 pairs and tests \(H_0:\mu_d=0\). What degrees of freedom should the paired t test use?
- For a two-sided test, a calculator reports \(t=0\). What is the p-value, and why?
- A test gives \(p=0.032\) at \(\alpha=0.05\). State the decision using AP-standard wording.
- Why does a significant paired t test of before-and-after scores not automatically prove that a program caused the change?