Tutorials › AP Statistics › Writing a Complete Paired t Test Solution

Paired data and paired t procedures · Tutorial 700 of 1000

Writing a Complete Paired t Test Solution

Practice connecting every part of a paired t test—from the difference definition and conditions to the p-value and a careful conclusion.

Intermediate 10 min read

What You'll Learn

  • Define the population mean difference using a clear subtraction order and units
  • Match paired t test hypotheses to the research question
  • Check the design, independence, 10% condition, and distribution of differences
  • Show the standard error, t statistic, degrees of freedom, and p-value
  • Write a conclusion that reflects the evidence and the study design

One Connected Response for a Paired t Test

A paired t test is a test about the population mean of within-pair differences. In “Writing the Full Four-Step One-Sample t Test,” you practiced organizing a test as State, Plan, Do, and Conclude. The same structure works for paired data, but the analysis must begin by defining the difference for each pair and the population mean of those differences.

This tutorial brings together skills from the earlier paired t tutorials. In particular, use the difference variable and parameter definitions from “Defining the Difference Variable” and “The Parameter \(\mu_d\) in Paired Inference,” and use the conditions discussed in “Conditions for a Paired t Procedure.” Here, the goal is to make the entire free-response solution clear, justified, and connected from start to finish.

Key idea: A complete paired t test identifies the population mean difference in context, states hypotheses that match the defined subtraction order, checks conditions using the differences, shows the test calculation and p-value, and concludes in context.

A Four-Step Structure for a Paired t Test

The four steps are not four disconnected boxes to fill in. Each one prepares the reader to understand the next: the parameter gives the hypotheses meaning, the conditions justify the procedure, and the calculation supports the conclusion.

1
State.
Define \(\mu_d\) in context, including the measurement, units, population, and subtraction order. State \(H_0\), \(H_a\), and the significance level \(\alpha\). For a no-average-difference claim, the null is \(H_0:\mu_d=0\).
2
Plan.
Name the paired t test and check its conditions. Refer to the actual study design, compare the number of pairs with the population size when the 10% condition applies, and assess the distribution of the differences.
3
Do.
Calculate the standard error, t statistic, degrees of freedom, and p-value. Use the alternative hypothesis to select the appropriate tail area.
4
Conclude.
Compare the p-value with \(\alpha\), state whether you reject or fail to reject \(H_0\), and describe the evidence about the mean difference in context.

For \(n\) pairs, the calculation uses the sample mean \(\bar d\) and sample standard deviation \(s_d\) of the \(n\) differences—not the two original measurement columns as if they were independent samples.

$$ SE=\frac{s_d}{\sqrt{n}}, \qquad t=\frac{\bar d-0}{s_d/\sqrt{n}}, \qquad df=n-1. $$

The numerator is \(\bar d-0\) because the usual null hypothesis says the population mean difference is zero. The p-value is calculated from a t distribution with \(n-1\) degrees of freedom, using the tail or tails specified by \(H_a\). As in “Writing a Conclusion for a One-Sample t Test,” a conclusion should say “convincing evidence” when the test rejects, and “do not provide convincing evidence” when it fails to reject. Do not say that a test proves the null hypothesis.

Worked Example: Change in Library Study Time

Worked Example: A Workshop and Weekly Study Time

A library randomly selects 16 members from its register of 600 members to take part in a study-skills workshop. Each person reports weekly hours spent studying before and after the workshop. Define \(d=\text{after}-\text{before}\), in hours. The sample has \(\bar d=1.5\) hours and \(s_d=2.0\) hours. A plot of the differences is roughly symmetric with no outliers. Test whether the population mean difference is positive, using \(\alpha=0.01\).

State. Let \(\mu_d\) be the true mean difference, after minus before in weekly study hours, for all members represented by the library’s register who would take part in this workshop. The research question is whether this mean difference is positive, so the hypotheses are

$$ H_0:\mu_d=0 \qquad\text{and}\qquad H_a:\mu_d>0. $$

Plan. Use a paired t test because each person contributes a before-and-after pair. The library randomly selected the 16 members, supporting inference to the population represented by that register. The 10% condition is met because \(16\leq 0.10(600)=60\), so it is reasonable to treat different sampled members’ differences as independent. The plot of the differences is roughly symmetric and has no outliers, supporting use of a t procedure for this sample size.

Do. The standard error is

$$ SE=\frac{s_d}{\sqrt{n}} =\frac{2.0}{\sqrt{16}} =\frac{2.0}{4} =0.5\text{ hours}. $$

The test statistic and degrees of freedom are

$$ t=\frac{\bar d-0}{s_d/\sqrt{n}} =\frac{1.5-0}{0.5} =3.000, \qquad df=16-1=15. $$

For the upper-tailed alternative, the p-value is the area to the right of \(t=3.000\) with 15 degrees of freedom. A calculator gives \(p\approx0.004486\). Since \(0.004486<0.01\), reject \(H_0\).

Conclude. The data provide convincing evidence that the mean after-minus-before difference in weekly study time is positive for the population represented by the library’s register. This is evidence of an average increase in reported study time. Because this is a before-and-after study rather than a randomized comparison of workshop exposure, the evidence does not by itself establish that the workshop caused the increase.

Worked Example: No Observed Mean Difference

Worked Example: Two Versions of a Scheduling Tool

Eight randomly selected staff members each complete a scheduling task with the current tool and with a redesigned tool. Define \(d=\text{redesigned-tool time}-\text{current-tool time}\), in minutes. The differences have \(\bar d=0\) minutes and \(s_d=2.4\) minutes. The differences appear roughly symmetric with no outliers. Test whether the population mean difference is nonzero at \(\alpha=0.05\). Assume the staff members’ order of using the tools was randomized.

Solution. Let \(\mu_d\) be the true mean difference in task completion time, redesigned tool minus current tool, for the staff population represented by the random sample. Because the question asks whether the mean difference is nonzero, use \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\).

A paired t test is appropriate because each staff member tried both tools and contributes one difference. The sample was randomly selected. If the population represented by the sample contains at least 80 staff members, \(8\leq0.10(80)\), so the 10% condition is satisfied; the study states that the staff members were selected from a population at least that large. The differences are roughly symmetric without outliers, supporting the t procedure.

The standard error is

$$ SE=\frac{2.4}{\sqrt{8}}\approx0.8485\text{ minutes}. $$

Thus,

$$ t=\frac{0-0}{2.4/\sqrt{8}}=0, \qquad df=8-1=7. $$

For a two-sided test, the p-value is the probability of a t statistic at least as far from zero as \(0\), in either direction. That probability is \(p=1.0000\). Because \(1.0000>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence of a nonzero mean difference in task completion time for the population represented by the sample. This does not prove that the tools perform identically; it says this study did not provide convincing evidence of a mean difference. Randomly assigning tool order helps address order effects and supports a causal comparison, assuming the study was otherwise conducted appropriately.

Worked Example: Check the Sign Before the Conclusion

Worked Example: Comparing Two App Interfaces

Ten randomly selected users complete a task with two app interfaces, and the order is randomized. The response is completion time in seconds. Define \(d=\text{new interface time}-\text{old interface time}\). The sample differences have \(\bar d=-1.2\) seconds and \(s_d=1.897\) seconds; their plot is roughly symmetric with no outliers. Test whether the new interface has a lower mean completion time, at \(\alpha=0.05\). The users are sampled from a population of 500.

Solution. Let \(\mu_d\) be the true mean difference in task completion time, new interface minus old interface, for the population represented by the random sample. A lower mean time with the new interface corresponds to a negative difference, so use \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\).

Use a paired t test because each user tries both interfaces. The sample is random, the 10% condition is met because \(10\leq0.10(500)=50\), and the differences are roughly symmetric without outliers. These checks support the procedure.

The standard error is

$$ SE=\frac{1.897}{\sqrt{10}}\approx0.600. $$

Then

$$ t=\frac{-1.2-0}{0.600}=-2.000, \qquad df=10-1=9. $$

The lower-tail p-value for \(t=-2.000\) with 9 degrees of freedom is approximately \(0.0383\). Since \(0.0383<0.05\), reject \(H_0\). The data provide convincing evidence that the mean completion time with the new interface is lower than with the old interface for the population represented by the sample. Because interface order was randomized, a causal comparison is supported, provided order effects and other features of the design were adequately handled.

Common Mistakes and AP Exam Tips

  • Leaving the parameter vague: “The mean change” does not specify what changed, for whom, or in which direction. Define \(\mu_d\) with the measurement, units, population, and subtraction order.
  • Reversing the difference but keeping the hypotheses: If you define \(d\) as after minus before, a positive mean difference means an increase. If you switch to before minus after, the signs and alternative must switch too.
  • Checking conditions on the original measurements instead of the differences: A paired t test analyzes the list of \(d_i\) values. Use the design and distribution of those differences when checking conditions.
  • Using the number of measurements as \(n\): With 16 people measured twice, there are 16 pairs, not 32 independent observations. The degrees of freedom are \(16-1=15\).
  • Omitting the evidence for a condition: “The conditions are met” is not a check. State how the sample was selected or treatments assigned, show the 10% comparison when relevant, and describe the differences’ shape or sample size.
  • Reporting a calculator result without showing the calculation: Include \(SE=s_d/\sqrt n\), the substituted values for \(t\), and \(df=n-1\). Then identify the correct tail for the p-value.
  • Writing “accept the null” or “prove no difference”: The appropriate wording is “fail to reject \(H_0\)” and, in context, “the data do not provide convincing evidence” for the alternative claim.
  • Claiming causation or generalization automatically: As emphasized in “Paired t Conclusions About Causation and Generalization,” random sampling supports generalization to the sampled-from population, while random assignment supports a causal conclusion. Pairing alone provides neither.

A full-credit response is precise at both ends: its parameter and hypotheses establish exactly what is being tested, and its conclusion says what the evidence supports without going beyond the design. Keep the four steps together as one argument.

Key takeaway: For a complete paired t test, define the population mean difference in context, check conditions using the paired design and the differences, show the standard error, t statistic, degrees of freedom, and p-value, then give a conclusion that matches the alternative and study design.

Check Your Understanding

For each question, focus on making the paired t solution complete and consistent.

  1. A study records pulse rate before and after a breathing exercise. If \(d=\text{after}-\text{before}\), what does a negative value of \(\mu_d\) mean in context?
  2. A random sample of 20 people is selected from a population of 150. Does the 10% condition hold? Show the comparison.
  3. For \(n=12\) pairs, \(\bar d=1.4\), and \(s_d=2.1\), calculate the standard error, t statistic for \(H_0:\mu_d=0\), and degrees of freedom.
  4. If the alternative is \(H_a:\mu_d<0\) and the observed statistic is negative, which tail area is the p-value?
  5. A paired t test has \(p=0.12\) at \(\alpha=0.05\). Write the decision and a careful conclusion in context, without claiming that the null hypothesis has been proved.