From Paired Differences to a t Statistic
In “Hypotheses for a Paired t Test,” you used \(\mu_d\) to describe the true mean of the paired differences and learned to state hypotheses about it. This tutorial develops the calculation that compares the sample mean difference, \(\bar d\), with the null value. For the usual test of no average change, that null value is zero.
The paired t procedure works with one difference per pair, not with the two original measurement columns as if they were independent samples. If there are twelve pairs, there are twelve differences. From those differences, calculate their mean \(\bar d\) and sample standard deviation \(s_d\). The standard error measures the estimated variability of the sample mean difference across repeated samples.
The numerator is the distance between the observed mean difference and the null value, measured in the original units. The denominator is the standard error in those same units. Thus, \(t\) tells how many estimated standard errors the observed \(\bar d\) lies above or below zero. A positive statistic means \(\bar d\) is positive; a negative statistic means it is negative. The sign must be interpreted using the defined subtraction order.
For twelve pairs, \(df=12-1=11\). The degrees of freedom come from the twelve difference values: the sample standard deviation \(s_d\) is calculated around their sample mean, which imposes one constraint on the deviations. This is the same \(n-1\) rule explained in “Why the t Test Uses n Minus 1 Degrees of Freedom,” applied to the list of paired differences.
Calculation Sequence
Keep the arithmetic organized. First form the differences using the stated subtraction order. Then find \(\bar d\), \(s_d\), and \(s_d/\sqrt{n}\). Divide the difference between \(\bar d\) and the null value by that standard error. Finally, report degrees of freedom from the number of pairs. Do not use the standard deviation of the before measurements, the after measurements, or all 24 original measurements in place of \(s_d\).
Define each \(d_i\) using the chosen subtraction order. Count pairs to obtain \(n\).
Calculate their sample mean \(\bar d\) and sample standard deviation \(s_d\).
Find \(s_d/\sqrt{n}\), then divide \(\bar d-\mu_{d,0}\) by that value. For a no-average-change test, \(\mu_{d,0}=0\).
Use \(df=n-1\), where \(n\) is the number of pairs and therefore the number of differences.
A convenient check on the sample standard deviation is the computational formula. For differences \(d_1,\ldots,d_n\), the sum of squared deviations from their mean is \(\sum d_i^2-(\sum d_i)^2/n\). Divide this result by \(n-1\), then take the square root to obtain \(s_d\). This provides a way to verify a calculator summary against the stated data.
As in “Conditions for One-Sample Versus Paired Data,” the inference is about the distribution of differences. Before using a paired t procedure, check the study design for randomness or appropriate random assignment, independence of differences across pairs, and a suitable shape for the differences. For a sample of twelve pairs, inspect the differences for strong skewness or pronounced outliers; the small-sample shape check is discussed in “What to Check With a Small Sample of 12 Observations.” If sampling without replacement, also consider the 10% condition.
Worked Examples
Worked Example: Change in a Practice Score
In a fictional study, twelve students complete a short skill practice program. Each student’s score is recorded before and after practice. Define \(d_i=\text{after}_i-\text{before}_i\), in points. The twelve differences are \(2,3,1,4,0,2,5,1,3,2,4,1\). Calculate the paired t statistic for \(H_0:\mu_d=0\) and find the degrees of freedom.
State. Let \(\mu_d\) be the true mean after-minus-before score difference, in points, for the population of students represented by this study. The null hypothesis is that the population mean difference is zero. The sample contains twelve paired differences.
Plan. Use a paired t statistic because each student supplies both measurements and the analysis uses one difference per student. The study description does not establish that the students were randomly sampled or randomly assigned, so that design condition must be verified before generalizing or drawing a causal conclusion. Assume for this calculation that the twelve student differences are independent and that a graph of the differences shows no strong skewness or pronounced outliers. If the students were sampled without replacement, the population should contain at least 120 students for the 10% condition. These checks concern whether inference is appropriate; they do not change the statistic formula.
Do. The sum of the differences is 28, so
The sum of their squares is \(90\). Therefore, the sum of squared deviations is \(90-28^2/12=74/3\approx24.6667\). The sample standard deviation is
The standard error is \(s_d/\sqrt{12}\approx1.4975/\sqrt{12}\approx0.4323\) points. Thus,
Conclude. The sample mean after-minus-before difference is about 5.398 estimated standard errors above the null value of zero, with 11 degrees of freedom. This reports the paired t statistic and its degrees of freedom; by itself it is not a test decision. The direction is positive, meaning the sample average score was higher after practice. Any population claim still depends on the study design, conditions, and the subsequent p-value assessment.
Worked Example: Difference in Appointment Wait Times
A fictional clinic records the wait time before and after a scheduling change for twelve returning patients. Let each difference be after minus before, in minutes. The differences are \(-1,-2,-3,0,-1,-4,-2,-3,1,-2,-1,-2\). Find the standard error, t statistic for \(H_0:\mu_d=0\), and degrees of freedom.
Solution. There are twelve patients and hence twelve differences, so \(n=12\). Their sum is \(-20\), giving \(\bar d=-20/12=-1.6667\) minutes. The sum of their squared values is \(54\). The sum of squared deviations from the mean is \(54-(-20)^2/12=54-33.3333=20.6667=62/3\).
Divide by \(n-1=11\) to get the sample variance, then take the square root:
The standard error is \(\sqrt{62/33}/\sqrt{12}=\sqrt{31/198}\approx0.3957\) minutes. Therefore,
The negative sign is consistent with the subtraction order: the sample average after-minus-before difference is negative, so the observed average wait was shorter after the scheduling change. The statistic is about 4.212 standard errors below zero. The twelve pairs, not the 24 recorded wait times, determine \(df=11\). Before making an inference about a broader population, the design and the shape and independence conditions for the differences must also be supported.
Worked Example: Small Positive Average Difference
A fictional environmental team measures a quantity at twelve matched locations before and after a local maintenance project. With \(d_i=\text{after}_i-\text{before}_i\), the differences, in the measurement’s units, are \(0,1,2,-1,3,0,-2,1,2,-1,0,2\). Calculate the paired t statistic and degrees of freedom for a null mean difference of zero.
Solution. There are \(n=12\) matched locations. The differences sum to \(7\), so \(\bar d=7/12\approx0.5833\). Their squared values sum to \(29\), and the sum of squared deviations is \(29-7^2/12=29-49/12=299/12\approx24.9167\).
The sample standard deviation and standard error are
Then
The sample mean difference is positive and is about 1.343 standard errors above zero. This statistic’s sign indicates the direction of the sample mean difference under the stated subtraction order; it does not alone establish that the population mean difference is positive. The degrees of freedom remain 11, regardless of whether the calculated statistic is large or small.
Accuracy, Rounding, and Interpretation
Carry extra digits through intermediate calculations and round the final statistic only at the end. If you round \(s_d\) or the standard error too early, the displayed quotient may not exactly match the value calculated from the original differences. For example, a standard error near \(0.4345\) should be used with enough unrounded digits when computing the final statistic; show a sensible rounded result such as \(t=1.343\).
Units help catch mistakes. The mean difference, standard deviation of differences, and standard error all have the same units as an individual difference. In the quotient, those units cancel, so \(t\) has no units. Degrees of freedom are also unitless. A reported statistic with units, or an SE calculated from one of the original measurement columns instead of the differences, signals a problem.
The size and sign of \(t\) describe the sample result relative to the null value in standard-error units. The sign does not determine which tail to use; that is set by the alternative hypothesis, as explained in “Test Statistic Sign and Direction of the Alternative.” A t statistic alone also does not determine whether to reject the null. The p-value and the decision are separate parts of a significance test.
Common Mistakes and AP Exam Tips
- Counting measurements instead of pairs. Twelve pairs give \(n=12\), not \(n=24\). There are twelve differences, so \(df=11\).
- Using the wrong spread. Use \(s_d\), the sample standard deviation of the differences. The standard deviation of before values or after values is not the paired procedure’s denominator.
- Forgetting the square root of \(n\). The standard error is \(s_d/\sqrt{n}\), not \(s_d/n\) and not \(s_d\).
- Reversing the subtraction order or sign. State how \(d_i\) is defined, then interpret the sign of \(\bar d\) and \(t\) using that same definition.
- Using the wrong null value in the numerator. For a no-average-change paired test, subtract zero: \(t=(\bar d-0)/(s_d/\sqrt n)\). If a problem specifies a different null mean difference, subtract that reference value instead.
- Rounding too early. Keep calculator precision for \(\bar d\), \(s_d\), and the standard error, then round the final \(t\) value. State rounded values consistently.
- Treating the statistic as the conclusion. A t statistic describes the sample in standard-error units. A test decision requires the p-value and significance level; a population conclusion also requires appropriate design and conditions.
For a strong AP response, show the difference-based inputs, identify \(n\) as the number of pairs, give the standard error calculation, substitute into the t formula, and state \(df=n-1\). Interpret the sign in context, but do not claim that the statistic alone proves an effect or settles the test.
Check Your Understanding
Use the paired-difference framework to answer each question.
- A study has twelve matched pairs. What sample size belongs in the paired t formula, and what are the degrees of freedom?
- If \(\bar d=3.6\), \(s_d=2.4\), and \(n=12\), calculate the standard error and t statistic for \(H_0:\mu_d=0\). Round the statistic to three decimals.
- A student uses the standard deviation of the before measurements to calculate the standard error. Explain why this is not the correct paired t calculation.
- With differences defined as after minus before, what does a negative t statistic say about the sample mean difference?
- Why does reporting \(t\) and \(df\) not, by itself, give a reject-or-fail-to-reject decision?