Use the Differences List as the Data
In “Paired t Test Statistic and Degrees of Freedom,” you learned that a paired t test is calculated from one difference per pair. This tutorial applies that idea on a calculator. The calculator does not automatically recognize that two columns of measurements are paired: you must first create a list of differences, then run a one-sample T-Test on that list.
Use the same subtraction order that you used to define \(d_i\). For example, if \(d_i=\text{after}_i-\text{before}_i\), enter each after-minus-before value. A positive difference then means the after measurement was higher. The calculator’s test statistic and p-value depend on both this list and the alternative hypothesis you select.
TI-84 Data-Mode Steps
On a TI-84, use the following sequence. Menu names may vary slightly on other calculator models, but the statistical choices are the same.
Press STAT, choose EDIT, and enter one \(d_i\) value for each pair in a list such as L1. If you have before and after columns already entered, calculate the differences in a third list using the defined order, such as L2-L1 for after minus before.
Press STAT, move to TESTS, and select T-Test. Choose Data, because you have entered individual difference values. Enter the null mean difference, usually \(\mu_{d,0}=0\), and set List to L1 and Freq to 1.
Choose \(<\mu_0\), \(>\mu_0\), or \(\ne\mu_0\) to match the research question. For a test of no average change, \(\mu_0=0\). Do not select a tail based on the sign of the calculated statistic.
Choose Calculate. Check the reported \(t\), \(p\), and \(df\), as well as \(\bar{x}\), \(Sx\), and \(n\). Here, \(\bar{x}\) is the mean of the entered differences, \(Sx\) is their sample standard deviation \(s_d\), and \(n\) is the number of pairs.
The display label \(\bar{x}\) does not mean the calculator has switched to analyzing the original measurements. Because the list contains differences, the calculator’s \(\bar{x}\) is \(\bar d\), its \(Sx\) is \(s_d\), and its \(n\) counts the pairs. Its degrees of freedom should be \(n-1\), as in “Paired t Test Statistic and Degrees of Freedom.”
Some calculators also offer a Stats option in T-Test. That option uses a supplied sample mean, sample standard deviation, and sample size rather than individual data. It can be useful when a problem gives only summaries. When you have the paired observations or their differences available, Data mode is the direct way to run the test and verify the entered list.
Worked Examples
Worked Example: Testing for a Reduction in Wait Time
In a fictional clinic study, four patients are randomly sampled from 80 eligible patients. Each patient’s wait time is recorded before and after a scheduling change. Define \(d_i=\text{after}_i-\text{before}_i\), in minutes. The differences are \(-4,-3,-2,-1\). Test whether the population mean after-minus-before difference is less than zero at \(\alpha=0.05\). Assume the population distribution of differences is approximately Normal.
State. Let \(\mu_d\) be the true mean after-minus-before wait-time difference, in minutes, for eligible patients represented by this study. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\), with \(\alpha=0.05\).
Plan. Use a paired t test because the same patient provides both measurements, and the analysis uses one difference per patient. The patients were randomly sampled, which supports generalizing to the eligible-patient population. The four patients are less than 10% of the 80 eligible patients, so the 10% condition is met. Different patients contribute separate differences, so independence between pairs is reasonable. Since the sample is small, the shape condition requires the differences to come from a population that is approximately Normal; this is stipulated in the study description. The observed list has no pronounced outlier, though four observations alone provide limited information about shape.
Do. Enter \(-4,-3,-2,-1\) in L1. In T-Test, choose Data, enter \(\mu_0=0\), set List to L1 and Freq to 1, and select \(<\mu_0\). The calculator reports \(\bar{x}=-2.5\), \(Sx\approx1.291\), \(n=4\), \(t\approx-3.873\), \(df=3\), and \(p\approx0.01523\) (rounded).
These output values can be checked from the difference list. Its mean is \((-4-3-2-1)/4=-2.5\). The sample variance is \(5/3\), so \(s_d=\sqrt{5/3}\approx1.291\). The standard error is \(s_d/\sqrt{4}\), giving
The lower-tail p-value is \(P(T\leq-3.873)\) for a t distribution with 3 degrees of freedom, approximately \(0.01523\). Because \(0.01523<0.05\), reject \(H_0\).
Conclude. The study provides convincing evidence that the true mean after-minus-before wait-time difference for eligible patients is less than zero. In context, the mean wait time after the scheduling change is lower. Because this study used a random sample rather than random assignment to the scheduling change, this conclusion does not by itself establish that the change caused the reduction.
Worked Example: Testing for a Positive Change in Practice Time
For a fictional sports-training exercise, five randomly selected athletes each complete a timed drill before and after a new practice routine. Define \(d_i=\text{after}_i-\text{before}_i\), in seconds. The differences are \(-1,0,1,2,3\). The athletes are sampled from a group of 100, their difference values are independent across athletes, and the population of differences is assumed approximately Normal. Test whether the true mean difference is greater than zero at \(\alpha=0.05\).
State. Let \(\mu_d\) be the true mean after-minus-before drill-time difference, in seconds, for athletes represented by the sample. Use \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\), with \(\alpha=0.05\).
Plan. This is paired data because each athlete is measured twice, so use the list of five differences. Random sampling supports generalizing to the represented athlete group. Five is less than 10% of 100, so the 10% condition is met. Separate athletes provide independent differences as described. With only five observations, the shape condition cannot be justified by a large sample; the problem states that the population differences are approximately Normal. The listed values show no pronounced outlier.
Do. Enter \(-1,0,1,2,3\) in L1. Choose T-Test in Data mode, set \(\mu_0=0\), List to L1, Freq to 1, and select \(>\mu_0\). The calculator reports \(\bar{x}=1\), \(Sx\approx1.581\), \(n=5\), \(t\approx1.414\), \(df=4\), and \(p\approx0.1151\) (rounded). To verify the statistic, the sum is 5, so \(\bar d=5/5=1\). The sum of squares is 15; the sum of squared deviations is \(15-5^2/5=10\). Thus \(s_d=\sqrt{10/4}=\sqrt{2.5}\approx1.581\), and
The upper-tail p-value is approximately \(0.1151\). Since \(0.1151>0.05\), fail to reject \(H_0\).
Conclude. The sample does not provide convincing evidence that the true mean after-minus-before drill-time difference is greater than zero. The sample mean difference is positive, but its direction alone is not enough to establish a positive population mean.
Worked Example: Create the Difference List From Two Columns
A fictional environmental team randomly selects six matched sites from 100 eligible sites and records a measurement before and after maintenance. Define \(d_i=\text{after}_i-\text{before}_i\). The before measurements are \(20,22,24,25,27,28\), and the after measurements are \(26,21,24,28,29,30\), in the same units. Test for any nonzero mean difference at \(\alpha=0.05\). Assume the population differences are approximately Normal.
State. Let \(\mu_d\) be the true mean after-minus-before measurement difference, in the measurement’s units, for eligible matched sites. Use \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\), with \(\alpha=0.05\).
Plan. The same sites are measured before and after, so this is a paired design. The sites were randomly selected, and six is less than 10% of 100, supporting generalization and the 10% condition. The six different sites contribute separate differences, so independence between pairs is reasonable. The sample is small, so the shape condition is not supplied by a large-sample rule; the problem stipulates an approximately Normal population of differences. Check the calculated differences for unusual values as well.
Do. Enter the before values in L1 and the after values in L2. On the calculator, create L3 by moving to the L3 heading and entering L2-L1, then press ENTER. The resulting differences are \(6,-1,0,3,2,2\). Run T-Test in Data mode with List set to L3, Freq set to 1, \(\mu_0=0\), and the alternative \(\ne\mu_0\).
The calculator reports \(\bar{x}=2\), \(Sx\approx2.449\), \(n=6\), \(t=2\), \(df=5\), and \(p\approx0.1019\) (rounded). Verify the summaries: the differences sum to 12, so \(\bar d=12/6=2\); their squared values sum to 54, so the sum of squared deviations is \(54-12^2/6=30\). Therefore \(s_d=\sqrt{30/5}=\sqrt{6}\approx2.449\), and the standard error is \(\sqrt{6}/\sqrt{6}=1\). Hence \(t=(2-0)/1=2\), with \(df=6-1=5\).
Conclude. Since the two-sided p-value \(0.1019\) is greater than \(0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence that the true mean after-minus-before measurement difference for eligible sites is nonzero.
Common Mistakes and AP Exam Tips
- Putting both measurement columns into T-Test. The paired test uses the difference list. Subtract within each pair first, keeping one difference per pair.
- Reversing the subtraction order. If differences are defined as after minus before, calculate L2-L1 when L1 holds before and L2 holds after. Reversing the order reverses the signs and changes which one-sided alternative is appropriate.
- Choosing the tail after seeing the statistic. Select \(<\), \(>\), or \(\ne\) to match the research question and hypotheses. The observed sign of \(t\) does not determine the alternative.
- Misreading \(\bar{x}\) and \(Sx\). With a difference list entered, these are \(\bar d\) and \(s_d\), not summaries of the original before or after measurements.
- Counting the measurements instead of the pairs. Six pairs give \(n=6\) differences and \(df=5\), not \(n=12\).
- Reporting only the calculator screen. A full AP response defines \(\mu_d\) in context, states both hypotheses and \(\alpha\), identifies the paired t procedure, checks the conditions, reports the test statistic, degrees of freedom and p-value, and concludes in context.
- Claiming causation from a paired pattern alone. Pairing determines how to analyze the data; it does not establish random assignment. Distinguish evidence about a population mean from evidence that a treatment caused a change.
Before accepting the output, compare the calculator’s \(n\) and \(\bar{x}\) with the list you intended to test, confirm that the alternative is correct, and check that \(df=n-1\). Keep the subtraction order and units visible in your written work. The calculator performs the arithmetic, but you are responsible for deciding whether the data and study design justify the procedure and what its result means.
Check Your Understanding
Use the paired-difference calculator method to answer each question.
- If L1 contains before measurements and L2 contains after measurements, what list expression produces after-minus-before differences?
- A study has 15 pairs. What should the calculator report for \(n\) and \(df\) when the difference list is entered correctly?
- In T-Test Data mode, what do \(\bar{x}\) and \(Sx\) represent when List is the paired-difference list?
- A student chooses \(>\mu_0\) because the calculator’s \(t\) value is positive. Explain why the alternative should instead be chosen from the research question.
- For a lower-tailed paired t test, the calculator reports \(p=0.032\) at \(\alpha=0.05\). State the decision and give the core conclusion wording in context.