Tutorials › AP Statistics › Common Mistakes With Paired t Procedures

Paired data and paired t procedures · Tutorial 698 of 1000

Common Mistakes With Paired t Procedures

Practice checking the paired design, calculating differences consistently, and carrying out a paired t test with clear, context-specific reasoning.

Intermediate 9 min read

What You'll Learn

  • Recognize when a paired t procedure is appropriate and why a two-sample test is not.
  • Define one difference variable and use the same subtraction order for every pair.
  • Use the number of pairs—not the number of measurements—for the standard error and degrees of freedom.
  • Match the paired t test hypotheses and p-value tail to the defined difference.
  • Identify common calculator and interpretation errors in paired t procedures.

Why Paired t Procedures Are Easy to Misapply

A paired t procedure is a one-sample t procedure applied to a list of within-pair differences. That sounds simple, but a few small choices can change the analysis: treating linked measurements as independent, subtracting in inconsistent directions, or using the number of measurements instead of the number of pairs.

In “Paired Data From Two Measurements on the Same Unit,” you learned to identify the link between two measurements and keep each unit’s values together. Here, the focus is on errors that can occur after recognizing that link. The central habit is to define the difference once, calculate one difference per pair, and use that single difference list throughout the analysis.

Key idea: For paired data, do not run a two-sample t procedure on the original measurement columns. Define one difference for each pair, then use a one-sample t procedure on those differences. If there are \(n\) pairs, the difference list has \(n\) values, the standard error is \(s_d/\sqrt{n}\), and the degrees of freedom are \(n-1\).

Mistake 1: Treating Paired Measurements as Independent Samples

A two-sample t procedure is designed to compare the means of two independent groups. A paired design is different: each value in one column is linked to a particular value in the other column. The analysis should retain that within-pair connection by calculating a difference for each unit.

For example, if the same sensor modules are measured before and after recalibration, each module provides a pair. Running a two-sample test on the before and after columns treats them as though they came from unrelated groups. That discards the pairings and uses a standard error based on the variability in the two separate columns, rather than the variability among the actual within-module differences. The paired method instead asks whether the population mean difference is zero, or is above or below zero as the research question specifies.

Worked Example: Two Readings on Each Sensor Module

In a fictional lab demonstration, six sensor modules are measured before and after recalibration. Let \(d=\text{after}-\text{before}\), in reading units.

ModuleBeforeAfter\(d=\text{after}-\text{before}\)
11001022
21201233
390911
41401444
51101111
61301333

Solution. The same six modules were measured twice, so these are six pairs, not two independent samples. The differences are \(2,3,1,4,1,3\), and their sum is \(14\). Thus \(\bar d=14/6\approx2.3333\) reading units.

The squared deviations from \(\bar d\) sum to approximately \(7.3333\), so the sample variance of the differences is \(7.3333/(6-1)\approx1.4667\), and \(s_d\approx1.2111\) reading units. For a paired t analysis, the standard error is \(s_d/\sqrt{6}\approx1.2111/\sqrt{6}\approx0.4944\) reading units, with \(df=6-1=5\).

The before-column mean is \(690/6=115\), and the after-column mean is \(704/6\approx117.3333\). Their difference, about \(2.3333\), matches \(\bar d\), as it should when the same pairs and the same subtraction order are used. But that match does not make a two-sample t test appropriate: the paired analysis uses the spread of the six differences, not the separate spreads of the before and after readings.

Mistake 2: Changing the Subtraction Order

A difference has meaning only in relation to its subtraction order. As in “Defining the Difference Variable,” state that order explicitly, such as \(d=\text{after}-\text{before}\). Then use it for every pair. If the order changes for one or more pairs, the resulting list is not a consistent set of differences for the parameter \(\mu_d\).

Worked Example: A Difference Sign That Gets Reversed

A fictional comparison records brightness under two display settings for six screens. Let \(d=\text{Setting Y}-\text{Setting X}\), in brightness units.

ScreenSetting XSetting YCorrect \(d=\text{Y}-\text{X}\)
110097-3
2105104-1
398980
41101122
51021053
61081091

Solution. The correct differences are \(-3,-1,0,2,3,1\). They sum to \(2\), so \(\bar d=2/6\approx0.3333\) brightness units. The squared deviations from this mean sum to approximately \(23.3333\). Therefore \(s_d=\sqrt{23.3333/5}\approx2.1602\) brightness units.

Suppose a student accidentally calculates Setting X minus Setting Y for screen 1, getting \(+3\), but keeps the defined order \(d=\text{Y}-\text{X}\) for all the other screens. The mistaken list is \(3,-1,0,2,3,1\). Its sum is \(8\), so its mean is \(8/6\approx1.3333\), not \(0.3333\). The mistake changes the center and spread of the difference list, and therefore can change the test statistic and conclusion.

Reversing the subtraction order for every screen would be consistent: the differences would all change sign, and so would \(\bar d\) and the t statistic. But the hypotheses and the interpretation would also need to reflect the new order. Reversing just one or a few differences is not a valid alternative convention; it is a calculation error.

Mistake 3: Using the Number of Measurements Instead of Pairs

Each pair produces one difference. If 12 people each contribute two measurements, there are 12 differences—not 24. The sample size for paired inference is the number of pairs, \(n\). That \(n\) determines both the standard error \(s_d/\sqrt{n}\) and the degrees of freedom \(n-1\). Using the total number of recorded measurements in either place gives the wrong calculations.

Another related error is calculating a standard error from the two original columns. For a paired t procedure, calculate \(s_d\) from the difference list. Do not substitute the standard deviation of either original column for \(s_d\). “Paired t Test Statistic and Degrees of Freedom” gives the paired statistic for testing \(H_0:\mu_d=0\):

$$ t=\frac{\bar d-0}{s_d/\sqrt{n}}, \qquad df=n-1. $$

If the null value is not zero, use \(t=(\bar d-\mu_{d,0})/(s_d/\sqrt{n})\), where \(\mu_{d,0}\) is the null value. In the common no-average-difference test, that value is zero.

Worked Example: A Complete Paired t Test

Worked Example: Comparing Two Soil-Moisture Probes

In an invented study, two probes measure soil moisture in each of 10 randomly selected garden plots. Let \(d=\text{Probe B}-\text{Probe A}\), in percentage points. The differences are \(2,3,1,4,2,3,0,5,2,3\). The question is whether Probe B reads higher on average.

1
State.
Let \(\mu_d\) be the true mean difference, Probe B minus Probe A in percentage points, for all garden plots in the population represented by the random sample. Test \(H_0:\mu_d=0\) against \(H_a:\mu_d>0\), using \(\alpha=0.05\).
2
Plan.
Use a paired t test, which is a one-sample t test on the differences. The design condition is met because both probes measure each plot. The 10 plots are a random sample. If the population contains 200 plots, then \(10\leq0.10(200)=20\), so the 10% condition is met. Different plots can reasonably be treated as independent. The differences are quantitative; a dotplot of the listed values is roughly symmetric, with no pronounced outlier, supporting use of a t procedure for this small sample.
3
Do.
The differences sum to \(25\), so \(\bar d=25/10=2.5\) percentage points. The sum of squared deviations from \(2.5\) is \(18.5\), giving \(s_d=\sqrt{18.5/9}\approx1.4337\) percentage points. Thus \(SE=\sqrt{18.5/90}\approx0.4534\) percentage points. The test statistic is \(t=(2.5-0)/\sqrt{18.5/90}\approx5.514\), with \(df=10-1=9\). For the upper-tailed alternative, the p-value is \(P(T\geq5.514)\approx0.0002\), rounded to four decimal places.
4
Conclude.
Because the p-value is less than \(\alpha=0.05\), reject \(H_0\). The data provide convincing evidence that the true mean Probe B minus Probe A reading is greater than zero for the population of garden plots represented by this sample.

Mistake 4: Choosing the Wrong Hypothesis or P-Value Tail

The alternative hypothesis must match both the research question and the defined difference. If \(d=\text{after}-\text{before}\), then \(H_a:\mu_d>0\) represents a higher mean after measurement, while \(H_a:\mu_d<0\) represents a lower mean after measurement. If the question asks whether the mean difference is nonzero in either direction, use \(H_a:\mu_d\ne0\).

The direction of \(H_a\) determines the tail used for the p-value, as in “Finding a P-Value Using tcdf.” A common mistake is to see a positive t statistic and automatically use an upper-tail probability. The alternative, not the sign alone, determines the p-value. A positive statistic in a lower-tailed test is not evidence in the direction claimed by that alternative.

If you consistently reverse the difference definition, you must also reverse the direction of a one-sided alternative. For instance, changing from \(d=\text{B}-\text{A}\) to \(d=\text{A}-\text{B}\) changes “B is higher on average” from \(H_a:\mu_d>0\) to \(H_a:\mu_d<0\). Keeping the original alternative after reversing the order would test the opposite claim.

Common Mistakes and AP Exam Tips

  • Running a two-sample t test on matched measurements: Identify the design link and analyze one difference per pair. Do not treat the original columns as independent groups.
  • Changing the difference order partway through: Write the definition of \(d\) before calculating. Check each row against it, including the sign.
  • Using the number of measurements as \(n\): Count pairs. Twelve units measured twice give \(n=12\) differences, not \(n=24\).
  • Using the wrong standard deviation: The standard error uses \(s_d\), the sample standard deviation of the differences, not the standard deviation of a measurement column.
  • Using the wrong degrees of freedom: For a paired t procedure with \(n\) pairs, use \(df=n-1\), because the procedure is applied to the \(n\) differences.
  • Using the wrong tail or conclusion: Choose the tail from \(H_a\), then compare the p-value with \(\alpha\). Say “reject” or “fail to reject” \(H_0\); do not claim that a large p-value proves the null hypothesis.
  • Giving a conclusion without context: Name the population mean difference and explain what its sign means using the stated subtraction order and units.

On a calculator, enter the \(n\) differences as one list and run a one-sample T-Test in Data mode, as described in “Running a Paired t Test on the Calculator.” Confirm that the displayed \(n\) is the number of pairs, the degrees of freedom are \(n-1\), the null value is correct, and the selected alternative matches the question. Choosing a two-sample test on the original columns is not a substitute for this setup.

Key takeaway: A paired t procedure is built from one consistently defined difference per pair. Use the differences’ mean and standard deviation, count pairs for \(n\), use \(df=n-1\), and let the defined subtraction order guide the hypotheses, p-value tail, and interpretation.

Check Your Understanding

For each question, focus on the design, the difference list, and how the calculation or interpretation follows from them.

  1. Six people each have a measurement recorded before and after an activity. Why is a two-sample t test on the two columns inappropriate?
  2. For 15 matched pairs, what sample size and degrees of freedom should a paired t procedure use?
  3. If \(d=\text{new}-\text{old}\), what does a negative \(d\) mean for an individual pair? What direction of alternative represents a lower mean new measurement?
  4. Why must \(s_d\), rather than the standard deviation of an original measurement column, be used in the paired t standard error?
  5. A student gets a positive t statistic for a lower-tailed alternative. Does the positive sign change which tail defines the p-value? Explain.