Why Compare the Spread of Differences?
In “Paired Design With Random Order of Treatments,” you saw how a crossover experiment can give each subject both treatments. Once the data are paired, the comparison is made within each subject: define a difference for every pair and study the list of differences. A key benefit can be that these differences vary less than the individual measurements do.
Why might that happen? Some subjects may tend to have high responses under either treatment, while others tend to have low responses under either treatment. Those subject-to-subject differences appear in both measurements. Subtracting one measurement from the other can cancel much of this shared variation, leaving a difference that is more consistent across subjects.
This benefit depends on how the paired measurements relate to each other. Pairing does not automatically make differences less variable. If a subject’s measurement under one condition gives little information about the measurement under the other condition—or if the two tend to move in opposite directions—the differences may remain quite variable or become more variable.
What “Less Variable” Means
The standard deviation describes the spread of observations around their mean. For paired data, \(s_d\) is the sample standard deviation of the \(n\) within-pair differences. It describes how much the individual differences tend to vary. It is not the standard deviation of either original measurement column.
When the goal is to estimate the population mean difference \(\mu_d\), the standard error \(s_d/\sqrt{n}\) describes the estimated variability of the sample mean difference \(\bar d\). This is the measure that tells us how the spread of the differences affects the precision of the mean comparison.
For comparison, if two groups are independent, the estimated standard error of the difference between their sample means is \(\sqrt{s_A^2/n_A+s_B^2/n_B}\). In a paired study, the relevant standard error is instead \(s_d/\sqrt{n}\), using one difference for each pair. As in “Matched Pairs Versus Independent Samples,” which calculation is appropriate depends on the study design: paired observations must be analyzed as pairs.
Here, \(s_A\) and \(s_B\) are the sample standard deviations in the two measurement columns, \(s_{AB}\) is their sample covariance, and \(r\) is their sample correlation. The equation helps explain the effect of the relationship between paired measurements. When \(r\) is positive, the last term is subtracted, which tends to make \(s_d^2\) smaller. When \(r\) is near zero, there is little reduction from that term. When \(r\) is negative, subtracting a negative quantity increases the variance of the differences.
You do not need to calculate covariance or correlation every time you analyze paired data. The practical lesson is to look at what pairing does: it compares like with like and can remove some of the variation associated with differences among the units. The worked examples show how to see that effect numerically.
Worked Examples: Pairing Can Reduce Variability
Worked Example: Comparing Two Task Interfaces
A fictional researcher records the time, in seconds, that six students need to complete a task using two interfaces. Each student tries both interfaces. Let \(A\) be the time with interface A and \(B\) the time with interface B, and define \(d=B-A\). The invented data are:
| Student | Time with A | Time with B | \(d=B-A\) |
|---|---|---|---|
| 1 | 70 | 74 | 4 |
| 2 | 80 | 83 | 3 |
| 3 | 90 | 94 | 4 |
| 4 | 100 | 102 | 2 |
| 5 | 110 | 115 | 5 |
| 6 | 120 | 123 | 3 |
Solution. The mean time with A is \(95\) seconds, since the six times total \(570\) seconds and \(570/6=95\). The mean time with B is \(98.5\) seconds, since those times total \(591\) seconds and \(591/6=98.5\). The differences total \(21\) seconds, so \(\bar d=21/6=3.5\) seconds. This also checks against the difference of the two means: \(98.5-95=3.5\) seconds.
The A times have squared deviations totaling \(1750\), so \(s_A^2=1750/(6-1)=350\) seconds squared and \(s_A=\sqrt{350}\approx18.708\) seconds. The B times have squared deviations totaling \(1745.5\), so \(s_B^2=1745.5/5=349.1\) seconds squared and \(s_B=\sqrt{349.1}\approx18.684\) seconds. The two columns each have a spread of about 19 seconds.
The differences are \(4,3,4,2,5,3\), with mean \(3.5\). Their squared deviations total \(5.5\), so \(s_d^2=5.5/5=1.1\) seconds squared and \(s_d=\sqrt{1.1}\approx1.049\) seconds. Thus, the individual task times vary widely across students, but the paired differences vary by only about one second. Students who take longer with A also tend to take longer with B, so much of their overall speed difference cancels in \(B-A\).
For a direct comparison of precision, the paired standard error is \(s_d/\sqrt{6}=\sqrt{1.1/6}\approx0.428\) seconds. If these had instead been two independent groups of six, the independent-groups benchmark would be \(\sqrt{350/6+349.1/6}=\sqrt{116.5167}\approx10.794\) seconds. This comparison illustrates the potential precision benefit of pairing; it does not mean that independent-group analysis is appropriate for these students, who each tried both interfaces.
The difference between the two means is \(3.5\) seconds, and the average paired difference is also \(3.5\) seconds. The major contrast here is not the estimated average difference, but the estimated variability around that difference: pairing makes the within-student comparison far less variable in this example.
When Pairing Does Not Reduce Variability
Worked Example: A Strong Negative Relationship
Consider an invented illustration with five matched units measured under conditions A and B. The A measurements are \(10,20,30,40,50\), and the B measurements are \(50,40,30,20,10\), in the same units. Define \(d=B-A\). This deliberately extreme pattern shows what can happen when paired measurements move in opposite directions.
Solution. Both columns have mean \(30\). In each column, the squared deviations from \(30\) total \(1000\), so \(s_A^2=s_B^2=1000/(5-1)=250\), and each standard deviation is \(\sqrt{250}\approx15.811\).
The differences are \(40,20,0,-20,-40\), with mean \(0\). Their squared deviations from \(0\) total \(4000\), so \(s_d^2=4000/4=1000\) and \(s_d=\sqrt{1000}\approx31.623\). The difference standard deviation is larger than either condition’s standard deviation.
The covariance equation verifies this result. The columns have a perfect negative sample association, so \(r=-1\). Therefore, \(s_d^2=250+250-2(-1)(15.811)(15.811)=500+500=1000\), matching the direct calculation. The negative relationship makes differences more spread out rather than less.
With five observations in each condition, the independent-groups standard error benchmark would be \(\sqrt{250/5+250/5}=\sqrt{100}=10\). The paired standard error is \(\sqrt{1000/5}=\sqrt{200}\approx14.142\). Pairing has not improved precision in this example. In an actual study with paired units, the paired calculation still reflects the design; the independent-groups number is a comparison, not a reason to ignore the pairs.
Using the Relationship to Predict the Effect
Worked Example: Estimating Variability From Summary Statistics
Suppose a fictional experiment has \(16\) matched participants. The sample standard deviations for the responses under conditions A and B are \(s_A=12\) units and \(s_B=10\) units, and the sample correlation between the paired responses is \(r=0.75\). Compare the standard error for a paired mean difference with the independent-groups benchmark for two groups of \(16\).
Solution. Using the sample correlation in the variance relationship gives \(s_d^2=12^2+10^2-2(0.75)(12)(10)=144+100-180=64\). Thus, \(s_d=8\) units. The paired standard error is \(s_d/\sqrt{16}=8/4=2\) units.
For two independent groups of \(16\), the benchmark is \(\sqrt{12^2/16+10^2/16}=\sqrt{9+6.25}=\sqrt{15.25}\approx3.905\) units. The paired standard error is smaller because the paired responses have a positive association. As a check on the paired variance calculation, the covariance is \(r s_A s_B=(0.75)(12)(10)=90\), and \(144+100-2(90)=64\), again giving \(s_d=8\).
The comparison assumes the same number of observations per condition and uses the stated sample summaries. It illustrates the role of within-pair association; it does not show that a particular experiment will have this correlation or this precision. If the correlation were near zero, the subtraction would not gain much from shared variation. If it were negative, the difference spread could increase.
How to Explain the Benefit Precisely
When describing why pairing may help, connect the design to the observed pattern. A strong explanation says that each unit contributes both measurements, the measurements tend to be positively associated, and the shared unit-to-unit variation tends to cancel when calculating within-pair differences. If the differences are less spread out, \(s_d\) is smaller; for the same number of pairs, that also makes \(s_d/\sqrt{n}\) smaller.
Do not say that pairing changes the actual responses or guarantees a smaller standard error. Pairing changes how the comparison is formed. Its benefit comes from the relationship between the paired values, and it can be absent when that relationship is weak or negative. Also keep the sample size straight: \(n\) in a paired calculation is the number of pairs, not the total number of measurements.
- Comparing unlike quantities: The standard deviations of the A and B measurements describe their separate spreads; \(s_d\) describes the spread of the differences. Use the paired standard error and independent-groups standard error to compare the estimated precision of a mean contrast.
- Claiming pairing always reduces variability: It often helps with positively associated measurements, but the association matters. The negative-association example shows that pairing can instead produce more variable differences.
- Ignoring the study design: When the same subjects or matched units provide both measurements, preserve the pairs. The independent-groups calculation is a useful benchmark for understanding variability, not a substitute for the analysis that respects the design.
- Confusing spread with average change: A smaller \(s_d\) says the individual differences are more consistent; it does not mean the average difference \(\bar d\) is smaller. The center and spread answer different questions.
For a clear AP response, state what each difference measures, identify \(s_d\) as the standard deviation of those differences, and explain how the relationship within pairs affects its spread. If comparing precision, name the paired and independent-groups standard errors and make clear which one matches the actual design.
Check Your Understanding
Use the relationship between paired measurements and their differences to answer each question.
- Why might the differences between two measurements on the same subjects be less variable than either measurement column?
- If paired measurements have a strong positive association, what does the covariance term suggest about the variance of the differences?
- Suppose \(s_A=8\), \(s_B=6\), and \(r=0\). Use the variance relationship to find \(s_d^2\).
- In a paired study with 20 subjects, what sample size belongs in the paired standard error \(s_d/\sqrt{n}\)?
- Does pairing guarantee a smaller standard error than treating groups as independent? Explain why or why not.