Why Design Can Change the Analysis
In “A Decision Flowchart for Mean Inference,” you learned to choose a procedure by identifying the response, the design-based parameter, and the goal of the question. This tutorial zooms in on one important comparison: the same two columns of numbers can call for different procedures if the study design links the observations in one situation but not in another.
A paired design might measure the same people under two conditions, or deliberately match one unit in one group to a unit in another. An independent-groups design has no such one-to-one links. The values in a table cannot tell you which design occurred by themselves. The research description does that.
This difference in procedure is not a matter of preference. Genuine pairing makes the within-pair differences the relevant data. If the groups are unlinked, making pairs after seeing the values does not create a paired design. As emphasized in “Two-Sample Versus Paired Test on the Same Numbers,” design-based links—not equal sample sizes or a two-column layout—determine the procedure.
Compare the Targets and the Data Used
Suppose two conditions produce quantitative measurements. With genuine pairs, first define a difference for each pair, such as condition A minus condition B. The paired t procedure then treats those differences as one sample and makes inference about \(\mu_d\). The sample size is the number of pairs, and the condition about shape applies to the differences.
With independent groups, the target is the difference between the population means, \(\mu_1-\mu_2\). The unpooled two-sample t procedure uses the two samples separately. Its standard error is based on each group’s sample standard deviation and sample size. For the formula and conditions of that procedure, see “Conditions for a Two-Sample t Test” and “How Sample Size and Spread Affect the Test Result.”
Both procedures compare an observed difference with a null value of zero in the test examples below, but the quantities in the denominators are not interchangeable. In the paired formula, \(s_d\) describes the sample-to-sample variation among the pairwise differences. In the independent formula, \(s_1\) and \(s_2\) describe variation within the separate groups.
Pairing can make differences less variable when linked measurements tend to move together, or more variable when the chosen pairings do not do that. The important point is not that pairing always makes a test stronger. It is that the pairwise differences—and their spread—belong to the paired design. The independent procedure does not use the row-by-row differences.
Worked Examples: One Set of Values, Two Designs
Worked Example: Same Scores From Matched or Independent Students
Imagine two practice formats, A and B, with these scores. In one version of the study, the same eight students take both formats, so each row represents one student. In another version, the A and B scores come from separate random samples of eight students each; the side-by-side layout is only for display, and rows do not link students.
| Row | Format A | Format B |
|---|---|---|
| 1 | 70 | 79 |
| 2 | 72 | 81 |
| 3 | 74 | 75 |
| 4 | 76 | 81 |
| 5 | 78 | 71 |
| 6 | 80 | 75 |
| 7 | 82 | 71 |
| 8 | 84 | 71 |
State: In both versions, consider a two-sided test at \(\alpha=0.05\). For the paired version, define each difference as A minus B, and let \(\mu_d\) be the true mean difference in scores for the population represented by the study. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\). For the independent version, let \(\mu_1\) and \(\mu_2\) be the population mean scores for formats A and B, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Plan and conditions: For the paired version, assume the eight students were randomly sampled from a population of at least 80 students and took both formats. The design provides genuine pairs; the differences from distinct students are independent, and \(8/80=0.10\) meets the 10% condition. Suppose a plot of the differences shows no strong skewness or outliers. The paired t test is appropriate.
For the independent version, assume the two separate random samples each contain eight students from populations of at least 80 students. The samples are independent by design, and each sample satisfies the 10% condition. Suppose plots for both groups show no strong skewness or outliers. The unpooled two-sample t test is appropriate. A paired procedure would not be justified for these unlinked students.
Do—paired version: The differences are \(-9,-9,-1,-5,7,5,11,13\). Their sum is 12, so \(\bar{x}_d=12/8=1.50\) points. The sum of squared deviations from 1.50 is 534, giving \(s_d=\sqrt{534/7}\approx8.734\) points and a standard error of \(8.734/\sqrt{8}\approx3.088\) points. With \(7\) degrees of freedom, \(t\approx1.50/3.088=0.486\), and the two-sided p-value is approximately \(0.642\), rounded.
Do—independent version: The means are \(\bar{x}_1=616/8=77.0\) and \(\bar{x}_2=604/8=75.5\), so the observed difference is \(1.50\) points. For group A, the sum of squared deviations from 77 is 168, so \(s_1=\sqrt{168/7}=\sqrt{24}\approx4.899\). For group B, the sum of squared deviations from 75.5 is 134, so \(s_2=\sqrt{134/7}\approx4.375\). The estimated standard error is \(\sqrt{24/8+(134/7)/8}\approx2.322\) points. The Welch degrees of freedom are approximately \(13.8\); the test statistic is \(1.50/2.322\approx0.646\), with two-sided p-value approximately \(0.529\), rounded.
Conclude: Both p-values exceed \(0.05\), so in each version we fail to reject the null hypothesis. The paired analysis does not provide convincing evidence of a nonzero mean A-minus-B score difference for the population represented by the paired study. The independent analysis does not provide convincing evidence of a difference in the population mean scores for the two independent groups. The conclusions refer to different design-based parameters, even though the two columns of values are identical.
Pairing Is About the Study, Not the Spreadsheet
In the paired version of the score example, row 1 means that one student earned both the A score of 70 and the B score of 79. That link makes the row difference \(-9\) meaningful. In the independent version, row 1 simply places one A score next to one B score for convenience. Subtracting those entries would not create a meaningful within-student difference.
This also explains why rearranging a column is not harmless when the data are paired. Reordering rows for display is fine only if the links are preserved. If the values are reassigned across pairs, the calculated differences and their variability can change. That would describe a different matching, not the observed study.
Worked Example: What Reversing the Matches Changes
Worked Example: Keep the Values, Change the Pairing
Return to the same A and B score columns. In the paired-study version, the actual row links give differences \(-9,-9,-1,-5,7,5,11,13\). Now imagine reversing the order of the B scores and pairing each A score with that reversed list. This is a different matching, not a permissible correction to the original rows when the original students were genuinely linked.
The reversed B list is \(71,71,75,71,81,75,81,79\). Subtracting it from the A values in order gives the differences \(-1,1,-1,5,-3,5,1,5\). They sum to 12, so their mean is \(12/8=1.50\). Their squared deviations from 1.50 sum to 70, so \(s_d=\sqrt{70/7}=\sqrt{10}\approx3.1623\). The standard error is \(\sqrt{10}/\sqrt{8}=\sqrt{10/8}\approx1.1180\), and the paired t statistic for a zero mean difference is \(1.50/1.1180\approx1.342\), with \(7\) degrees of freedom.
This reversed matching would give a two-sided p-value of about \(0.2216\), rounded. It differs from the paired result using the actual row links because the sample of differences has a different spread. The values in each column and their means have not changed. If the design had genuinely matched students in this reversed way, these would be the differences to analyze. But when the original same-student links are real, reversing them discards the study design and produces an invalid analysis.
Worked Example: A Paired Interval Uses the Differences
Worked Example: Estimate a Mean Change in Running Time
A random sample of five runners records a short course time before and after following a training plan. Define improvement as before minus after, so a positive difference means a faster time afterward. The times, in minutes, are:
| Runner | Before | After | Before − after |
|---|---|---|---|
| 1 | 30 | 29 | 1 |
| 2 | 32 | 30 | 2 |
| 3 | 29 | 27 | 2 |
| 4 | 35 | 32 | 3 |
| 5 | 31 | 29 | 2 |
Identify and plan: Each runner is measured twice, so the observations are paired. Let \(\mu_d\) be the true mean improvement in course time for the population represented by the random sample. The question asks for an estimate, so use a paired t interval. Assume at least 50 runners are in the population; \(5/50=0.10\) meets the 10% condition. The differences \(1,2,2,3,2\) show no strong skewness or outliers, supporting t inference for this small sample, and distinct runners’ differences are independent.
Calculate: The mean difference is \(\bar{x}_d=10/5=2\) minutes. The squared deviations from 2 sum to 2, so \(s_d=\sqrt{2/4}\approx0.7071\) minutes. The standard error is \(0.7071/\sqrt{5}\approx0.3162\) minutes. For a 95% confidence interval with \(4\) degrees of freedom, \(t^*\approx2.776\). The margin of error is \(2.776(0.3162)\approx0.878\) minutes.
Interpret: We are 95% confident that the true mean improvement in course time for the population represented by these runners is between about 1.12 and 2.88 minutes. The interval uses the five within-runner differences. Treating the before and after columns as independent groups would ignore the repeated measurements and answer a different design-based question.
Common Mistakes and Full-Credit Communication
- Choosing from the table layout: Two columns do not prove that observations are paired. Identify the actual link from the study description.
- Pairing independent observations after the fact: Equal sample sizes do not authorize row-by-row subtraction. Without genuine links, use an unpooled two-sample t procedure.
- Using the wrong data for the conditions check: For paired t inference, assess the differences. For a two-sample t procedure, assess each group separately.
- Assuming paired inference always gives a smaller standard error: The spread of the differences depends on the actual links. Pairing can change the standard error; its effect is not guaranteed to go in one direction.
- Leaving subtraction order unstated: Define differences, for example, as A minus B. This tells the reader what positive and negative differences mean and identifies \(\mu_d\).
- Reporting a conclusion without the parameter: A paired conclusion concerns a mean difference. An independent-groups conclusion concerns a difference in population means. Name the appropriate population and variable in context.
A clear procedure choice can be brief but specific: “Because the same students took both formats, I will analyze the A-minus-B differences with a paired t procedure. If the samples instead came from separate, unlinked students, I would use an unpooled two-sample t procedure.” This makes the design—not just the arithmetic—visible.
Check Your Understanding
For each situation, decide whether the observations are paired or independent, identify the target parameter, and name the procedure that fits the stated goal.
- A random sample of 15 phones is tested for battery life with two settings, and each phone is tested under both settings. The question asks whether mean battery life differs between settings. What is the design and procedure?
- Two separate random samples of 20 households compare monthly water use under two billing plans. No household is matched to another. The question asks for a confidence interval. What parameter and procedure fit?
- A table lists the test scores of 12 students in one column and 12 different students in another. The columns have the same number of rows. Does that make the observations paired? Explain.
- For a genuine paired study, what data should be checked for compatibility with t inference: the two original columns separately, or the pairwise differences?
- In a before-minus-after study, a positive difference is observed. What does its sign mean, and why should the subtraction order be stated?