Why the Design Determines the Analysis
Two studies can both compare measurements under two conditions and still require different analyses. If each person, object, or deliberately matched pair contributes two related measurements, the comparison is paired. If the two groups consist of unrelated units, the comparison is between independent samples. The number of observations in each group does not decide which design was used.
In “What Makes Data Paired,” you learned to look for a meaningful design link, such as measuring the same person twice or matching two participants before assigning treatments. This tutorial adds a practical next step: connect that design choice to the analysis it requires. A paired design reduces the comparison to one within-pair difference per pair. An independent-groups design compares the two groups’ means directly.
Two Designs, Two Comparisons
In a paired design, the two measurements in a pair are linked. A person’s before-and-after measurements share that person; two matched participants share a planned matching link. The analysis focuses on how much the two measurements differ within each pair. For a population of pairs, the parameter of interest is the population mean difference, often denoted \(\mu_d\). The paired t procedure analyzes the sample of pairwise differences as a one-sample t procedure.
In an independent-groups design, each unit contributes an observation to just one group, and units in one group are not linked to units in the other group. The question concerns the difference between two population means, \(\mu_1-\mu_2\). A two-sample t procedure uses information from both groups to estimate that difference.
These formulas reflect different sources of variation. The paired standard error uses the variability of the within-pair differences. The independent-samples standard error uses the variability in each group separately. When paired measurements tend to move together, their differences may vary less than the separate measurements do. Pairing can then make a comparison more precise. It does not guarantee a smaller standard error in every study, and it does not justify treating independent observations as paired.
A Design-First Decision Process
When a description might fit either analysis, pause before looking at the means or calculating anything. A short design audit helps prevent a common error: trying to create pairs from observations that were never linked.
Determine who or what contributes each observation, and which condition or group each unit receives.
Ask whether the same unit contributes both measurements or whether two different units were deliberately matched by the study design.
If there is a meaningful link, compare within pairs. If there is no link and the groups are separate, compare the two group means.
For paired data, the pairs should be independent of one another. For independent samples, check independence within each group and between the groups.
As discussed in “Checking Independence Between Two Groups,” independence is a design question, not a visual property of a data table. In paired data, the two observations inside a pair are expected to be related; that is why they are paired. What matters for paired inference is that different pairs are independent. In a two-sample analysis, observations should not be linked across groups through repeated units or deliberate matching.
Worked Examples
Worked Example: Classifying Several Study Descriptions
For each fictional study, decide whether the comparison is paired or independent, and name the design feature that determines the choice.
| Study description | Design classification |
|---|---|
| A clinic records each patient’s sleep duration before and after a schedule change. | ? |
| A researcher compares test scores from two randomly selected, unrelated groups of students taught with different review plans. | ? |
| Researchers match cyclists by prior experience, then assign one cyclist in each pair to each training plan. | ? |
| A lab tests one of two materials on each separate sample of a manufactured product; each sample receives only one material. | ? |
Solution. The clinic study is paired because each patient provides both measurements. The patient is the shared unit, so the before and after values are linked.
The two student groups are independent if the study has no matching or shared units across groups. Having two groups to compare does not by itself create pairs. The cyclists form matched pairs because the researchers deliberately match them by experience and assign one cyclist in each pair to each plan. The product samples are independent groups because each separate sample receives only one material and no matching plan links samples across conditions.
The analysis follows those classifications: use a paired comparison for the clinic and cyclist studies, and an independent two-sample comparison for the students and product samples. If the lab instead tested both materials on each sample, that would change the design to paired, because each sample would then contribute both measurements.
Worked Example: How Pairing Changes the Standard Error
A fictional equipment team measures the operating time, in hours, of four devices before and after a maintenance adjustment. Each device is measured in both conditions. For this comparison, calculate the paired standard error and compare it with the standard error that would result from incorrectly treating the condition measurements as independent groups.
| Device | Before adjustment | After adjustment | After minus before |
|---|---|---|---|
| 1 | 10 | 12 | 2 |
| 2 | 12 | 16 | 4 |
| 3 | 14 | 20 | 6 |
| 4 | 16 | 24 | 8 |
Solution. This is a paired design: each device supplies both measurements. There are four pairs, and the within-pair differences are 2, 4, 6, and 8 hours. Their mean is \((2+4+6+8)/4=5\) hours. The sample standard deviation of the differences is
So the standard error for the paired mean difference is
For comparison, the before measurements have mean 13 hours and sample standard deviation \(s_1=\sqrt{20/3}\approx2.5820\) hours. The after measurements have mean 18 hours and sample standard deviation \(s_2=\sqrt{80/3}\approx5.1640\) hours. If these same data were treated as independent groups, the calculated standard error would be
The paired standard error is smaller here because the measurements move together across devices, while the differences are comparatively consistent. This numerical comparison explains why the design matters: the two methods estimate uncertainty using different information. It is not a reason to choose whichever method produces the smaller standard error. The study design determines the analysis. Also, the value 2.8868 is a hypothetical calculation using an inappropriate independent-groups model for these linked measurements; it is not a valid alternative analysis of this paired study.
Worked Example: Equal Group Sizes Do Not Create Pairs
A fictional food-testing lab compares the fill mass, in grams, of packages made on two separate production lines. Five packages are randomly selected from each line. No package is measured on both lines, and packages are not matched. The observed masses are:
| Line A (grams) | Line B (grams) |
|---|---|
| 10 | 12 |
| 10 | 12 |
| 12 | 15 |
| 14 | 18 |
| 14 | 18 |
Solution. These are independent groups, even though there are five packages in each. The packages are distinct units, and the design gives no reason to link the first Line A package with the first Line B package, or any other specific combination. The target comparison is the difference in the population mean fill masses for the two lines.
For Line A, \(\bar{x}_1=12\) grams. Its deviations from the mean are \(-2,-2,0,2,2\), so \(s_1=\sqrt{16/4}=2\) grams. For Line B, \(\bar{x}_2=15\) grams. Its deviations are \(-3,-3,0,3,3\), so \(s_2=\sqrt{36/4}=3\) grams. The standard error for the difference in sample means is
The observed difference in sample means is \(\bar{x}_1-\bar{x}_2=12-15=-3\) grams. The negative sign indicates that the Line A sample mean is 3 grams lower than the Line B sample mean. This example identifies the appropriate comparison and its standard error; it does not by itself establish a population difference. The design, conditions, and inferential question would all matter for a test or interval.
It would be unjustified to subtract the table entries row by row and analyze those five differences as if they were pairs. Their row positions are merely a way to display two samples of equal size; the study did not establish those links. Pairing them after data collection would invent a design feature that never existed.
Consequences of Choosing the Wrong Analysis
The choice is not just a calculator setting. A paired analysis estimates the mean of the within-pair differences, so its conclusion addresses the average change or average within-match contrast. A two-sample analysis estimates the difference between two population means. Those are related questions, but they arise from different designs and use different measures of variability.
If genuinely paired data are analyzed as independent, the calculation ignores the link between the observations. The resulting standard error may be larger or smaller than the paired standard error, depending on how the measurements vary together. The procedure may therefore give an inaccurate account of uncertainty. Conversely, if independent observations are artificially paired, the calculated differences depend on arbitrary pairings and can produce misleading variability and conclusions.
Pairing also does not make a study randomized or representative. As covered in “Random Assignment Versus Random Sampling Conditions,” random assignment and random sampling have different roles: assignment can support a cause-and-effect conclusion for the study units, while sampling can support generalizing to the population sampled. The presence of pairs alone establishes neither. For a paired t procedure, use the conditions discussed in “Conditions for One-Sample Versus Paired Data”; for independent samples, use those in “Conditions for Two-Sample t Procedures.” The condition checks must match the design.
Common Mistakes and AP Exam Tips
- Using equal sample sizes as evidence of pairing. Equal counts are not a matching rule. Full-credit reasoning identifies a shared unit or a deliberate matching plan—or states that no such link is described.
- Pairing by row number or collection order. Rows can be used to display data, but a row position does not establish a meaningful relationship between two units.
- Choosing the analysis from the result you want. Do not select the paired or independent procedure based on which gives a smaller standard error or a more favorable test result. The design determines the analysis.
- Calling the two observations within a pair independent. They are related by design. For paired inference, check independence between pairs, not independence of the two measurements inside each pair.
- Interpreting the wrong parameter. A paired analysis concerns a population mean difference; an independent two-sample analysis concerns a difference between population means. State which comparison fits the study.
- Claiming that pairing guarantees a more precise result. Pairing can reduce variability when within-pair differences are less variable than the separate measurements, but the actual standard error depends on the data.
For a strong AP response, state the design evidence and connect it to the procedure. For example: “The observations are paired because each device was measured under both conditions. I would analyze one difference per device with a paired t procedure.” Or: “The groups are independent because separate packages were sampled from each line and no matching or repeated measurement links them; I would use a two-sample t procedure.” These explanations justify the method rather than merely naming it.
Check Your Understanding
For each situation, choose paired or independent analysis and give the design reason.
- A mechanic records fuel use for each of eight vehicles using two engine settings, testing both settings on every vehicle. What makes this a paired design?
- Two unrelated groups of 20 students use different study apps. Is the data paired because each group has 20 students? Explain.
- Researchers match participants by baseline fitness, then assign one participant in each pair to each exercise plan. What link supports paired analysis?
- In a paired study, why are the two observations within a pair not treated as independent, and at what level should independence be checked?
- A study genuinely uses paired measurements, but a student analyzes the two columns as independent samples. What feature of the design has the analysis ignored?