Tutorials › AP Statistics › Same Data, Different Design, Different Procedure

Choosing a mean-inference procedure · Tutorial 769 of 1000

Same Data, Different Design, Different Procedure

Learn how study design—not the numbers or table layout—determines whether mean comparisons use paired differences or two independent samples.

Intermediate 10 min read

What You'll Learn

  • Distinguish genuine pairs from two unlinked groups, even when both designs produce the same columns of numbers.
  • Identify whether the target parameter is a mean of pairwise differences or a difference between population means.
  • Compare how pairing changes the standard error and t statistic.
  • Check conditions for paired and unpooled two-sample t procedures using the appropriate data.
  • Explain why changing or inventing pairings can change an analysis without changing either group’s values.

Why Design Can Change the Analysis

In “A Decision Flowchart for Mean Inference,” you learned to choose a procedure by identifying the response, the design-based parameter, and the goal of the question. This tutorial zooms in on one important comparison: the same two columns of numbers can call for different procedures if the study design links the observations in one situation but not in another.

A paired design might measure the same people under two conditions, or deliberately match one unit in one group to a unit in another. An independent-groups design has no such one-to-one links. The values in a table cannot tell you which design occurred by themselves. The research description does that.

Key distinction: In a paired design, the target is \(\mu_d\), the true mean of the pairwise differences defined in a stated order. In an independent-groups design, the target is \(\mu_1-\mu_2\), the difference between the two population means. The same observed columns can yield the same observed difference in means, but the estimated standard error—and therefore the t statistic and p-value—can differ.

This difference in procedure is not a matter of preference. Genuine pairing makes the within-pair differences the relevant data. If the groups are unlinked, making pairs after seeing the values does not create a paired design. As emphasized in “Two-Sample Versus Paired Test on the Same Numbers,” design-based links—not equal sample sizes or a two-column layout—determine the procedure.

Compare the Targets and the Data Used

Suppose two conditions produce quantitative measurements. With genuine pairs, first define a difference for each pair, such as condition A minus condition B. The paired t procedure then treats those differences as one sample and makes inference about \(\mu_d\). The sample size is the number of pairs, and the condition about shape applies to the differences.

With independent groups, the target is the difference between the population means, \(\mu_1-\mu_2\). The unpooled two-sample t procedure uses the two samples separately. Its standard error is based on each group’s sample standard deviation and sample size. For the formula and conditions of that procedure, see “Conditions for a Two-Sample t Test” and “How Sample Size and Spread Affect the Test Result.”

$$ \text{Paired:}\qquad t=\frac{\bar{x}_d-0}{s_d/\sqrt{n}} $$
$$ \text{Independent groups:}\qquad t=\frac{\bar{x}_1-\bar{x}_2} {\sqrt{s_1^2/n_1+s_2^2/n_2}} $$

Both procedures compare an observed difference with a null value of zero in the test examples below, but the quantities in the denominators are not interchangeable. In the paired formula, \(s_d\) describes the sample-to-sample variation among the pairwise differences. In the independent formula, \(s_1\) and \(s_2\) describe variation within the separate groups.

Pairing can make differences less variable when linked measurements tend to move together, or more variable when the chosen pairings do not do that. The important point is not that pairing always makes a test stronger. It is that the pairwise differences—and their spread—belong to the paired design. The independent procedure does not use the row-by-row differences.

Worked Examples: One Set of Values, Two Designs

Worked Example: Same Scores From Matched or Independent Students

Imagine two practice formats, A and B, with these scores. In one version of the study, the same eight students take both formats, so each row represents one student. In another version, the A and B scores come from separate random samples of eight students each; the side-by-side layout is only for display, and rows do not link students.

RowFormat AFormat B
17079
27281
37475
47681
57871
68075
78271
88471

State: In both versions, consider a two-sided test at \(\alpha=0.05\). For the paired version, define each difference as A minus B, and let \(\mu_d\) be the true mean difference in scores for the population represented by the study. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\). For the independent version, let \(\mu_1\) and \(\mu_2\) be the population mean scores for formats A and B, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).

Plan and conditions: For the paired version, assume the eight students were randomly sampled from a population of at least 80 students and took both formats. The design provides genuine pairs; the differences from distinct students are independent, and \(8/80=0.10\) meets the 10% condition. Suppose a plot of the differences shows no strong skewness or outliers. The paired t test is appropriate.

For the independent version, assume the two separate random samples each contain eight students from populations of at least 80 students. The samples are independent by design, and each sample satisfies the 10% condition. Suppose plots for both groups show no strong skewness or outliers. The unpooled two-sample t test is appropriate. A paired procedure would not be justified for these unlinked students.

Do—paired version: The differences are \(-9,-9,-1,-5,7,5,11,13\). Their sum is 12, so \(\bar{x}_d=12/8=1.50\) points. The sum of squared deviations from 1.50 is 534, giving \(s_d=\sqrt{534/7}\approx8.734\) points and a standard error of \(8.734/\sqrt{8}\approx3.088\) points. With \(7\) degrees of freedom, \(t\approx1.50/3.088=0.486\), and the two-sided p-value is approximately \(0.642\), rounded.

$$ t=\frac{\bar{x}_d-0}{s_d/\sqrt{n}} =\frac{1.50}{\sqrt{534/7}/\sqrt{8}} \approx0.486 $$

Do—independent version: The means are \(\bar{x}_1=616/8=77.0\) and \(\bar{x}_2=604/8=75.5\), so the observed difference is \(1.50\) points. For group A, the sum of squared deviations from 77 is 168, so \(s_1=\sqrt{168/7}=\sqrt{24}\approx4.899\). For group B, the sum of squared deviations from 75.5 is 134, so \(s_2=\sqrt{134/7}\approx4.375\). The estimated standard error is \(\sqrt{24/8+(134/7)/8}\approx2.322\) points. The Welch degrees of freedom are approximately \(13.8\); the test statistic is \(1.50/2.322\approx0.646\), with two-sided p-value approximately \(0.529\), rounded.

$$ t=\frac{77.0-75.5}{\sqrt{24/8+(134/7)/8}} \approx0.646 $$

Conclude: Both p-values exceed \(0.05\), so in each version we fail to reject the null hypothesis. The paired analysis does not provide convincing evidence of a nonzero mean A-minus-B score difference for the population represented by the paired study. The independent analysis does not provide convincing evidence of a difference in the population mean scores for the two independent groups. The conclusions refer to different design-based parameters, even though the two columns of values are identical.

Pairing Is About the Study, Not the Spreadsheet

In the paired version of the score example, row 1 means that one student earned both the A score of 70 and the B score of 79. That link makes the row difference \(-9\) meaningful. In the independent version, row 1 simply places one A score next to one B score for convenience. Subtracting those entries would not create a meaningful within-student difference.

This also explains why rearranging a column is not harmless when the data are paired. Reordering rows for display is fine only if the links are preserved. If the values are reassigned across pairs, the calculated differences and their variability can change. That would describe a different matching, not the observed study.

AP exam tip: State what makes observations paired: the same units measured twice or deliberate one-to-one matching. Then name the parameter and procedure. “There are two columns” or “both samples have eight observations” is not evidence of pairing.

Worked Example: What Reversing the Matches Changes

Worked Example: Keep the Values, Change the Pairing

Return to the same A and B score columns. In the paired-study version, the actual row links give differences \(-9,-9,-1,-5,7,5,11,13\). Now imagine reversing the order of the B scores and pairing each A score with that reversed list. This is a different matching, not a permissible correction to the original rows when the original students were genuinely linked.

The reversed B list is \(71,71,75,71,81,75,81,79\). Subtracting it from the A values in order gives the differences \(-1,1,-1,5,-3,5,1,5\). They sum to 12, so their mean is \(12/8=1.50\). Their squared deviations from 1.50 sum to 70, so \(s_d=\sqrt{70/7}=\sqrt{10}\approx3.1623\). The standard error is \(\sqrt{10}/\sqrt{8}=\sqrt{10/8}\approx1.1180\), and the paired t statistic for a zero mean difference is \(1.50/1.1180\approx1.342\), with \(7\) degrees of freedom.

$$ t=\frac{1.50}{\sqrt{10/8}}\approx1.342 $$

This reversed matching would give a two-sided p-value of about \(0.2216\), rounded. It differs from the paired result using the actual row links because the sample of differences has a different spread. The values in each column and their means have not changed. If the design had genuinely matched students in this reversed way, these would be the differences to analyze. But when the original same-student links are real, reversing them discards the study design and produces an invalid analysis.

Worked Example: A Paired Interval Uses the Differences

Worked Example: Estimate a Mean Change in Running Time

A random sample of five runners records a short course time before and after following a training plan. Define improvement as before minus after, so a positive difference means a faster time afterward. The times, in minutes, are:

RunnerBeforeAfterBefore − after
130291
232302
329272
435323
531292

Identify and plan: Each runner is measured twice, so the observations are paired. Let \(\mu_d\) be the true mean improvement in course time for the population represented by the random sample. The question asks for an estimate, so use a paired t interval. Assume at least 50 runners are in the population; \(5/50=0.10\) meets the 10% condition. The differences \(1,2,2,3,2\) show no strong skewness or outliers, supporting t inference for this small sample, and distinct runners’ differences are independent.

Calculate: The mean difference is \(\bar{x}_d=10/5=2\) minutes. The squared deviations from 2 sum to 2, so \(s_d=\sqrt{2/4}\approx0.7071\) minutes. The standard error is \(0.7071/\sqrt{5}\approx0.3162\) minutes. For a 95% confidence interval with \(4\) degrees of freedom, \(t^*\approx2.776\). The margin of error is \(2.776(0.3162)\approx0.878\) minutes.

$$ \bar{x}_d\pm t^*\frac{s_d}{\sqrt{n}} =2\pm2.776\left(\frac{\sqrt{2/4}}{\sqrt{5}}\right) \approx(1.122,\ 2.878) $$

Interpret: We are 95% confident that the true mean improvement in course time for the population represented by these runners is between about 1.12 and 2.88 minutes. The interval uses the five within-runner differences. Treating the before and after columns as independent groups would ignore the repeated measurements and answer a different design-based question.

Common Mistakes and Full-Credit Communication

  • Choosing from the table layout: Two columns do not prove that observations are paired. Identify the actual link from the study description.
  • Pairing independent observations after the fact: Equal sample sizes do not authorize row-by-row subtraction. Without genuine links, use an unpooled two-sample t procedure.
  • Using the wrong data for the conditions check: For paired t inference, assess the differences. For a two-sample t procedure, assess each group separately.
  • Assuming paired inference always gives a smaller standard error: The spread of the differences depends on the actual links. Pairing can change the standard error; its effect is not guaranteed to go in one direction.
  • Leaving subtraction order unstated: Define differences, for example, as A minus B. This tells the reader what positive and negative differences mean and identifies \(\mu_d\).
  • Reporting a conclusion without the parameter: A paired conclusion concerns a mean difference. An independent-groups conclusion concerns a difference in population means. Name the appropriate population and variable in context.

A clear procedure choice can be brief but specific: “Because the same students took both formats, I will analyze the A-minus-B differences with a paired t procedure. If the samples instead came from separate, unlinked students, I would use an unpooled two-sample t procedure.” This makes the design—not just the arithmetic—visible.

Key takeaway: Identical numerical columns do not determine the procedure. Genuine links call for inference on pairwise differences; unlinked groups call for inference on the difference in population means. Preserve the study’s actual links, and check the conditions for the procedure those links justify.

Check Your Understanding

For each situation, decide whether the observations are paired or independent, identify the target parameter, and name the procedure that fits the stated goal.

  1. A random sample of 15 phones is tested for battery life with two settings, and each phone is tested under both settings. The question asks whether mean battery life differs between settings. What is the design and procedure?
  2. Two separate random samples of 20 households compare monthly water use under two billing plans. No household is matched to another. The question asks for a confidence interval. What parameter and procedure fit?
  3. A table lists the test scores of 12 students in one column and 12 different students in another. The columns have the same number of rows. Does that make the observations paired? Explain.
  4. For a genuine paired study, what data should be checked for compatibility with t inference: the two original columns separately, or the pairwise differences?
  5. In a before-minus-after study, a positive difference is observed. What does its sign mean, and why should the subtraction order be stated?