Tutorials › AP Statistics › Why the Wrong Procedure Gives Misleading Results

Choosing a mean-inference procedure · Tutorial 777 of 1000

Why the Wrong Procedure Gives Misleading Results

Compare paired and incorrectly independent calculations to see how ignoring links between observations can change the standard error and the conclusion.

Intermediate 11 min read

What You'll Learn

  • Explain why repeated measurements from the same units are not independent samples.
  • Calculate the paired standard error from the standard deviation of the differences.
  • Compare a valid paired t test with an invalid unpooled two-sample t test on the same data.
  • Show how the incorrect procedure can change a test decision.
  • Explain why misusing two-sample t does not always make the standard error larger.

The Link Between Measurements Matters

In “Choosing Procedures From Summary Tables,” you learned to choose a mean procedure by tracing the study design and identifying the target parameter. This tutorial examines what can go wrong when that step is skipped. In particular, treating repeated measurements on the same units as if they came from independent groups can produce a very different standard error—and potentially a different conclusion.

A two-sample t procedure is built for two independent groups. When the same people or other units are measured twice, each unit’s two responses are linked. The analysis should preserve that link by calculating one difference per unit and using a paired t procedure. If the two measurements tend to move together, ignoring their relationship can discard useful information about the change within each unit.

Key idea: The paired t standard error is \(s_d/\sqrt{n}\), where \(s_d\) is the standard deviation of the within-pair differences. The unpooled two-sample t standard error is \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). These standard errors answer to different designs; using the second formula on paired data does not make it a valid analysis.

The difference between the formulas is not just a technical detail. A standard error measures the estimated variability of a statistic across repeated samples under the model for the design. When observations are linked, the variability of the differences depends on how the two measurements within each pair vary together. The separate group standard deviations do not capture that link.

Compare the Calculations on the Same Paired Data

For paired observations, define the difference in a stated order, such as \(d=\text{after}-\text{before}\). The target is the population mean difference, \(\mu_d\). As in “Identifying the Parameter in a Mean Problem,” the order matters: reversing it reverses the sign of the sample mean difference and the test statistic, but it does not change the standard error or the strength of evidence in a two-sided test.

If the same measurements are incorrectly treated as independent samples, the estimated difference in means is still \(\bar{x}_{\text{after}}-\bar{x}_{\text{before}}\). But its standard error is calculated from the two marginal standard deviations as if the two samples were unlinked. That assumption is contrary to the design when each person appears in both groups.

Worked Example: A Conclusion Changes When Pairing Is Ignored

A fictional clinic randomly selects 8 clients from a population of 120 clients and records the number of minutes each needs to complete a routine before and after a program. The same clients are measured twice. The differences, defined as after minus before, are:

$$ -4,\ -3,\ -2,\ -2,\ -2,\ -2,\ -1,\ 0 $$

The differences are symmetric around \(-2\), with no apparent outliers. The question is whether the program reduces the population mean completion time. Use \(\alpha=0.05\).

State: Let \(\mu_d\) be the true mean change in completion time, in minutes, for the population represented by the sampled clients, where \(d=\text{after}-\text{before}\). A reduction corresponds to a negative difference. Test \(H_0:\mu_d=0\) against \(H_a:\mu_d<0\).

Plan: Use a paired t test because the same clients contribute both measurements. The clients were randomly selected, and \(8\leq0.10(120)=12\), so the 10% condition is satisfied. The differences from distinct clients are treated as independent. With 8 differences, the shape condition is supported by their symmetry and lack of apparent outliers.

Do—the valid paired analysis: The sum of the differences is \(-16\), so \(\bar{x}_d=-16/8=-2\) minutes. Deviations from \(-2\) are \(-2,-1,0,0,0,0,1,2\); their squared deviations sum to \(10\). Therefore,

$$ s_d=\sqrt{\frac{10}{8-1}}=\sqrt{\frac{10}{7}}\approx1.195 \qquad SE_{\text{paired}}=\frac{s_d}{\sqrt{8}} =\sqrt{\frac{10}{56}}\approx0.423 $$

The test statistic is

$$ t=\frac{\bar{x}_d-0}{s_d/\sqrt{n}} =\frac{-2}{\sqrt{10/56}} \approx-4.733 $$

The degrees of freedom are \(8-1=7\). The left-tailed p-value is approximately \(0.0010\), rounded. Since \(0.0010<0.05\), reject \(H_0\). The data provide convincing evidence that the mean completion time after the program is lower than the mean completion time before the program for the population represented by the sampled clients.

Now calculate the misleading result: The before measurements are \(20,22,24,26,28,30,32,34\), with mean \(27\) and sample variance \(24\). The after measurements are \(16,19,22,24,26,28,31,34\), with mean \(25\). Their deviations from \(25\) are \(-9,-6,-3,-1,1,3,6,9\), whose squared values sum to \(254\). Thus the after sample variance is \(254/7\). If these repeated measurements are wrongly entered as two independent groups, the unpooled standard error is

$$ SE_{\text{wrong}} =\sqrt{\frac{24}{8}+\frac{254/7}{8}} =\sqrt{\frac{211}{28}} \approx2.745 $$

The incorrect test statistic is \((25-27)/2.745\approx-0.729\). The Welch degrees of freedom are approximately \(13.44\), giving a left-tailed p-value of approximately \(0.239\), rounded. Since this p-value is greater than \(0.05\), that invalid analysis would fail to reject \(H_0\).

The paired standard error is about \(0.423\) minutes, while the wrongly calculated independent-groups standard error is about \(2.745\) minutes. The latter treats the before and after values as if they came from unrelated clients, overlooking the within-client comparisons. Here, that change makes the test statistic much closer to zero and reverses the decision. The two-sample p-value is not valid evidence for this study: its calculation relies on a design that did not occur.

Why the Standard Errors Can Differ

With equal sample sizes and the same units measured twice, the paired standard error uses the variability of the differences, while the independent-groups calculation uses the variability of each set of measurements separately. If clients with higher before values also tend to have higher after values, much of the variation in their separate measurements can cancel when each client’s change is calculated. In that situation, the differences may be less variable than either set of measurements.

This is why matching can improve precision: each unit serves as its own comparison. But do not memorize the rule “the two-sample standard error is always bigger.” The direction depends on how the measurements within pairs relate. The essential point is that only the procedure matching the design gives a justified standard error and test.

Worked Example: A Second Paired Study and a Different Wrong Result

In another fictional study, 10 randomly selected volunteers are measured under two conditions. The same volunteers participate in both, and researchers define \(d=\text{condition B}-\text{condition A}\). The summary statistics are \(\bar{x}_d=3\), \(s_d=2\), and the separate condition standard deviations are both \(5\). Assume the differences are approximately symmetric with no apparent outliers, and the volunteers were sampled from a population of 300. Test whether the population mean difference is nonzero at \(\alpha=0.05\).

State and plan: Let \(\mu_d\) be the true mean difference, in the response’s units. Test \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\). Use a paired t test because each volunteer contributes a measurement under both conditions. The volunteers were randomly selected, \(10\leq0.10(300)=30\), and the differences satisfy the stated shape check; treat differences from distinct volunteers as independent.

Do—the paired calculation:

$$ SE_{\text{paired}}=\frac{2}{\sqrt{10}}\approx0.632 \qquad t=\frac{3-0}{2/\sqrt{10}}\approx4.743 $$

The degrees of freedom are \(9\), and the two-sided p-value is approximately \(0.0011\), rounded. Reject \(H_0\); the data provide convincing evidence that the population mean response differs between the two conditions.

Compare with the incorrect calculation: Treating the measurements as independent would give

$$ SE_{\text{wrong}} =\sqrt{\frac{5^2}{10}+\frac{5^2}{10}} =\sqrt{5}\approx2.236 \qquad t_{\text{wrong}}=\frac{3}{\sqrt{5}}\approx1.342 $$

With equal group sizes and standard deviations, the Welch degrees of freedom are \(18\). The two-sided p-value is approximately \(0.196\), rounded, so this invalid calculation would fail to reject \(H_0\). Again, the conclusion changes because the wrong standard error fails to use the paired design. The correct inference is based on the sample of 10 differences, not on two independent groups of 10.

Worked Example: The Wrong Standard Error Is Not Always Larger

A fictional lab applies two measurement settings to each of 8 randomly selected specimens. The paired differences, defined as setting B minus setting A, are \(14,10,6,2,-2,-6,-10,-14\). They are symmetric, with no apparent outliers. The sample mean difference is \(0\). Suppose the sample variances for both settings are \(24\); the sample variance of the differences is \(96\). The specimens come from a population of 200, so \(8\leq0.10(200)=20\). Treat differences from distinct specimens as independent.

Correct paired calculation: The standard error is \(\sqrt{96}/\sqrt{8}=\sqrt{12}\approx3.464\). The test statistic for \(H_0:\mu_d=0\) is \(t=(0-0)/3.464=0\), with \(7\) degrees of freedom and a two-sided p-value of \(1.000\).

Incorrect independent-groups calculation: The separate sample variances would give \(SE_{\text{wrong}}=\sqrt{24/8+24/8}=\sqrt{6}\approx2.449\), which is smaller than the correct paired standard error. Its test statistic is still \(0\), but the calculation remains invalid: it uses an independence assumption the paired design does not satisfy.

This example illustrates an important limit on the comparison. Misusing two-sample t can make the standard error smaller or larger, depending on the relationship between measurements in each pair. Either way, the wrong procedure does not provide a sound basis for the p-value or conclusion.

Common Mistakes and AP Exam Tips

  • Choosing two-sample t because there are two columns: Two columns may contain repeated measurements on the same units. Identify the observational units and the link between columns before choosing a procedure.
  • Using separate standard deviations for paired data: The paired procedure needs \(s_d\), the standard deviation of the individual differences. The before and after standard deviations do not replace it.
  • Calling the two-sample result an alternative valid answer: If the same units were measured twice, the independent-groups assumption is not met. A calculator’s p-value cannot correct a procedure that does not match the design.
  • Claiming the wrong standard error must be larger: Its direction depends on how the paired measurements vary together. Say that the wrong calculation ignores the pairing; do not claim it always overstates uncertainty.
  • Reporting only that the procedures differ: Show the numerical effect when asked. State both standard errors, the corresponding test statistics and p-values, and whether the decision changes.
  • Making a conclusion about individuals from a mean test: A paired t test addresses the population mean difference, not whether every individual improved or whether a particular client’s change is typical.

A full-credit explanation identifies the repeated-measures design, defines the difference in a stated order, and gives the paired standard error based on \(s_d\). When comparing with a mistaken two-sample calculation, label it clearly as invalid and explain which independence assumption fails. State the decision in context using the valid test.

Key takeaway: Paired data must be analyzed through one difference per pair. Treating the two measurements as independent can change the standard error, test statistic, p-value, and even the apparent conclusion; the design—not the table layout—determines which calculation is valid.

Check Your Understanding

Use the design and calculations in this tutorial to answer each question.

  1. Why is the unpooled two-sample t procedure inappropriate when the same clients are measured before and after a program?
  2. For paired data with \(n=16\) and \(s_d=3.2\), calculate the paired standard error.
  3. In the first worked example, which standard error and p-value support the valid conclusion, and why?
  4. Can the standard error from a mistaken two-sample calculation be assumed to exceed the paired standard error? Explain.
  5. What information about the design should you state when explaining why paired t is the appropriate procedure?