Tutorials › AP Statistics › Common Errors in Two-Sample t Intervals

Two-sample t confidence intervals · Tutorial 718 of 1000

Common Errors in Two-Sample t Intervals

Learn to avoid common two-sample t interval errors by checking the standard error, using the interval for the difference, and interpreting results in context.

Intermediate 10 min read

What You'll Learn

  • Calculate the standard error using each sample’s variance and sample size separately.
  • Spot why dividing a combined variance by the total sample size gives the wrong standard error.
  • Explain why overlap between two separate confidence intervals does not decide whether means differ.
  • Construct and interpret a two-sample t interval for the difference in population means.
  • Check the two-sample t interval conditions and communicate conclusions without overclaiming.

Errors Can Change the Answer

A two-sample t interval estimates the difference between two population means. As in “Using an Interval to Judge a Difference Between Groups,” zero is the reference value for no difference, and the order \(\mu_1-\mu_2\) determines what positive and negative values mean. But a correct interpretation depends on first calculating the interval correctly.

Two common errors are using the wrong standard error and trying to decide whether the means differ by comparing two separate confidence intervals. The standard error must reflect the variability and sample size of each group separately. And when the question is about a difference between means, use an interval constructed for that difference—not overlap between intervals constructed for the individual means.

Key reminder: For independent samples, the standard error of \(\bar{x}_1-\bar{x}_2\) is \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). Use an unpooled two-sample t interval to estimate \(\mu_1-\mu_2\) and judge whether zero is plausible.

Use Each Group’s Variance Contribution

The variability of the difference in sample means comes from both samples. Group 1 contributes \(s_1^2/n_1\), and group 2 contributes \(s_2^2/n_2\). Add those contributions and then take the square root to get the estimated standard deviation of \(\bar{x}_1-\bar{x}_2\), called its standard error.

$$ SE_{\bar{x}_1-\bar{x}_2} = \sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} $$

This is the unpooled standard error used in the AP Statistics two-sample t interval. As covered in “Pooled Versus Unpooled Two-Sample Intervals,” do not combine the sample standard deviations into a pooled estimate for the standard AP procedure. Also, do not replace the two sample sizes with their total. The groups can have different sample sizes and different sample variability, so their contributions need not be equal.

A frequent incorrect calculation is to add the sample variances and divide by the total number of observations, such as \(\sqrt{(s_1^2+s_2^2)/(n_1+n_2)}\). That does not give the standard error of the difference in sample means. Another error is to add \(s_1^2/n_1\) and \(s_2^2/n_2\) but forget the square root; that result is a variance, not a standard error, and its units are squared.

The degrees of freedom are also based on the two samples’ separate variance contributions. Use Welch’s formula or the conservative method, as introduced in “Computing Two-Sample Degrees of Freedom With the Welch Formula.” Do not automatically use \(n_1+n_2-2\), which is associated with a different, pooled approach.

Worked Examples

Worked Example: Don’t Divide a Combined Variance by the Total Sample Size

In an invented study of commute times, independent random samples of 16 workers from town 1 and 25 workers from town 2 report their one-way commute in minutes. The sample means are 72 and 66 minutes, and the sample standard deviations are 8 and 10 minutes, respectively. Both town populations are large compared with their samples. The sample distributions are roughly symmetric, with no apparent outliers. Construct a 95% confidence interval for \(\mu_1-\mu_2\).

State: Let \(\mu_1\) be the true mean one-way commute time for workers in town 1, and let \(\mu_2\) be the true mean one-way commute time for workers in town 2. The parameter \(\mu_1-\mu_2\) is measured in minutes.

Plan: Use an unpooled two-sample t interval. The samples are independent random samples of different workers, so observations are not paired. Each sample is less than 10% of its town’s worker population, satisfying the 10% condition. The sample distributions are roughly symmetric without apparent outliers, supporting the t procedure. Use each sample’s variance contribution separately.

Do: The observed difference in sample means is \(72-66=6\) minutes. The correct standard error is:

$$ SE = \sqrt{\frac{8^2}{16}+\frac{10^2}{25}} = \sqrt{4+4} = \sqrt{8} \approx 2.828 $$

For comparison, the incorrect calculation \(\sqrt{(8^2+10^2)/(16+25)}=\sqrt{4}\) would give \(2\) minutes. It does not account for how each group contributes to the variability of the difference. Welch’s degrees of freedom are about 36.9, and the 95% critical value is about \(t^*=2.026\). The margin of error is \(2.026(2.828)\approx5.730\) minutes. Thus:

$$ 6\pm5.730 \quad\Longrightarrow\quad (0.270,\ 11.730) $$

Conclude: We are 95% confident that the mean one-way commute time for workers in town 1 is about 0.27 to 11.73 minutes greater than for workers in town 2. The interval is entirely above zero, so it suggests a difference between the population means at the corresponding two-sided 5% significance level. The random samples support generalizing to the two worker populations, subject to the study’s design and conditions. This observational comparison does not establish that living in one town causes a longer commute.

Worked Example: Overlapping Individual Intervals, but a Difference Interval Excluding Zero

In an invented study, independent random samples of 25 plants from each of two large greenhouse sections are measured for height in centimeters. Section 1 has a sample mean of 100 cm and a sample standard deviation of 10 cm. Section 2 has a sample mean of 107.5 cm and a sample standard deviation of 10 cm. Assume the random sampling, 10% condition, and t-procedure shape conditions are met. Some students look at separate 95% intervals for the two means and conclude that overlap means there is no evidence of a difference. Check that reasoning by calculating an interval directly for \(\mu_1-\mu_2\).

State: Let \(\mu_1\) and \(\mu_2\) be the true mean heights of plants in sections 1 and 2, respectively. The parameter \(\mu_1-\mu_2\) is measured in centimeters.

Plan: Use an unpooled two-sample t interval for the difference in means. The samples consist of different plants, not matched pairs. They are independent random samples, each is less than 10% of its population, and the stated conditions support the t procedure. Separate intervals for \(\mu_1\) and \(\mu_2\) are not a substitute for an interval for their difference.

Do: Each separate 95% interval has standard error \(10/\sqrt{25}=2\) cm. With \(df=24\), \(t^*\approx2.064\), so each margin of error is \(2.064(2)=4.128\) cm. The intervals are approximately \((95.872,\ 104.128)\) cm for section 1 and \((103.372,\ 111.628)\) cm for section 2. They overlap from about 103.372 to 104.128 cm.

Now calculate the interval for the difference. The sample mean difference is \(100-107.5=-7.5\) cm, and its standard error uses both groups’ variance contributions:

$$ SE = \sqrt{\frac{10^2}{25}+\frac{10^2}{25}} = \sqrt{4+4} = \sqrt{8} \approx2.828 $$

Welch’s degrees of freedom are 48, giving \(t^*\approx2.011\) for a 95% interval. The margin of error is \(2.011(2.828)\approx5.688\) cm. Therefore:

$$ -7.5\pm5.688 \quad\Longrightarrow\quad (-13.188,\ -1.812) $$

Conclude: We are 95% confident that the mean height in section 1 is about 1.812 to 13.188 cm less than the mean height in section 2. The difference interval excludes zero, so it suggests a difference in population means at the corresponding two-sided 5% significance level, even though the separate 95% intervals overlap. The overlap is not the correct decision rule for this question; use the interval for \(\mu_1-\mu_2\).

Worked Example: Different Sample Sizes Need Different Contributions

In an invented comparison of two types of reusable water filters, independent random samples of 12 filters of type 1 and 30 filters of type 2 are tested for the volume of water, in liters, filtered before replacement. Type 1 has a sample mean of 54 liters and standard deviation of 6 liters. Type 2 has a sample mean of 50 liters and standard deviation of 9 liters. The populations are large relative to the samples, and the sample distributions are reasonably symmetric without strong outliers. Construct and interpret a 95% confidence interval for the difference in mean filtered volume.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean volumes filtered by types 1 and 2, respectively. The parameter \(\mu_1-\mu_2\) is measured in liters.

Plan: Use an unpooled two-sample t interval. The two groups are independent samples of different filters, with no pairing. They are random samples, and each sample is less than 10% of its population. The distributions are reasonably symmetric without strong outliers, supporting use of the t procedure.

Do: The difference in sample means is \(54-50=4\) liters. Since the sample sizes differ, the two variance contributions differ too:

$$ SE = \sqrt{\frac{6^2}{12}+\frac{9^2}{30}} = \sqrt{3+2.7} = \sqrt{5.7} \approx2.387 $$

Welch’s degrees of freedom are about 30.4. Using \(df\approx30\), \(t^*\approx2.042\) for a 95% interval. The margin of error is \(2.042(2.387)\approx4.875\) liters. Therefore:

$$ 4\pm4.875 \quad\Longrightarrow\quad (-0.875,\ 8.875) $$

Conclude: We are 95% confident that the mean volume filtered by type 1 is between about 0.875 liters less and 8.875 liters more than the mean volume filtered by type 2. Because the interval includes zero, it does not provide convincing evidence of a difference between the population means at the corresponding two-sided 5% significance level. It does not prove that the means are equal; the interval also includes nonzero differences.

Common Mistakes and AP Exam Tips

  • Dividing a combined variance by the total sample size: For the difference in sample means, calculate \(s_1^2/n_1\) and \(s_2^2/n_2\) separately, add them, and take the square root. Do not use \((s_1^2+s_2^2)/(n_1+n_2)\).
  • Adding standard deviations instead of variance contributions: Use \(s_1^2/n_1+s_2^2/n_2\) inside the square root. Adding \(s_1\) and \(s_2\) directly does not produce the standard error.
  • Forgetting the square root: The sum \(s_1^2/n_1+s_2^2/n_2\) is a variance estimate. Take its square root to obtain a standard error in the original measurement units.
  • Using overlap between separate intervals as a decision rule: Overlap of two individual confidence intervals does not establish that a difference interval includes zero. Construct the interval for \(\mu_1-\mu_2\) to assess a difference between means.
  • Assuming that an interval including zero proves equality: Say that the data do not provide convincing evidence of a difference at the corresponding two-sided significance level. Zero is plausible, but nonzero values may be plausible too.
  • Using pooled degrees of freedom automatically: For the unpooled two-sample t interval, use Welch degrees of freedom or the specified conservative method, not \(n_1+n_2-2\) by default.
  • Leaving out the context or subtraction order: State which group is first, give the units, and interpret both endpoints. A negative interval for \(\mu_1-\mu_2\) means the plausible mean for population 1 is lower than the plausible mean for population 2.

A strong AP response names the parameter, identifies the unpooled two-sample t interval, checks the design and distribution conditions, shows the separate variance contributions, and interprets the result in context. If judging evidence of a difference, compare zero with the interval for the difference—not with a visual comparison of two separate intervals.

Key takeaway: Calculate the standard error as \(\sqrt{s_1^2/n_1+s_2^2/n_2}\), using each sample’s variance and size separately. To judge a difference between population means, use the interval for \(\mu_1-\mu_2\); overlap between separate intervals does not decide whether that difference interval contains zero.

Check Your Understanding

For each question, focus on the calculation or interpretation error described.

  1. Two independent samples have \(s_1=5,\ n_1=10,\ s_2=8,\ n_2=20\). Write the correct standard error expression for \(\bar{x}_1-\bar{x}_2\). Do not round it.
  2. A student calculates \(s_1^2/n_1+s_2^2/n_2\) and reports that number as the standard error. What calculation is missing, and why?
  3. Two separate 95% confidence intervals for population means overlap. Does that fact alone show that a 95% interval for the difference includes zero? Explain.
  4. A 95% interval for \(\mu_1-\mu_2\) is \((-4.2,\ -0.8)\) hours. Interpret its direction in context if \(\mu_1\) and \(\mu_2\) are mean hours of sleep for groups 1 and 2.
  5. A two-sample t interval includes zero. Write a careful conclusion that does not claim the population means have been proven equal.