Tutorials › AP Statistics › How Variability Affects Width of a Two-Sample Interval

Two-sample t confidence intervals · Tutorial 714 of 1000

How Variability Affects Width of a Two-Sample Interval

Use numerical comparisons to see why more variable samples widen a two-sample interval and larger samples generally make it narrower.

Intermediate 9 min read

What You'll Learn

  • Explain how increasing either sample standard deviation changes its contribution to the standard error.
  • Calculate a two-sample t interval’s margin of error and total width.
  • Compare interval widths while holding other quantities fixed.
  • Describe how larger sample sizes affect standard error and the t critical value.
  • Distinguish a change in interval width from a change in its center.
  • Communicate the assumptions behind a numerical width comparison.

From Standard Error to Interval Width

In “Effect of Unequal Sample Sizes on Standard Error,” you saw how the two groups contribute separately to the standard error. This tutorial connects those contributions to the width of a two-sample t confidence interval. The key is to compare intervals while holding the confidence level and other relevant quantities fixed.

Recall the unpooled interval from “The Form of a Two-Sample t Interval.” Its center is the observed difference in sample means, and its margin of error is the critical value times the standard error. The full width is twice the margin of error.

$$ (\bar{x}_1-\bar{x}_2)\ \pm\ t^*SE_{\bar{x}_1-\bar{x}_2}, \qquad SE_{\bar{x}_1-\bar{x}_2} = \sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} $$
$$ \text{Margin of error}=t^*SE_{\bar{x}_1-\bar{x}_2}, \qquad \text{Width}=2t^*SE_{\bar{x}_1-\bar{x}_2} $$

The sample standard deviations \(s_1\) and \(s_2\) estimate variability within their respective groups. If one of them gets larger while sample sizes and the other standard deviation stay fixed, its squared contribution \(s_i^2/n_i\) increases. That makes the standard error larger and tends to make the interval wider.

Increasing a sample size \(n_i\), with the standard deviations held fixed, reduces that group’s contribution \(s_i^2/n_i\). The standard error therefore gets smaller. Also, the Welch degrees of freedom usually increase as sample sizes increase, which tends to reduce \(t^*\). Both changes generally narrow the interval at the same confidence level.

Key idea: At a fixed confidence level, interval width is \(2t^*SE\). Greater sample variability increases standard error; larger sample sizes reduce standard error. The critical value also matters, so compare both \(SE\) and \(t^*\) when calculating actual widths.

What a Width Comparison Holds Fixed

An interval’s width describes its precision, not the value at its center. If the sample means are held fixed, changing a standard deviation or sample size changes the margin of error and spreads the endpoints farther apart or brings them closer together. If the sample means change too, the center can move as well as the width.

To isolate the effect of variability, keep \(n_1\), \(n_2\), the confidence level, and \(s_2\) fixed while changing \(s_1\), for example. To isolate the effect of sample size, keep the standard deviations, confidence level, and other sample size fixed while changing one or both \(n_i\). These controlled comparisons help explain the formula; actual samples may not keep their standard deviations fixed when their sizes change.

Comparison checklist: State what is being changed and what is held fixed. Calculate each group’s variance contribution, find the standard error, use the appropriate Welch degrees of freedom to get \(t^*\), and calculate width as \(2t^*SE\).

Worked Examples

Worked Example: Increasing One Sample Standard Deviation

An invented study compares the mean time, in minutes, that two independent groups take to complete a task. Each sample has 20 observations. First suppose \(s_1=5\) minutes and \(s_2=5\) minutes. Then compare that interval with one in which \(s_1\) increases to 8 minutes while \(s_2\), both sample sizes, and the 95% confidence level remain fixed. The sample means are assumed to stay the same, so only the width is being compared.

For the first case, the standard error is:

$$ SE = \sqrt{\frac{5^2}{20}+\frac{5^2}{20}} = \sqrt{1.25+1.25} = \sqrt{2.5} \approx 1.5811\text{ minutes} $$

The Welch degrees of freedom are 38 because the two variance contributions are equal and the sample sizes are both 20. For 95% confidence, \(t^*\approx2.0244\). Therefore, the width is:

$$ \text{Width} = 2(2.0244)(1.5811) \approx 6.402\text{ minutes} $$

Now increase only \(s_1\) to 8 minutes. The new standard error is:

$$ SE = \sqrt{\frac{8^2}{20}+\frac{5^2}{20}} = \sqrt{3.2+1.25} = \sqrt{4.45} \approx 2.1095\text{ minutes} $$

The Welch degrees of freedom are now approximately:

$$ df = \frac{(3.2+1.25)^2} {\frac{3.2^2}{19}+\frac{1.25^2}{19}} = \frac{19.8025}{11.8025/19} \approx31.88 $$

For 95% confidence and about 31.88 degrees of freedom, \(t^*\approx2.037\). The new width is:

$$ \text{Width} = 2(2.037)(2.1095) \approx8.594\text{ minutes} $$

Increasing \(s_1\) from 5 to 8 minutes increased the width from about 6.402 to 8.594 minutes, with the other stated quantities held fixed. The result follows the formula: group 1’s contribution increased from \(25/20=1.25\) to \(64/20=3.2\). The critical value changed slightly too, because the Welch degrees of freedom changed.

Worked Example: Increasing Both Sample Sizes

An invented environmental study compares mean daily water use, in liters, for two independent neighborhoods. For a controlled comparison, suppose both groups have a sample standard deviation of 6 liters. Compare 95% interval widths with 10 observations per group and 25 observations per group. Keep the standard deviations and confidence level fixed.

With 10 observations in each group:

$$ SE = \sqrt{\frac{6^2}{10}+\frac{6^2}{10}} = \sqrt{3.6+3.6} = \sqrt{7.2} \approx2.6833\text{ liters} $$

Here, the Welch degrees of freedom are 18. For 95% confidence, \(t^*\approx2.1009\), so the width is:

$$ \text{Width} = 2(2.1009)(2.6833) \approx11.275\text{ liters} $$

With 25 observations in each group:

$$ SE = \sqrt{\frac{6^2}{25}+\frac{6^2}{25}} = \sqrt{1.44+1.44} = \sqrt{2.88} \approx1.6971\text{ liters} $$

The Welch degrees of freedom are 48, giving \(t^*\approx2.0106\) for 95% confidence. Thus:

$$ \text{Width} = 2(2.0106)(1.6971) \approx6.824\text{ liters} $$

The width falls by about \(11.275-6.824=4.451\) liters. In this comparison, both groups’ variance contributions became smaller, and the critical value decreased slightly as the degrees of freedom increased. The calculation shows the combined effect on actual interval width, rather than considering standard error alone.

Worked Example: A Complete Interval and a Larger-Sample Comparison

An invented study compares mean completion times, in minutes, for two independent versions of a digital task. A random sample of 16 users testing version 1 has \(\bar{x}_1=27\) minutes and \(s_1=4\) minutes. An independent random sample of 16 users testing version 2 has \(\bar{x}_2=24\) minutes and \(s_2=4\) minutes. We will construct a 95% interval, then compare its width with a hypothetical study that has 36 users in each group while keeping the sample standard deviations fixed.

State: The parameter is \(\mu_1-\mu_2\), the true mean completion time for all users of version 1 minus the true mean completion time for all users of version 2, in minutes.

Plan: Use an unpooled two-sample t interval, as in “The Form of a Two-Sample t Interval.” The groups are independent because each user tests one version and there is no pairing. The samples are described as random. If sampling is without replacement, each population should contain at least 160 users for the original samples to meet the 10% condition. The samples are small, so assume plots of the completion times show no strong skewness or outliers. These conditions support using the procedure.

Do: The observed difference in sample means is \(27-24=3\) minutes. The standard error is:

$$ SE = \sqrt{\frac{4^2}{16}+\frac{4^2}{16}} = \sqrt{1+1} = \sqrt{2} \approx1.4142\text{ minutes} $$

The Welch degrees of freedom are 30. For a 95% interval, \(t^*\approx2.0423\), so the margin of error and interval are:

$$ ME = 2.0423(1.4142) \approx2.8882\text{ minutes} $$
$$ 3\pm2.8882 = (0.1118,\ 5.8882)\text{ minutes} $$

The width is \(5.8882-0.1118=5.7764\) minutes, or about 5.776 minutes. For the hypothetical comparison with 36 users in each group, using the same standard deviations:

$$ SE = \sqrt{\frac{16}{36}+\frac{16}{36}} = \sqrt{\frac{8}{9}} \approx0.9428\text{ minutes} $$

The Welch degrees of freedom for equal sample sizes and equal standard deviations are 70. For 95% confidence, \(t^*\approx1.9944\), so the hypothetical width is:

$$ \text{Width} = 2(1.9944)(0.9428) \approx3.761\text{ minutes} $$

For this larger-sample comparison, the 10% condition would require at least 360 users in each population if sampling without replacement. The larger sample sizes also make the sample-size condition less concerning, but the data must still be checked for distributional features that would undermine the method.

Conclude: We are 95% confident that the true mean completion time for version 1 minus the true mean completion time for version 2 is between about 0.112 and 5.888 minutes. Under the stated comparison assumptions, increasing both sample sizes from 16 to 36 per group reduces the interval width from about 5.776 to 3.761 minutes. Its center remains 3 minutes because the sample means were held fixed.

Common Mistakes and AP Exam Tips

  • Calling the standard error the interval width: The width is \(2t^*SE\), not just \(SE\) or \(t^*SE\). The latter is the margin of error.
  • Changing several quantities without saying so: A comparison of widths is hard to interpret unless you state what stays fixed. To show the effect of \(s_1\), hold the sample sizes, confidence level, and \(s_2\) fixed.
  • Ignoring the other group’s contribution: Increasing \(s_1\) affects \(s_1^2/n_1\), while \(s_2^2/n_2\) remains in the standard-error formula. Show both terms in your calculation.
  • Assuming the critical value never changes: \(t^*\) depends on the confidence level and degrees of freedom. When sample sizes or variance contributions change, recompute the Welch degrees of freedom and critical value for a numerical interval-width comparison.
  • Saying a wider interval has a different center: Width concerns the distance between endpoints. A change in sample means affects the center; changes in variability or sample size affect the margin of error.
  • Presenting a controlled comparison as a real study result: If you hold standard deviations fixed while changing sample sizes, explain that this isolates a mathematical effect. Actual samples can have different standard deviations.

For full credit, show the standard error formula with both separate contributions, identify the relevant \(t^*\), and calculate the margin of error or total width as requested. Explain the direction of the change in context, and avoid claiming that a comparison proves anything about real populations beyond its stated assumptions.

Key takeaway: At a fixed confidence level, larger \(s_1\) or \(s_2\) increases a variance contribution and generally widens a two-sample t interval. Larger sample sizes reduce the variance contributions and generally narrow it. Calculate \(2t^*SE\) to compare actual widths.

Check Your Understanding

Use the width formula and state what is held fixed in each comparison.

  1. With \(n_1=n_2=25\), \(s_1=s_2=5\), and \(t^*=2.01\), calculate the standard error and 95% interval width.
  2. If \(s_1\) increases while \(n_1\), \(n_2\), and \(s_2\) stay fixed, which term in the standard-error formula changes, and in what direction?
  3. With both standard deviations fixed, why does increasing both sample sizes generally reduce the interval width?
  4. What is the difference between a confidence interval’s margin of error and its full width?
  5. Why should you recalculate \(t^*\) when comparing actual interval widths after changing sample sizes?