Overlap Does Not Answer the Difference Question
Suppose you want to know whether two populations have different means. You might see two separate confidence intervals—one for each mean—and compare them visually. If they overlap, it can be tempting to conclude that there is no convincing evidence of a difference. That conclusion does not follow from the overlap.
The intervals answer different questions. An interval for \(\mu_1\) estimates the true mean for population 1; an interval for \(\mu_2\) estimates the true mean for population 2. The question “How far apart are the population means?” concerns \(\mu_1-\mu_2\). To address it, use an interval constructed directly for that difference, as covered in “The Form of a Two-Sample t Interval.”
Why the Intervals Can Tell Different Stories
Each separate interval has its own standard error and margin of error. For example, the standard error for \(\bar{x}_1\) is \(s_1/\sqrt{n_1}\), while the standard error for \(\bar{x}_2\) is \(s_2/\sqrt{n_2}\). Their intervals describe uncertainty about their respective population means.
For independent samples, the standard error for the difference in sample means combines the two samples’ variance contributions. As introduced in “Standard Error for a Difference in Means,” it is not the sum of the two individual standard errors.
The unpooled two-sample t interval uses this standard error and an appropriate \(t^*\) for the difference. In contrast, the individual intervals use their own standard errors and critical values. Their widths are not designed to provide a confidence interval for \(\mu_1-\mu_2\).
Here is a useful intuition. If the critical values were about the same, two separate intervals would overlap when the observed difference in means is smaller than the sum of their margins of error. But the margin of error for the difference is based on the square root of the sum of the squared standard errors. For positive values \(a\) and \(b\), \(a+b\) is larger than \(\sqrt{a^2+b^2}\). So there can be a range of differences for which separate intervals overlap even though the interval for the difference excludes zero. In actual calculations, the critical values can differ, so this comparison is an explanation—not a substitute for calculating the interval.
As explained in “Using an Interval to Judge a Difference Between Groups,” an interval for \(\mu_1-\mu_2\) that excludes zero suggests a difference in population means at the corresponding two-sided significance level. An interval that includes zero does not prove the means are equal. Either way, the interval for the difference—not overlap between two individual intervals—is the relevant result.
Worked Examples
Worked Example: Overlap but the Difference Interval Excludes Zero
In an invented comparison of two large distribution centers, independent random samples of 16 delivery drivers from each center record the number of minutes they spend loading their vehicles each morning. Center 1 has a sample mean of 68 minutes and a sample standard deviation of 12 minutes. Center 2 has a sample mean of 79 minutes and a sample standard deviation of 12 minutes. The populations are each more than ten times the sample size, and the sample distributions are roughly symmetric without apparent outliers. Compare the two individual 95% intervals with a 95% interval for the difference.
State: Let \(\mu_1\) be the true mean loading time for drivers at center 1, and let \(\mu_2\) be the true mean loading time for drivers at center 2. The parameter of interest is \(\mu_1-\mu_2\), measured in minutes.
Plan: Use an unpooled two-sample t interval for \(\mu_1-\mu_2\). The samples are random, come from different drivers, and are independent rather than paired. Each sample is less than 10% of its population, satisfying the 10% condition. The stated sample shapes support the t procedure. Individual intervals can illustrate overlap, but the difference interval will answer the question about how far apart the population means are.
Do: Each individual interval has standard error \(12/\sqrt{16}=3\) minutes. With \(df=15\), \(t^*\approx2.131\), so each margin of error is \(2.131(3)\approx6.394\) minutes. The intervals are approximately \((61.606,\ 74.394)\) minutes for center 1 and \((72.606,\ 85.394)\) minutes for center 2. They overlap from about 72.606 to 74.394 minutes.
The sample mean difference is \(68-79=-11\) minutes. Its standard error is:
Welch’s degrees of freedom are 30, giving \(t^*\approx2.042\) for a 95% interval. The margin of error is \(2.042(4.243)\approx8.665\) minutes. Thus, the interval is:
Conclude: We are 95% confident that the mean loading time at center 1 is about 2.335 to 19.665 minutes less than the mean loading time at center 2. The interval for \(\mu_1-\mu_2\) excludes zero, suggesting a difference between the population means at the corresponding two-sided 5% significance level. The overlap between the separate 95% intervals does not contradict that conclusion.
Worked Example: Overlap and a Difference Interval That Includes Zero
In an invented study of two large community gardens, independent random samples of 12 raised beds from each garden are used to measure the weekly mass of harvested tomatoes, in kilograms. Garden 1 has a sample mean of 42 kg and a sample standard deviation of 10 kg. Garden 2 has a sample mean of 47 kg and a sample standard deviation of 10 kg. Each garden has more than ten times as many beds as sampled, and the sample distributions are reasonably symmetric without strong outliers. Examine whether the visual overlap alone gives the answer.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean weekly tomato harvests per bed for gardens 1 and 2. The parameter \(\mu_1-\mu_2\) is measured in kilograms per bed.
Plan: Use an unpooled two-sample t interval. The samples are random samples of different beds, so there is no pairing. Each sample is less than 10% of its garden’s beds, and the stated sample shapes support the t procedure.
Do: For each individual interval, the standard error is \(10/\sqrt{12}\approx2.887\) kg. With \(df=11\), \(t^*\approx2.201\), giving a margin of error of \(2.201(2.887)\approx6.354\) kg. The individual intervals are approximately \((35.646,\ 48.354)\) kg for garden 1 and \((40.646,\ 53.354)\) kg for garden 2. They overlap substantially.
For the difference, the sample estimate is \(42-47=-5\) kg. The standard error is:
Welch’s degrees of freedom are 22, giving \(t^*\approx2.074\). The margin of error is \(2.074(4.082)\approx8.467\) kg. Therefore:
Conclude: We are 95% confident that the mean weekly harvest per bed in garden 1 is between about 13.467 kg less and 3.467 kg more than in garden 2. The interval includes zero, so the data do not provide convincing evidence of a difference between the population means at the corresponding two-sided 5% significance level. This conclusion comes from the interval for the difference; overlap of the individual intervals did not determine it. The result also does not prove that the two means are equal.
Worked Example: Separate Intervals That Do Not Overlap
In an invented quality-control comparison, independent random samples of 20 rechargeable cells from each of two large production lines are tested for operating time, in hours. Line 1 has a sample mean of 31 hours and a sample standard deviation of 8 hours. Line 2 has a sample mean of 41 hours and a sample standard deviation of 8 hours. Each line’s population is more than ten times the sample size, and the sample distributions are roughly symmetric with no strong outliers. Calculate the difference interval rather than relying only on the visual comparison.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean operating times for cells from lines 1 and 2. The parameter \(\mu_1-\mu_2\) is measured in hours.
Plan: Use an unpooled two-sample t interval. The random samples are independent because different cells are tested, each sample satisfies the 10% condition, and the sample shapes support the t procedure.
Do: For each individual interval, the standard error is \(8/\sqrt{20}\approx1.789\) hours. With \(df=19\), \(t^*\approx2.093\), so the margin of error is \(2.093(1.789)\approx3.744\) hours. The individual intervals are approximately \((27.256,\ 34.744)\) hours for line 1 and \((37.256,\ 44.744)\) hours for line 2. These intervals do not overlap.
The difference in sample means is \(31-41=-10\) hours. The standard error for that difference is:
Welch’s degrees of freedom are 38, giving \(t^*\approx2.024\). The margin of error is \(2.024(2.530)\approx5.121\) hours. The 95% interval is:
Conclude: We are 95% confident that the mean operating time for cells from line 1 is about 4.879 to 15.121 hours less than for cells from line 2. This difference interval excludes zero and suggests a difference between the population means at the corresponding two-sided 5% significance level. The separate intervals also do not overlap, but the interval for \(\mu_1-\mu_2\) is the direct answer to the question.
Common Mistakes and AP Exam Tips
- Concluding that overlap means “no difference”: That conclusion is not justified. Calculate an interval for \(\mu_1-\mu_2\) and check whether it includes zero.
- Using the amount of overlap as a test statistic: The overlap amount is not the standard error or margin of error for the difference. The two-sample interval uses \(\sqrt{s_1^2/n_1+s_2^2/n_2}\).
- Adding the individual margins to get a difference margin: The standard error for the difference combines variance contributions; do not add individual margins of error. Use the appropriate two-sample \(t^*\) as well.
- Treating separate intervals as intervals for the same parameter: One interval concerns \(\mu_1\), another concerns \(\mu_2\), and the difference interval concerns \(\mu_1-\mu_2\). Name the parameter you are interpreting.
- Claiming that an interval including zero proves equality: Say that the data do not provide convincing evidence of a difference at the corresponding two-sided significance level. The interval may still include meaningful nonzero differences.
- Forgetting subtraction order and units: If the interval is for \(\mu_1-\mu_2\), negative values mean the population 1 mean is lower than the population 2 mean. State the context and measurement units.
For full credit, state the parameter in context, identify the unpooled two-sample t interval, check the design and distribution conditions, show the standard error and interval, and interpret both endpoints. When the question asks about a difference, do not substitute a visual comparison of two separate intervals for the direct interval about \(\mu_1-\mu_2\).
Check Your Understanding
Use the distinction between individual means and their difference in each response.
- Two separate 95% intervals for \(\mu_1\) and \(\mu_2\) overlap. What additional interval should you calculate to assess a difference between the population means?
- For independent samples with \(s_1=6,\ n_1=9,\ s_2=8,\ n_2=16\), write the standard error expression for \(\bar{x}_1-\bar{x}_2\).
- Why is the sum of the two individual margins of error not the margin of error for the unpooled two-sample t interval?
- A 95% interval for \(\mu_1-\mu_2\) is \((1.4,\ 6.8)\) liters. Interpret its direction in context if the parameter compares mean water use in group 1 minus mean water use in group 2.
- A 95% interval for a difference in population means includes zero. Give a careful conclusion that avoids claiming the means have been proven equal.