Zero as the Reference for No Difference
A two-sample t interval estimates the difference between two population means. As in “Reading the Sign of the Interval for a Difference,” the subtraction order determines what positive and negative values mean. This tutorial uses the interval to answer a related question: Do the data provide evidence that the population means differ?
For a difference defined as \(\mu_1-\mu_2\), no difference between the population means corresponds to \(\mu_1-\mu_2=0\). So zero is the reference value. If zero is outside the interval, the interval’s plausible values are all nonzero. If zero is inside the interval, no difference remains among the plausible values.
This is a way to judge evidence, not a new calculation. Construct the interval using the unpooled two-sample t procedure from “The Form of a Two-Sample t Interval,” and then compare its endpoints with zero. Keep the interval’s confidence level, parameter, subtraction order, and units in view.
What the Interval Says About a Difference
For a 95% confidence interval, the matching two-sided test uses a significance level of \(\alpha=0.05\). If the 95% interval excludes zero, the data would lead to rejecting \(H_0:\mu_1-\mu_2=0\) in that matching two-sided test. If the interval includes zero, the data would lead to failing to reject that null hypothesis. The confidence-interval approach is often a direct way to communicate the evidence because it shows not only whether zero is plausible, but also which nonzero differences are plausible.
The matching matters: the interval and test must use the same two-sided question and corresponding confidence level and significance level. A 95% interval does not, by itself, answer a one-sided question at the 5% level. Nor should you treat the result as a test at a different significance level without checking the corresponding interval.
When zero is outside the interval, its position also tells you the direction of the difference. An interval entirely above zero suggests \(\mu_1>\mu_2\); an interval entirely below zero suggests \(\mu_1<\mu_2\). Describe the size of the difference in context by interpreting both endpoints. When zero is inside, the interval includes both zero and nonzero differences, so the data do not establish a clear direction at that confidence level.
- Entirely above zero: The interval suggests that population 1 has a higher mean than population 2.
- Entirely below zero: The interval suggests that population 1 has a lower mean than population 2.
- Includes zero: Zero is a plausible difference, so the interval does not provide convincing evidence of a difference at the corresponding two-sided significance level.
“Includes zero” does not mean the population means have been shown to be equal. It means that the interval includes no difference along with other plausible differences. Those nonzero possibilities may be small or substantial; read the endpoints to understand what the interval allows.
Likewise, excluding zero does not automatically mean the difference matters in practice. An interval can suggest a difference while containing only small differences, or it can suggest a difference that might be important in context. Statistical evidence and practical importance are separate judgments.
Worked Examples
Worked Example: An Interval Entirely Above Zero
In an invented study, researchers take independent random samples of 20 adults from each of two communities and record minutes spent walking per day. Community 1 has a sample mean of 18.6 minutes and a sample standard deviation of 4.0 minutes. Community 2 has a sample mean of 15.0 minutes and a sample standard deviation of 4.0 minutes. Both populations are large compared with their sample sizes, and the sample distributions are roughly symmetric with no apparent outliers. Construct and interpret a 95% confidence interval for \(\mu_1-\mu_2\), then decide whether it suggests a difference in population mean walking time.
State: Let \(\mu_1\) be the true mean daily walking time for adults in community 1, and let \(\mu_2\) be the true mean daily walking time for adults in community 2. The parameter is \(\mu_1-\mu_2\), measured in minutes per day.
Plan: Use an unpooled two-sample t interval. The adults in the two samples are different people, so the samples are independent and unpaired. Each sample is a random sample. The populations are large compared with the sample sizes, so the 10% condition is met for each sample. With 20 observations in each group, the roughly symmetric distributions and lack of apparent outliers support using the t procedure.
Do: The sample mean difference is \(18.6-15.0=3.6\) minutes per day. The standard error is calculated separately from the two groups’ variance contributions:
Welch’s degrees of freedom are 38. For a 95% confidence interval, \(t^*\approx2.024\). The margin of error is \(2.024(1.265)\approx2.561\) minutes per day. Therefore:
Conclude: We are 95% confident that the mean daily walking time in community 1 is about 1.04 to 6.16 minutes greater than in community 2. The interval is entirely above zero, so it suggests a difference between the population means at the corresponding two-sided 5% significance level, with community 1’s mean higher. The random samples support generalizing to the two communities, subject to the study’s design and conditions. Because this is not a randomized experiment, the result does not establish that community membership causes a difference in walking time.
Worked Example: An Interval That Includes Zero
In an invented survey, independent random samples of 16 students from each of two large schools report the number of minutes they spend reading for pleasure on a typical weekday. School 1’s sample mean is 42 minutes, and school 2’s is 40 minutes. Both sample standard deviations are 8 minutes. The distributions are roughly symmetric without apparent outliers. Construct a 95% confidence interval for \(\mu_1-\mu_2\) and use it to assess whether the data suggest a difference in population mean reading time.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean weekday reading times for students at school 1 and school 2, respectively. The parameter \(\mu_1-\mu_2\) is measured in minutes per weekday.
Plan: Use an unpooled two-sample t interval. The students come from independent random samples, and each population is large enough that its sample is less than 10% of the population. The samples are unpaired, and the roughly symmetric distributions without apparent outliers support the t procedure for these sample sizes.
Do: The sample mean difference is \(42-40=2\) minutes. The standard error is:
Welch’s degrees of freedom are 30. For a 95% confidence interval, \(t^*\approx2.042\). The margin of error is \(2.042(2.828)\approx5.776\) minutes. Thus:
Conclude: We are 95% confident that the mean weekday reading time at school 1 is between about 3.78 minutes lower and 7.78 minutes higher than the mean at school 2. Because zero is in the interval, these data do not provide convincing evidence of a difference between the population means at the corresponding two-sided 5% significance level. The interval does not prove the means are equal: it also includes nonzero differences in either direction.
Worked Example: An Interval Entirely Below Zero
In an invented environmental survey, independent random samples of 25 homes in each of two large regions are used to estimate weekly household water use in hundreds of liters. Region 1 has a sample mean of 31.8 and a sample standard deviation of 5.0. Region 2 has a sample mean of 36.0 and a sample standard deviation of 5.0. Both sample distributions are reasonably symmetric and show no strong outliers. Construct and interpret a 95% confidence interval for \(\mu_1-\mu_2\), and decide what it suggests.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean weekly household water use in regions 1 and 2, respectively. The difference \(\mu_1-\mu_2\) is measured in hundreds of liters per household per week.
Plan: Use an unpooled two-sample t interval. The homes are from independent random samples, not matched pairs. The regions contain far more than 250 homes each, so the 10% condition is met. The sample distributions are reasonably symmetric without strong outliers, supporting the procedure for samples of 25 homes each.
Do: The sample mean difference is \(31.8-36.0=-4.2\) hundreds of liters. The standard error is:
Welch’s degrees of freedom are 48. For a 95% confidence interval, \(t^*\approx2.011\). The margin of error is \(2.011(1.414)\approx2.843\) hundreds of liters. The interval is:
Conclude: We are 95% confident that region 1’s mean weekly household water use is about 1.357 to 7.043 hundreds of liters lower than region 2’s mean. Equivalently, the interval for \(\mu_1-\mu_2\) is entirely below zero, suggesting a difference in population means at the corresponding two-sided 5% significance level, with region 1’s mean lower. The interval describes a difference in means, not a difference that applies to every household.
Common Mistakes and AP Exam Tips
- Calling an interval that includes zero proof of equality: Zero is one plausible difference, but nonzero values may also be plausible. Say the data do not provide convincing evidence of a difference at the corresponding level; do not say the population means are proven equal.
- Ignoring the subtraction order: For an interval for \(\mu_1-\mu_2\), a positive endpoint means population 1’s mean is greater than population 2’s mean by that amount. A negative difference means population 1’s mean is lower. Interpret both endpoints using the order stated.
- Reporting only whether zero is included: Explain the direction and size of the plausible differences when the interval excludes zero. If it includes zero, describe the nonzero values it also allows rather than stopping at “there is no difference.”
- Mixing confidence levels and significance levels: A 95% interval corresponds to a two-sided test at the 5% level. Do not use that equivalence to claim a result for a one-sided test or a different significance level without a matching analysis.
- Equating evidence with importance: Excluding zero suggests evidence of a difference, but the endpoints and context determine whether the plausible sizes could matter in practice.
- Making causal claims from random samples alone: Random sampling can support generalization to a population. Cause-and-effect conclusions require an appropriate randomized experiment, not simply an interval that excludes zero.
For a full-credit response, name the parameter and subtraction order, interpret the interval in context with units, and explicitly compare the interval with zero. Then state whether the interval suggests a difference at the matching two-sided significance level. If zero is included, explain that this is not proof of equality; if zero is excluded, identify the direction and plausible size of the difference.
Check Your Understanding
For each question, use the stated difference in population means and interpret the interval in context.
- A 95% interval for mean battery life in brand A minus mean battery life in brand B is \((1.2,\ 3.8)\) hours. Does it suggest a difference? What is the direction?
- A 95% interval for the mean commute time in town X minus that in town Y is \((-2.5,\ 4.1)\) minutes. What does the interval say about evidence of a difference, and what does it not prove?
- Why does a 95% confidence interval correspond to a two-sided test at the 5% significance level only when the hypotheses and reference value match?
- An interval excludes zero but contains only small differences in a context where even a small change may not matter. What two separate judgments should you make?
- A random sample study finds an interval entirely below zero for the difference in mean daily screen time between two groups. What does the sign tell you, and does it alone establish a cause of the difference?