One Difference, Two Ways to Make a Decision
A two-sample t test asks whether the data provide convincing evidence of a difference between two population means. A two-sample t confidence interval estimates the size and direction of that difference. When the question is whether the means differ, the interval can also make the test decision: check whether it includes zero.
In “Does an Interval for \(\mu_d\) Containing Zero Matter,” we used this idea for paired data. Here we apply the same decision logic to independent samples and an interval for \(\mu_1-\mu_2\). The key is to match the confidence level to the test’s significance level and to use the same unpooled two-sample t procedure for both.
Why the Decisions Match
The interval and test use the same observed difference in sample means, \(\bar{x}_1-\bar{x}_2\), and the same estimated standard error, \(SE_{\bar{x}_1-\bar{x}_2}\). They also use the same degrees of freedom. The interval places a margin of error on either side of the observed difference. For a two-sided test, the test compares the absolute value of the test statistic with the matching critical t-value.
The critical value \(t^*\) comes from the same t distribution and degrees of freedom used by the test. For a 5% two-sided test, for example, use a 95% interval. The interval excludes zero precisely when the observed difference is farther than \(t^*\) standard errors from zero. That is the same evidence represented by a test statistic beyond either critical value, \(t^*\) or \(-t^*\).
This connection is about a two-sided test. A one-sided test does not generally have its decision represented by checking whether a two-sided \(1-\alpha\) interval excludes zero. As covered in “Choosing One-Sided or Two-Sided for Two Means,” the alternative hypothesis determines whether the question concerns a difference in either direction or a difference in one specified direction.
An interval gives more information than the reject-or-fail-to-reject decision alone: it shows plausible values for the population mean difference, with units. If the entire interval is positive, the plausible differences favor a larger mean for population 1. If the entire interval is negative, they favor a smaller mean for population 1. If the interval includes zero, the data do not provide convincing evidence of a difference at the matching test level; that does not prove the population means are equal.
A Practical Interval-to-Test Process
For a two-sided test at significance level \(\alpha\), identify the interval with confidence level \(1-\alpha\).
For \(H_0:\mu_1-\mu_2=0\), determine whether zero lies inside the interval.
If zero is outside, reject \(H_0\); if zero is inside, fail to reject \(H_0\). Keep the endpoint convention in mind if zero is exactly an endpoint.
Describe the evidence about the difference between the population means, keeping the order \(\mu_1-\mu_2\) and the variable’s units clear.
This is a decision shortcut, not a replacement for checking whether the procedure is appropriate. In a complete inference response, identify the populations and parameter, check the design and two-sample t conditions, and interpret the result in context. The interval and test must be based on the same data, group order, unpooled standard error, and degrees of freedom to match.
Worked Examples
Worked Example: Interval Excludes Zero
A hypothetical experiment randomly assigns 16 greenhouse sensors to Firmware A and 16 to Firmware B. Researchers measure each sensor’s average daily operating time, in hours, during a fixed test period. Firmware A has \(\bar{x}_1=52\) hours and \(s_1=4\) hours; Firmware B has \(\bar{x}_2=47\) hours and \(s_2=4\) hours. Is there evidence that the population mean operating times differ? Use \(\alpha=0.05\).
State: Let \(\mu_1\) and \(\mu_2\) be the true mean operating times for sensors using Firmware A and Firmware B, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).
Plan: Use an unpooled two-sample t procedure. The experiment randomly assigns sensors to the two settings, and each sensor receives only one firmware, so the treatment groups are independent. The sample sizes of 16 are small enough that the response distributions should be checked; assume plots show no strong skewness or outliers in either group. These design and distribution features support the procedure. No 10% condition is needed for sampling without replacement because these are randomized treatment groups, not samples drawn from a finite population.
Do: The observed difference is \(52-47=5\) hours. The standard error is
The Welch degrees of freedom are 30. A 95% interval uses \(t^*\approx2.0423\), so the margin of error is \(2.0423(1.4142)\approx2.8882\) hours. The interval is
Zero is outside this interval. To check the matching test decision directly, \(t=5/1.4142\approx3.5355\), which is greater than the 95% interval’s critical value of 2.0423 in absolute value. The two-sided p-value is approximately 0.0013, rounded to four decimal places, so the test also rejects \(H_0\) at \(\alpha=0.05\).
Conclude: There is convincing evidence that the true mean operating times differ between sensors using Firmware A and Firmware B. The interval suggests that the mean for Firmware A is greater, with plausible differences from about 2.11 to 7.89 hours. This is a conclusion about population means, not a claim that every sensor using Firmware A operates longer.
Worked Example: Interval Includes Zero
A hypothetical study takes independent random samples of 12 garden plots from each of two large neighborhoods and measures water absorbed by the soil after a standard test, in centimeters. Neighborhood 1 has a sample mean of 6.5 centimeters and a sample standard deviation of 3 centimeters. Neighborhood 2 has a sample mean of 5.5 centimeters and a sample standard deviation of 3 centimeters. At \(\alpha=0.05\), do the data provide convincing evidence of different population means?
Let \(\mu_1\) and \(\mu_2\) be the true mean water absorption for plots in Neighborhoods 1 and 2. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
The plots were selected using independent random samples, and no plot appears in both groups, so the samples are independent. Each sample is less than 10% of its neighborhood’s plots, satisfying the 10% condition. With 12 plots per group, inspect each distribution; assume neither shows severe skewness or outliers. These conditions support an unpooled two-sample t procedure.
The sample mean difference is \(6.5-5.5=1.0\) centimeter. The standard error is
The Welch degrees of freedom are 22, and the 95% critical value is \(t^*\approx2.0739\). The margin of error is \(2.0739(1.2247)\approx2.5400\) centimeters, giving
Zero is inside the interval. The matching test statistic is \(t=1.0/1.2247\approx0.8165\), which is less than 2.0739 in absolute value. The two-sided test therefore fails to reject \(H_0\) at \(\alpha=0.05\).
There is not convincing evidence that the true mean water absorption differs between plots in the two neighborhoods. The interval includes both negative and positive differences, as well as zero. It does not demonstrate equality; it shows that this study does not establish a difference at the chosen level.
Worked Example: Confirming a Corrected Test Result
In a hypothetical randomized experiment, 12 routers are assigned to each of two cooling designs. The response is average operating temperature in degrees Celsius. Design 1 has a sample mean of 31.5 degrees and Design 2 has a sample mean of 27.0 degrees. Both sample standard deviations are 3 degrees. Researchers test whether the population means differ at \(\alpha=0.05\).
Let \(\mu_1\) and \(\mu_2\) be the true mean operating temperatures for routers using Designs 1 and 2. Use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). Random assignment and one design per router support independent groups. With 12 routers per group, assume inspection finds no severe skewness or outliers in either response distribution; these features support the two-sample t procedure.
The observed difference is \(31.5-27.0=4.5\) degrees. The standard error is
The test statistic is \(t=4.5/1.2247\approx3.6742\), with 22 Welch degrees of freedom. The two-sided p-value is approximately 0.00133043 (about 0.0013 rounded to four decimal places). Since \(0.00133043<0.05\), reject \(H_0\).
For the matching 95% interval, use \(t^*\approx2.0739\). Its margin of error is \(2.0739(1.2247)\approx2.5400\) degrees, so the interval is
The interval excludes zero, matching the test’s rejection. There is convincing evidence that the true mean operating temperatures differ between the two designs. The interval indicates that Design 1’s population mean is higher, with plausible differences of about 1.96 to 7.04 degrees.
Common Mistakes and AP Exam Tips
The interval-to-test connection is simple once the levels and null value are matched. These are the details that most often lead to a mismatched decision or an incomplete interpretation.
- Using the wrong confidence level: A 95% interval matches a two-sided test at \(\alpha=0.05\), not at \(\alpha=0.10\). Always use confidence level \(1-\alpha\).
- Using an interval for the wrong difference: If the test is about \(\mu_1-\mu_2\), use the interval in that same order. Reversing the group order reverses the signs and switches the endpoints.
- Checking whether the interval is “mostly” positive: The decision rule is whether zero is inside the interval, not whether most of its length lies on one side of zero.
- Treating an interval that includes zero as proof of equality: The appropriate conclusion is that there is not convincing evidence of a difference at the matching significance level. Failing to reject \(H_0\) does not establish that it is true.
- Using a one-sided alternative with the two-sided shortcut: This connection applies to a two-sided test. A claim about which mean is greater requires following the one-sided test procedure and its matching tail area.
- Forgetting the design and conditions: The interval’s decision does not verify that the inference method is appropriate. A full-credit response still addresses random sampling or assignment, independence, the 10% condition when applicable, and the distribution condition.
At the exact boundary, a closed interval can have zero as an endpoint. The usual test rule rejects when \(p\leq\alpha\), whereas a closed interval with zero at an endpoint technically includes zero. Exact equality is unusual with unrounded data, but rounded calculator output can make it appear to happen. When a result is near the boundary, use unrounded values and state the decision rule being applied rather than forcing a conclusion from rounded endpoints.
Check Your Understanding
For each question, use a two-sided test and its matching confidence interval.
- What confidence level matches a two-sided test with \(\alpha=0.10\)?
- A 95% interval for \(\mu_1-\mu_2\) is \((-4.2,-0.8)\) minutes. What decision matches a two-sided test at \(\alpha=0.05\), and what direction does the interval suggest?
- A 99% interval for \(\mu_1-\mu_2\) is \((-1.5,3.1)\) kilograms. What is the matching decision at \(\alpha=0.01\), and does the interval establish equality?
- Why must the interval and test use the same group order and degrees of freedom for their decisions to match?
- A two-sided test has a p-value reported as 0.050 after rounding, and the matching interval endpoint is reported as zero. What should you check before deciding that the results conflict?