One Interval Can Decide a Two-Sided Test
A two-sided t test asks whether a population mean differs from a specified value, in either direction. A 95% t interval for that same population mean can answer the corresponding test question: check whether the null value falls inside the interval. If it does, fail to reject the null hypothesis at \(\alpha=0.05\). If it does not, reject it.
This shortcut works only when the test and interval use the same sample, the same one-sample t procedure, and the same conditions. It also depends on matching a two-sided test at \(\alpha=0.05\) with a 95% confidence interval. A 95% interval does not automatically decide every test about a mean; the confidence level and test significance level must correspond.
Why the Decisions Match
As in “The One-Sample t Test Statistic,” the test statistic measures how many standard errors the sample mean is from the null value. The 95% t interval extends a critical number of standard errors in both directions from the sample mean. Both calculations use \(s/\sqrt{n}\) and \(df=n-1\).
For a two-sided test at \(\alpha=0.05\), the rejection regions place 0.025 in each tail of the t distribution. The 95% interval uses the critical value \(t^*\) that leaves the same 0.025 in each tail. So the test rejects when the absolute value of the observed statistic is greater than the critical value, while the interval excludes \(\mu_0\) under exactly the same circumstances.
If \(\mu_0\) is inside the interval, then the observed test statistic is not far enough from 0 to enter either rejection region. The test decision is to fail to reject \(H_0\). This does not prove that \(\mu=\mu_0\); it means the sample does not provide convincing evidence of a difference at the 0.05 level.
If the null value is exactly at an endpoint, the test statistic is exactly at the corresponding critical value. With the usual rule to reject when the p-value is less than or equal to \(\alpha\), this boundary case is a rejection. In practice, calculator rounding can make an endpoint look equal to the null value when the unrounded values are slightly different, so avoid making a boundary decision from heavily rounded endpoints.
A Practical Decision Process
First check that the interval is appropriate for the data. The one-sample t conditions discussed in “Writing the Full Four-Step t Interval Solution” and “Why Conditions Matter in Mean Inference” still matter here: a shortcut for making the decision does not remove the need to justify the procedure. Then calculate the interval with \(df=n-1\), using unrounded values where possible. Finally, compare the stated null value with the two endpoints.
Define \(\mu\), write \(H_0:\mu=\mu_0\) and \(H_a:\mu\ne\mu_0\), and identify \(\alpha=0.05\).
Verify the Random, 10%, and Normal/Large Sample conditions for one-sample mean inference.
Use \(\bar{x}\pm t^*(s/\sqrt{n})\) with \(df=n-1\) and a 95% confidence level.
If \(\mu_0\) is outside the interval, reject \(H_0\); if it is inside, fail to reject \(H_0\). State what the result says about the population mean in context.
On a TI-84 or similar calculator, a 95% TInterval in Stats mode can use \(\bar{x}\), \(s\), and \(n\); alternatively, find the critical value with \(\operatorname{invT}(0.975,n-1)\) and calculate the endpoints. The 0.975 input leaves 0.025 to the left of the positive critical value, giving 0.025 in each tail overall. A two-sided T-Test on the same summary statistics should give the matching reject-or-fail-to-reject decision.
Worked Examples
Worked Example: Mean Repair Time Compared With a Standard
A fictional bicycle shop randomly selects 16 repair jobs from 160 jobs completed during a period. The repair times have a sample mean of \(\bar{x}=52\) minutes and a sample standard deviation of \(s=8\) minutes. The population distribution of repair times is approximately Normal. Use a 95% t interval to decide the two-sided test of whether the true mean repair time differs from 48 minutes, at \(\alpha=0.05\).
State. Let \(\mu\) be the true mean repair time, in minutes, for all 160 jobs completed during the period. The hypotheses are \(H_0:\mu=48\) minutes and \(H_a:\mu\ne48\) minutes. Use \(\alpha=0.05\).
Plan and check conditions. Use a one-sample t procedure for a population mean. The jobs were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(160)=16\), and the sample size is \(n=16\), so the sample is no more than 10% of the population. Since \(n=16<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. Use \(df=16-1=15\).
Do. The standard error is \(8/\sqrt{16}=2\) minutes. For 95% confidence with 15 degrees of freedom, \(t^*\approx2.131\). The interval is
The null value, 48 minutes, is inside the interval. Therefore, fail to reject \(H_0\) at \(\alpha=0.05\). As a check using the test statistic, \(t=(52-48)/2=2.00\), which is less than the two-sided critical value \(2.131\) in absolute value.
Conclude. The sample does not provide convincing evidence that the true mean repair time for these jobs differs from 48 minutes.
Worked Example: Battery Runtime Compared With a Claim
A fictional electronics tester randomly selects 25 batteries from a production lot of 500. Their mean runtime is \(\bar{x}=106\) minutes, and their sample standard deviation is \(s=10\) minutes. The runtime distribution is approximately Normal. Use a 95% t interval to decide whether the true mean runtime differs from the 100-minute claim, at \(\alpha=0.05\).
State. Let \(\mu\) be the true mean runtime, in minutes, of batteries in this production lot. Test \(H_0:\mu=100\) minutes against \(H_a:\mu\ne100\) minutes at \(\alpha=0.05\).
Plan and check conditions. Use a one-sample t procedure. Random selection supports the Random condition. For the 10% condition, \(0.10(500)=50\), and \(25\leq50\), so the sample is no more than 10% of the lot. Since \(n=25<30\), the large-sample route is not met; the stated approximately Normal distribution supports the Normal/Large Sample condition. The degrees of freedom are \(25-1=24\).
Do. The standard error is \(10/\sqrt{25}=2\) minutes. The 95% critical value for 24 degrees of freedom is \(t^*\approx2.064\):
The null value of 100 minutes is below the interval, so it is outside. Reject \(H_0\) at \(\alpha=0.05\). The equivalent test-statistic check gives \(t=(106-100)/2=3.00\); since \(|3.00|>2.064\), the statistic falls in a two-sided rejection region.
Conclude. The sample provides convincing evidence that the true mean runtime of batteries in this production lot differs from 100 minutes. The interval is entirely above 100, so the sample’s estimated mean is higher than the claim; the two-sided test itself addresses whether the mean differs in either direction.
Worked Example: Apartment Electricity Use Near a Reference Value
A fictional housing cooperative randomly selects 12 apartments from 180 apartments. The sample mean monthly electricity use is \(\bar{x}=74\) kilowatt-hours, with \(s=6\) kilowatt-hours. The population distribution is approximately Normal. Use a 95% t interval to decide whether the true mean differs from 78 kilowatt-hours, at \(\alpha=0.05\).
State. Let \(\mu\) be the true mean monthly electricity use, in kilowatt-hours, for all apartments in the cooperative. The hypotheses are \(H_0:\mu=78\) kilowatt-hours and \(H_a:\mu\ne78\) kilowatt-hours.
Plan and check conditions. Use a one-sample t procedure for a population mean. The apartments were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(180)=18\), and \(12\leq18\), so the sample is no more than 10% of the population. Because \(n=12<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. Use \(df=12-1=11\).
Do. The standard error is \(6/\sqrt{12}\approx1.732\) kilowatt-hours. For 95% confidence with 11 degrees of freedom, \(t^*\approx2.201\). Thus,
The null value of 78 kilowatt-hours is above the upper endpoint, so it is outside the interval. Reject \(H_0\) at \(\alpha=0.05\). The test-statistic check is \(t=(74-78)/(6/\sqrt{12})\approx-2.309\), whose absolute value is greater than the 95% critical value \(2.201\).
Conclude. The sample provides convincing evidence that the true mean monthly electricity use for apartments in the cooperative differs from 78 kilowatt-hours. The sample mean is lower than 78, but the stated alternative was two-sided.
What the Interval Decision Does—and Does Not—Say
A 95% interval contains all null mean values that would not be rejected by the corresponding two-sided t test at \(\alpha=0.05\), apart from boundary and rounding details. For instance, after calculating one interval, you could compare more than one proposed reference mean with its endpoints. A value inside would lead to failing to reject the matching test; a value outside would lead to rejecting it. Each comparison uses the same sample and the same conditions.
This does not mean every value inside the interval is equally plausible, or that there is a 95% probability the fixed population mean is inside this particular interval. As covered in “Common Errors in Interpreting t Intervals,” the confidence level describes the long-run success rate of the interval method. For the test decision here, the practical question is simply whether the specified null value is included.
Common Mistakes and AP Exam Tips
- Using the wrong confidence level. A 95% interval matches a two-sided test at \(\alpha=0.05\), not a one-sided test at that level. A different \(\alpha\) requires a matching confidence level.
- Comparing the sample mean with the null value instead of checking the interval. A sample mean can be above or below \(\mu_0\) without being far enough away to reject. Compare \(\mu_0\) with both interval endpoints.
- Reversing the decision. A null value outside the interval means reject \(H_0\); a null value inside means fail to reject \(H_0\).
- Ignoring conditions because the interval is already calculated. A calculator output does not establish that the inference is justified. Check randomness, independence (including the 10% condition when relevant), and the Normal/Large Sample condition.
- Calling a failure to reject proof that the null is true. Say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence of a difference. Do not say that the mean has been proved equal to the null value.
- Making a one-sided claim from a two-sided test. If the interval excludes \(\mu_0\), the test supports a difference. The sample mean shows which direction the observed difference takes, but the alternative \(H_a:\mu\ne\mu_0\) did not specify a direction in advance.
Check Your Understanding
Assume each interval is a valid 95% one-sample t interval and the test uses \(\alpha=0.05\).
- A 95% interval for a population mean is \((18.4,\ 24.7)\). For \(H_0:\mu=20\) versus \(H_a:\mu\ne20\), what is the test decision?
- A 95% interval is \((31.2,\ 35.6)\), and the null value is 36. Describe the decision and the evidence in general terms.
- A sample has \(n=20\), \(\bar{x}=44\), and \(s=10\). What are the degrees of freedom and standard error for its one-sample t interval?
- Why does a 95% t interval match a two-sided test at \(\alpha=0.05\), but not a one-sided test at \(\alpha=0.05\)?
- An interval contains the null value. Explain why “fail to reject” is appropriate but “prove the null is true” is not.