Tutorials › AP Statistics › Connecting Confidence Intervals to Test Decisions

P-values and mean-inference conclusions · Tutorial 752 of 1000

Connecting Confidence Intervals to Test Decisions

Use a confidence interval to judge a hypothesized population mean and connect the decision to the test’s significance level, including the endpoint case.

Intermediate 11 min read

What You'll Learn

  • Match a two-sided confidence interval to a test’s significance level.
  • Decide whether a hypothesized population mean is inside or outside the interval.
  • State the correct reject-or-fail-to-reject conclusion in context.
  • Handle a hypothesized value that equals an interval endpoint.
  • Distinguish two-sided interval decisions from one-sided test decisions.

One Interval, One Test Decision

A confidence interval gives a range of plausible values for a population mean. A hypothesis test asks whether the data provide convincing evidence against a particular value of that mean. When the interval and test use the same inference procedure, the interval can make the test decision visible: check whether the hypothesized value is in the interval.

As in “Linking Two-Sample Intervals and Tests,” this connection depends on matching the interval’s confidence level to the test’s significance level. Here we focus on a one-sample t interval and a two-sided test of a population mean. A \(100(1-\alpha)\%\) two-sided t interval matches a two-sided t test at significance level \(\alpha\), provided both use the same data and conditions.

Key rule: For a two-sided test of \(H_0:\mu=\mu_0\) at level \(\alpha\), use the matching \(100(1-\alpha)\%\) confidence interval for \(\mu\). Reject \(H_0\) if \(\mu_0\) is at or beyond either endpoint. Fail to reject \(H_0\) only if \(\mu_0\) is strictly between the endpoints.

The endpoint wording matters. A confidence interval is conventionally written with its endpoints included. If the null value equals an endpoint, the test statistic is exactly at the critical value, so the two-sided p-value equals \(\alpha\). Since the test rule is to reject when \(p\le\alpha\), equality belongs with rejection—not with failure to reject.

Why the Confidence Level and Alpha Must Match

A two-sided t interval has the form sample mean plus or minus a critical value times the standard error. A two-sided test rejects when the test statistic is sufficiently far from zero in either direction. The interval and test therefore use the same critical t value when the confidence level is \(1-\alpha\).

$$ \bar{x}\pm t^*\frac{s}{\sqrt{n}} \qquad\text{and}\qquad t_{\text{obs}}=\frac{\bar{x}-\mu_0}{s/\sqrt{n}}. $$

For example, a two-sided test at \(\alpha=0.05\) matches a 95% confidence interval. A 90% interval matches a two-sided test at \(\alpha=0.10\). Do not pair a test with an interval at an unrelated confidence level and assume the interval settles that test’s decision.

1
Match the levels.
For a two-sided test at level \(\alpha\), select the \(100(1-\alpha)\%\) two-sided interval for the same population mean.
2
Locate the null value.
Compare \(\mu_0\) with the interval’s lower and upper endpoints. Use unrounded values for a decision when available.
3
Make the test decision.
If \(\mu_0\) is at or beyond an endpoint, reject \(H_0\). If it is strictly between the endpoints, fail to reject \(H_0\).
4
Conclude in context.
Describe whether the data provide convincing evidence for the alternative claim about the population mean.

The interval does not change the hypotheses or replace the need for conditions. It is another way to express the decision from the matching test. As in earlier tutorials on one-sample t procedures, identify the population mean, check the study design and t-procedure conditions, and keep the conclusion tied to the stated alternative.

Worked Example: A Hypothesized Mean Outside the Interval

Worked Example: A Hypothesized Mean Outside the Interval

A fictional greenhouse randomly selects 25 tomato plants from a large growing operation and records the mass of fruit harvested per plant, in kilograms, over a defined harvest period. The sample mean is 42 kilograms and the sample standard deviation is 10 kilograms. The sample is less than 10% of the plants in the operation, and a plot of the sample values shows no severe skewness or extreme outliers. Test whether the population mean harvest is different from 46 kilograms at \(\alpha=0.05\), using the matching confidence interval.

1
State.
Let \(\mu\) be the true mean harvest mass, in kilograms per plant, for tomato plants in this greenhouse operation during the defined harvest period. The hypotheses are \(H_0:\mu=46\) kilograms and \(H_a:\mu\ne46\) kilograms.
2
Plan and check conditions.
Use a one-sample t procedure for a population mean. The plants were randomly selected, meeting the random condition. The sample is less than 10% of the operation’s plants, supporting independence under the 10% condition. The plot shows no severe skewness or extreme outliers, so using a t procedure is reasonable. A 95% two-sided t interval matches the two-sided test at \(\alpha=0.05\).
3
Do.
The degrees of freedom are \(25-1=24\). For a 95% interval with 24 degrees of freedom, \(t^*\approx2.064\). The standard error is \(s/\sqrt{n}=10/\sqrt{25}=10/5=2\) kilograms. The interval is
\(42\pm2.064(2)=42\pm4.128\),
so the 95% confidence interval is approximately \((37.872,\ 46.128)\) kilograms. The null value, 46 kilograms, is strictly between the endpoints, so fail to reject \(H_0\) at \(\alpha=0.05\).
4
Conclude in context.
The data do not provide convincing evidence that the true mean harvest mass per plant in this greenhouse operation differs from 46 kilograms during the defined harvest period.

The interval contains values below and above 46 kilograms, so this test does not find convincing evidence of a difference at the 0.05 level. That is not proof that the population mean is exactly 46 kilograms. The interval still allows a range of plausible mean values.

Worked Example: A Hypothesized Mean Beyond an Endpoint

Worked Example: A Hypothesized Mean Beyond an Endpoint

A fictional community garden randomly samples 16 plots and measures the mass of cucumbers harvested per plot, in kilograms, during one month. The sample mean is 31 kilograms and the sample standard deviation is 8 kilograms. The sample is less than 10% of all plots in the garden, and the sample data show no severe skewness or extreme outliers. Test whether the population mean differs from 36 kilograms at \(\alpha=0.05\).

State: Let \(\mu\) be the true mean cucumber harvest mass, in kilograms per plot, for plots in this garden during the month. The hypotheses are \(H_0:\mu=36\) kilograms and \(H_a:\mu\ne36\) kilograms.

Plan and check conditions: Use a one-sample t procedure. The plots were randomly sampled, and the sample is less than 10% of all plots, supporting the random and 10% conditions. The data show no severe skewness or extreme outliers, so the t procedure is reasonable. Use the matching 95% two-sided interval for a test at \(\alpha=0.05\).

Do: The degrees of freedom are \(16-1=15\), and the critical value is \(t^*\approx2.131\). The standard error is \(8/\sqrt{16}=8/4=2\) kilograms. The interval is

$$ 31\pm2.131(2) =31\pm4.262 \approx(26.738,\ 35.262)\text{ kilograms}. $$

The null value of 36 kilograms is greater than the upper endpoint, so reject \(H_0\) at \(\alpha=0.05\). As a check, the test statistic is

$$ t_{\text{obs}} =\frac{31-36}{8/\sqrt{16}} =\frac{-5}{2} =-2.50. $$

Its absolute value exceeds the critical value \(2.131\), which gives the same decision.

Conclude: The data provide convincing evidence that the true mean cucumber harvest mass per plot in this garden during the month differs from 36 kilograms. The interval is entirely below 36 kilograms, suggesting the mean is lower; the stated two-sided test, however, is specifically evidence of a difference in either direction.

Worked Example: The Null Value at an Endpoint

Worked Example: The Null Value at an Endpoint

Consider a fictional random sample of 10 garden beds used to estimate mean plant height. Suppose the sample mean is 20 centimeters and the sample standard deviation is \(\sqrt{10}\) centimeters. The sample is less than 10% of the beds in the target population, and the sample values have no severe skewness or extreme outliers. This example sets the hypothesized mean equal to the upper endpoint of the matching 95% interval to examine the boundary rule.

State and plan: Let \(\mu\) be the true mean plant height, in centimeters, for plants in the target population of garden beds. Test \(H_0:\mu=\mu_0\) against \(H_a:\mu\ne\mu_0\) at \(\alpha=0.05\), where \(\mu_0\) is defined to be the upper endpoint of the 95% t interval. The sample is random and below 10% of the population; the data show no severe skewness or extreme outliers. These conditions support a one-sample t procedure.

Do: There are \(10-1=9\) degrees of freedom. The standard error is

$$ \frac{s}{\sqrt{n}} =\frac{\sqrt{10}}{\sqrt{10}} =1\text{ centimeter}. $$

For 9 degrees of freedom, the 95% interval uses \(t^*\approx2.262\). Its upper endpoint is \(20+t^*(1)\), or approximately 22.262 centimeters. Define \(\mu_0=20+t^*\) centimeters using the unrounded critical value. The test statistic is therefore

$$ t_{\text{obs}} =\frac{20-(20+t^*)}{1} =-t^*. $$

The statistic is exactly at the two-sided critical value, so the two-sided p-value is exactly \(\alpha=0.05\). The null value is at the interval endpoint; therefore reject \(H_0\), because the test rule rejects when \(p\le\alpha\). The displayed endpoint is rounded, so the exact equality in this example is defined using the unrounded critical value.

Conclude: At the 0.05 significance level, the data provide convincing evidence that the true mean plant height in the target population differs from this hypothesized endpoint value. This boundary example is why “inside means fail to reject” must mean strictly inside: equality with an endpoint leads to rejection under the usual test rule.

One-Sided Tests Need a Different Match

The interval rule above is for a two-sided alternative. Do not automatically use a standard two-sided confidence interval to make an exact decision for a one-sided test at the same \(\alpha\). For a one-sided test at level \(\alpha\), the matching interval is a one-sided confidence bound with confidence level \(1-\alpha\). For example, a right-tailed test asks whether \(\mu>\mu_0\); its matching lower confidence bound must be above \(\mu_0\) to reject.

A central two-sided 95% interval uses 2.5% in each tail, while a one-sided test at \(\alpha=0.05\) uses 5% in one tail. Those are different cutoffs. A central 90% two-sided interval has the same one-tail critical cutoff as a one-sided 95% bound, but it represents a two-sided test at \(\alpha=0.10\), not a two-sided test at \(\alpha=0.05\). Always match the alternative and the confidence procedure, not just the printed percentage.

Common Mistakes and AP Exam Tips

  • Pairing the wrong confidence level with alpha: A 95% two-sided interval matches a two-sided test at \(\alpha=0.05\); a 90% two-sided interval matches \(\alpha=0.10\). State the match explicitly.
  • Treating the endpoint as strictly inside: If \(\mu_0\) equals an endpoint, reject under the usual \(p\le\alpha\) rule. Failure to reject applies only when \(\mu_0\) is strictly between the endpoints.
  • Using rounded endpoints to settle a borderline decision: Rounding can make a value appear equal to an endpoint when it is slightly inside or outside. Use unrounded calculator values or the test statistic and critical value when the decision is close.
  • Claiming the interval proves the null: If the null value is inside, fail to reject; do not say the null has been proven or that there is no difference. Say the data do not provide convincing evidence for the alternative at the stated level.
  • Forgetting the direction of the conclusion: If the null value is outside a two-sided interval, the test supports a difference. Describe a lower or higher direction only when the interval’s location supports it, and keep the conclusion in context.
  • Applying the two-sided rule to a one-sided test: A one-sided alternative requires its matching one-sided bound for an exact interval decision. Do not assume a usual two-sided interval at confidence level \(1-\alpha\) is equivalent.

A full-credit response identifies the population mean and hypotheses, names the matching interval and significance level, checks the conditions, shows the interval calculation, and makes the reject-or-fail-to-reject decision using the correct endpoint rule. It then states what the data show about the population mean in context.

Key takeaway: For a two-sided test of a population mean at level \(\alpha\), use the matching \(100(1-\alpha)\%\) confidence interval. Reject if the null value is at or beyond an endpoint; fail to reject only if it is strictly between the endpoints. A null value exactly at an endpoint corresponds to \(p=\alpha\) and is rejected under the usual rule.

Check Your Understanding

Use the relationship between a two-sided interval and its matching test to answer each question.

  1. A two-sided test uses \(\alpha=0.10\). What confidence level should its matching two-sided interval have?
  2. A 95% confidence interval for \(\mu\) is \((12.4,\ 18.1)\). For a two-sided test at \(\alpha=0.05\), what decision would you make about \(H_0:\mu=19\)?
  3. A 90% interval is \((7.2,\ 11.8)\), and the null value is \(\mu_0=10\). What is the matching two-sided test’s significance level, and what decision follows?
  4. For a two-sided test at \(\alpha=0.05\), the null value equals the unrounded upper endpoint of the matching interval. Should you reject or fail to reject? Explain using the p-value rule.
  5. Why does a standard two-sided 95% interval not automatically give the exact decision for a one-sided test at \(\alpha=0.05\)?