Using an Interval to Assess a Claimed Mean
A confidence interval gives a range of plausible values for a population mean \(\mu\). That range can also help assess a specific claim about \(\mu\). If the claimed value lies outside an appropriate interval, the sample provides evidence against that claim at a level connected to the confidence level. If the value lies inside, the sample does not provide convincing evidence against it at that level.
This connection works when the interval and test use the same data, the same one-sample t procedure, and a two-sided test. As described in “The Form of a One-Sample t Interval,” the interval is centered at \(\bar{x}\) and uses \(s/\sqrt{n}\) to account for uncertainty in estimating \(\mu\). The new idea here is to use its endpoints to make a test decision about a hypothesized mean.
For example, a 95% confidence interval matches a two-sided test at \(\alpha=0.05\), and a 90% confidence interval matches a two-sided test at \(\alpha=0.10\). This is an exact connection for the corresponding t interval and t test, apart from minor differences that can result from rounding. A hypothesized value exactly on an endpoint is not outside the interval; with unrounded calculations, it corresponds to a test result at the significance-level boundary.
The rule is for a two-sided alternative, such as \(H_a:\mu\ne\mu_0\). It does not, by itself, make a decision for a one-sided alternative such as \(H_a:\mu>\mu_0\). The matching interval is not a substitute for checking whether the study design and data meet the conditions for inference.
How to Make and Explain the Decision
Begin by translating the claim into a null hypothesis. If a manufacturer claims that the mean battery life is 8 hours, for instance, the null hypothesis is \(H_0:\mu=8\) hours, where \(\mu\) is the true mean battery life for the population of interest. For a two-sided assessment, the alternative is \(H_a:\mu\ne 8\) hours.
Then identify the confidence level and compare the claimed value with the interval endpoints. If it is excluded, reject the null hypothesis at the matching significance level. If it is included, fail to reject the null hypothesis. In either case, state the result in context and name the population mean, rather than describing only a number or saying that the claim has been “proved.”
Define \(\mu\), state \(H_0:\mu=\mu_0\), and use a two-sided alternative \(H_a:\mu\ne\mu_0\) when assessing whether the mean differs from the claim.
Pair a confidence level of \(100(1-\alpha)\%\) with a two-sided test at significance level \(\alpha\). Check the conditions for the one-sample t procedure.
If \(\mu_0\) is outside the interval, reject \(H_0\). If \(\mu_0\) is inside, fail to reject \(H_0\).
Describe whether the sample provides convincing evidence that the population mean differs from the claimed value. Do not say that an included claim has been proved true.
The conditions are the ones used for the one-sample t interval, as reviewed in “Building a One-Sample t Interval.” Check that the data come from a random sample or a suitable randomized process; check independence, including the 10% condition when sampling without replacement from a finite population; and check that the data’s shape is suitable for a t procedure. For a small sample, the population should be approximately Normal or the sample data should be roughly symmetric with no strong outliers. A larger sample can make the procedure more robust to departures from Normality, but it does not correct a biased sample or dependent observations.
The interval decision is convenient, but it does not give the exact p-value. It tells you whether the p-value for the matching two-sided test is below the chosen \(\alpha\), or at or above it. A value outside the interval corresponds to a p-value below \(\alpha\); a value inside corresponds to a p-value at least \(\alpha\). If an exact p-value is requested, the interval alone is not enough to report one.
Worked Examples
Worked Example: Assessing a Claim About Sensor Response Time
A fictional laboratory randomly selects 25 sensors from a population of more than 2,500 sensors and measures their response times. The sample has \(\bar{x}=72\) milliseconds and \(s=10\) milliseconds. A researcher claims the population mean response time is 68 milliseconds. A plot of the sample shows no strong skewness or outliers. Use a 95% confidence interval to assess the claim.
State. Let \(\mu\) be the true mean response time, in milliseconds, for the population of sensors. The hypotheses are \(H_0:\mu=68\) milliseconds and \(H_a:\mu\ne68\) milliseconds.
Plan. A 95% one-sample t interval matches a two-sided test with \(\alpha=0.05\). The sensors were randomly selected. The observations are independent under the sampling design, and the 10% condition is met because \(25<0.10(2500)=250\); the population is larger than 2,500, so the sample is less than 10% of it. With \(n=25\), the plot’s lack of strong skewness or outliers supports using a one-sample t procedure.
Do. The degrees of freedom are \(df=n-1=25-1=24\). For a 95% interval with 24 degrees of freedom, \(t^*\) is about 2.064. The estimated standard error is:
The margin of error is \(2.064(2)=4.128\) milliseconds. The interval is:
The claimed mean, 68 milliseconds, lies inside the interval because \(67.872<68<76.128\).
Conclude. Fail to reject \(H_0\) at the 0.05 significance level. The sample does not provide convincing evidence that the population’s mean sensor response time differs from 68 milliseconds. This result does not establish that the true mean is exactly 68 milliseconds; values in the interval are plausible given the data and procedure.
Worked Example: Testing a Claim About Reusable Bottle Capacity
A fictional outdoor equipment shop randomly samples 16 reusable bottles from a large shipment. Their measured capacities have a mean of 42.5 fluid ounces and a standard deviation of 6 fluid ounces. The sample data are roughly symmetric with no apparent outliers. The shop wants to assess a claim that the population mean capacity is 47 fluid ounces, using a 90% confidence interval.
State. Let \(\mu\) be the true mean capacity, in fluid ounces, of bottles in the shipment. The hypotheses are \(H_0:\mu=47\) fluid ounces and \(H_a:\mu\ne47\) fluid ounces.
Plan. The 90% interval matches a two-sided test at \(\alpha=0.10\). The sample is random. The shipment is large enough that \(16<0.10N\), where \(N\) is the number of bottles, so the 10% condition is met. With a sample of 16, the roughly symmetric data and lack of apparent outliers support using a one-sample t procedure.
Do. The degrees of freedom are \(df=16-1=15\). For a 90% interval with 15 degrees of freedom, \(t^*\) is about 1.753. The standard error is:
The margin of error is \(1.753(1.5)=2.6295\) fluid ounces. Thus, the interval is:
The claimed value of 47 fluid ounces is above the upper endpoint of 45.1295 fluid ounces, so it is outside the interval.
Conclude. Reject \(H_0\) at the 0.10 significance level. The sample provides convincing evidence that the population mean bottle capacity differs from 47 fluid ounces. The interval decision supports a difference, but does not by itself establish the direction or size of that difference as a test conclusion.
Worked Example: How the Confidence Level Can Change the Decision
A fictional community garden randomly selects 36 bags of a particular soil mixture and records their masses. The sample has \(\bar{x}=18.4\) kilograms and \(s=3.6\) kilograms. The data show no strong skewness or outliers, and the garden wants to assess the claim that the population mean mass is 20 kilograms. Compare decisions using 95% and 99% confidence intervals.
State. Let \(\mu\) be the true mean mass, in kilograms, of bags of this soil mixture. For each interval, the hypotheses are \(H_0:\mu=20\) kilograms and \(H_a:\mu\ne20\) kilograms.
Plan. The 95% interval matches a two-sided test at \(\alpha=0.05\); the 99% interval matches a two-sided test at \(\alpha=0.01\). The sample is random. If the bags come from a population of at least 360 bags, then \(36<0.10(360)=36\) would not hold as a strict inequality at 360, so the 10% condition requires a population larger than 360. The garden confirms that the population contains more than 360 bags, so the condition is met. The sample size of 36 and the absence of strong skewness or outliers support the t procedure.
Do. The degrees of freedom are \(df=36-1=35\). The standard error is:
For 95% confidence, \(t^*\) is about 2.030. The margin of error is \(2.030(0.6)=1.218\) kilograms, so the interval is:
The claimed mean of 20 kilograms is outside this interval. At the 0.05 level, reject \(H_0\).
For 99% confidence, \(t^*\) is about 2.724. The margin of error is \(2.724(0.6)=1.6344\) kilograms, giving:
Now 20 kilograms lies inside the interval. At the 0.01 level, fail to reject \(H_0\).
Conclude. The 95% interval gives convincing evidence at the 0.05 level that the mean bag mass differs from 20 kilograms, while the 99% interval does not give convincing evidence at the 0.01 level. This is not a contradiction: the higher confidence level produces a wider interval and corresponds to a more demanding test for rejecting the claim.
Common Mistakes and AP Exam Tips
- Saying “accept the null hypothesis.” If the claimed value is inside the interval, say “fail to reject \(H_0\).” The data have not shown that the claim is true; they have not provided sufficient evidence against it at the chosen level.
- Matching the wrong confidence level and significance level. A 95% interval matches a two-sided test at \(\alpha=0.05\), not at \(\alpha=0.95\). Use \(C=1-\alpha\), where \(C\) is the confidence level written as a proportion.
- Using an interval for a one-sided question without checking the match. The inclusion rule described here applies to a two-sided alternative. Do not use it to claim a one-sided test decision.
- Ignoring the conditions. A numerical comparison cannot fix a biased sample, dependence, or an unsuitable distribution for the t procedure. State and check the randomization, independence, 10% condition when applicable, and shape conditions.
- Overstating what an included value means. Inclusion means the result is not statistically significant at the matching level. It does not mean that the claimed value is the only plausible mean, or that it has a particular probability of being true.
- Deciding from rounded endpoints near the claim. If the claimed value is very close to an endpoint, compare it with unrounded endpoints when possible. Rounding may make a value appear to be inside or outside when it is actually on the other side.
- Reporting an exact p-value from the interval alone. The interval tells you whether the matching p-value is below \(\alpha\), not its exact value. Use an appropriate test calculation if the question requests the p-value.
A full-credit response identifies the hypotheses and parameter, matches the interval to a two-sided significance level, checks the inference conditions, compares the claimed value with the interval, and gives a contextual conclusion using “reject” or “fail to reject.” The conclusion should say whether there is convincing evidence that the population mean differs from the claimed value.
Check Your Understanding
Use the matching two-sided test and state each decision in context.
- A 95% confidence interval for the mean commute time is 18 to 24 minutes. A claim says the population mean is 25 minutes. What is the test decision at \(\alpha=0.05\)?
- A 90% confidence interval for the mean daily water use is 110 to 130 liters. A claim says the population mean is 120 liters. What can you conclude at \(\alpha=0.10\), and what can you not conclude?
- Which confidence interval matches a two-sided test at \(\alpha=0.01\)?
- Why is “accept the claim” not an appropriate conclusion when the claimed mean lies inside the interval?
- A claimed mean is nearly equal to a rounded interval endpoint. What should you check before deciding whether to reject the null hypothesis?