What a Confidence Interval Does—and Does Not—Describe
A confidence interval can be easy to misread because it is calculated from sample data but used to estimate a population value. As in Interpreting a Confidence Interval for a Proportion Using Percentages, the interval concerns the population proportion \(p\), not a percentage of the sample. This tutorial focuses on two especially common mistakes: treating the interval as a range for individual data, and saying there is a 95% chance that \(p\) is in the particular interval we calculated.
For example, a 95% confidence interval from 41% to 51% estimates one population quantity: the proportion of people in a specified population who have a specified characteristic. It is not a range of values describing 95% of those people, and it does not say that 95% of the sample responses fall between 41% and 51%.
The distinction is between a parameter and individual outcomes. The parameter \(p\) is a single proportion for the population. Each individual either has the characteristic or does not. The interval estimates the population’s overall proportion; it does not give a range of individual values.
Misinterpretation 1: “95% of the Data Lie in the Interval”
A statement such as “95% of the data lie between the interval endpoints” confuses a confidence interval with a range for individual observations. For a proportion, the endpoints are possible values of \(p\), not cutoffs for individual people’s responses. A person’s response might be yes or no, but the interval endpoints—such as 41% and 51%—are proportions for the population.
Similarly, an interval does not say that 95% of the sampled people had the characteristic. The sample proportion \(\hat p\) reports the fraction in the sample who had it. A confidence interval uses that sample result to estimate \(p\). The confidence level describes how the interval method performs over repeated samples, as explained in Interpreting a 95% Confidence Level Correctly.
Misinterpretation 2: “There Is a 95% Chance That \(p\) Is in This Interval”
It is tempting to say, “There is a 95% chance that \(p\) is between the endpoints.” But in the usual AP Statistics interpretation, \(p\) is a fixed, though unknown, population proportion. After the sample is selected and the interval is calculated, its endpoints are fixed too. The interval either captures \(p\) or it does not; the confidence level does not describe a probability assigned to \(p\) being in this particular interval.
The 95% refers instead to the method: if we repeatedly took appropriate random samples from the same population and calculated an interval in the same way each time, about 95% of those intervals would capture the fixed population proportion \(p\). The interval calculated from one sample is one of those intervals, but we do not know whether it is among the ones that capture \(p\).
This interpretation does not make the observed interval unhelpful. Its endpoints are still reasonable estimates based on the sample and method. You can say, “We are 95% confident that \(p\) is between the endpoints,” using the standard confidence-interval wording. Just do not translate that confidence statement into a probability claim about the fixed parameter.
Worked Examples
Worked Example: Correct a Claim About Individual Residents
A fictional random survey of 400 adult residents in a large town finds that 184 used a shared bicycle at least once in the past month. A student says, “We are 95% confident that 95% of the town’s adults used the service between 41.1% and 50.9% of the time.” Identify the errors and give an appropriate interpretation of the interval.
First, the sample proportion is \(\hat p=184/400=0.46\), or 46%. To check the reported interval, use the one-proportion \(z\)-interval formula with \(z^*=1.96\) for 95% confidence. The sample has 184 successes and \(400-184=216\) failures, both at least 10. Assume the survey is a random sample, and that the town has at least 4,000 adults; then \(400\leq0.10(4{,}000)\), so the 10% condition is met.
The estimated standard error is \(\sqrt{0.46(0.54)/400}=\sqrt{0.000621}\approx0.02492\). The margin of error is \(1.96(0.02492)\approx0.04884\). Therefore, the interval is \(0.46\pm0.04884\), or approximately \((0.4112,0.5088)\). As percentages, the endpoints are about 41.1% and 50.9%.
The student’s statement wrongly suggests that the interval describes the amount of use by individual adults, and “95% of the time” changes the characteristic. The interval is for the proportion of adults who used a shared bicycle at least once in the past month. A suitable interpretation is: “We are 95% confident that between 41.1% and 50.9% of all adult residents of the town used a shared bicycle at least once in the past month.”
Worked Example: Correct a Claim About the Sample Data
A fictional community survey asks randomly selected households whether they compost food scraps. Of 300 households, 168 say yes. A 90% confidence interval is reported as approximately 51.3% to 60.7%. Someone interprets this as, “We are 90% confident that between 51.3% and 60.7% of the sampled households compost.” What is wrong, and how should the interval be described?
The sample proportion is \(\hat p=168/300=0.56\), or 56%. The statement describes the sample, but the sample result is already known: exactly 168 of the 300 sampled households, or 56%, said yes. The interval estimates the proportion for all households in the population being studied.
The reported interval is consistent with a 90% one-proportion \(z\)-interval. There are 168 yes responses and \(300-168=132\) no responses, so both observed counts are at least 10. Assume a random sample and a population of at least 3,000 households; then \(300\leq0.10(3{,}000)\), meeting the 10% condition. With \(z^*=1.645\), the estimated standard error is \(\sqrt{0.56(0.44)/300}=\sqrt{0.00082133}\approx0.02866\). The margin of error is \(1.645(0.02866)\approx0.04714\). Thus, \(0.56\pm0.04714\approx(0.5129,0.6071)\), which rounds to 51.3% to 60.7%.
A suitable interpretation is: “We are 90% confident that between 51.3% and 60.7% of all households in the community compost food scraps.” The sample percentage is 56%; the interval’s endpoints estimate plausible values for the population proportion. The 90% confidence level refers to the method’s long-run capture rate, not to a percentage of sampled households.
Worked Example: Correct a Probability Claim About \(p\)
A fictional random sample of 250 eligible customers finds that 150 would recommend a new online help service. A student says, “There is a 95% chance that the true proportion \(p\) of eligible customers who would recommend it is between 53.9% and 66.1%.” Explain why the wording is not appropriate and write a careful interpretation.
The sample proportion is \(\hat p=150/250=0.60\). There are 150 successes and \(250-150=100\) failures, each at least 10. Assume the customers were randomly selected from a population of at least 2,500 eligible customers; then \(250\leq0.10(2{,}500)\), so the 10% condition is met. The one-proportion \(z\)-interval conditions are supported under these assumptions.
For 95% confidence, use \(z^*=1.96\). The estimated standard error is \(\sqrt{0.60(0.40)/250}=\sqrt{0.00096}\approx0.03098\). The margin of error is \(1.96(0.03098)\approx0.06073\). The interval is \(0.60\pm0.06073\), or approximately \((0.5393,0.6607)\), which is 53.9% to 66.1% after converting to percentages and rounding.
The issue is not the endpoints; it is the probability claim. The fixed population proportion \(p\) does not change from sample to sample. The interval does. A careful interpretation is: “We are 95% confident that between 53.9% and 66.1% of all eligible customers would recommend the new online help service.” In repeated sampling, about 95% of intervals constructed by this method would capture the fixed population proportion.
A Quick Routine for Checking an Interpretation
When you evaluate an interpretation, ask what the interval estimates, what the confidence level describes, and whether the wording matches the population and characteristic. This routine helps distinguish an appropriate confidence statement from a claim about individual data or a probability for \(p\).
Identify \(p\): the proportion in the specified population with the specified characteristic.
They are plausible values for that population proportion—not a range for individual responses or a percentage of the sample.
The stated confidence level describes the long-run capture rate of the interval method under appropriate conditions.
Name the confidence level, population, characteristic, and both endpoints. Do not turn the confidence level into a probability that the fixed \(p\) is in this observed interval.
Common Mistakes and AP Exam Tips
- Calling the endpoints a range for individual people. An interval for \(p\) estimates a population proportion. It does not describe the range of individual responses or the amount of time each person spent on an activity.
- Saying that a percentage of the sample lies in the interval. The sample statistic \(\hat p\) is already calculated from the sample. The interval uses it to estimate the population parameter \(p\).
- Giving the fixed parameter a 95% chance of being in the observed interval. The usual AP interpretation treats \(p\) as fixed and the interval as varying across samples. Say that you are 95% confident, then connect 95% to the method’s long-run capture rate.
- Leaving out the context. “The proportion is between 41% and 51%” is not complete if it does not identify the population and characteristic. A full-credit response specifies both.
- Changing the characteristic while interpreting. “Used a service at least once” is not the same as “used it regularly” or “used it a certain percentage of the time.” Preserve the question’s exact characteristic.
Key Takeaway
A confidence interval for a proportion estimates one fixed population value, \(p\). Its endpoints are not a range for individual data or a statement about what fraction of the sample has the characteristic. A 95% confidence level describes how often the interval method captures \(p\) over repeated appropriate samples; it does not mean there is a 95% chance that \(p\) lies in the interval calculated from the one sample.
Check Your Understanding
For each item, decide what the interval estimates and identify any mistaken interpretation.
- A 95% confidence interval for the proportion of local households with a rain barrel is 18% to 26%. Explain why “95% of households have between 18% and 26% of a rain barrel” is incorrect.
- A student says, “There is a 90% chance that the fixed population proportion is in this observed 90% confidence interval.” What does the 90% confidence level describe instead?
- A sample of 200 students includes 84 who walk or bike to school. Explain why a confidence interval based on this sample does not describe what percentage of these 200 students walk or bike.
- Write a suitable interpretation for a 95% confidence interval from 32% to 40% for the proportion of all eligible voters in a district who support a proposed community project.
- In your own words, distinguish a plausible range for \(p\) from a range containing a stated percentage of individual responses.