Which Counts Matter for a Confidence Interval?
Before using a one-proportion \(z\)-interval, check whether the observed sample contains enough successes and failures to support the Normal-based method. This is the Large Counts condition for a confidence interval. The key is that the check uses the observed sample proportion \(\hat{p}\), not a hypothesized population proportion.
In Checking Normality of \(\hat{p}\) with \(np\) and \(n(1-p)\), you used a population proportion \(p\) to check whether a Normal model for \(\hat{p}\) was reasonable when \(p\) was specified. For a confidence interval, the population proportion is unknown—that is what the interval is meant to estimate. So the interval procedure estimates the success and failure counts using the sample proportion.
The condition is therefore easy to check directly from the data. Count the observations with the characteristic of interest, then count the observations without it. Do not round \(\hat{p}\) and use the rounded value to reconstruct the counts; the actual counts are already available and are exact.
Why the Interval Uses the Sample Counts
A one-proportion confidence interval estimates an unknown population proportion \(p\) using the sample proportion \(\hat{p}\). As discussed in Structure of a One-Proportion \(z\)-Interval, the interval uses \(\hat{p}\) to estimate the standard error. The Large Counts condition checks whether that estimate is based on enough observed successes and failures for the Normal-based interval to be appropriate.
A test asks a different question. In a one-proportion test, the null hypothesis supplies a specific value \(p_0\), so the test’s Large Counts condition uses counts based on \(p_0\), such as \(np_0\) and \(n(1-p_0)\). For a confidence interval, there is no null value \(p_0\) to use. The relevant counts are the observed ones, \(n\hat{p}\) and \(n(1-\hat{p})\).
These expressions answer related but distinct questions. For an interval, the observed counts show whether the sample data support using the interval procedure. For a test, the counts calculated from \(p_0\) help assess the sampling model under the null hypothesis. Keep the two checks separate, even when working with the same sample.
Meeting the Large Counts condition does not, by itself, establish that an interval is appropriate. As covered in Why Inference Procedures Need Conditions, the sampling process also matters. The sample should be random or otherwise support treating observations as random, and a random sample taken without replacement from a finite population should satisfy the 10% condition. Each condition addresses a different aspect of the inference.
A Direct Check
Use the counts rather than decimal approximations whenever possible. If \(x\) of the \(n\) observations are successes, then \(\hat{p}=x/n\). Substituting this definition shows why the check is simply a count:
The estimated failure count works the same way:
The condition passes only when both counts are at least 10. One count cannot make up for the other: a sample with 80 successes and 2 failures does not meet the condition, even though its total sample size is large.
Record \(n\), the full number of observations, and \(x\), the number classified as successes.
The success count is \(x=n\hat{p}\); the failure count is \(n-x=n(1-\hat{p})\).
The condition is met only if the success count and the failure count are each at least 10.
If both counts meet the threshold, the Large Counts condition for the interval is met. If either is below 10, this condition does not support using the usual one-proportion \(z\)-interval.
Worked Examples
Worked Example: 44 of 75 Households Report a Feature
A fictional city survey randomly selects 75 households without replacement from 12,000 eligible households. Forty-four households report having a particular energy-saving feature. Check the conditions for using a one-proportion \(z\)-interval for the population proportion of eligible households with the feature.
State. Let \(p\) be the proportion of all 12,000 eligible households in this fictional city that have the feature. The sample size is \(n=75\), and the number of successes is \(x=44\). The sample proportion is $$ \hat{p}=\frac{x}{n}=\frac{44}{75}\approx0.5867. $$
Plan: check the sampling conditions. The problem states that the households were randomly selected, so the Random condition is met. The sample was selected without replacement from a finite population, so check the 10% condition: $$ 75\leq0.10(12{,}000)=1{,}200. $$ This is true, so the 10% condition is met. For the Large Counts condition, use the observed success and failure counts: $$ n\hat{p}=75\left(\frac{44}{75}\right)=44, \qquad n(1-\hat{p})=75\left(1-\frac{44}{75}\right)=31. $$ Both counts are at least 10, so the Large Counts condition is met.
Do. The condition check is based on 44 observed successes and \(75-44=31\) observed failures. No rounding is needed to decide whether either count reaches 10. If constructing a 95% interval after these checks, the estimated standard error is $$ SE_{\hat{p}} =\sqrt{\frac{\hat{p}(1-\hat{p})}{n}} =\sqrt{\frac{(44/75)(31/75)}{75}} \approx0.0569. $$ Using \(z^*=1.96\), the interval is $$ \frac{44}{75}\pm1.96(0.0569) \approx(0.475,\ 0.698). $$ The endpoints are rounded to three decimal places.
Conclude. The random, 10%, and Large Counts conditions are met for these data, so using a one-proportion \(z\)-interval is supported. We are 95% confident that the true proportion of eligible households in this fictional city with the feature is between about 0.475 and 0.698.
Worked Example: A Sample With Too Few Successes
A fictional school surveys 30 randomly selected students about whether they use a particular study-planning tool. Three students say they use it. Is the Large Counts condition for a one-proportion confidence interval met?
Here \(n=30\) and \(x=3\), so there are \(30-3=27\) failures. Equivalently, $$ n\hat{p}=30\left(\frac{3}{30}\right)=3, \qquad n(1-\hat{p})=30\left(1-\frac{3}{30}\right)=27. $$ The failure count is at least 10, but the success count is not. Because both counts must be at least 10, the Large Counts condition is not met. The usual one-proportion \(z\)-interval is not supported by this condition for these data. This result does not show that the true proportion is zero or that no interval method could ever be used; it says this particular Normal-based condition has not been satisfied.
Worked Example: Both Counts Are Just Large Enough
A fictional environmental group randomly samples 20 garden plots and records whether each plot contains a particular plant species. The species is found in 10 plots. Check the Large Counts condition for a one-proportion confidence interval.
There are \(n=20\) plots and \(x=10\) successes. The number of failures is \(20-10=10\). In the condition’s notation, $$ n\hat{p}=20\left(\frac{10}{20}\right)=10, \qquad n(1-\hat{p})=20\left(1-\frac{10}{20}\right)=10. $$ Both counts equal 10, and the rule is “at least 10,” so the Large Counts condition is met. The fact that the counts are exactly at the threshold does not make the condition fail.
How to Report the Check
A good condition check gives the evidence, not just a yes-or-no conclusion. For the 44-out-of-75 sample, a concise AP-style statement is: “There are 44 observed successes and \(75-44=31\) observed failures. Both counts are at least 10, so the Large Counts condition for a one-proportion \(z\)-interval is met.”
If a count is below 10, say so directly and identify which one. For the sample with 3 successes and 27 failures, an appropriate statement is: “The sample has only 3 observed successes, which is less than 10. Therefore, the Large Counts condition is not met, and the usual one-proportion \(z\)-interval is not supported by this condition.” This is clearer than saying only that the sample is “too small,” because a small sample can have enough successes and failures, while a larger sample with a proportion near 0 or 1 can still have too few observations in one category.
When the data give \(x\) and \(n\), calculate \(n-x\) for the failures. When the data instead give \(\hat{p}\) and \(n\), you can calculate \(n\hat{p}\) and \(n(1-\hat{p})\), but remember that these quantities represent estimated numbers of successes and failures. For actual sample data, using the original counts avoids rounding confusion.
Common Mistakes and AP Exam Tip
- Using \(np\) instead of \(n\hat{p}\) for an interval. The population proportion \(p\) is unknown in a confidence interval. Use the observed counts, \(n\hat{p}\) and \(n(1-\hat{p})\).
- Checking only one count. Both successes and failures must be at least 10. A large success count does not offset too few failures, or vice versa.
- Using the total sample size as the condition. Having \(n\geq10\) is not enough. For example, \(n=30\) with 3 successes still fails the condition.
- Rounding a proportion before checking. Use \(x\) and \(n-x\) directly. This avoids a rounded \(\hat{p}\) producing an imprecise count near the threshold.
- Confusing an unmet condition with proof that the parameter is extreme. A failed check means the usual interval’s Large Counts condition is not supported; it does not prove a specific value of \(p\).
- Forgetting the other conditions. Meeting Large Counts alone does not establish randomness or the 10% condition. Check each requirement separately, as in the complete four-step procedure.
Key Takeaway
For a one-proportion confidence interval, the Large Counts condition uses the observed sample proportion. Since \(n\hat{p}=x\) and \(n(1-\hat{p})=n-x\), the check is simply whether the sample has at least 10 observed successes and at least 10 observed failures. In the 44-out-of-75 example, those counts are 44 and 31, so the condition is met.
Check Your Understanding
For each situation, check the Large Counts condition for a one-proportion confidence interval and explain your conclusion.
- A random sample of 75 households includes 44 successes. What are the observed success and failure counts, and is the condition met?
- In a sample of 28 people, 6 have a specified characteristic. Which count determines whether the condition fails?
- A sample of 24 observations has 12 successes and 12 failures. Does equality with 10 matter to the decision?
- For a confidence interval, should you check \(n\hat{p}\) and \(n(1-\hat{p})\), or \(np_0\) and \(n(1-p_0)\)? Explain briefly.
- A student says, “The sample size is 50, so the Large Counts condition is met.” What additional information is needed to assess the condition?