Conditions Protect the Meaning of an Inference
In Exam-Style Free Response on a One-Proportion Interval, you practiced checking conditions before calculating a confidence interval. Those checks are not formalities. A one-proportion interval or test uses a mathematical model for how \(\hat{p}\) would vary from sample to sample. If the real sampling process does not behave sufficiently like that model, the procedure’s reported uncertainty and conclusions may not be trustworthy.
Two important ideas underlie the formulas: the sample proportion should be centered at the population proportion, and its sampling distribution should be approximately Normal. Random sampling and, when applicable, the 10% condition help support the sampling model and independence assumptions. The Large Counts condition helps support the Normal approximation. Each condition addresses a different part of the reasoning; meeting one does not automatically meet the others.
Why the Center of the Sampling Distribution Matters
As established in Mean of the Sampling Distribution of p-hat, \(\hat{p}\) is an unbiased estimator of \(p\) when the sampling process is random and represents the population of interest: over repeated samples, the mean of the sampling distribution of \(\hat{p}\) is \(p\). “Unbiased” describes the center of results over many repetitions. It does not mean that every individual sample proportion equals \(p\), or that an estimate from one sample cannot be far from \(p\).
This matters for inference because a confidence interval is built around the observed \(\hat{p}\), while a test compares the observed \(\hat{p}\) with a value specified by the null hypothesis. The formulas quantify ordinary sample-to-sample variation under an assumed model. If the sampling method systematically favors certain individuals or responses, the resulting \(\hat{p}\) may be centered away from the target population proportion. A standard error can describe random variation around that wrong center, but it cannot remove the systematic shift.
Random selection is not the same thing as a guarantee that a sample perfectly mirrors a population. Chance can produce an unrepresentative sample, and problems such as undercoverage or nonresponse can still matter. The point is that a suitable random sampling process supports the model’s claim about the center; a large number of responses by itself does not establish that claim. The role of the random condition is the subject of the next tutorial.
Why the Normal Approximation Matters
The familiar one-proportion \(z\)-interval and \(z\)-test use Normal-based calculations. Those calculations work well when the sampling distribution of \(\hat{p}\) is approximately Normal. As covered in Checking Normality of p-hat with np and n(1-p), the Large Counts condition is a practical check for that approximation.
For a confidence interval, the procedure estimates the standard error using the observed sample proportion. Its Large Counts check uses the observed number of successes and failures, \(x\) and \(n-x\). For a significance test, the calculation is made under the null hypothesis. Its Large Counts check therefore uses the expected counts under the null, \(np_0\) and \(n(1-p_0)\), where \(p_0\) is the null value. These are different checks because the two procedures use different assumptions to calculate spread.
- Representative random process: Supports treating \(\hat{p}\) as centered at the population proportion being studied.
- Independence: For sampling without replacement from a finite population, the 10% condition supports treating observations as independent.
- Approximately Normal sampling distribution: For an interval, check \(x\geq10\) and \(n-x\geq10\). For a test of \(H_0:p=p_0\), check \(np_0\geq10\) and \(n(1-p_0)\geq10\).
If the Normal approximation is poor, the usual \(z\)-based tail areas or interval endpoints may not accurately represent the sampling behavior. If the center is biased, even a very good approximation to a Normal curve would describe results centered at the wrong value. Thus, “approximately Normal” and “unbiased” are separate requirements, not interchangeable descriptions.
Worked Examples
Worked Example: Conditions Support a Confidence Interval
A fictional wildlife team takes a random sample of 80 nesting boxes from 1,200 boxes in a conservation area. It finds signs of use in 24 boxes. The team wants to estimate the proportion of all nesting boxes in the area that show signs of use.
Check the sampling model. The sample was randomly selected, supporting the use of \(\hat{p}\) as an unbiased estimator for the area’s proportion, provided the sampling frame covers the target boxes. Because selection is without replacement, check the 10% condition: \(0.10(1{,}200)=120\), and \(80\leq120\). This supports treating the observations as independent.
Check the Normal approximation for an interval. The observed success count is \(x=24\), and the observed failure count is \(80-24=56\). Both are at least 10, so the Large Counts condition for the one-proportion interval is met.
Calculate and interpret. The sample proportion is \(\hat{p}=24/80=0.30\). For a 95% interval, \(z^*=1.96\). The estimated standard error is \(\sqrt{0.30(0.70)/80}=\sqrt{0.002625}\approx0.05123\). The margin of error is \(1.96\sqrt{0.30(0.70)/80}\approx0.10042\). Therefore,
We are 95% confident that between about 19.96% and 40.04% of all nesting boxes in the conservation area show signs of use. The conditions support using this interval to quantify sampling uncertainty for the stated population. They do not rule out every possible source of measurement or coverage error.
Notice that the condition checks justify interpreting the calculation as an interval for the population proportion, rather than merely reporting arithmetic endpoints. If the sample had not represented the target population, the calculation could still be performed, but the intended inference would not be justified.
Worked Example: The Test Check Uses the Null Proportion
A fictional school nutrition committee tests whether more than 20% of students bring fruit as a snack. It takes a random sample of 40 students and finds that 13 brought fruit. The proposed hypotheses are \(H_0:p=0.20\) and \(H_a:p>0.20\), where \(p\) is the proportion of all students at the school who bring fruit as a snack.
State. We are testing whether the proportion of all students at this school who bring fruit as a snack is greater than 0.20.
Plan and check conditions. The random sample supports the sampling model for the school population. If the school has at least 400 students, then \(40\leq0.10N\), so the 10% condition is met for sampling without replacement. For the test’s Large Counts condition, use the null value \(p_0=0.20\): \(np_0=40(0.20)=8\) expected successes and \(n(1-p_0)=40(0.80)=32\) expected failures. The expected success count is less than 10, so the Large Counts condition for a Normal-based one-proportion \(z\)-test is not met.
Do not proceed as if the test were justified. The observed sample proportion is \(\hat{p}=13/40=0.325\). If one mechanically substitutes into the usual test statistic, the result is
The arithmetic is correct, but the failed Large Counts check means the usual Normal-based \(z\)-test is not supported by the conditions. We should not report its Normal-curve p-value as though the approximation had been justified or draw a conclusion from it using the standard \(z\)-test. This example shows why checking conditions before interpreting a calculator output matters.
The interval and test can have different Large Counts outcomes for the same sample because they ask different modeling questions. In this example, an interval’s observed counts would be 13 successes and \(40-13=27\) failures, both at least 10. That check could support the interval’s Normal approximation. The test, however, evaluates the sampling distribution under the null proportion of 0.20, where only 8 successes are expected. Always match the check to the procedure.
Worked Example: A Small Standard Error Cannot Fix a Biased Sample
Imagine a fictional town wants to estimate the proportion of all adult residents who support a proposed community garden. Instead of selecting residents at random, a group posts an optional online poll on a gardening discussion page. Suppose the poll receives 1,000 responses, with \(\hat{p}=0.62\) supporting the proposal.
Examine the center. The large response count does not make the poll a random sample of town adults. People who visit a gardening discussion page and choose to respond may differ systematically from other residents. Therefore, there is no sound basis for treating this poll’s \(\hat{p}\) as unbiased for the proportion among all town adults. Repeating the same self-selected poll could consistently overrepresent people interested in gardening.
See what the formula alone would say. If someone mechanically used the one-proportion interval formula with \(\hat{p}=0.62\), the estimated standard error would be \(\sqrt{0.62(0.38)/1{,}000}=\sqrt{0.0002356}\approx0.01535\). Using \(z^*=1.96\), the nominal margin of error would be \(1.96\sqrt{0.62(0.38)/1{,}000}\approx0.03008\). The resulting arithmetic interval would be
Explain why that is not a valid population interval. These endpoints do not have the usual 95% confidence interpretation for all town adults, because the data-collection method does not support the required sampling model. The small calculated standard error reflects the formula’s assumed sampling variability; it does not measure how far a biased poll may be from the town’s true proportion. A narrow interval can be precisely centered on an unrepresentative estimate.
Common Mistakes and Full-Credit Reasoning
A complete AP response connects each condition to the claim it supports. It is not enough to write “conditions are met” or to list counts with no explanation. State the evidence and explain its role: randomness supports inference to the population, the 10% condition supports independence when sampling without replacement, and Large Counts supports the Normal approximation.
- Checking the wrong counts for a test. A one-proportion test uses the null value \(p_0\), so check \(np_0\) and \(n(1-p_0)\), not just the observed success and failure counts.
- Assuming a large \(n\) guarantees validity. A large sample may help the Normal approximation, but it does not repair a biased selection method or nonrepresentative sample.
- Treating unbiased as exact. Unbiasedness means the sampling distribution is centered at \(p\) over repetitions. It does not mean one sample must equal \(p\), nor does it eliminate sampling variability.
- Confusing precision with accuracy. A small standard error indicates limited modeled sampling variability. It does not establish that the estimate is close to the target when the assumptions fail.
- Using a failed check without qualification. If a required condition fails, do not present the standard \(z\)-procedure’s result as fully justified. Explain which condition failed and what part of the model it was intended to support.
Key Takeaway
One-proportion inference relies on more than inserting numbers into a formula. The sampling process must support a sample proportion centered at the population proportion, and the sampling distribution must be sufficiently close to Normal for the \(z\)-based calculations to work well. Checking conditions before calculating protects the meaning of the interval or test; a calculator cannot substitute for that reasoning.
Check Your Understanding
For each question, identify what a condition supports and distinguish calculation from a justified inference.
- In a one-proportion confidence interval, a sample has 18 successes and 42 failures. Which Large Counts check applies, and is it met?
- A test uses \(H_0:p=0.35\) with \(n=50\). Calculate the null expected success and failure counts. Is the Large Counts condition met?
- Explain why an unbiased estimator can still produce an estimate that differs from the true population proportion in one particular sample.
- A self-selected poll has 2,000 responses and a very small calculated standard error. Explain why this alone does not justify a confidence interval for the whole population.
- For a test, the observed counts are both at least 10, but one null expected count is 7. Which count check governs the Normal-based test, and what should the analyst conclude about using it?