Read the Study Description Before Checking the Numbers
A condition check is not a box-ticking exercise. To decide whether a one-proportion inference procedure is supported, you need to read how the data were produced, identify the population the study is about, and match the numerical evidence to the intended procedure. A study description may show that a condition fails, show that it passes, or leave too little information to decide.
In Conditions Check for a Proportion Confidence Interval Scenario, you applied the Random, 10%, and Large Counts conditions to an interval. This tutorial focuses on spotting evidence in short passages: who was selected or assigned, how large the sample is relative to its source, and which success and failure counts the procedure requires. As in earlier tutorials on random sampling and study design, the word “random” matters only when the description tells you what was randomized.
A useful approach is to read the passage in layers. First identify the intended population and the way participants or observations entered the data. Then look for a sample size and a source-population size. Finally, identify whether the proposed procedure is an interval or a test, because the Large Counts check uses different counts for those two procedures.
A Diagnostic Map for the Three Conditions
The Random condition asks whether the data came from a random sample or an appropriate randomized process. A random sample can support inference to the population represented by its sampling frame. In an experiment, random assignment can support inference about a treatment effect for the experimental units; it does not, by itself, make those units representative of a broader population. Randomly inviting people to a survey is not the same as randomly selecting the people who answer it.
The 10% condition applies when observations are sampled without replacement from a finite population. Compare the sample size \(n\) with 10% of the actual source population of size \(N\). It is not a measure of whether a sample “seems small.” The source population is the group from which the sample was drawn, not a larger group chosen after the fact to make the comparison pass.
The Large Counts condition checks whether a Normal approximation is reasonable for the intended one-proportion \(z\)-procedure. For a confidence interval, use the observed success and failure counts. For a test of \(H_0:p=p_0\), use the expected counts under the null proportion \(p_0\). The earlier tutorial Interval Versus Test Condition Checks Compared explains why the two checks differ.
These checks answer different questions. Randomness concerns how the data were obtained and what population or process they represent. The 10% condition supports treating observations as approximately independent when sampling without replacement. Large Counts supports the Normal approximation. Passing one check does not compensate for failing another.
A Short-Passage Reading Strategy
Identify the population proportion being studied and whether the passage proposes a confidence interval or a test. This determines what the Large Counts check should use.
Look for random selection, random assignment, convenience sampling, or voluntary response. Say who was actually observed, not only who was invited.
If sampling without replacement, locate the size \(N\) of the population from which the sample was actually selected and compare \(n\) with \(0.10N\).
For an interval, check the observed successes and failures. For a test, check the expected success and failure counts under \(H_0\).
Name the condition, cite the passage or calculation that supports your decision, and explain the implication. If the passage omits necessary information, say that the condition cannot be verified from the description.
A short answer can be rigorous without repeating every definition. For example: “The Random condition is not met for estimating the proportion among all subscribers because the 80 respondents chose whether to answer the invitation. The description does not establish that respondents represent the full subscriber group.” That diagnosis identifies the evidence and limits the claim.
Worked Examples
Worked Example: Optional App Survey
A transit agency wants a 95% confidence interval for the proportion of all monthly-pass holders who use a new trip-planning feature. It emails 1,200 pass holders an optional survey. The 240 people who respond include 186 who report using the feature and 54 who do not. A report says the respondents are “a manageable fraction” of the pass-holder population but does not give the number of pass holders. Identify which conditions are violated, met, or not established.
The target is the proportion of all monthly-pass holders who use the feature, and the proposed procedure is a one-proportion \(z\)-interval. The Random condition is not met for this target: pass holders chose whether to respond. Randomly sending invitations does not make the group of respondents a random sample of all pass holders. The response method could favor people who use the feature or who have stronger opinions about it.
For the 10% condition, the description does not give the size of the pass-holder population. It also describes voluntary responses rather than a random sample drawn without replacement. Therefore, the 10% condition is not established by the stated facts; the phrase “a manageable fraction” is not evidence that \(n\leq0.10N\). Even a valid numerical comparison would not fix the self-selection problem.
For Large Counts in an interval, use the observed counts: \(x=186\) successes and \(n-x=240-186=54\) failures. Both are at least 10, so the Large Counts condition is met. This does not repair the failed Random condition. The data support describing the respondents—\(186/240=0.775\), or 77.5%—but the usual interval is not supported as an estimate for all pass holders by these condition checks.
Worked Example: A Random Sample That Is Too Large for the 10% Check
A city parks department wants a one-proportion \(z\)-interval for the proportion of registered community-garden members who compost food scraps. It randomly selects 84 members without replacement from the program’s list of 700 members. Of the selected members, 59 compost and 25 do not. Diagnose the conditions.
Let \(p\) be the proportion of the 700 registered members who compost food scraps. The Random condition is met because the 84 members were randomly selected from the list representing that population.
For the 10% condition, the source population is the 700 members on the program list. Ten percent of 700 is \(0.10(700)=70\), and \(84\not\leq70\). The sample is 12% of the source population, so the 10% condition is not met. Do not compare the sample with all city residents: they were not the source from which these members were selected.
For the interval’s Large Counts condition, there are 59 observed successes and \(84-59=25\) observed failures. Both counts are at least 10, so Large Counts is met. The diagnosis is therefore specific: Random and Large Counts are met, but the 10% condition fails. The usual one-proportion \(z\)-interval is not supported by all the checks as stated.
Worked Example: A Test With Too Few Expected Successes
A seed company claims that 8% of a large shipment of packets contain a labeling error. An auditor takes a random sample of 150 packets and plans a one-proportion \(z\)-test of \(H_0:p=0.08\) against \(H_a:p>0.08\). The shipment contains 8,000 packets. The sample finds 18 packets with errors and 132 without errors. Which condition is violated?
The Random condition is met: the packets were randomly sampled from the shipment. For the 10% condition, the source population is the 8,000 packets in the shipment. Since \(0.10(8{,}000)=800\) and \(150\leq800\), the 10% condition is met.
Because the proposed procedure is a test, check Large Counts using the null proportion, not the observed count of 18. Under \(H_0\), the expected number of packets with errors is \(np_0=150(0.08)=12\). The expected number without errors is \(n(1-p_0)=150(0.92)=138\). Both expected counts are at least 10, so the Large Counts condition is met. No condition is violated according to the information given.
The observed 18 errors may differ from the null expectation of 12, but that difference is not itself a condition failure. The condition concerns expected counts under the null model. Confusing observed counts with expected counts would incorrectly apply the interval check to a test.
Worked Example: A Description That Does Not Establish Randomness
A school district wants a confidence interval for the proportion of its high-school students who bring lunch from home. A staff member surveys 100 students who happen to be in the library during lunch. The district has 2,400 high-school students, and 62 surveyed students bring lunch from home. Assess the conditions and distinguish a violation from missing information.
The target population is all 2,400 high-school students. The Random condition is not met for inference to that population: the students were chosen because they happened to be in the library, not by a random selection process. Students in the library at lunchtime may differ from other students in ways related to bringing lunch.
The description identifies a finite source population of 2,400 students and says 100 were surveyed. The numerical comparison is \(0.10(2{,}400)=240\), and \(100\leq240\). This satisfies the 10% size comparison. However, that comparison does not make the convenience sample random or remove possible selection bias.
For the interval’s observed Large Counts check, there are \(x=62\) successes and \(100-62=38\) failures. Both are at least 10, so Large Counts is met. The Random condition fails even though the other two checks pass. The data describe the 100 students surveyed, but the usual interval is not justified for all district high-school students by these conditions.
Common Mistakes and AP Exam Tips
- Equating an invitation with random selection. State whether the people who supplied data were randomly selected. If they decided whether to respond, explain that voluntary response does not establish the Random condition.
- Calling every unknown detail a definite violation. If the passage gives no source-population size, say the 10% condition cannot be checked from the description. Do not invent \(N\) or assume it is large enough.
- Using the wrong source population. The 10% comparison uses the population the sample was drawn from. A larger regional or national population is irrelevant if the actual sampling frame is a smaller program list.
- Using observed counts for a test. For a one-proportion \(z\)-test, use \(np_0\) and \(n(1-p_0)\). For an interval, use \(x\) and \(n-x\).
- Assuming one failed condition makes the others fail. Diagnose each condition separately. A random sample can exceed 10% of its source population; a voluntary sample can have large observed counts.
- Writing only “conditions are not met.” Name the specific condition and cite evidence. Full-credit communication explains what the evidence means for the proposed inference and its target population.
Key Takeaway
To spot a violated condition, connect each requirement to the part of the study description that addresses it. The sampling or assignment method informs Random; the actual source population and sample size inform the 10% condition; and the procedure type determines which counts to use for Large Counts. Keep those judgments separate, and be honest when the description leaves a check unresolved.
Check Your Understanding
For each description, identify any condition that is violated or cannot be verified, and explain what evidence supports your diagnosis.
- A community center randomly selects 70 members without replacement from a list of 500. For an interval, 52 report using the fitness room and 18 do not. Which conditions are met?
- A school emails a survey to 900 students selected at random, but only 120 choose to respond. Can the respondents automatically be treated as a random sample? Explain.
- A one-proportion test uses a random sample of 80 observations and \(p_0=0.05\). What expected counts should be checked, and does Large Counts pass?
- A report says that 40 patients were randomly selected, but gives neither the source-population size nor the observed success and failure counts. Which checks can and cannot be assessed from this information?
- A random sample is taken without replacement from a list of 1,200 customers. The sample size is 100. Does the 10% condition pass, and what population size should be used in the comparison?