Why Sample Size Matters When Sampling Without Replacement
In Recognizing a Binomial Setting and The BINS Checklist for Binomial Conditions, you learned that independent trials and a constant probability of success are part of a binomial setting. But when you sample people or objects from a finite population without replacement, the selections are not truly independent: after each selection, the population left to sample from has changed.
The 10% condition gives a practical rule for deciding when this dependence is small enough to ignore for a binomial model. It applies to a random sample taken without replacement from a population whose size is known. The condition does not make the selections independent; it supports treating them as approximately independent.
What Changes After Each Selection?
Suppose a population contains \(N\) individuals, of whom \(K\) have the characteristic labeled success. Before sampling, the probability that a randomly selected individual is a success is \(K/N\). If sampling is without replacement, that probability can change after the first selection. For example, after selecting a success, there is one fewer success in the population; after selecting a failure, there is one fewer failure.
Here is a small illustration. Imagine a population of 1,000 items, 200 of which are successes. The first selection has success probability \(200/1000=0.20\). After selecting one success, the next success probability is \(199/999\approx0.1992\). After selecting one failure, the next success probability is \(200/999\approx0.2002\). These probabilities are not exactly the same, and the selections are not truly independent.
If the sample is small relative to the population, however, each selection changes only a small part of the population. The success probability and the dependence between selections then change only slightly, which makes a binomial model a reasonable approximation in many settings. The 10% condition is a practical guideline for this purpose; it is not a claim that dependence disappears.
When the condition holds, the binomial model uses \(n\) as the number of sampled individuals and \(p=K/N\) as the population proportion of successes. The random variable \(X\) counts the successes in the sample, so the approximate model is \(X\sim B(n,p)\). As discussed in Defining \(n\), \(p\), and \(X\) in Context, define the success and the count in the situation before using this notation.
The 10% condition addresses the independence part of BINS and helps support the same-probability part. It does not replace the other checks. You still need two possible outcomes for each trial, a fixed number of trials, and a random selection process that makes the stated probability model sensible. A sample that is not selected randomly can be biased, even if it is less than 10% of the population.
How to Check the Condition
First identify the population being sampled and count its members to get \(N\). Then identify the number selected, \(n\). Multiply the sample size by 10 and compare that result with the population size. If \(N\geq10n\), the 10% condition is satisfied. If \(N<10n\), it is not.
You can also calculate the sample’s fraction of the population, \(n/N\), and check whether it is at most 0.10. These two methods are equivalent. For example, if \(n=40\), the population must have at least \(10(40)=400\) members. A population of 400 meets the condition exactly; a population of 399 does not.
Be precise about what counts as the population. If a survey samples from eligible registered participants, \(N\) is the number of eligible participants in that sampling frame—not everyone who lives in the area or belongs to a broader group. Likewise, use the sample size actually drawn, not the number who later respond, unless the problem specifically defines the sample that way.
Worked Example: Inspecting a Small Fraction of an Inventory
Worked Example: Inspecting a Small Fraction of an Inventory
A warehouse has 900 sealed packages. A quality inspector randomly selects 60 packages without replacement and records whether each package is damaged. There are 108 damaged packages in the warehouse. Is a binomial model reasonable for the number of damaged packages selected?
State. Let \(X\) be the number of damaged packages among the 60 selected. The population size is \(N=900\), the sample size is \(n=60\), and the population proportion of damaged packages is \(p=108/900\).
Plan. Check BINS, using the 10% condition to assess whether the without-replacement selections can be treated as approximately independent. The outcome for each package is binary: damaged or not damaged. The sample size is fixed at 60. The packages are randomly selected, and we will check whether the population is at least 10 times the sample size. The population success proportion supplies the model’s \(p\).
Do. The population proportion of damaged packages is:
For the 10% condition, compare the population size with 10 times the sample size:
Equivalently, the sample is \(60/900\approx0.0667\), or about 6.67% of the population, which is no more than 10%. The sample is randomly selected, each package is classified as damaged or not damaged, and the number selected is fixed. Because the 10% condition holds, treating the selections as approximately independent and using a constant success probability of 0.12 is reasonable for a binomial approximation.
The model is therefore \(X\approx B(60,0.12)\). The approximation symbol matters: the actual selections without replacement are dependent, even though the condition supports using a binomial model.
Conclude. A binomial model is reasonable as an approximation for the number of damaged packages in the sample. The 10% condition is satisfied because 900 is at least 10 times 60, and the success probability used is the population damaged-package proportion, 0.12.
Worked Example: A Sample That Is Too Large Relative to the Population
Worked Example: A Sample That Is Too Large Relative to the Population
A community center has 4,800 eligible members. A staff member randomly selects 500 members without replacement for a survey. In the full membership list, 1,440 members use a particular facility. Is it reasonable to use a binomial model for the number of facility users selected?
State. Let \(X\) be the number of facility users among the 500 selected members. Here \(N=4{,}800\), \(n=500\), and the population proportion of facility users is \(p=1{,}440/4{,}800\).
Plan. The response is binary: uses the facility or does not. The sample size is fixed, and the selection is random. Check the 10% condition before treating the without-replacement selections as approximately independent. A failure of this condition means that this rule does not support the binomial approximation.
Do. Calculate the population proportion and compare \(N\) with \(10n\):
The population size is \(4{,}800\), which is less than \(5{,}000\). Equivalently, the sample fraction is:
The sample is about 10.42% of the eligible membership, which exceeds 10%. The 10% condition is not satisfied. Although the selection is random and the outcome is binary, the condition does not support treating the selections as approximately independent for a binomial model. The value \(p=0.30\) is the population proportion, but knowing that proportion does not make the independence condition hold.
Conclude. Under the AP Statistics 10% guideline, we should not justify a binomial model for this sample as approximately independent. The sample is too large relative to the population for this condition to support that approximation.
Worked Example: Meeting the Condition Exactly
Worked Example: Meeting the Condition Exactly
A library has 250 books in a particular section, including 30 books that are overdue. A librarian randomly selects 25 books without replacement. Let \(X\) count the overdue books selected. Does the 10% condition support a binomial approximation?
State. The population size is \(N=250\), the sample size is \(n=25\), and the success proportion is the proportion of books in the section that are overdue. Success means that a selected book is overdue.
Plan. Books have two outcomes for this count—overdue or not overdue—and the sample size is fixed. The selection is random. Check whether \(N\geq10n\). Equality satisfies the condition because the sample is allowed to be exactly 10% of the population.
Do. The population success proportion is:
The population size is exactly 10 times the sample size:
The equivalent sample fraction is \(25/250=0.10\), exactly 10%. Thus the 10% condition is satisfied at its boundary. A binomial model with \(n=25\) and \(p=0.12\) is reasonable as an approximation, given the random selection and the binary outcome.
The boundary result does not mean that the selections become independent. For instance, after selecting one overdue book, the fraction of overdue books remaining is \(29/249\approx0.1165\), rather than the initial \(30/250=0.12\). The 10% condition allows an approximate binomial model; it does not turn the sampling process into sampling with replacement.
Conclude. The 10% condition is satisfied exactly, so it supports treating the number of overdue books in the sample as approximately binomial: \(X\approx B(25,0.12)\). The model remains an approximation because the books were selected without replacement.
What the Condition Does—and Does Not—Tell You
The condition is a check on how much the sample can change the population as sampling proceeds. If a small fraction of the population is sampled, removing one selected member has a relatively small effect on the mix of successes and failures left. If a large fraction is sampled, the remaining mix can change more noticeably, so the selections are more strongly dependent.
Passing the 10% condition does not prove that a binomial model is perfect or guarantee a particular level of accuracy for every probability. It is a practical AP Statistics guideline for treating sampling without replacement as approximately independent. The actual process is still without replacement, and so its selection outcomes remain dependent.
Also, the condition is not a substitute for checking that the sample was randomly selected. A convenience sample of 5% of a population can satisfy the numerical inequality while still failing to represent that population well. The condition addresses the effect of sampling without replacement; random selection addresses how the sample was obtained.
If the population size is not given or cannot be determined, you cannot verify \(N\geq10n\) from the available information. Do not assume that the condition holds just because the sample sounds small. State what additional population-size information is needed, or explain that the binomial approximation is not justified by the information provided.
Common Mistakes and AP Exam Tips
- Reversing the inequality. The condition is \(N\geq10n\), not \(n\geq10N\). Compare population size with ten times the sample size.
- Checking only whether the sample seems small. Use the stated population and sample sizes. For \(n=500\), for example, the required population size is at least \(5{,}000\).
- Calling the selections independent after the check passes. Sampling without replacement creates dependence. Say that the condition supports treating the selections as approximately independent.
- Assuming the 10% condition verifies all of BINS. It addresses dependence in a without-replacement sample. You must still identify two outcomes, a fixed number of selections, and a reasonable probability model.
- Using the wrong population size. Use the population from which the sample was actually drawn, such as all eligible members on a list—not a larger group that includes people who could not have been selected.
- Ignoring how the sample was selected. A small sample fraction does not make a nonrandom sample representative. Describe the selection process as well as checking the numerical condition.
For full-credit communication, identify \(N\) and \(n\), show the comparison \(N\geq10n\) or calculate \(n/N\), and state whether the condition is satisfied. Then connect that result to the binomial model in context. If sampling is without replacement, describe the model as an approximation rather than claiming the selections are truly independent.
Key Takeaway
The 10% condition connects finite-population sampling to the independence requirement in BINS. It gives a clear numerical check, but the conclusion is always about a reasonable approximation—not exact independence.
Check Your Understanding
For each situation, use the stated population size and sample size to assess the 10% condition. When appropriate, describe the binomial model as an approximation.
- A random sample of 35 bulbs is selected without replacement from 420 bulbs. Does the 10% condition hold? Show the comparison.
- A survey samples 240 people without replacement from a population of 2,100. Calculate \(n/N\) and decide whether the condition holds.
- Explain why selections without replacement are dependent, even when \(N\geq10n\).
- A random sample is 4% of a population, but people were selected because they were easy to contact. Does meeting the 10% condition alone justify a binomial model? Explain.
- A population has 600 members, of whom 90 are successes. A random sample of 50 is selected without replacement. Identify \(N\), \(n\), and \(p\), check the 10% condition, and state whether a binomial approximation is reasonable.