More Data Can Mean More Precision—but Not Less Bias
In Identifying the Source of Bias in a Scenario, you learned to trace how people are selected, who may be missing, and whether answers could be influenced. This tutorial connects those ideas to sample size. A larger random sample generally gives a statistic that varies less from sample to sample, making it more precise. But increasing the number of observations does not repair a biased selection process.
A sample statistic is calculated from the selected sample, while a parameter describes the population, as explained in Parameters and Statistics. Because a different random sample may produce a different statistic, one sample result is not guaranteed to match the population value. Here, precision refers to how little a statistic tends to vary across samples collected in the same way. It is not a promise that one particular statistic is close to the parameter.
Imagine repeatedly selecting random samples from the same population and calculating the same statistic each time. The statistics would not all be identical. With a larger sample, they generally cluster more closely together. This is the connection between sample size and sampling variability, introduced in Sources of Variability in Collected Data. The next tutorial will examine variability across repeated samples in more detail.
“Generally” matters. A larger sample does not guarantee that its statistic will be closer to the population value than the statistic from a smaller sample. Random variation can make a particular large-sample statistic farther away. The benefit of a larger sample is about the typical variation across samples, not a certainty about every result.
What a Larger Sample Changes
Suppose a town wants to estimate the proportion of residents who support a park renovation. If the town selects a random sample, the sample proportion \(\hat{p}\) may differ from the population proportion \(p\). A larger random sample usually gives more information about the population and reduces the amount that \(\hat{p}\) tends to fluctuate across repeated samples.
For a sample proportion, the possible values also become more finely spaced as the sample size increases. If \(n=20\), each additional person who answers “yes” changes \(\hat{p}\) by \(1/20=0.05\), or 5 percentage points. If \(n=100\), each additional “yes” changes it by \(1/100=0.01\), or 1 percentage point. Finer possible steps do not themselves guarantee accuracy, but they help explain why estimates from larger samples can be more precise.
The same broad idea applies to other statistics, such as a sample mean \(\bar{x}\): increasing the size of a random sample generally reduces the statistic’s sampling variability. The exact amount depends on the population and the sampling method. In this tutorial, focus on the direction of the effect rather than a formula for measuring that variability.
A fair comparison holds the rest of the sampling plan as constant as possible. Compare random samples from the same population, using the same selection process and measuring the same variable. If one study uses a different frame, a different question, or a different selection method, sample size alone cannot explain why its result differs.
Worked Example: Comparing Two Random Polls
Worked Example: Support for a Community Garden
Suppose 40% of adults in a fictional city support a proposed community garden. Two separate simple random samples are selected from the same population. In a sample of 50 adults, 17 support the plan. In a sample of 200 adults, 81 support it. Compare the estimates with the population proportion.
State: The parameter is \(p=0.40\), the proportion of all adults in the city who support the garden. The statistics are the sample proportions from the samples of 50 and 200 adults.
Plan: Calculate each sample proportion, then find its difference from the known population proportion. The samples use the same random method and target population, so the comparison illustrates how two sample sizes can produce different estimates. The result of these particular samples does not establish what will happen in every pair of samples.
Do: For the sample of 50 adults,
This estimate is 0.06, or 6 percentage points, below 0.40. For the sample of 200 adults,
This estimate is 0.005, or 0.5 percentage points, above 0.40. The absolute differences are \(|0.34-0.40|=0.06\) and \(|0.405-0.40|=0.005\).
Conclude: In these particular samples, the estimate from 200 adults is closer to the population proportion than the estimate from 50 adults. This is consistent with the general expectation that a larger random sample tends to give a more precise estimate. It does not prove that every sample of 200 will be closer than every sample of 50.
Precision Does Not Guarantee a Close Result
A common misunderstanding is that a larger random sample must produce an estimate closer to the population value. Random selection does not guarantee this. A sample of 200 could, by chance, contain a higher or lower proportion of supporters than a sample of 50. The larger sample tends to vary less across repeated samples, but its result can still be on either side of the population value.
Another way to see the distinction is to consider the possible values of a sample proportion. With \(n=20\), the count of “yes” responses can only be a whole number, so \(\hat{p}\) changes in increments of \(1/20\). With \(n=100\), the increments are \(1/100\). For example, 7 “yes” responses out of 20 gives 35%, while 8 gives 40%. With 100 people, 35, 36, or 37 “yes” responses give 35%, 36%, or 37%. The larger sample allows more finely spaced proportions, but a finely spaced estimate can still be inaccurate.
Worked Example: More Possible Levels of Detail
Worked Example: A Sample Estimate of Bike Commuting
A transportation planner wants to estimate the proportion of employees at a large fictional workplace who commute by bicycle. Consider two possible sample sizes, \(n=20\) and \(n=100\). Find how much the sample proportion changes when one additional sampled employee reports biking to work.
State: The statistic is the sample proportion \(\hat{p}\) of employees who commute by bicycle. We are comparing the size of one step between possible sample proportions for the two sample sizes.
Plan: Each additional “yes” response adds one to the numerator while the sample size remains fixed. Therefore, the step size is \(1/n\).
Do: With 20 employees, the step size is
With 100 employees, the step size is
For instance, in a sample of 20, 6 bicycle commuters give \(6/20=0.30\), while 7 give \(7/20=0.35\). In a sample of 100, 30 give \(30/100=0.30\), while 31 give \(31/100=0.31\).
Conclude: The possible sample proportions from 100 employees are spaced more closely than those from 20 employees. This is one way a larger sample can support a more finely detailed estimate. It does not establish which sample proportion is closer to the true workplace proportion; that depends on the particular people selected.
Why a Bigger Sample Cannot Fix Bias
Sample size addresses sampling variability, not every weakness in data collection. If a sampling frame leaves out part of the population, selecting more people from that incomplete frame does not give the missing group a chance to be selected. If people with a particular opinion are more likely to answer, adding more responses through the same process may preserve that imbalance. These concerns connect to Undercoverage Bias, Nonresponse Bias, and Voluntary Response Bias.
A very large biased sample can produce a result that is stable but systematically different from the population value. In that case, the result may be precise in the sense that repeated use of the same flawed method produces similar statistics, yet still be inaccurate for the population of interest. Precision is not the same as freedom from bias.
Worked Example: A Large Sample From an Incomplete Frame
Worked Example: Support for a Rural Transit Route
A fictional county wants to estimate the proportion of all adult residents who support a new transit route. There are two groups: 70% of adults live in urban areas, where 60% support the route; 30% live in rural areas, where 10% support it. A survey frame includes only urban residents. A random sample from that frame finds 79 supporters among 120 respondents; a later, much larger sample from the same frame finds 721 supporters among 1,200 respondents.
State: The target population is all adult residents of the county. The variable is whether an adult supports the route. The sampling frame omits rural residents, so it does not cover the full population of interest.
Plan: First calculate the county’s overall support proportion using the two group proportions and population shares. Then calculate each sample proportion and compare it with the county value. Finally, identify whether increasing the sample size corrected the missing coverage.
Do: The county’s population proportion is
So 45% of all county adults support the route. For the sample of 120 urban residents,
Its difference from the county proportion is about \(65.83\%-45\%=20.83\) percentage points. For the sample of 1,200 urban residents,
Its difference from the county proportion is about \(60.08\%-45\%=15.08\) percentage points. The frame’s omission of rural residents makes the expected sample proportion 60%, above the county value of 45%. The particular sample proportions also reflect random variation.
Conclude: The larger sample’s estimate happens to be closer to 45% than the smaller sample’s estimate, but it still overestimates support among all county adults. Increasing the sample size did not include rural residents or remove the undercoverage. The sampling method must be improved to address that bias.
Choosing a Sample Size Is Only Part of Planning
When planning a study, first define the population and the variable, as in Defining the Population of Interest and Deciding Which Variables to Measure. Then choose a selection method that gives the intended population appropriate representation. Only after considering these design choices should you treat a larger sample as a way to reduce sampling variability.
A larger sample can also take more time, money, and staff effort. If the method has serious bias, collecting many more observations the same way may be a poor use of resources. Improving the sampling frame, reducing nonresponse, or changing how participants are recruited may matter more for the credibility of the result than simply increasing \(n\).
If the study is a census, it seeks information from everyone in the defined population, as discussed in Census Versus Sample. A census removes sampling variability for that population at that time, but it does not automatically remove nonresponse, response bias, measurement problems, or errors in defining the population. Collecting data from everyone is not a guarantee of perfect information.
Common Mistakes and AP Exam Tips
- Claiming that a larger sample is always closer. Say that larger random samples generally have less sampling variability. Do not claim that every larger sample must produce a more accurate statistic.
- Using “precision” to mean “unbiased.” Precision concerns how much statistics vary across samples. Bias concerns a systematic tendency to differ from the population value. A method can be consistent and still biased.
- Assuming sample size repairs undercoverage. A large random sample selected from a frame that omits part of the population still cannot include the omitted people. Identify and fix the frame problem.
- Ignoring the sampling method when comparing sample sizes. The general sample-size comparison assumes the same target population and a comparable random selection method. A larger convenience sample is not automatically better than a smaller random sample.
- Confusing one observed comparison with a general rule. If one larger sample is closer in an example, report that as what happened in those samples. Explain separately that the broader claim concerns typical variability across repeated random samples.
- Calling every large sample representative. A high response count does not show that respondents represent the target population. Consider selection, coverage, and response, as practiced in Identifying the Source of Bias in a Scenario.
A strong AP response names the sample size and method, distinguishes variability from bias, and limits its claim. For example: “Using a larger random sample from the same population generally reduces sampling variability, so the sample proportion tends to be more precise. It does not guarantee that this particular estimate is closer to the population proportion, and it does not correct undercoverage in the sampling frame.”
Check Your Understanding
Answer each question using the distinction between sample size, sampling variability, and bias.
- A random sample of 400 residents is selected using the same method as a random sample of 40 residents. What is generally expected to happen to sampling variability, and what is not guaranteed for these particular samples?
- A sample proportion is calculated from 25 people and then from 100 people. How large is one step between possible proportions for each sample size?
- A survey uses a random sample from a list that excludes residents without internet access. The survey team triples the sample size but keeps using the same list. Does this fix the undercoverage? Explain.
- One sample of 500 people gives an estimate farther from the population value than a sample of 50. Does this disprove the general relationship between sample size and precision? Explain.
- A large open online poll receives thousands of responses. What does the large number of responses suggest about sample size, and why does it not by itself establish that the result is unbiased?