Tutorials › AP Statistics › Sampling Variability Versus Bias in Proportions

Sampling distributions for proportions · Tutorial 416 of 1000

Sampling Variability Versus Bias in Proportions

Distinguish ordinary sample-to-sample variation from systematic bias in polls, and see why a larger sample cannot automatically repair a flawed sampling method.

Intermediate 9 min read

What You'll Learn

  • Explain sampling variability as the changes in sample proportions across repeated samples.
  • Describe what it means for a sampling method to be unbiased for a population proportion.
  • Identify how coverage, nonresponse, and opt-in polls can systematically misrepresent a population.
  • Calculate the expected proportion among respondents when groups have different response rates.
  • Compare how sample size affects variability and bias.
  • Communicate why one poll result alone cannot diagnose the source of a difference.

Two Reasons a Poll Can Miss the Population Proportion

In Sampling Distribution of \(\hat{p}\) from a Claimed Percentage, you used a sampling distribution to describe how sample proportions vary when a population proportion and a sampling model are specified. This tutorial separates two ideas that can both make a poll’s \(\hat{p}\) differ from the population proportion \(p\): random sample-to-sample variation and bias from a flawed sampling method.

Suppose a random poll estimates the share of residents who support a transit change. Even if the sampling method is sound, different random samples will usually include different people. Their sample proportions will not all be identical. That ordinary fluctuation is sampling variability. A different problem arises if the poll systematically misses or underrepresents some kinds of residents. Then the method may tend to produce sample proportions that are too high or too low, even across repeated polls.

Definition: Sampling variability is the natural change in a statistic, such as \(\hat{p}\), from one random sample to another. Bias is a systematic tendency for a statistic produced by a method to differ from the population parameter it is intended to estimate.

The key distinction is about the pattern across repeated samples. A sound random sampling method can produce an estimate above \(p\) in one sample and below \(p\) in another, while its sampling distribution is centered at \(p\). A biased method tends to place its estimates away from \(p\). A single sample can be far from \(p\) by chance, so one result alone does not establish that the method is biased.

Variation Has a Center; Bias Shifts the Center

As established in Mean of the Sampling Distribution of \(\hat{p}\), for an appropriate random sample, the mean of the sampling distribution of \(\hat{p}\) is \(p\). This is what it means for \(\hat{p}\) to be unbiased for \(p\) under that sampling method: across repeated samples, the estimates are centered at the population proportion. Unbiased does not mean every sample gives the exact population proportion.

The standard deviation \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\), from Standard Deviation of \(\hat{p}\) Formula, describes the amount of sample-to-sample variation under the stated random sampling model. Increasing \(n\) reduces this spread. But that formula describes variation around the model’s center; it does not show that a poll’s selection method represents the population well.

Key distinction: Sampling variability describes how spread out estimates are around their center. Bias describes whether that center is systematically displaced from the population proportion. A larger sample can reduce variability, but it does not automatically remove bias.

Bias can enter in several ways. A poll may use a list that excludes some members of the population (a coverage problem). Some selected people may not respond, and response rates may differ among groups (nonresponse bias). Or a poll may invite people to choose whether to participate, as in an open online poll (voluntary response). If people with stronger opinions are more likely to answer, the respondents may not reflect the population.

These problems are systematic when they consistently favor some groups or opinions. They are different from sampling variability, which remains even with a well-designed random sample. A poll can also have both: estimates may vary from sample to sample and cluster around a biased center.

Worked Example: Variation in a Random Transit Poll

Worked Example: Different Random Samples, Same Target

Suppose 40% of residents in a town support changing a bus route. A polling team takes a random sample of 250 residents without replacement from 10,000 residents. One sample has 110 supporters. How much sample-to-sample variation would be expected under the random-sampling model, and does this one result demonstrate bias?

State. Let \(\hat{p}\) be the proportion of sampled residents who support the change. The population proportion is \(p=0.40\), and this sample gives \(\hat{p}=110/250=0.44\). We will use the sampling distribution’s standard deviation to describe ordinary variation under the stated random-sampling model.

Plan and check conditions. The poll uses a random sample. Because sampling is without replacement, check the 10% condition: \(250\leq0.10(10{,}000)=1{,}000\), so the condition is met. The expected numbers of supporters and nonsupporters are \(np=250(0.40)=100\) and \(n(1-p)=250(0.60)=150\). Both are at least 10, so the Large Counts condition supports an approximate Normal model if one is needed.

Do. The standard deviation and approximate two-standard-deviation range are:

$$ \sigma_{\hat{p}} =\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{0.40(0.60)}{250}} =\sqrt{0.00096} \approx0.03098 $$
$$ p\pm2\sigma_{\hat{p}} =0.40\pm2(0.03098) \approx(0.3380,\ 0.4620) $$

The observed \(\hat{p}=0.44\) is within this approximate range. Its distance above the center is \(0.44-0.40=0.04\), or about \(0.04/0.03098\approx1.29\) standard deviations.

Conclude. Under this random-sampling model, sample proportions will vary from sample to sample, with a standard deviation of about 0.031. The result of 0.44 is a plausible random-sample fluctuation around 0.40. It does not, by itself, show that the polling method is biased; assessing bias requires examining whether the method systematically represents residents.

How Different Response Rates Shift a Poll

A poll’s respondent proportion can differ systematically from the population proportion when response rates vary by group. To see why, track both the number in each group and the number expected to respond. The respondent proportion is calculated among respondents, not among everyone invited.

The following example uses response rates as part of a hypothetical model. In a real poll, actual response behavior may be more complicated. The calculation illustrates how a difference in response rates can shift the center of the respondent results even before sample-to-sample variation is considered.

Worked Example: An Open Poll Overrepresents Supporters

Worked Example: A Bus-Route Poll on a News Website

Suppose 40% of residents support changing a bus route and 60% oppose it. A news website posts an open poll. Assume that 80% of supporters respond and 40% of opponents respond. What proportion of respondents would be expected to support the change under these assumptions, and how does it compare with the population proportion?

State. Let \(p=0.40\) be the population proportion who support the change. We want the expected proportion of supporters among those who respond. The response rate for supporters is \(0.80\), while the response rate for opponents is \(0.40\). Thus, supporters are twice as likely to respond as opponents, since \(0.80/0.40=2\).

Plan. Work with a hypothetical population of 10,000 residents. Find the expected number of respondents in each group, then divide the number of responding supporters by the total number of respondents. This shows the respondent mix implied by the assumed response rates.

Do. Of 10,000 residents, 4,000 support the change and 6,000 oppose it. The expected respondent counts are:

$$ \text{Supporting respondents}=4{,}000(0.80)=3{,}200 \qquad \text{Opposing respondents}=6{,}000(0.40)=2{,}400 $$

The total expected number of respondents is \(3{,}200+2{,}400=5{,}600\). Therefore, the expected support proportion among respondents is:

$$ \frac{3{,}200}{5{,}600} =\frac{4}{7} \approx0.5714 $$

This is about 17.14 percentage points higher than the population proportion: \(0.5714-0.40=0.1714\), rounded. The shift occurs because supporters are more likely to respond, not because the population itself has changed.

Conclude. Under the stated response-rate assumptions, about 57.14% of respondents would support the change, even though only 40% of residents support it. This is a systematic shift caused by the response process. It is not ordinary random sample-to-sample variation around 40%.

A Larger Sample Does Not Automatically Fix Bias

A larger sample is valuable when the sampling method represents the population: it reduces the standard deviation of \(\hat{p}\), as discussed in How Sample Size Changes the Spread of \(\hat{p}\). But increasing the number of respondents to a biased poll does not make the respondents representative. If the response pattern continues to favor one group, a larger poll can estimate the respondent mix more precisely while remaining centered away from the population proportion.

For illustration, suppose the open-poll respondent mix in the previous example is treated as a large pool with support proportion \(q=4/7\). If responses were independent random draws from that pool, then for \(n=1{,}000\) responses the standard deviation around \(q\) would be approximately 0.01565. For \(n=4{,}000\), it would be approximately 0.00782. Both describe spread around \(q\), not around the town’s true support proportion of 0.40. This hypothetical calculation does not make an open poll a random sample of residents; it only illustrates how variability within the respondent pool can shrink while the selection bias remains.

Worked Example: More Responses, a Shifted Center

Worked Example: Support for Extended Library Hours

Suppose 55% of residents support extending library hours. In a poll, assume 80% of supporters respond but only 50% of opponents respond. Consider what would happen if the poll collected 1,000 respondents. Under a model that treats these respondents as independent draws from the resulting respondent mix, find the expected support proportion and its standard deviation. Explain what a larger response count would and would not accomplish.

State. The population proportion is \(p=0.55\). The assumed response rates are \(0.80\) for supporters and \(0.50\) for opponents. We first find the support proportion \(q\) among respondents, then calculate the variability for a sample proportion from that respondent mix.

Plan and check conditions. For every 1,000 residents, the expected supporter response share is \(0.55(0.80)=0.44\), and the expected opponent response share is \(0.45(0.50)=0.225\). The overall response share is \(0.44+0.225=0.665\). For the variability calculation, assume responses act like independent draws from a very large respondent pool with proportion \(q\). The Large Counts condition for 1,000 responses is met because the expected counts \(1000q\) and \(1000(1-q)\) are both greater than 10.

Do. Among respondents, the expected support proportion is:

$$ q=\frac{0.44}{0.665} =\frac{88}{133} \approx0.6617 $$

The expected respondent proportion is about 66.17%, around 11.17 percentage points above the population proportion of 55%. Its standard deviation for \(n=1{,}000\) is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{q(1-q)}{n}} =\sqrt{\frac{(88/133)(45/133)}{1{,}000}} \approx0.01496 $$

If the response count were increased to 4,000 while the same respondent mix and assumptions remained, the standard deviation would be about half as large:

$$ \sqrt{\frac{(88/133)(45/133)}{4{,}000}} \approx0.00748 $$

Conclude. Under this illustrative model, more responses reduce random variation around the respondent proportion of about 0.6617. They do not move that center to the population proportion of 0.55. The calculation assumes independent responses from the respondent mix; it cannot establish that an actual poll has that structure or that its respondents represent all residents.

What a Poll Result Can—and Cannot—Tell You

When a poll result differs from a known or claimed population proportion, do not immediately label the difference as bias. A single random sample can land above or below \(p\). The sampling distribution helps describe how much variation is expected under a specified random sampling model. A difference may be unusual under that model, but the calculation does not identify its cause.

A careful evaluation asks how people were selected, who could be reached, who responded, and whether response was related to the opinion being measured. If those details point to systematic underrepresentation, the poll may be biased. If the sampling design is sound, an observed difference may instead be ordinary sample-to-sample variation. In practice, both explanations can matter, and the poll’s method should be described before making a broad claim about the population.

Bias is also not the same as a mistake in arithmetic. A poll can calculate \(\hat{p}\) correctly from its respondents and still fail to estimate the target population proportion well. Conversely, a random sample may happen to produce an estimate far from \(p\) without the method being biased. Keep the quality of the calculation separate from the quality of the sampling method.

Common Mistakes and AP Exam Communication

  • Calling every difference bias. A single sample proportion that differs from \(p\) may reflect sampling variability. Bias refers to a systematic shift across repeated use of a method.
  • Assuming an unbiased method gives the right answer every time. Unbiased means the sampling distribution is centered at \(p\), not that every \(\hat{p}\) equals \(p\).
  • Assuming a bigger poll is automatically better. A larger \(n\) reduces random variability under a suitable model. It does not correct undercoverage, voluntary response, or unequal nonresponse by itself.
  • Confusing the respondent proportion with the population proportion. A poll’s \(\hat{p}\) describes the respondents. To generalize to the population, the sampling and response methods must support that step.
  • Claiming that a small standard deviation rules out bias. A standard deviation measures spread around a center. It does not establish that the center equals the population proportion.
  • Using response rates without weighting by group size. The respondent mix depends on both each group’s population share and its response rate. Calculate the expected respondent counts or shares for every group.
AP Exam Tip: State whether the issue is sample-to-sample variation or a systematic feature of the sampling method. For a poll-bias claim, identify who may be overrepresented or underrepresented and explain how that could shift \(\hat{p}\). Do not infer bias solely from one sample result.

Key Takeaway

Sampling variability is the ordinary spread of sample proportions across samples. Bias is a systematic shift in the center of results produced by a method. A representative random sample can be unbiased yet variable; a flawed poll can have both variability and bias. Larger samples reduce variability, but they cannot by themselves repair a method that systematically favors some people or opinions.

Key takeaway: Ask two separate questions about a poll: How much would \(\hat{p}\) vary from sample to sample under its model, and does the sampling method systematically shift the estimates away from the population proportion? A small spread does not guarantee little bias.

Check Your Understanding

Use the distinction between variation and bias to explain each situation. When a numerical response-rate calculation is requested, show how the respondent mix is obtained.

  1. A random sample of 400 residents from a town of 20,000 is used to estimate the proportion who support a new recycling program. Explain why two such samples might produce different \(\hat{p}\) values even if the method is unbiased.
  2. In a population, 30% support a proposal and 70% oppose it. If 90% of supporters and 30% of opponents respond to a poll, find the expected proportion of supporters among respondents. Is it above or below the population proportion?
  3. Explain why increasing the number of responses in an open online poll may reduce variation among poll results without removing selection bias.
  4. A poll’s sample proportion is 0.08 higher than a trusted population proportion. Explain why this difference alone does not prove that the poll’s method is biased.
  5. In your own words, distinguish an estimate being unbiased from an estimate being exactly correct in every sample.