Trace the Survey to Find the Source
In Response Bias and Question Wording, you learned that answers can be influenced by the survey process. In this tutorial, you will practice identifying that and other sources of bias in short scenarios, then explaining how the bias could affect the result. The key is to trace what happened from the population of interest to the reported answers.
Start by naming the population and the variable being measured. Then ask how people entered the survey, whether the sampling frame covered the population, whether selected people responded, and whether the survey process could have influenced their answers. This order helps separate problems that can sound similar. For example, a person who was never included in the selection process was not a nonrespondent; a selected person who did not answer may be.
The direction of bias describes the likely direction of error in the statistic compared with the population value. A survey might overestimate the proportion of residents who support a proposal, for instance, if the people whose answers are counted are more likely to support it than residents overall. It might underestimate that proportion if those counted are less likely to support it. The direction depends on the particular population, variable, and selection or response process—not just the name of the bias.
Separate the Main Sources
A convenience sample includes people who are easy to reach. Ask whether ease of access is connected to the variable being studied. If it is, the sample may lean toward particular answers. In Convenience Sampling and Its Bias, you saw why availability is not a substitute for chance-based selection.
A voluntary response sample forms when people choose themselves in response to an open invitation. People with strong opinions or experiences may be especially likely to participate. A voluntary response sample is not simply a random sample with a few selected people failing to respond: the invitation lets people decide whether to enter the sample in the first place. See Voluntary Response Bias for the distinction.
Undercoverage occurs when the sampling frame omits some members of the population or represents them inadequately. Random selection from that frame cannot select people who are missing from it. The direction depends on how the omitted or underrepresented group differs on the variable, as discussed in Undercoverage Bias.
Nonresponse bias is a concern when people selected for a sample do not provide information and their answers may differ from those who do respond. The important comparison is between respondents and nonrespondents on the variable of interest, not merely whether the response rate is low. This builds on Nonresponse Bias.
Response bias concerns answers that are systematically influenced by question wording, social expectations, interviewer behavior, or another feature of the survey process. It can occur even when selection was appropriate and everyone responds. A low response rate, by contrast, points to a possible nonresponse problem; it does not, by itself, show that the answers of respondents were influenced.
Worked Example: A Convenient Location
Worked Example: Support for a New Protected Bike Lane
A town wants to estimate the proportion of its adult residents who support building a protected bike lane. To save time, a survey team asks people leaving a cycling shop whether they support the plan. The people approached are mainly customers who ride bikes regularly. The survey team reports that support among those asked is high.
State: The population is all adult residents of the town. The variable is whether a resident supports the proposed bike lane. The sample was chosen because its members were easy to approach at a cycling shop, not by chance from a frame covering town residents.
Plan: Identify the selection method, then consider whether people selected in that way may differ from the full population on the variable. This is a convenience sample. To discuss direction, compare the likely views of regular cyclists with those of town residents overall.
Do: Regular cyclists may be more likely than residents overall to favor a protected bike lane. If that is true, their answers would make support appear more common in the survey than it is among all adult residents. The likely direction is an overestimate of the town-wide support proportion.
Conclude: Because the survey only approached people leaving a cycling shop, the sample may overrepresent residents who ride bikes and favor the proposal. The result could overestimate support among all adult residents. The direction is plausible given the scenario, but the survey alone does not establish how every resident—or even every cyclist—feels.
Worked Example: An Open Online Invitation
Worked Example: Opinions About Longer Library Hours
A public library posts an open online poll asking whether it should stay open later on weekdays. The poll is shared on the library’s social media account. People who see the post can click a link and submit an answer. Many of the responses come from frequent library users who strongly want more evening access.
State: The population is the library’s community, and the variable is whether a community member supports longer weekday hours. The poll does not select a chance-based sample; people decide for themselves whether to respond.
Plan: Classify how people entered the poll and ask whether willingness to respond may relate to the opinion being measured. This is a voluntary response sample. The fact that the poll is online is not, by itself, the key diagnosis; the defining feature is that anyone who encounters the invitation chooses whether to participate.
Do: People who strongly want later hours may be especially motivated to answer. If supporters are more likely to participate than community members who are indifferent or opposed, supporters will make up too large a share of the responses. The poll could overestimate support for longer hours in the community.
Conclude: This open poll may overestimate community support because people with strong interest in later hours may be more likely to submit an answer. The poll describes the responses received, but those responses do not provide a dependable estimate for the whole community.
Worked Example: A Missing Group in the Frame
Worked Example: A Mail Survey About Access to Health Appointments
A fictional county wants to estimate the proportion of adult residents who can get a routine health appointment within two weeks. The survey team randomly selects addresses from a residential mailing list and mails a questionnaire. The list omits residents who do not have a stable mailing address. In this county, those residents are more likely than other adults to have difficulty getting an appointment.
State: The population is all adult residents of the county. The variable is whether a resident can get a routine appointment within two weeks. The sampling frame is the residential mailing list, which does not adequately cover adults without stable mailing addresses.
Plan: Separate the selection method from the frame’s coverage. The team selected addresses at random, but random selection cannot include adults who are absent from the list. This is undercoverage. Use the stated difference between the omitted group and other adults to reason about the direction.
Do: Adults without stable mailing addresses are more likely to have difficulty getting an appointment. Leaving them out means the survey underrepresents people who cannot get an appointment within two weeks. The estimated proportion able to get an appointment promptly could therefore be too high.
Conclude: The mailing list’s undercoverage may cause the survey to overestimate the proportion of all county adults who can get a routine appointment within two weeks. Random selection from the list does not fix the omission, because the missing adults had no chance to be selected.
Worked Example: Selected People Who Do Not Reply
Worked Example: Satisfaction With a Clinic Visit
A fictional clinic randomly selects recent patients from its appointment records and mails them a survey about satisfaction with their visit. Patients who had an especially smooth experience are likely to return the survey. Patients who had a frustrating experience may be less willing to spend time on it. The team summarizes the answers from returned surveys as an estimate of patient satisfaction.
State: The population is the clinic’s recent patients, and the variable is whether a patient was satisfied with the visit. Some patients were selected but did not respond, so the concern is whether respondents and nonrespondents differ in satisfaction.
Plan: The sample was selected from patient records, but not everyone selected provided an answer. This suggests possible nonresponse bias, rather than voluntary response: the patients who did not answer had been selected first. Use the scenario’s description of likely response behavior to assess direction.
Do: If satisfied patients are more likely to return the survey than dissatisfied patients, the responses will include a higher proportion of satisfied patients than the selected group or patient population. The estimated satisfaction proportion could be too high.
Conclude: The survey may overestimate satisfaction because dissatisfied selected patients may be less likely to respond. The direction follows from the stated difference in response likelihood; it would not be justified from the existence of nonresponse alone.
When the Direction Is Unclear
A source of bias does not always tell you whether the result is too high or too low. Suppose a survey uses a convenience sample of people at a train station to estimate support for a new local parking rule. The scenario does not say whether people who use the station are more or less likely to support the rule than residents overall. You can identify convenience sampling, but you cannot confidently determine the direction from the information given.
In that case, say what is known and what is missing: “This is a convenience sample because participants were approached at the train station. The result may not represent all residents, but the direction of bias cannot be determined without information about how station users’ views differ from residents’ views.” This is a stronger answer than guessing. It explains both the concern and the limit of the evidence.
Keep bias separate from ordinary sampling variability, introduced in Sources of Variability in Collected Data. Different random samples can produce different statistics even when the method is not systematically biased. A scenario that says only that one sample result is high does not establish bias. Look for a feature of selection, coverage, response, or measurement that could systematically shift results.
Common Mistakes and AP Exam Tips
- Naming a source without linking it to the variable. “This is nonresponse bias” is incomplete. Explain which selected people are less likely to answer and how their answers might differ on the variable.
- Calling all self-selection nonresponse. If an open invitation lets people choose whether to join, identify voluntary response. If people were selected first and some of them did not answer, consider nonresponse.
- Assuming random selection eliminates every problem. Random selection from an incomplete frame can still leave undercoverage. Selected people can also fail to respond, and respondents can still be influenced by the survey process.
- Claiming a direction just from the bias label. Convenience sampling does not automatically overestimate a result, and nonresponse does not automatically make it too low. State how the affected group is likely to differ on the measured variable.
- Mixing up who is missing and who is answering inaccurately. Undercoverage concerns people missing or poorly represented in the frame; nonresponse concerns selected people who do not answer; response bias concerns how answers are given.
- Forgetting to name the comparison. Say whether the statistic may be too high or too low compared with the value for the population of interest, and specify the population and variable in context.
A full-credit explanation usually follows a simple pattern: name the source, point to the detail in the scenario, identify the group or answers affected, and connect that difference to the variable. For example: “Because patients who were dissatisfied are less likely to return the survey, respondents may be more satisfied than all recent patients, causing the estimated satisfaction proportion to be too high.” If the scenario does not establish how the affected group differs, state that the direction cannot be determined.
Check Your Understanding
For each scenario, name the most relevant possible source of bias and explain the direction, if the information supports one.
- A student newspaper posts a poll asking whether the school should add more sports teams. Readers can choose whether to submit a response, and students with strong opinions share the poll widely.
- A random sample is drawn from a list of all current phone numbers, but the survey aims to estimate the views of all town residents. Residents without listed phone numbers are more likely to favor the proposal being studied. What possible source of bias is present, and what direction might result?
- People selected for a survey about weekly exercise do not respond at equal rates. The scenario does not say whether responders exercise more or less than nonresponders. What can you identify, and what can you not determine?
- A survey interviewer asks a neutral question about following a school rule, but praises students who say they follow it. Identify the possible source of bias and explain how answers might shift.
- A researcher asks people leaving a grocery store about how often they shop at that store. What selection method is being used? Can you determine the direction of bias from the information given? Explain.