Tutorials › AP Statistics › Identifying Bias and Confounding in Scenarios

Statistical practices and exam synthesis · Tutorial 1006 of 1020

Identifying Bias and Confounding in Scenarios

Use who was selected, how answers were gathered, and what else could explain a comparison to identify bias and confounding.

Intermediate 9 min read

What You'll Learn

  • Distinguish sampling bias from inaccurate or influenced survey responses
  • Separate nonresponse from response bias
  • Identify a lurking variable that could confound an observed association
  • Explain how a flaw could distort a specific conclusion
  • State when a scenario does not establish the direction of distortion
  • Check whether a study has more than one limitation

Look for the Pathway From Study Design to Distorted Conclusion

In “Reading Study Descriptions for Design Flaws,” you checked who could enter a study, who responded, and how groups were compared. Now sharpen that scan: identify whether a scenario raises a concern about who is represented, what people report, or what else might explain an observed association. These questions help distinguish sampling bias, response bias, and confounding.

All three can weaken a conclusion, but they do so in different ways. Sampling bias concerns the process used to select or reach individuals. Response bias concerns answers that are systematically inaccurate or influenced. Confounding concerns a third variable that makes it difficult to separate the relationship between an explanatory variable and a response variable from another relationship.

Key idea: Do more than attach a label. Identify the feature in the scenario, explain how it could affect the data or comparison, and say which conclusion is therefore limited. If the scenario does not tell you whether the result would be too high or too low, do not invent a direction.

Three Different Questions to Ask

A short study description can be read with three questions in mind. The questions focus on different stages of the data process, so a study may raise more than one concern.

1
Who had a chance to be represented?
Look at how individuals were selected or reached. Ask whether the method systematically favors some members of the target population or leaves others out.
2
Could the answers be systematically inaccurate?
Consider the wording, setting, interviewer, memory demands, or social pressure involved in reporting the measured information.
3
Could another variable help explain the comparison?
For an association between an explanatory variable and a response variable, look for a third variable related to both that could account for some or all of the observed association.
4
Connect the concern to the claim.
Say whether the issue limits a population estimate, the accuracy of reported information, or an explanation of why groups differ.

Sampling Bias: Who Is Overrepresented or Missing?

Sampling bias occurs when the way a sample is selected or recruited systematically favors some members of the target population over others. The resulting sample may differ from the population in ways related to the question being studied. A biased sample can distort a population estimate or a comparison intended to represent a broader group.

A convenience sample, such as asking only people leaving one particular store, may overrepresent people who shop there. A voluntary-response sample, such as an open online poll, may attract people with especially strong opinions. Undercoverage, discussed in the previous tutorial, is another way the sampling process can fail to represent the target population: some members have little or no chance to be included.

Do not confuse sampling bias with ordinary sampling variability. Different probability samples can produce different estimates by chance. Sampling bias is a systematic concern about who is more likely to be included. Also, the fact that a sample is not perfectly representative does not tell you the direction of the resulting error unless the scenario provides evidence about which kinds of people are overrepresented and how they differ.

Response Bias: Could the Answer Itself Be Distorted?

Response bias occurs when the answers people give are systematically inaccurate or influenced by how information is collected. For example, a leading question may encourage a particular answer. People may underreport behavior they consider embarrassing, overstate behavior they consider admirable, or recall past events inaccurately. An interviewer’s presence or the setting in which questions are asked may also affect responses.

Response bias is different from nonresponse. Nonresponse concerns selected or contacted individuals who do not provide answers; response bias concerns answers that are provided but may not accurately reflect the information being measured. As in the earlier tutorial on study-design flaws, nonresponse can make the respondents differ from the people who do not reply. Keep these concerns separate when describing a scenario.

A response-bias explanation should name a plausible pathway. For instance, “the survey may be biased” is vague. “Students may underreport missed homework because they answer in front of their teacher” explains how the setting could influence the recorded responses. Unless the scenario establishes that the answers actually changed, describe this as a risk rather than a proven effect.

Confounding: Could a Third Variable Explain the Association?

Confounding occurs when the relationship between an explanatory variable and a response variable is mixed up with the effect of another variable. The third variable, often called a confounding variable or a lurking variable, is associated with the explanatory variable and is also related to the response variable. Because the influences overlap, the study cannot clearly separate the relationship of interest from the third variable’s relationship with the outcome.

For example, suppose students who choose to attend a tutoring program have higher scores later than students who do not attend. If students who were already scoring higher are also more likely to choose tutoring, prior achievement could be related to both program attendance and later scores. The observed difference does not by itself show that tutoring caused the higher scores.

Confounding is not simply “there might be another variable.” Explain how that variable connects to both variables in the comparison. Also, do not claim that a suggested confounder definitely explains the result if the scenario only makes it plausible. In an observational study, a confounding variable can make a causal conclusion less secure even when the observed association is real.

Keep the categories straight: Sampling bias is about which individuals enter the data. Response bias is about whether recorded answers are systematically inaccurate or influenced. Confounding is about whether another variable makes an observed relationship difficult to interpret. More than one category can apply to the same study.

Worked Examples: Trace the Possible Distortion

Worked Example: An Open Poll About Transit

Situation. A fictional regional website posts an open poll asking, “Do you support expanding bus service in your area?” Anyone who visits the page can vote. The poll’s summary is presented as evidence of what all residents want.

Identify the target and selection method. The target is all residents in the region, but the poll includes only people who visit the website and choose to vote. Residents who do not use the site cannot take part through this poll, and visitors with especially strong views may be more likely to respond. This is a voluntary-response method and raises a risk of sampling bias.

Explain how the conclusion could be distorted. If people who favor expansion are especially motivated to vote, support could be overrepresented. If opponents are more motivated, opposition could be overrepresented instead. The scenario does not say which group is more likely to vote, so it does not establish the direction of possible distortion.

Conclude in context. The poll describes the opinions of the people who chose to vote on that website. Because residents were not selected through a method that represents all regional residents, the result does not reliably establish the level of support among all residents. A carefully selected probability sample would better address the question about the region’s population.

Worked Example: A Sensitive School Survey

Situation. A fictional school asks students to report how many days they skipped breakfast in the past month. Students complete the form in class, and their teacher collects it with names visible. The school uses the answers to estimate how common skipping breakfast is among its students.

Identify the concern. The central concern is response bias. Students may not remember the number of days accurately, and some may report fewer days because they feel uncomfortable giving an answer their teacher might judge. The visible names and collection setting provide a specific reason answers might not fully reflect students’ actual behavior.

Explain the possible distortion. If students who skipped breakfast more often underreport their behavior, the survey could underestimate how common skipping breakfast is. This direction is plausible given the setting, but the description does not prove that students changed their answers.

Conclude and suggest an improvement. The collected answers can be summarized as students’ reported breakfast habits, but they may not accurately estimate actual behavior if students underreport. An anonymous survey collected without the teacher observing individual responses could reduce pressure to give socially acceptable answers. It would not guarantee perfect recall or remove every source of error.

Worked Example: Sunscreen Use and Sunburn

Situation. In a fictional observational study, researchers compare sunburn reports among beachgoers who say they use sunscreen regularly and those who do not. The sunscreen users report more sunburns. The researchers conclude that sunscreen increases the chance of sunburn.

Identify the association and proposed conclusion. The explanatory variable is regular sunscreen use, and the response variable is reported sunburn. The study observes an association, but it does not randomly assign sunscreen use. Therefore, the comparison alone does not establish a cause-and-effect relationship.

Look for a possible confounding variable. Time spent outdoors could be related to both sunscreen use and sunburn. People who spend more time at the beach may be more likely to use sunscreen, while also having more exposure to sunlight and more opportunity to get sunburned. Outdoor time could therefore help explain why the sunscreen users report more sunburns.

Conclude without reversing the evidence. The study found more reported sunburns among the observed sunscreen users, but the comparison does not show that sunscreen caused more sunburns. Differences in outdoor exposure could confound the association. This possibility does not prove that outdoor time explains the entire difference, nor does the study establish that sunscreen has no protective effect.

Worked Example: Choosing a Tutoring Program

Situation. At a fictional high school, students choose whether to attend an after-school mathematics tutoring program. At the end of the term, students who attended have higher average test scores than students who did not attend. A coordinator says the program raised scores.

Check group formation. Students chose whether to attend; the description does not report random assignment. This is an observational comparison, so students in the two groups may have differed before the term began.

Identify a plausible confounder. Prior achievement could be related to both program attendance and end-of-term scores. For example, students who were already doing well might be more motivated to seek extra practice, and prior achievement would also be related to later scores. Conversely, students who are struggling might be more likely to seek tutoring. The scenario does not tell us which pattern occurred, so do not assume the direction.

Conclude in context. Students who attended tutoring had higher average test scores at the end of the term, but the observed difference does not establish that tutoring caused higher scores. Because students selected their own participation, prior achievement or other differences could help explain the comparison. Random assignment, if practical and appropriate, would help make the groups comparable at the start.

Common Mistakes and AP Exam Tips

  • Using “bias” without identifying its source. State whether the concern is selection, inaccurate answers, or a third variable, then describe the specific feature of the scenario.
  • Calling every survey problem response bias. A group that is systematically more likely to be selected points to sampling bias. A selected person who does not reply raises a nonresponse concern. An answer that is influenced or inaccurate raises response bias.
  • Claiming a direction without evidence. If a poll may attract people with strong opinions but the scenario does not say which opinion is stronger, say the result could be distorted. Do not claim support must be overestimated or underestimated.
  • Naming a confounder without explaining its connections. A full explanation says how the third variable could be related to the explanatory variable and to the response variable. A variable that is unrelated to one side does not explain confounding in the stated comparison.
  • Treating confounding as proof that an association is false. Confounding limits how the association can be interpreted; it does not show that the observed difference disappears or that the explanatory variable has no effect.
  • Mixing up association and causation. In an observational comparison, use wording such as “was associated with” or “the groups differed.” Do not say a factor caused the outcome unless the design supports that claim.
  • Listing several flaws without ranking or linking them. Identify the concern that most directly affects the claim, then add another only if the scenario supports it. Explain each one’s separate pathway to distortion.

A strong short answer often follows this pattern: “Because [specific design feature], [specific group or measurement] may be distorted or another explanation may remain. Therefore, the study does not establish [the particular population or causal claim].” If the direction is supported by the scenario, state it; otherwise, describe the risk without guessing.

Key takeaway: Sampling bias concerns who is represented, response bias concerns the accuracy or influence of answers, and confounding concerns an alternative explanation for an association. For each, name the pathway and limit only the conclusion that the scenario does not support.

Check Your Understanding

For each scenario, identify the most relevant concern and explain how it could affect the conclusion. If more than one concern is plausible, distinguish them.

  1. A town asks people leaving its largest fitness center whether residents exercise regularly, then uses the answers to describe all adults in town. What selection concern should be considered?
  2. A survey asks students whether they have ever copied homework, and students complete it while their names are visible to the teacher. What kind of bias could affect their answers, and in what plausible direction?
  3. People who buy gardening classes have more plants at home than people who do not. Could a third variable help explain the association? Name one and explain its possible connection to both variables.
  4. A randomly selected group receives a survey, but many selected people do not reply. Is this the same as response bias? Explain the distinction.
  5. A news report says a voluntary online poll proves that most residents oppose a proposed park. What can the poll describe, and what conclusion about residents is not well supported?