When Selected People Do Not Respond
In Undercoverage Bias, you learned that a sampling frame can leave some members of the population out before a sample is selected. Nonresponse is different: people are selected for the sample, but some do not provide the requested information. They might refuse, be unavailable, or start but not complete the survey.
A random sample is a strong starting point, but the results can still be biased if the people who respond differ from those who do not on the variable being measured. A survey about sleep, for example, might miss selected people who work overnight and are hard to reach during the survey team’s calling hours. The key question is not simply how many people declined; it is whether responding and not responding are related to the answers the survey seeks.
A refusal is one kind of nonresponse. It is not the same as a voluntary response sample. In voluntary response, people choose themselves into the sample, often by answering an open invitation. In a survey with nonresponse, people are selected first, and some of those selected do not participate. As discussed in Undercoverage Bias, undercoverage happens earlier still: people are missing or inadequately represented in the frame used to select the sample.
Response Rate: Useful, but Not a Bias Measure
The response rate describes what fraction of selected people completed the survey. For a simple survey in which everyone selected is eligible, calculate it as the number of completed responses divided by the number selected. A survey protocol should specify how to handle people who are ineligible or whose eligibility cannot be determined; the examples here assume that everyone selected is eligible.
A low response rate means that many selected people did not provide data, so there is greater reason to investigate possible nonresponse bias. But the response rate alone cannot tell you whether bias occurred or which way an estimate may be off. A low-response survey could have little bias if respondents and nonrespondents are similar on the variable of interest. A high-response survey could still be biased if the few who do not respond differ substantially.
This is similar to the caution about undercoverage: identify the people who are missing from the results and consider how they might differ on the measured variable. With nonresponse, those missing people were in the selected sample, but their answers are unavailable. Unless the scenario supplies relevant information, avoid claiming that refusals definitely raised or lowered an estimate.
Worked Example: Refusals in a Park Survey
Worked Example: Do Refusals Change the Estimated Support?
Suppose a fictional city selects a random sample of 200 adult residents to ask whether the city should add evening hours at a public park. Of those selected, 120 complete the survey and 80 refuse or cannot be reached. Among the 120 respondents, 54 support the proposal. To illustrate the effect of nonresponse, suppose a later, successful follow-up establishes that 12 of the 80 initial nonrespondents support it.
State: The population of interest is the city’s adult residents. The variable is whether a resident supports adding evening park hours. The concern is that answers from the 120 respondents might not represent answers from all 200 selected residents if willingness to respond is related to support.
Plan: Calculate the response rate and compare support among initial respondents with support among the people initially missing from the results. The follow-up information in this constructed example lets us examine the difference; ordinarily, the answers of nonrespondents are unknown.
Do: The response rate is
Among the initial respondents, the proportion supporting evening hours is
Among the 80 initial nonrespondents, the proportion supporting the proposal is
Using the follow-up information, support among all 200 selected residents is
The initial respondent estimate of 0.45 is 0.12, or 12 percentage points, higher than the 0.33 proportion for all 200 selected residents. The same total can be checked by counting the 66 supporters and 134 residents who do not support the proposal: \(66+134=200\).
Conclude: In this constructed example, the initial respondents are more supportive than the initial nonrespondents, so using respondents alone overestimates support among the selected residents by 12 percentage points. Without information about the nonrespondents, we could identify a possible risk of bias but could not establish its direction or size.
Practical Ways to Reduce Nonresponse
Reducing nonresponse starts with making participation clear, respectful, and manageable. A well-designed contact plan gives selected people reasonable opportunities to respond without pressuring them. The goal is not to obtain answers at any cost; participation should remain voluntary and consistent with the ethical principles discussed in Ethical Considerations in Data Collection.
- Explain the request clearly. Briefly state who is conducting the survey, why the person was selected, how long participation is expected to take, and how the information will be used.
- Make responding convenient. Offer an accessible format, such as a short online survey or a phone option when appropriate. Provide a reasonable range of times to respond.
- Follow up respectfully. Send or make a limited number of reminders to people who have not responded. Different contact times may reach people who were unavailable earlier.
- Protect privacy honestly. Explain confidentiality protections accurately. Do not promise anonymity or security that the survey cannot guarantee.
- Consider a modest incentive. A small, appropriate incentive may encourage participation, but it should not be coercive or presented as a guaranteed way to eliminate bias.
A follow-up can improve participation and provide clues about nonresponse. For example, a second contact might reach people who were unavailable the first time. But people who respond after several attempts might still differ from those who never respond. Replacing nonrespondents with easier-to-reach people is not a reliable fix: replacements may have different characteristics and do not supply the missing selected individuals’ answers.
Worked Example: Improving Response to a Student Survey
A fictional school randomly selects 300 students to ask about access to quiet study spaces. After one online invitation, 180 students complete the survey, and 120 do not. Among the first 180 respondents, 72 report that they have access to a quiet study space. The school sends a neutral reminder and offers a brief phone option. An additional 54 students respond; 18 of those students report having access.
State: The population of interest is the school’s students. The variable is whether a student has access to a quiet study space. The school wants to increase participation among the students selected and consider what the first round of nonresponse could mean for its estimate.
Plan: Calculate the response rate after the first invitation and after follow-up. Then calculate the proportion reporting access among the combined respondents. The follow-up may reduce nonresponse, but the answers of the remaining nonrespondents are still unknown.
Do: After the first invitation, the response rate is
After 54 additional students respond, there are \(180+54=234\) responses. The new response rate is
The total number of respondents reporting access is \(72+18=90\). The proportion reporting access among the 234 respondents is
So about 38.5% of the combined respondents report access. As a check, the 234 respondents consist of 90 who report access and \(234-90=144\) who do not; \(90/234\) is the same proportion.
Conclude: The reminder and phone option raised the response rate from 60% to 78%, making the results available from more of the selected students. The estimate of about 38.5% describes those who responded, not automatically all students. The school should still consider whether the 66 remaining nonrespondents might differ in access to quiet study spaces.
Use Known Differences Carefully
Sometimes a survey team knows that response rates vary across groups—for example, by grade, region, or a characteristic already recorded for everyone in the sample. Comparing these rates can identify where nonresponse is concentrated. If those groups may also differ on the survey variable, the respondent results could overrepresent one group and underrepresent another.
When a group’s size in the sample is known, an analyst may adjust a respondent estimate to reflect the sample’s group proportions rather than the respondents’ proportions. This approach can help if the groups are meaningfully related to the response variable and if respondents within each group are reasonably representative of nonrespondents in that same group. It cannot correct for unknown differences within groups, so an adjusted result still needs cautious interpretation.
Worked Example: Different Response Rates Among Commuters
A fictional transportation survey selects 100 commuters: 60 who regularly use public transit and 40 who rarely use it. Of the regular transit users, 48 respond and 36 support a proposed bus-lane change. Of the commuters who rarely use transit, 16 respond and 4 support the change.
State: The variable is whether a commuter supports the bus-lane change. The concern is that the respondents include a different mix of regular and infrequent transit users than the selected sample does.
Plan: Compare response rates by group, calculate the support proportion among all respondents, and then calculate an illustrative adjustment using the groups’ proportions in the selected sample. The adjustment assumes that respondents within each group represent nonrespondents in that group.
Do: The response rate among regular transit users is
Among commuters who rarely use transit, the response rate is
There are \(48+16=64\) respondents, including \(36+4=40\) supporters. The unadjusted support proportion among respondents is
Regular transit users make up \(60/100=0.60\) of the selected sample, and infrequent users make up \(40/100=0.40\). The group support proportions among respondents are \(36/48=0.75\) and \(4/16=0.25\), respectively. Applying the selected-sample group proportions gives
The adjusted value is 0.55, compared with the unadjusted respondent proportion of 0.625. As a check on the group weighting, the weights add to \(0.60+0.40=1.00\), and the weighted support counts per selected person are \(0.60(0.75)=0.45\) and \(0.40(0.25)=0.10\), totaling 0.55.
Conclude: Regular transit users responded at a higher rate and were more supportive among respondents, so they make up a larger share of respondents than of the selected sample. The illustrative adjustment reduces the estimate to 0.55 by restoring the sample’s original group proportions. It would be appropriate only if respondents within each group provide a reasonable account of nonrespondents in that group; the calculation does not prove that assumption or remove all possible nonresponse bias.
Common Mistakes and AP Exam Tips
- Treating the response rate as the amount of bias. A response rate tells you how many selected people answered, not how different their answers are from nonrespondents’ answers. Explain what the rate suggests and what information is still missing.
- Claiming a direction without evidence. Say that nonresponse could bias a result when you do not know how nonrespondents differ. State a likely direction only when the scenario supports it.
- Confusing nonresponse with voluntary response or undercoverage. A full-credit answer identifies the mechanism: selected people do not respond, people choose themselves into a sample, or the frame leaves people out.
- Assuming reminders eliminate the problem. Reminders may increase participation, but those who continue not to respond may still differ from respondents. Describe the improvement without claiming that it guarantees an unbiased result.
- Reporting respondents’ answers as the population’s answers. Say whom the observed proportion describes. Generalizing to the population requires considering both how the sample was selected and whether nonresponse could matter.
- Using an adjustment without stating its assumption. If you reweight groups, explain that the adjustment relies on respondents representing nonrespondents within the groups used. It cannot account for differences the groups do not capture.
For a strong AP response, identify the selected sample, describe who did not respond, and connect possible differences to the variable being measured. Use precise, cautious language: “The refusals may bias the estimate if those who refused differ from respondents in their support for the proposal.” If the scenario establishes how they differ, explain the likely direction; otherwise, do not invent one.
Check Your Understanding
Answer each question using the distinction between selected people who respond and those who do not.
- A random sample of 250 residents is selected for a survey. If 175 complete it, calculate the response rate and state what population the resulting answers directly describe.
- In a survey, people who refuse are known to be especially likely to work night shifts. Explain why that fact alone does not establish whether the estimated average sleep time is too high or too low.
- Explain how nonresponse differs from undercoverage and from voluntary response.
- A survey team sends one reminder and increases participation. Give one reason why this may reduce concern about nonresponse without proving that nonresponse bias is absent.
- In a sample, two age groups have different response rates. What additional assumption would be needed to use an adjustment based on the age groups to address nonresponse bias?