Random Selection Is Not the Same as Getting Responses
In Writing Condition Checks in Complete Sentences, you practiced naming the evidence for a condition, showing any relevant calculation, and stating what it means. A survey with nonresponse needs one more distinction: researchers may randomly select people to contact, but not everyone selected will reply. The response process can affect whether the people who provide data still represent the target population.
This matters when using a mail survey to estimate a population proportion. A random sample can support inference about the population from which it was selected, as discussed in Random Condition for Proportion Inference. But if some selected people do not respond, the actual data come only from those who replied. Those respondents may differ from nonrespondents in ways related to the survey question.
The key question is not simply, “Was the original sample random?” Also ask, “Do the people who responded provide a reasonable basis for inference about the target population?” In a mail survey, each person chooses whether to return the questionnaire. That choice is not random assignment to the respondent group.
Calculate the Response Rate, Then Assess Its Meaning
The response rate is the proportion of the people contacted who completed and returned the survey. The nonresponse rate is the proportion who did not respond. If everyone selected was eligible and there are no other outcomes to account for, the rates add to 100%.
For a 30% nonresponse rate, the response rate is 70%. That percentage describes how many people replied, not how similar respondents are to nonrespondents. A low response rate can make concern about nonresponse more serious, but a high response rate does not guarantee that the respondents are representative. Even a smaller group of nonrespondents could differ substantially on the characteristic being measured.
Nonresponse threatens the Random condition when response is connected to the survey characteristic, or to factors associated with it, so that the respondents systematically differ from the intended population. For example, if people who use a community service are more likely to return a survey about that service, the respondent results could overstate how common its use is. If nonresponse is unrelated to the characteristic, it might reduce the number of observations without producing this particular bias. The response rate alone does not tell us which situation applies.
The target population also matters. A mail survey sent to a random sample from a complete list can support inference to the people on that list if the response process does not undermine representativeness. It does not automatically represent people missing from the list, people outside the stated population, or all respondents in some broader group.
A response rate is therefore evidence to report and interpret, not a pass-or-fail cutoff that settles the Random condition. For a survey with 30% nonresponse, say that the initial sample was random if the description supports that claim, calculate the response rate, and explain that the nonresponse may introduce bias if response is related to the characteristic of interest. Do not declare the condition fully satisfied just because the original selection was random.
Worked Examples
Worked Example: A Random Mail Sample With 30% Nonresponse
A town’s research team randomly selects 500 households from its current residential mailing list and mails each a questionnaire about whether anyone in the household used a public trail during the past month. The team receives 350 completed questionnaires. Assess how the response rate affects the Random condition for estimating the proportion of households on the list that used a trail.
First calculate the response rate and nonresponse rate:
The initial selection was random, so that part of the survey design supports the Random condition for inference about households on the mailing list. However, the observed data come from the 350 households that returned the questionnaire, not from all 500 selected households. Whether trail use is related to the likelihood of returning the survey is unknown.
A complete condition assessment is: “The 500 households were randomly selected from the town’s residential mailing list, which supports the Random condition for that list. However, only 350 responded, so the nonresponse rate is \(150/500=30\%\). If households’ likelihood of responding is related to trail use, the respondents may not represent all households on the list. The Random condition is therefore a concern, and the response rate alone does not establish that the survey respondents are representative.”
This assessment neither asserts that the estimate is biased nor treats random selection as irrelevant. It states what the sampling method supports and identifies the uncertainty created by nonresponse. A careful report would limit its population claim to the mailing list and acknowledge that nonresponse may weaken the basis for inference to that population.
Worked Example: Comparing Response Rates Across Groups
A fictional college randomly selects 200 students from its enrollment list for a mailed survey about whether students have a campus meal plan. The selected sample contains 100 students living on campus and 100 living off campus. Of the on-campus students, 80 respond; of the off-campus students, 40 respond. What does this information suggest about nonresponse?
Calculate each group’s response rate:
There are 120 responses in all, so the overall response rate is \(120/200=0.60\), or 60%. Although the selected sample had equal numbers from the two housing groups, the respondents include 80 on-campus students and 40 off-campus students. Thus, two-thirds of the responses are from on-campus students:
The different response rates show that respondents do not preserve the selected sample’s equal split between these two groups. If housing status is related to having a meal plan, the respondent results could misrepresent the proportion of all selected students who have one. This information does not show whether the survey’s meal-plan estimate is too high or too low; it gives no meal-plan results for the nonrespondents.
A complete assessment is: “The students were randomly selected from the enrollment list, but response differed by housing status: 80% of on-campus students responded compared with 40% of off-campus students. The responding group is therefore not balanced across these groups in the same way as the selected sample. If housing status is related to meal-plan use, nonresponse could bias the estimate. The Random condition is a concern for inference about all students, and the direction or size of any bias cannot be determined from these response counts alone.”
Worked Example: Four-Step Assessment for a One-Proportion Inference
A fictional regional agency randomly selects 600 residents from its mailing list and mails a questionnaire asking whether residents support adding a neighborhood recycling pickup. The agency receives 420 replies, and 252 respondents support the proposal. The agency wants to use the results to estimate the proportion of all residents on the list who support the proposal. Assess the Random condition in a four-step response.
Let \(p\) be the proportion of all residents on the agency’s mailing list who support adding the recycling pickup. The target is residents on that list, not automatically every person in the region.
For a one-proportion inference, assess whether the process supports treating the observed responses as a random sample from the target population. The agency randomly selected residents, but the mail survey also depends on those selected choosing to respond. Calculate the response and nonresponse rates, then consider what nonresponse means for representativeness.
The response rate is \(420/600=0.70\), or 70%. The nonresponse rate is \((600-420)/600=180/600=0.30\), or 30%. The random selection supports the Random condition for the mailing list at the contact stage. However, only the 420 respondents provided support information. If opinions about recycling pickup are related to the likelihood of returning the survey, the respondents may differ systematically from all residents on the list. The response rate does not show whether that relationship exists.
The original random selection is a strength, but the 30% nonresponse creates a concern about the Random condition for inference to all residents on the list. The data description alone does not establish that nonresponse caused bias, nor does it show its direction or size. Any population conclusion should acknowledge that limitation.
The value \(252/420=0.60\) describes support among the people who responded. It is not automatically an unbiased estimate of the proportion among all 600 selected residents or all residents on the list. The nonrespondents’ opinions are unknown, so the observed respondent proportion cannot by itself resolve the concern.
What Additional Information Can and Cannot Establish
Information about response patterns can make an assessment more specific. For example, response rates by region, age group, or another known characteristic may reveal that some parts of the sampling frame were less likely to reply. If those characteristics are relevant to the survey response, that difference raises a concrete representativeness concern. As in the housing example, however, a group difference in response rates does not by itself tell us how the survey estimate is shifted; that requires information about the characteristic being studied.
A follow-up effort that reaches some initial nonrespondents may provide useful evidence about their views. Still, a follow-up response is not proof that every nonrespondent would answer similarly, and it does not automatically eliminate nonresponse concerns. In an AP response, use only the evidence given. Do not invent explanations for why people did not reply or claim that a particular response pattern corrected the survey.
Also keep nonresponse separate from sampling variability. A larger number of replies can reduce variability under a suitable sampling model, but it does not automatically remove systematic differences between respondents and nonrespondents. As explained in Sampling Variability Versus Bias in Proportions, variability and bias are different issues. A calculation of a margin of error describes sampling uncertainty under the procedure’s assumptions; it does not measure the possible effect of nonresponse.
Common Mistakes and AP Exam Tips
- Calling the respondents a random sample just because selection was random. State that the initial contact sample was random, then address the response process separately. Respondents chose whether to reply.
- Treating 30% nonresponse as proof of bias. It creates a concern, but the rate alone does not establish that respondents differ on the survey characteristic. Say “may introduce bias” or “raises a concern,” not “proves the estimate is biased.”
- Assuming 70% response guarantees representativeness. A majority response does not show that respondents and nonrespondents have similar views. The relationship between response and the characteristic matters.
- Claiming a direction without evidence. A response rate does not reveal whether the estimate is too high or too low. To claim a direction, the description must provide information connecting response patterns with the survey characteristic.
- Ignoring the target population. Name the population actually represented by the sampling frame. A random selection from a mailing list does not automatically represent people absent from that list.
- Confusing respondents with everyone selected. In the recycling example, 60% is the proportion supporting the proposal among respondents. Do not silently describe it as the proportion among all residents.
Key Takeaway
A mail survey’s random selection and its response rate describe different parts of how the data were produced. With 30% nonresponse, report that 70% replied, recognize that the observed data come from respondents, and explain why differences between respondents and nonrespondents could threaten the Random condition. The rate alone cannot prove bias, establish representativeness, or determine the direction of any effect.
Check Your Understanding
For each situation, distinguish what the response rate establishes from what it does not establish.
- A random mail sample includes 800 people, and 560 return a survey. Calculate the response and nonresponse rates. What concern should be mentioned when assessing the Random condition?
- A survey has a 90% response rate. Does that rate alone prove the respondents represent the target population? Explain why or why not.
- In a random sample of 300 students, 120 live in residence halls and 180 live off campus. Replies come from 96 residence-hall students and 72 off-campus students. Calculate each group’s response rate and explain what the difference may suggest.
- A random mail survey finds that 65% of respondents support a proposal, but 30% do not respond. Which group does the 65% directly describe, and why should you be cautious about generalizing it?
- Write a complete condition assessment for a randomly selected mail sample with 30% nonresponse. Include a statement about the initial selection, possible nonresponse bias, and what cannot be concluded from the rate alone.