Why the Selection Method Matters
In Population, Sampling Frame, and Sample, we distinguished the group a study wants to learn about, the list used to find people, and the people actually selected. Now consider how those people are selected. A sample can be large and still give a misleading picture if the selection method favors certain members of the population.
Suppose a community center wants to learn how often residents use its walking paths. Asking people who happen to be on a path would be convenient, but it would tend to reach path users and miss residents who rarely or never go there. The issue is not just sample size. The selection method has made some perspectives much easier to include than others.
Selection bias occurs when the way individuals enter a sample systematically favors some members or types of members over others. A chance-based selection process helps prevent a researcher from choosing only people who are easy to reach, seem typical, or are expected to give a particular answer. It does not make every possible sample equally informative, and it does not guarantee that the selected sample will match the population perfectly.
This distinction connects to Generalizing Results to a Population and Scope of Inference: Four Combinations. Random selection supports generalizing from a sample to the population from which the sample was selected. It is separate from random assignment, which is used in experiments and supports cause-and-effect conclusions. A random sample does not, by itself, show that one variable causes another.
Chance Reduces Favoritism, Not All Error
A chance-based selection method gives the selection process a rule that does not depend on a person choosing the individuals they prefer. For example, a school could use a random-number generator to select student IDs from an accurate roster, then invite the selected students to report their weekly screen time. The method is different from asking students who happen to be in the library or inviting students to answer an open online poll.
Random selection matters even when nobody intends to be unfair. People who are easiest to contact may have different experiences from people who are difficult to reach. A researcher might unknowingly choose people who are nearby, available, or familiar. Chance offers a systematic alternative to those judgment-based choices.
However, a random selection process does not promise a sample that mirrors every population feature. Different chance-based samples from the same population can include different people and produce different statistics. This ordinary sample-to-sample variation is sampling variability, as introduced in Sources of Variability in Collected Data. Random selection makes that variability part of a chance process rather than a result of deliberate or predictable favoritism.
A useful way to think about the benefit is that chance makes the selection process defensible. If the frame is suitable and the selection is genuinely random, the sample is not simply a group chosen because it was convenient or appealing. That supports generalization, but the conclusion must still be limited to the population the frame can represent.
Worked Example: A Random Sample of Households
Worked Example: Solar Panels in a Fictional Community
Imagine a fictional community with 240 households, all listed once on an accurate, complete household roster. In this illustrative population, 96 households have solar panels. A researcher selects 40 households from the roster using a chance-based method that gives every household an equal probability of selection, and 17 of the selected households have solar panels. Compare the sample and population percentages. What does the difference tell us?
State: The population is all 240 households on the roster, and the sample is the 40 households selected from it. The variable is whether a household has solar panels. The goal is to use the sample to learn about the population percentage.
Plan: Compare the sample proportion with the population proportion. Because the sample was selected by chance from a stated complete frame, it can support generalization to these 240 households. The sample proportion is still a statistic, so it may not equal the population proportion.
Do: The sample proportion is the number of sampled households with panels divided by the sample size of 40:
The population proportion is the number of households with panels divided by the population size of 240:
The difference is \(42.5\%-40.0\%=2.5\) percentage points. The sample percentage is 2.5 percentage points higher than the population percentage in this example.
Conclude: The random sample’s 42.5% is close to, but not identical to, the population’s 40.0%. The difference is an example of sampling variability: chance selected 17 households with panels out of 40, rather than exactly 16. Random selection supports using the sample to learn about the listed households, but it does not guarantee an exact match.
The sample size is explicitly 40, so its percentage is \(17/40\), not \(17/240\). The population percentage uses the population count and population size, \(96/240\). Keeping the two denominators attached to their correct groups is essential.
Random Selection Cannot Repair a Poor Frame
As explained in Population, Sampling Frame, and Sample, the sampling frame is the list or source used to select individuals. A random selection from a frame can only choose from people or units that the frame includes. If some members of the population are missing, they have no chance of being selected from that frame.
This is called undercoverage: part of the population of interest is not adequately represented in the sampling frame. Randomly selecting from an undercovered frame does not make the missing people appear in the sample. In the same way, a chance-based sample from a list that includes ineligible entries can select those entries. First ask whether the frame fits the population; then consider how selection was carried out.
Worked Example: A Random Draw from an Incomplete List
Worked Example: Residents and an Online Portal
A city wants to estimate how many hours per week all adult residents spend using public parks. It selects a random sample from the city’s online recreation-portal account list. The list includes 8,000 account holders, but many residents do not have an account. Is random selection from this list enough to generalize to all adult residents?
Identify the population: The population is all adult residents of the city, because that is the group the question asks about.
Identify the frame: The sampling frame is the list of 8,000 recreation-portal account holders. It is not the same as all adult residents, because residents without accounts are absent.
Assess the selection: The city uses chance to select people from the account list. That helps prevent staff from choosing particular account holders, but it does not give non-account holders a chance to be selected from this frame.
Conclude in scope: The random selection supports generalizing to the account holders represented by the list, subject to participation and measurement issues. It does not, by itself, support generalizing to all adult residents. If account holders differ from non-account holders in park use, the estimate for the listed group could systematically differ from the citywide value.
The concern is not that random selection was useless. It was useful for choosing from the frame without judgment. The limitation is that the frame does not cover the full population of interest. A better chance-based draw from a more complete frame would be needed to support the broader citywide conclusion.
Nonresponse Is Another Limit
Selection and response are separate steps. A person can be randomly selected but not answer the survey. As in Population, Sampling Frame, and Sample, the selected sample and the respondents are not always the same group. If response rates or willingness to answer are related to the variable being measured, the responses may systematically differ from what would have been learned from all selected people.
For instance, suppose a random sample of residents is selected for a survey about park use, but residents who rarely visit parks are less likely to respond. The initial selection was chance-based, yet the completed responses could overrepresent frequent visitors. Random selection supports generalization only when the selection process and the data collected allow a reasonable connection to the population. A researcher should report nonresponse rather than quietly treating respondents as if they were the entire selected sample.
- The population of interest is clearly defined.
- The sampling frame covers that population adequately and does not include many ineligible units.
- Individuals are selected by a genuine chance process, not by convenience or preference.
- Nonresponse does not create a major difference between selected individuals and respondents, or its effect is addressed and reported.
Worked Example: Randomly Selected, but Not All Responding
Worked Example: A Community Garden Survey
A neighborhood association has a current list of 500 households in its service area. It uses a chance-based method to select 50 households and asks how many times each household used the community garden last month. Thirty-five households respond. The association wants to describe garden use among all 500 households. What can it say about selection and generalization?
Selection: The 50 households selected by chance form the sample. The fact that only 35 respond does not change the number selected. The 35 households are the respondents.
Response count: The number of selected households that did not respond is \(50-35=15\). The response proportion among those selected is:
Assess the frame and method: The list is stated to be current and to cover the 500 households in the service area, and selection was by chance. Those features support generalization beyond the respondents. But the 15 nonrespondents may differ from the 35 respondents in garden use; the response rate alone cannot show whether they do.
Conclude in context: The association can report the responses and explain that 35 of the 50 selected households responded. Random selection supports the plan to learn about households on the list, but the association should be cautious about treating the 35 responses as a complete representation of all 500 households unless it has reason to believe nonresponse did not distort the results.
This example separates three numbers: 500 households in the population and frame, 50 households selected, and 35 respondents. Reporting the correct group for each number makes the limits of the conclusion clear.
What Random Selection Does—and Does Not—Support
A clear conclusion matches its scope to the way the sample was obtained. If a random sample is selected from a frame that represents a defined population and the response process does not seriously undermine that connection, sample findings can be generalized to that population. If the frame covers only a subgroup, the conclusion should be limited to that subgroup.
Random selection does not mean that every selected person will be typical, that the sample statistic equals the parameter, or that every possible source of bias has disappeared. It also does not turn an observational study into an experiment. As covered in Observational Studies and Their Limits and Association Versus Causation in Study Conclusions, generalizing a result and claiming a cause are different kinds of conclusions.
In the next tutorial, Simple Random Sample Defined, we will give a more specific name and definition to one chance-based sampling method. For now, the central point is that random selection is a way to let chance—not convenience or preference—guide who enters the sample.
Common Mistakes and Full-Credit Communication
- Assuming random selection guarantees a representative sample. Say that chance-based selection reduces selection bias and supports generalization, not that it guarantees the sample will match the population exactly. Acknowledge sampling variability.
- Confusing random selection with random assignment. Selection concerns who enters a study and can support generalizing to a population. Assignment concerns which treatment experimental units receive and can support a cause-and-effect conclusion.
- Using “random” to describe an arbitrary choice. A researcher choosing whoever is nearby has not used random selection. Full-credit wording identifies a chance-based method, such as a random-number draw from a stated frame.
- Ignoring the frame. A random sample from an incomplete list cannot represent people absent from that list merely because the draw itself was random. Name the frame and the population, then explain any mismatch.
- Forgetting nonrespondents. Report how many were selected and how many responded. If the response process may be related to the measured variable, explain why that could limit generalization.
- Using the wrong denominator. A sample percentage uses the sample size; a population percentage uses the population size. State both groups and counts when comparing the two.
- Claiming a cause from a random sample. Random selection does not establish cause and effect. A careful conclusion describes the population the results may represent and avoids unsupported causal language.
A strong short explanation might say: “The 50 households were selected by chance from a current list of households in the service area, so the selection process reduces the risk of staff choosing a preferred group and can support generalizing to households represented by that list. Because 15 selected households did not respond, the association should consider whether their garden use may differ from that of respondents.” This states what random selection contributes and identifies a specific limitation.
Check Your Understanding
For each situation, explain what random selection does and does not allow the researcher to conclude.
- A school uses a random-number generator to select 30 students from a complete roster of 300. Why is this preferable to surveying students who happen to be in the cafeteria?
- A randomly selected sample has a larger percentage of households with solar panels than the population. Does that alone show that the selection was not random? Explain.
- A city randomly selects people from a list of recreation-portal users but wants to generalize to every adult resident. Identify the frame problem.
- A random sample of 60 people is selected, and 42 respond. How many were selected, how many responded, and why should those numbers be reported separately?
- Explain the difference between the conclusion supported by random selection and the conclusion supported by random assignment.