“Random” Can Describe Two Different Design Choices
In Why Inference Procedures Need Conditions, you saw that the random condition helps connect a sample proportion \(\hat{p}\) to a population proportion \(p\). But study descriptions use the word “random” in more than one way. A researcher might randomly select people to participate, randomly assign participants to treatments, do both, or do neither. These choices support different kinds of conclusions.
The key question is not simply, “Was the study random?” Ask: What was randomized, and what is the conclusion meant to be about? Random selection concerns which individuals enter the sample. Random assignment concerns which treatment the study participants receive. A convenience sample, by contrast, includes individuals who are easy to reach or who choose to take part, without random selection from the target population.
For one-proportion inference about a population, the random condition asks whether the data come from a random sample or another sampling process that supports treating the observations as representative of the population named in the conclusion. Random assignment is not the same thing as random sampling: assigning treatments at random does not randomly select participants from a larger population.
Random Sample: A Basis for Generalizing to a Population
A random sample is selected using a chance process from a defined population. When the sampling frame covers the population of interest and the selection process is carried out as planned, every individual has a known chance of selection. A suitable random sample supports using \(\hat{p}\) to estimate the proportion \(p\) in that population.
This support is about generalization, not cause and effect. A survey that randomly samples residents and asks whether they use a public bus can estimate the proportion of residents who report using one. It cannot establish that using the bus caused a particular outcome, because the researcher did not assign people to use or not use the bus.
Random sampling does not guarantee that every sample will resemble the population perfectly. As covered in Sampling Variability Versus Bias in Proportions, chance variation remains from sample to sample. A random method supports the inference model; it does not erase sampling variability or automatically solve problems such as a poorly constructed sampling frame, nonresponse, or inaccurate responses.
Randomized Experiment: A Basis for Studying Cause and Effect
In a randomized experiment, researchers assign experimental units to treatments using a chance process. Random assignment helps make the treatment groups comparable, on average, with respect to other characteristics. If the groups’ outcomes differ, the design can support a causal conclusion about the effect of the treatments for the experimental units, subject to the study’s conduct and other limitations.
Random assignment does not, by itself, make the participants a random sample from a larger population. If volunteers are recruited from one school and then randomly assigned to two study conditions, the assignment can support a treatment comparison among the study participants. It does not automatically show that the same treatment effect applies to every student in the region or country.
Sometimes a study uses both steps: researchers select a random sample from a population and then randomly assign those sampled individuals to treatments. The random sample supports generalizing to the population represented by the sampling frame; random assignment supports a causal comparison of treatments. The two design features answer different questions.
- Who entered the study? Random selection may support generalizing to a population.
- Who received each treatment? Random assignment may support a causal conclusion about treatment effects.
Convenience Sample: Easy to Collect, Hard to Generalize
A convenience sample consists of individuals selected because they are easy to reach, available at a particular place or time, or willing to respond to an open invitation. A poll posted on a club’s social-media page, a survey of shoppers leaving one store, and a questionnaire answered only by people who choose to click a link are common examples.
The concern is not merely that a convenience sample might be small. The people who are easiest to reach or most motivated to respond may differ systematically from those who are missed. That can shift the sample proportion away from the population proportion. Increasing the number of convenience responses may make the estimate more numerically stable for that selection process, but it does not turn the process into random sampling or remove the risk of bias.
A convenience sample can still describe the people who responded. For example, it is accurate to report the percentage of respondents who selected “yes.” But without a suitable random sampling process, that percentage does not automatically justify a confidence interval or test about all members of a broader population. This is the distinction between describing observed data and making an inference.
A Design-Based Check Before One-Proportion Inference
Before using a one-proportion \(z\)-interval or \(z\)-test, identify the target parameter: \(p\), the proportion of a clearly named population with a specified characteristic. Then examine how the observations were obtained. As in Why Inference Procedures Need Conditions, a formula may still produce an answer when a condition is not supported; the issue is whether that answer justifies the intended inference.
State whose proportion \(p\) is being estimated or tested and what counts as a success.
Was a sample selected by chance, were treatments assigned by chance, or did neither happen?
A random sample may support generalizing to its represented population. Random assignment may support a causal treatment comparison. Do not use one as evidence for the other.
For sampling without replacement, apply the 10% condition when appropriate. For a Normal-based interval, check the observed success and failure counts; for a test, check the null expected counts.
Passing the random-condition check does not replace the 10% condition or the Large Counts condition. These address different parts of the inference model, as explained in earlier tutorials on those conditions. Also consider whether the sampling frame covers the stated population and whether nonresponse could undermine representativeness.
Worked Examples
Worked Example: A Random Sample of Residents
A fictional city has 2,400 households on a complete residential list. A researcher uses a computer to select 120 households at random and asks whether each household composts food scraps. Of the 120 sampled households, 51 say yes. The researcher wants to estimate the proportion of all households on the list that compost.
Identify the design. Households were randomly selected from the stated list, so this is a random sample, not a randomized experiment. No treatment was assigned. The sampling process supports generalizing to the households represented by the list, assuming the selected households are contacted and their responses are recorded appropriately.
Check the sampling conditions relevant to an interval. The 10% condition is met because \(0.10(2{,}400)=240\), and the sample size \(120\) is no more than 240. The observed success count is 51 and the failure count is \(120-51=69\); both are at least 10, so the Large Counts condition for a one-proportion \(z\)-interval is met.
Summarize what the design supports. The sample proportion is \(\hat{p}=51/120=0.425\). The random selection supports estimating the proportion of households on the list that compost. It does not support a causal claim, because the researcher did not assign households to compost or not compost. If the researcher wants to generalize to all households in the city, the list must adequately cover that population; a random sample from an incomplete list cannot represent households the list omits.
Worked Example: Volunteers Randomly Assigned to Treatments
A fictional health educator recruits 200 volunteers from a community center to compare two reminder messages. Using a random process, the educator assigns 100 volunteers to receive a text reminder and 100 to receive a printed reminder. After a week, 64 people in the text group and 48 people in the printed group report completing a planned activity.
Identify each random feature. The participants volunteered; they were not randomly selected from all community-center members or all residents. However, the educator randomly assigned the volunteers to the two reminder treatments.
Judge the possible inference. The random assignment supports comparing the treatments for these experimental participants and, with an appropriate analysis of the two groups, can support a causal conclusion about the effect of the reminder method for the volunteers. The observed proportions are \(64/100=0.64\) in the text group and \(48/100=0.48\) in the printed group. Those sample results alone do not establish that the text reminder caused a higher completion rate; chance variation must be considered through an appropriate inference procedure.
State the limit on generalization. Random assignment did not make the volunteers a random sample. The study therefore does not automatically support generalizing the estimated effect to all community-center members or all residents. The participants may differ from people who did not volunteer. The design has a basis for causal inference because of assignment, but not broad population generalization based on random selection.
Worked Example: A Large Convenience Poll
A fictional city-news website asks readers to click a poll answering, “Do you support adding protected bike lanes downtown?” By the end of the week, 880 readers have responded, and 600 select “yes.” The editor wants to report the proportion of all city residents who support the proposal.
Calculate the sample result. Among respondents, \(\hat{p}=600/880\approx0.6818\), or about 68.18%. This describes the 880 people who chose to answer the website poll.
Evaluate the sampling method. This is a convenience sample with voluntary response, not a random sample of city residents. Readers who visit that website and choose to answer may differ from residents who do not visit it or choose not to respond. The sample size of 880 does not fix that selection problem.
Decide what inference is justified. The editor should not use the usual one-proportion \(z\)-interval or \(z\)-test to make a claim about the proportion of all city residents. The poll’s result can be reported as a description of the respondents, with the selection method made clear. It cannot be presented as a statistically justified estimate for all residents just because many people responded.
Worked Example: Random Selection with Substantial Nonresponse
A fictional school district randomly selects 300 families from its enrollment list to ask whether they favor a proposed change to the school calendar. Only 210 families respond, and 126 responding families favor the change. The district wants to estimate the proportion of all families in the district who favor it.
Recognize the sampling design. The district began with a random sample, which is stronger than a convenience sample and supports a connection to the listed families. The response rate is \(210/300=0.70\), so 30% of the selected families did not respond.
Explain the limitation. The 126 “yes” responses are 60% of the 210 respondents, since \(126/210=0.60\). But the original random selection does not guarantee that the respondents remain representative if response is related to opinion. For instance, families with strong views might be more likely to reply. The district should consider whether nonresponse could systematically change the observed proportion before treating it as a reliable estimate for all families.
Make a careful judgment. Random selection provides a sound starting design, but the nonresponse creates a possible source of bias that a standard margin of error does not measure. The district should not claim that random selection alone proves the responses represent every family. It should report the response situation and qualify any population inference accordingly.
Common Mistakes and Full-Credit Reasoning
A strong response identifies the design feature, names the inference it supports, and limits the conclusion to an appropriate target. Simply saying “the study is random” is not enough. State whether people were randomly selected or treatments were randomly assigned, and connect that detail to the claim.
- Equating random assignment with random sampling. Assignment can support a causal treatment comparison; it does not make the study participants representative of a larger population.
- Assuming a random sample proves causation. A survey can estimate a population proportion, but observing a characteristic does not show that it caused another outcome.
- Calling a poll random because many people responded. A large voluntary-response poll can still systematically overrepresent some viewpoints. Sample size does not repair the selection method.
- Ignoring who the sample can represent. A random sample from a list supports generalization to the population covered by that list, not necessarily to groups missing from it.
- Overlooking nonresponse. Randomly selecting people is a useful design feature, but people who respond may differ from those who do not. Mention this concern when the description gives evidence of nonresponse.
- Claiming too much from a randomized experiment. Say that random assignment supports a causal comparison for the experimental units. Do not automatically extend the conclusion to a broader population without a suitable sampling basis.
Key Takeaway
The random condition is about how the study’s data were produced, not about whether a researcher used the word “random.” Random selection and random assignment are distinct tools: one can support population generalization, and the other can support cause-and-effect reasoning. A convenience sample may describe its respondents, but it does not gain population representativeness from having a large sample.
Check Your Understanding
For each situation, identify the design and state what kind of inference it supports, if any.
- A researcher randomly selects 160 registered library members and asks whether they borrowed an e-book this month. Is this random sampling, random assignment, or convenience sampling? What population might the result represent?
- A group of volunteers is randomly assigned to use one of two study apps. What does random assignment support, and what does it not establish about all students?
- A website poll receives 5,000 optional responses. Explain why the large response count does not by itself justify a confidence interval for all residents.
- A random sample is drawn from a list that excludes families without internet access. Identify the target-population concern.
- A random sample is selected, but many selected people do not respond. Explain why the original selection method may not be enough to guarantee representative responses.