Two Questions Determine the Scope of a Conclusion
In Drawing Conclusions About Causation From Test Results, you learned that random assignment can support a cause-and-effect conclusion when a well-conducted experiment provides convincing evidence of a difference. A related question is: To whom does that conclusion apply? The answer depends on how the study’s participants were chosen, as well as how they were assigned.
Two uses of chance play different roles. Random sampling uses chance to select individuals from a population. It can support generalizing results from the sample to the population the sample was drawn from. Random assignment uses chance to place experimental units into treatment groups. It can support a causal conclusion about the effect of the assigned treatment. Neither process automatically does the other’s job.
Keep the two questions separate: Was there random sampling from a population? and Was there random assignment to treatments? A study can use either, both, or neither. A p-value addresses how unusual the sample results would be under a null model; it does not expand the scope allowed by the study design.
| How the data were collected | What the design can support | What it does not establish by itself |
|---|---|---|
| Random sample, no random assignment | Generalizing an association or estimate to the population sampled, subject to study limitations | A cause-and-effect claim |
| Random assignment, no random sample | A cause-and-effect conclusion about the experimental units studied, if the experiment is well conducted | Generalizing to a broader population |
| Both random sample and random assignment | A causal conclusion that may be generalized to the population sampled, subject to study limitations | Claims about populations or settings beyond those represented |
| Neither | A description of the observed participants, or a cautious association within the studied group | Broad generalization or a causal claim |
Start by Naming the Population and the Units
Before interpreting a result, identify the people or units actually represented in the study. The sample is the set of individuals from whom data were collected. The population is the larger group the researchers want to understand. These groups may not match: a study may recruit volunteers from one school while its researchers hope to draw conclusions about all students in a state.
A random sample supports generalizing to the population from which the individuals were randomly selected—not automatically to every population of interest. For example, a random sample of people on a city library’s membership list represents that list’s population, not necessarily all residents of the city. The sample frame, or list from which the sample is selected, matters.
Even a random sample does not remove every concern. Undercoverage can leave some members of the intended population out of the sampling frame, and nonresponse can mean that selected individuals do not provide data. A random selection is useful evidence about how a sample was obtained, but it does not guarantee that these other problems are absent. State the population that the design actually represents, and be cautious if the sampling or response process limits that representation.
For a treatment comparison, also name the experimental units: the individuals or objects assigned to treatments. Random assignment helps make the treatment groups comparable, on average, and supports a causal interpretation for those units. If the units were volunteers or a convenience group rather than a random sample, that does not erase the value of random assignment—but it does limit how confidently the result can be generalized to people who were not in the experiment.
Worked Examples: Match the Conclusion to the Design
Worked Example: A Random Sample in an Observational Study
A fictional city department wants to know whether residents who bike to work differ from residents who do not in the proportion who regularly wear a helmet. The department takes a random sample of 500 adults from a current city-resident registry. In the sample, 84 of 140 bike commuters regularly wear a helmet, compared with 126 of 360 other sampled adults. A two-proportion test reports convincing evidence of a difference. Residents chose their commuting behavior; the department did not assign it.
Identify the groups: The two observed groups are sampled bike commuters and sampled adults who do not bike to work. The comparison concerns the proportions who regularly wear a helmet. Since residents chose their behavior, this is an observational comparison, not a randomized experiment.
Consider generalization: The department randomly sampled from its registry of city adults. If the registry adequately covers the intended population and response does not introduce important bias, the result can be generalized to the city adults represented by that registry. The random sample supports generalization of the association; it does not change how the commuting groups were formed.
Conclude in scope: The data provide convincing evidence that the proportion who regularly wear a helmet differs between city adults who bike to work and those who do not. The result supports an association in the population represented by the registry, subject to coverage and response limitations. It does not show that biking to work causes residents to wear helmets more or less often. Other differences between the groups could help explain the association.
Worked Example: Random Assignment Among Volunteers
A fictional health team recruits 160 adults who volunteer for a four-week hydration program. The volunteers are randomly assigned: 80 receive daily text reminders and 80 receive no reminders. At the end, the team compares the proportions who meet a hydration goal. A two-proportion test reports convincing evidence that the reminder group has a higher proportion meeting the goal.
Identify the groups: The groups are volunteers randomly assigned to receive reminders and volunteers randomly assigned to receive no reminders. Chance determined treatment assignment, so this is a randomized experiment.
Consider causation: Assuming the experiment was carried out appropriately, the result supports a causal conclusion that assignment to the text reminders increased the proportion meeting the goal for these 160 volunteers. Random assignment helps support that conclusion because the treatment groups were formed by chance rather than by participants choosing their treatment.
Consider generalization: The team recruited volunteers; it did not randomly sample adults from a defined population. Random assignment therefore does not, on its own, justify generalizing the result to all adults, all residents of the area, or all people who might receive such reminders. Volunteers may differ from people who did not volunteer.
Conclude in scope: The experiment provides convincing evidence that the reminders increased the proportion of these participating volunteers who met the hydration goal. The conclusion is causal but limited to the experimental units studied. More evidence, such as a random sample from a clearly defined population, would be needed to support broad generalization.
Worked Example: Random Sampling and Random Assignment
A fictional school district has 1,200 students in its after-school program. Researchers randomly select 240 students from the program roster and invite them to take part in a study. All 240 take part. The researchers randomly assign 120 to receive a new study-planning routine and 120 to continue their usual routine. A comparison of the proportions who complete a weekly study goal provides convincing evidence of a difference between the groups.
Identify the target population: The students were randomly selected from the after-school program roster, so the population represented by the sample is the district’s 1,200 students in that program. The result does not automatically represent students who are not in the program.
Consider causation: The researchers randomly assigned the selected students to the two routines. If the study was carried out appropriately, the evidence supports a causal conclusion about the effect of assignment to the new routine on the participating students’ goal-completion rate.
Combine the design features: Random assignment supports the causal comparison, while random sampling supports generalizing to the population sampled. Because all selected students participated, this scenario avoids nonresponse among those invited; researchers would still need to consider whether the roster covers the intended population and whether the experiment was conducted appropriately.
Conclude in scope: The study provides convincing evidence that the new routine changed the proportion of students who completed the weekly goal, and the design supports generalizing that causal conclusion to students in this district’s after-school program. It does not establish the same effect for all district students or for students in other settings.
When Neither Chance Process Was Used
Suppose a researcher recruits students who happen to be in the cafeteria and asks which of two study apps they already use. The researcher then compares the proportions who meet a study goal. The students were not randomly sampled from a school population, and the researcher did not assign app use at random. A test could describe evidence of a difference between the observed app-user groups, but the design alone does not support generalizing that difference to all students or claiming that an app caused it.
This does not mean every study without random sampling or assignment is useless. A carefully described study can provide information about the participants observed, suggest a question for further research, or identify an association worth investigating. The important step is to keep the conclusion proportionate to the evidence: do not claim more than the study design supports.
Likewise, a study that uses random sampling but no random assignment has an advantage for generalization, not causation. A study that uses random assignment but no random sampling has an advantage for causation, not broad generalization. Those two forms of chance solve different problems, so naming the design feature precisely is essential.
Common Mistakes and AP Exam Tips
- Equating random sampling with random assignment: A random sample can support generalization to the population sampled. It does not establish that a treatment caused an outcome.
- Equating random assignment with a representative sample: Randomly assigning volunteers makes the treatment groups comparable on average. It does not make those volunteers representative of everyone else.
- Generalizing to a population broader than the sampling frame: Name the source population. A random sample from one school, roster, or registry does not automatically represent people outside it.
- Claiming causation from an observational difference: A small p-value does not eliminate confounding. If people chose their groups or treatments, describe an association or difference rather than a cause.
- Ignoring the evidence of the test: A design tells you what kinds of conclusions may be supported; the test result tells you whether the data provide convincing evidence of a difference. A sound design does not turn a non-significant result into evidence of an effect.
- Writing a conclusion without naming its limits: Specify the group, outcome, and—in an experiment—the treatment. Then say whether the conclusion is causal, generalizable, both, or neither.
Key Takeaway
A statistical result does not decide its own scope. Random sampling can support generalizing findings to the population sampled, and random assignment can support a causal conclusion about the experimental units. When a study uses both, it may support both kinds of inference; when it uses neither, keep claims closely tied to the participants and groups observed.
Check Your Understanding
For each situation, identify whether generalization, a causal conclusion, both, or neither is supported by the design.
- Researchers randomly select residents from a town registry and observe whether they own a bicycle and regularly exercise. What population might the results represent, and can the researchers conclude that bicycle ownership causes exercise?
- Researchers recruit volunteers and randomly assign them to two versions of a stretching routine. A test finds convincing evidence of a difference in the proportion who reach a flexibility goal. What conclusion is supported, and to whom should it be limited?
- A random sample of eligible students is selected, then those students are randomly assigned to two study programs. What does each chance process contribute to the scope of the conclusion?
- A researcher compares people who chose to use a meal-planning app with people who did not, after recruiting participants at a shopping center. Name two reasons the conclusion should be limited.
- Why does a very small p-value not make a convenience sample representative or turn an observational study into an experiment?