Who Can a Conclusion Describe?
A test about population means is not only about the calculation. The way the data were collected determines which population the evidence can describe. A small p-value may provide convincing evidence of a difference between means, but it cannot make a sample represent people who had no chance of being selected.
As in “Statistical Significance in Observational Studies,” distinguish how units were selected from how they were assigned to groups. Random sampling uses chance to select units from a population. It can support generalizing results to the population represented by that selection process. Random assignment uses chance to place study units into groups. It can support a cause-and-effect conclusion, but does not by itself make those units representative of a larger population.
The relevant population is often narrower than a broad label such as “adults,” “students,” or “patients.” A random sample from one clinic may represent that clinic’s eligible patients, not all patients in the region. A random sample from a list of current customers may represent those customers, not everyone who might become a customer.
Follow the Sampling Method
Begin by asking: What units were eligible to be selected, and from what list or group were they selected? That list or group is the sampling frame. Then ask whether chance selected the sample from that frame. If the selection was random and the sampling process was carried out appropriately, an inference may extend to the population represented by the frame.
Random selection does not guarantee that every sample will perfectly mirror the population. It does provide a basis for accounting for sampling variability. It also does not automatically remove undercoverage, nonresponse, or other sources of bias. For example, a random sample from an outdated customer list may still miss recent customers. Consider the population and the actual selection process, not just the word “random.”
A convenience sample—such as people who happen to be nearby or who volunteer after seeing an advertisement—does not use chance to select from the intended population. Its results can describe the people who took part, but the sample alone does not justify a formal generalization to everyone in the intended population. A large sample or a small p-value does not fix this limitation.
This distinction applies whether a study compares groups or estimates a single mean. For a two-sample t test, define the population means in terms of the groups and setting actually represented by the sampling process. As in “Defining Both Population Means in Context,” identify the quantitative response, units, and target setting. Then make sure the conclusion does not silently expand those populations.
Worked Example: Random Samples from Two Defined Groups
Worked Example: Random Samples from Two Defined Groups
A fictional electric utility wants to compare daily winter electricity use by households with and without programmable thermostats. From complete lists of households in its service area, it randomly selects 25 households that currently use programmable thermostats and 25 that do not. The first sample has a mean use of 31.2 kilowatt-hours per day and a standard deviation of 5.0 kilowatt-hours. The second sample has a mean of 34.2 kilowatt-hours and a standard deviation of 5.0 kilowatt-hours. Assume there are at least 2,000 households in each group. The samples are independent, and plots show no severe skewness or extreme outliers. Test for a difference at \(\alpha=0.05\).
Let \(\mu_1\) be the true mean daily winter electricity use, in kilowatt-hours per day, for households in this utility’s service area that currently use programmable thermostats. Let \(\mu_2\) be the true mean for households in the same service area that do not. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).
Use an unpooled two-sample t test. Each group was randomly sampled from its own complete list, so the samples support inference about those two groups of households in the utility’s service area. Each sample of 25 is less than 10% of its group, since each group has at least 2,000 households. The groups contain different households, so the samples are independent. The plots show no severe skewness or extreme outliers, supporting the t procedure. Thermostat use was not randomly assigned, so a difference would not establish that thermostat use caused a difference in electricity use.
The observed difference is \(31.2-34.2=-3.0\) kilowatt-hours per day. The standard error and test statistic are
The data provide convincing evidence of a difference in mean daily winter electricity use between households with and without programmable thermostats in this utility’s service area. This conclusion applies to the two household populations represented by the random samples. Because the study was observational, it does not establish that using a programmable thermostat caused the difference.
The sampling method gives this result a defined reach: households in the utility’s service area in the two thermostat-use groups. It does not justify extending the conclusion to households served by other utilities or to all households in the country. The small p-value addresses evidence about a difference between the specified means; it does not broaden the populations those means describe.
Worked Example: A Large Volunteer Sample Still Has Limited Reach
Worked Example: A Large Volunteer Sample Still Has Limited Reach
A student research team posts a survey link on its school’s social media account, asking people to report their average nightly sleep. Forty-eight students respond voluntarily. Their reported mean is 6.3 hours per night, with a standard deviation of 1.2 hours. The team hopes to describe the mean sleep time of all students in the district.
Identify the sampling method: The students were not randomly selected from a district-wide list. They chose whether to notice the post and complete the survey. This is a volunteer sample, so students who responded may differ from those who did not—for example, students especially interested in sleep may have been more likely to participate.
Identify what the data describe: The sample mean is \(6.3\) hours per night for the 48 respondents. That is a valid descriptive summary of those responses. The team cannot use this selection process alone to conclude that the mean for all district students is \(6.3\) hours, or to make a well-supported inference about that population.
Limit the conclusion: A careful report would say, “Among the 48 students who voluntarily completed the survey, the mean reported sleep time was 6.3 hours per night.” It would not say, “District students average 6.3 hours per night.” Even if a test or interval based on the responses produced a striking result, it would not correct the lack of random selection or the possibility of volunteer bias.
Worked Example: Random Assignment Does Not Make Volunteers Representative
Worked Example: Random Assignment Does Not Make Volunteers Representative
A fictional research team recruits 40 volunteers from one community college to test a study-planning app. The team randomly assigns 20 volunteers to use the app and 20 to use their usual planning method. At the end of the term, the app group has a mean course score of 78 points, while the usual-method group has a mean of 72 points. Suppose a correctly conducted two-sample t test reports \(p=0.023\) for a two-sided test at \(\alpha=0.05\).
Interpret the test result: Since \(0.023<0.05\), the data provide convincing evidence of a difference in mean course scores between the app and usual-method groups in this experiment. The observed difference is \(78-72=6\) points, favoring the app group.
Consider what random assignment supports: The researchers used chance to assign the volunteers to the two methods. As discussed in “Randomized Experiment Versus Observational Comparison,” random assignment can support a cause-and-effect conclusion about the treatments for the study units, provided the experiment was carried out appropriately. It helps make the two treatment groups comparable at the start.
Consider what the recruitment method does not support: The participants volunteered from one college; they were not randomly sampled from all college students. Therefore, random assignment does not establish that the observed treatment difference would also occur for students at other colleges, nonvolunteers, or all students. A suitably limited conclusion is that the experiment provides evidence of a difference in mean course scores caused by assignment to the app rather than the usual method for these volunteers. Generalizing that effect to a wider population would require additional justification.
The example separates two questions: “Did the assigned treatment make a difference for these study units?” and “To which population does that result extend?” Random assignment addresses the first question; the recruitment and sampling method informs the second.
Common Mistakes and AP Exam Tips
- Generalizing to everyone named in the research question: A broad research question does not make a narrow sample representative. Name the population from which units could actually be selected.
- Treating random assignment as random sampling: Assignment determines which group study units enter; sampling determines which units enter the study. State what each process supports.
- Assuming a larger sample removes selection bias: A large volunteer sample can estimate the volunteer group’s characteristics precisely while still differing systematically from the intended population.
- Letting a small p-value expand the conclusion: The p-value describes evidence against a null hypothesis under the test model. It does not repair an unrepresentative sampling method or justify a broader target population.
- Using a vague population label: “People,” “students,” or “patients” may be too broad. Specify the group, location, and relevant time or setting represented by the sampling frame.
- Claiming random sampling eliminates all bias: Random selection supports inference to the population sampled, but coverage problems, nonresponse, or poorly measured responses may still limit the conclusion.
For a full-credit conclusion, connect the statistical result to the specific population means in the hypotheses, then check that the sampling method represents those populations. If the sample came from volunteers or a convenience group, state that the results describe the participants and avoid claiming that they represent a wider population. Keep this question separate from whether random assignment permits a causal interpretation.
Check Your Understanding
For each situation, identify the population a conclusion can reasonably describe and explain which design feature matters.
- A random sample of 60 registered members is drawn from one neighborhood’s community garden. What population does the sample represent, and can the result automatically be generalized to all gardeners in the city?
- A survey about study habits is completed by students who choose to click a link in a class chat. Why does having 500 responses not, by itself, justify generalizing to all students at the school?
- Researchers recruit volunteers from one clinic and randomly assign them to two treatments. Which feature supports a causal comparison, and what limits generalization to patients at other clinics?
- A random sample is drawn from a complete list of current customers of a small online shop. Name the population the sample represents and one broader group the result may not represent.
- Explain why a p-value below 0.05 cannot repair a sample that was selected by convenience rather than randomly from the target population.