Tutorials › AP Statistics › Generalizing Results to a Population

Investigative questions and data collection · Tutorial 153 of 1000

Generalizing Results to a Population

See how random selection provides a basis for extending sample findings to the population the sample represents—and how to show that scope with a diagram.

Beginner 8 min read

What You'll Learn

  • Distinguish a study’s sample from its population of interest.
  • Explain how random selection supports generalizing findings to a defined population.
  • Draw and interpret a scope-of-inference diagram showing a sample-to-population link.
  • Identify why convenience samples and voluntary-response samples may not represent a population.
  • Limit a generalization to the population the sampling process actually covered.
  • Explain why random assignment and random selection support different conclusions.

From the Sample to the Population

A study often collects information from a sample because reaching every member of a population is difficult, costly, or impractical. The goal may be to use what the sample reveals to learn about the larger group. Whether that extension is justified depends in part on how the sample was chosen.

In Defining the Population of Interest and Parameters and Statistics, you learned to identify the population a question is about and distinguish a population parameter from a sample statistic. Here, the key question is: Does the way the sample was selected give us a sound basis for generalizing the findings to that population?

Definition: To generalize a study’s findings is to use results from a sample to draw a conclusion about a broader population. A generalization is most defensible when the sample was selected at random from the population of interest and the sampling process adequately reached that population.

Random selection uses chance to choose the sample from a population. Because the selection process is not based on a person’s preference, availability, or a researcher’s judgment, it can provide a sound basis for representing the population. It does not guarantee that a particular sample will match the population perfectly. Instead, it makes the method of selection fair in a way that supports generalizing from the sample, while allowing for chance differences between a sample and the population.

A random sample also does not automatically fix every problem in a study. For example, selected people may not respond, questions may be misunderstood, or the list used to select participants may leave out part of the population. These issues can weaken how well the sample represents the population, even if the initial selection used chance.

Scope-of-Inference Diagrams

A scope-of-inference diagram is a compact way to show what a study’s design supports. For generalization, it makes the relationship between the population and the sample visible. The important link is random selection from the population to the sample.

$$ \text{Population of interest} \ \xrightarrow{\text{random selection}}\ \text{Sample} \ \xrightarrow{\text{analyze data}}\ \text{Sample findings} \ \xrightarrow{\text{generalize cautiously}}\ \text{Population findings} $$

Read the diagram from left to right. The first arrow indicates how the sample was formed. If chance was used to select the sample from the stated population, the design supports using sample findings to learn about that population. The final arrow represents the reasoning from a statistic calculated from the sample to a conclusion about a population parameter. It is a supported inference, not a claim that the sample and population are identical.

The diagram’s scope depends on the population actually covered by the selection process. If a sample is selected from one school’s enrollment list, for example, the design does not by itself support a conclusion about all students in the region. A population should be described with clear boundaries, and the selection method should be checked against those boundaries.

Key distinction: Random selection supports generalizing from a sample to the population from which it was selected. It does not automatically support generalizing to other populations, places, or time periods that the sampling process did not cover.

A diagram can also help you spot when a generalization is not supported. If people volunteer after seeing an open invitation, or if researchers choose the easiest people to contact, the sample-to-population arrow should not be labeled “random selection.” The results still describe the people who responded, but the method may not justify extending those results to everyone the study hopes to represent.

A Practical Decision Process

To judge whether a study supports generalization, first name the population the researchers want to learn about. Then trace how the sample was obtained. Ask whether chance was used to select sample members from that population, and whether the selection process reached the relevant members. Finally, state the conclusion no more broadly than the design allows.

1
Name the population.
Specify the people or units, location, and time period the conclusion is meant to describe.
2
Trace the selection method.
Determine how the sample was chosen and whether chance was used to select it from the population.
3
Check coverage and participation.
Consider whether the selection list covered the population and whether nonresponse or other issues could affect who provided data.
4
State the scope.
Say which population the findings can reasonably describe, and identify limits when the design does not support a broader claim.

The sampling frame is the list or source from which a sample is selected. If the frame leaves out members of the stated population, even a random selection from that frame may not represent the full population. For example, a random sample drawn from a list of people with internet access may not represent all residents if some residents are absent from that list and internet access is related to the study’s question.

Nonresponse requires similar care. If some randomly selected people do not provide data, the final set of responses may differ from the selected sample. If responders and nonresponders differ in ways connected to the variable being studied, the results may be less representative. Random selection supports generalization, but it is not a guarantee against these problems.

Worked Example: A Random Sample of City Residents

Worked Example: A Random Sample of City Residents

A city wants to estimate the percentage of its adult residents who use public transit at least once a week. The study team has a current list of adult residents and randomly selects 500 people from it. Of those selected, 420 complete the survey; 168 respondents report using public transit at least once a week. What does this design support?

The population of interest is the city’s adult residents. The sample initially selected consists of 500 adults chosen at random from the city’s resident list. The 420 completed surveys are the responses used to calculate the reported result. The sample percentage among respondents is \(168/420=0.40\), or 40%.

A scope-of-inference diagram for the selection is:

$$ \text{City adult residents} \ \xrightarrow{\text{random selection from resident list}}\ \text{500 selected adults} \ \xrightarrow{\text{420 respond}}\ \text{420 survey responses} \ \xrightarrow{\text{sample finding: 40\%}}\ \text{estimate for city adults, with nonresponse caution} $$

Because the team randomly selected adults from a list intended to cover the city’s adult residents, the design provides a basis for generalizing the findings to that population. However, only 420 of the 500 selected adults responded. The 40% is the percentage among respondents; nonresponse could affect how well that percentage represents all city adults. A careful conclusion is that the survey estimates the percentage of adult city residents who use public transit weekly, while noting that nonresponse may limit the generalization.

The conclusion should not automatically extend to children, residents of nearby towns, or city residents at some different time. Those groups were not the stated population represented by this selection.

Worked Example: A Random Sample from a Narrow Population

Worked Example: A Random Sample from a Narrow Population

A fictional company wants to estimate how many hours per month its current warehouse employees spend using a new inventory system. The company randomly selects 80 employees from its complete list of 640 warehouse employees. The sample’s mean is 6.5 hours per month. A manager says, “Warehouse workers across the country spend an average of 6.5 hours per month using this system.” Is that claim supported?

The random selection supports generalizing the sample finding to the company’s 640 current warehouse employees, assuming the list is complete and the selected employees’ data were collected appropriately. The sample mean of 6.5 hours is a statistic describing the 80 employees; the corresponding population quantity of interest is the mean for all 640 employees.

The scope diagram is:

$$ \text{Company's 640 current warehouse employees} \ \xrightarrow{\text{random selection}}\ \text{80 sampled employees} \ \xrightarrow{\text{sample mean } 6.5\text{ hours/month}}\ \text{estimate for those 640 employees} $$

The manager’s national claim is too broad. The random sample was drawn from one company’s employees, not from warehouse workers across the country. A suitable conclusion is: “The sample provides an estimate of average monthly system use for this company’s current warehouse employees.” The selection design does not establish that the same mean applies to other companies or to all warehouse workers.

This example shows why naming the population is essential. “Random sample” is not a complete justification on its own; you must also say what group the sample was randomly drawn from.

Worked Example: An Open Online Poll

Worked Example: An Open Online Poll

A neighborhood library posts a link on its public website asking residents whether the library should extend its evening hours. In all, 612 people complete the poll, and 70% favor extending the hours. The library director reports that “70% of neighborhood residents support the change.” Does the poll support that generalization?

The intended population is neighborhood residents, but the poll did not randomly select residents from that population. People decided for themselves whether to visit the website and respond. Residents who noticed the link or felt strongly about the hours may have been more likely to participate. The 70% describes the people who completed the poll; it does not, on its own, estimate the percentage of all neighborhood residents who favor the change.

The diagram shows the problem:

$$ \text{Neighborhood residents} \ \xrightarrow{\text{open invitation; self-selection}}\ \text{612 poll respondents} \ \xrightarrow{\text{70\% favor extension}}\ \text{description of respondents, not a supported population estimate} $$

The large number of responses does not turn the poll into a random sample. The appropriate conclusion is: “Among the 612 people who completed the online poll, 70% favored extending evening hours.” To support a generalization to neighborhood residents, the library could use a chance-based method to select residents and make a serious effort to obtain responses from those selected.

Keep Generalization Separate from Causation

As you learned in Association Versus Causation in Study Conclusions, random selection and random assignment have different roles. Random selection helps make a sample representative of a population and supports generalizing results to that population. Random assignment helps make treatment groups comparable and supports cause-and-effect conclusions about the treatments studied.

A study can have one of these features without the other. In this tutorial, the diagram’s sample-to-population arrow concerns generalization. It does not, by itself, show that a treatment caused an outcome. Conversely, random assignment in an experiment does not automatically mean that the participants represent a wider population. Keep the conclusion tied to the design feature that supports it. The next tutorial, Scope of Inference: Four Combinations, will bring these two design questions together.

Common Mistakes and AP Exam Tips

  • Using “random” without saying what was random. State whether people were randomly selected from a population or randomly assigned to groups. For generalization, identify random selection.
  • Generalizing to a group the sample did not represent. A random sample from one school does not automatically represent a district or a state. Name the population from which the sample was selected.
  • Treating a large sample as a random sample. A large convenience sample or voluntary-response poll can still be unrepresentative. Sample size does not replace random selection.
  • Claiming random selection guarantees a perfect match. Random selection provides a basis for generalizing, but chance differences, nonresponse, and gaps in the sampling frame can still matter.
  • Ignoring who actually responded. If many selected people do not respond, explain that nonresponse may limit how well the responses represent the intended population.
  • Confusing population description with causal explanation. Random selection supports generalization; random assignment supports cause-and-effect reasoning. Use the design feature that matches the claim.

For full credit, name the population and sample, identify whether selection was random, and connect that design to the scope of the conclusion. For example: “Because the researchers randomly selected participants from the city’s adult residents, the findings can be generalized to that population, although nonresponse may limit how representative the completed surveys are.” If selection was not random, describe the sample results without claiming they represent the full population.

Key takeaway: A random sample can support generalizing findings to the population from which it was selected. Draw the scope-of-inference diagram from the stated population through the sampling method to the sample findings, and do not extend the conclusion beyond the population the design represents.

Check Your Understanding

For each situation, identify the population and decide whether the sample design supports generalizing the findings to it.

  1. A school randomly selects 60 students from its current enrollment list to estimate the percentage of its students who bring lunch from home. What population does the sample represent?
  2. A sports website asks visitors to click a poll about a new team logo. Explain why the poll does not automatically represent all fans.
  3. A researcher randomly selects 200 households from a town’s complete address list, but only 90 respond. Name one concern to include when describing the scope of the findings.
  4. A random sample of employees at one hospital is used to make a claim about all hospital employees in the country. What is too broad about the claim?
  5. In a scope-of-inference diagram, what does the random-selection arrow connect, and what kind of conclusion does it support?