Tutorials › AP Statistics › Population Scope and Generalizing Results

Regression and context · Tutorial 931 of 1000

Population Scope and Generalizing Results

Use the way a regression sample was chosen to identify the population the observed relationship can reasonably describe.

Intermediate 10 min read

What You'll Learn

  • Distinguish the observed sample, sampling frame, and target population in a regression setting.
  • Explain how random selection supports generalizing a relationship to the population represented by the sampling process.
  • Limit conclusions from convenience and volunteer samples to the people actually observed.
  • Identify why a random sample does not justify generalizing to every related group.
  • Separate the population scope of a conclusion from causal claims and model reliability.

Who Does a Regression Conclusion Describe?

A regression line describes an association among the cases whose \(X\) and \(Y\) values were recorded. But a conclusion about those cases is not automatically a conclusion about a larger group. To decide how far a regression conclusion can reach, ask how the sample was chosen.

In “Common Response and Third Variables,” you considered possible explanations for an association between two variables. Here the question is different: Which group does the observed association describe? Even a clearly described pattern in a sample may not represent people or objects who had no chance of being selected.

Definition: The population scope of a regression conclusion is the group of individuals or objects to which the conclusion can reasonably be generalized. The scope depends in part on how the observed sample was selected.

Keep three groups distinct. The sample is the set of cases with recorded paired values of \(X\) and \(Y\). A sampling frame is the list or source from which cases could be selected. The target population is the larger group the study aims to describe. A sample may represent a target population well only if the selection process gives members of that population a fair chance of inclusion and the frame covers that population appropriately.

A random sample selected from a suitable frame provides a basis for generalizing a pattern from the sample to the population represented by that frame. This does not mean the sample must produce exactly the same regression results as the whole population. It means the selection process supports using the sample to draw conclusions about that population, while allowing for sample-to-sample variation.

By contrast, a convenience sample consists of cases that were easy to reach, and a voluntary-response sample consists of people who chose to participate. These methods may overrepresent particular kinds of individuals. A regression pattern in such a sample describes the observed cases, but the selection method alone does not justify generalizing the pattern to a wider population.

Key distinction: Random selection concerns which individuals enter the sample and supports generalization to a population. Random assignment concerns which treatment individuals receive and can support a causal conclusion in a well-designed experiment. One does not substitute for the other.

As in “Association Versus Causation in Regression,” an association does not by itself establish causation. A random sample can support generalizing an observed association to the population sampled, but it does not show that changing \(X\) causes a change in \(Y\). Likewise, a convenience sample may contain a real and useful pattern, but it does not automatically represent people who were not sampled.

A Practical Scope Check

Use a short “sample-to-population” check before writing a regression conclusion. Start with the cases actually measured, then identify who could have been selected and who the study hopes to describe. The result should name the population precisely, not leap to a larger group because it seems similar.

1
Name the observed cases.
State who or what supplied the paired measurements and identify the setting and time period when relevant.
2
Trace the selection process.
Determine whether cases were randomly selected from a frame, chosen for convenience, recruited as volunteers, or selected in another way.
3
Set the boundary.
For a random sample, name the population represented by the frame. For a nonrandom sample, keep the conclusion focused on the observed cases unless additional evidence supports broader generalization.
4
Write a limited conclusion.
Describe the regression relationship in context and state exactly which group it may describe. Do not extend it to unrepresented groups or turn association into causation.

The sampling frame matters because a random selection from an incomplete list can still leave part of the intended population out. For example, a random sample from a list of current customers can represent customers on that list; it cannot automatically represent people who have never been customers. Also consider the stated time and eligibility rules. “Adults in the region” is broader than “adults registered with this clinic during the spring.”

Nonresponse can also narrow how confidently a sample represents its intended population. If selected people do not provide measurements, the cases with complete data may differ from those who did not respond. The original random selection is relevant, but it does not make every possible source of representation problem disappear. For this tutorial, the central habit is to say what the sample selection supports and avoid claiming more than that.

Worked Examples: Matching Conclusions to Samples

Worked Example: Household Water Use in a City

A fictional city has 4,800 active residential water accounts. Researchers use that account list to randomly select 150 households and obtain complete measurements from each selected household. For each one, they record average daily outdoor watering time in July, \(X\), and total July water use, \(Y\), in gallons. The sample scatterplot shows a positive, roughly linear association.

State. Decide which population a conclusion about the positive association can reasonably describe.

Plan. Identify the observed cases, the sampling frame, and the selection method. Then compare the population represented by the frame with any broader group someone might want to discuss.

Do. The observed cases are the 150 households with complete paired measurements. The sampling frame is the list of 4,800 active residential accounts, and the researchers randomly selected households from that list. Because the sample was randomly selected and the frame covers those active residential accounts, the selection process supports generalizing the observed association to the city’s households represented by the list, for the July setting studied. It does not directly represent businesses, households missing from the account list, other cities, or water use in a different season.

Conclude. “Among the randomly selected households represented by the city’s list of 4,800 active residential accounts, average outdoor watering time and total July water use were positively associated in the sample. The sampling method supports generalizing this relationship to the city’s residential accounts represented by that list. It does not show that additional watering causes higher total use, and it does not establish the same relationship for businesses or households outside the frame.” The conclusion connects the scope to the frame rather than treating “the city” as an unlimited population.

Worked Example: Training Hours and Resting Heart Rate

A fictional researcher stands near the entrance to a neighborhood fitness center and invites arriving adults to join a study. Seventy-five people agree. Each reports weekly exercise-training hours, \(X\), and has resting heart rate, \(Y\), measured. In these participants, more reported training hours tend to occur with lower resting heart rates.

State. Decide whether the negative association can be generalized to all adults in the neighborhood.

Plan. Examine whether the 75 participants were randomly selected from a frame covering neighborhood adults. Consider who had a chance to be included and who chose to participate.

Do. These are volunteers recruited at a fitness center, not a random sample of all adults in the neighborhood. Adults who visit that center may differ from people who do not visit it, and people willing to participate may differ from those who decline. The observed association describes the 75 participants. The selection method does not provide a basis for generalizing the relationship to all neighborhood adults.

Conclude. “Among the 75 adults who volunteered at the fitness center, reported weekly training hours and resting heart rate were negatively associated. Because these participants were recruited at one fitness center and volunteered, the result should not be generalized to all adults in the neighborhood on the basis of this sample.” This limitation concerns population scope; it does not claim that the association among participants is absent or meaningless.

Worked Example: Commute Distance and Tardiness at One School

A fictional high school has a roster of all 620 enrolled ninth-grade students for the fall term. Staff randomly select 90 students from that roster and obtain commute distance, \(X\), in miles, and number of tardy arrivals during a specified four-week period, \(Y\), for every selected student. The sample shows a positive association.

State. Identify the group to which the observed association can reasonably be generalized, and explain why the conclusion should not be expanded to all teenagers.

Plan. Match the random selection to the population on the roster. Check whether other schools or age groups were included in that frame.

Do. The 90 students were randomly selected from the roster of enrolled ninth graders at one high school. The frame therefore represents that school’s enrolled ninth graders for the fall term, not all teenagers. Students at other schools, teenagers who are not enrolled, and students in other grades were not part of this sampling frame. The positive association can be discussed for the sampled group and, with the random-selection basis, generalized to the school’s enrolled ninth graders represented by the roster during the stated period.

Conclude. “For enrolled ninth-grade students at this high school during the fall term, commute distance and the number of tardy arrivals in the four-week period were positively associated. Because the students were randomly selected from the school’s ninth-grade roster, this relationship may be generalized to that roster population. The sample does not justify extending the conclusion to all teenagers or to students at other schools.” Random selection improves the basis for generalization within the represented population; it does not make that population limitless.

How Far Is Too Far?

A useful way to check a proposed conclusion is to compare its nouns with the sampling frame. If the data come from selected customers at one store, “customers represented by that store’s customer list” is closer to the evidence than “all consumers.” If the data come from a random sample of one school’s students, “students at that school” is more appropriate than “students everywhere.”

Be precise about what a regression conclusion describes. The slope and other summaries come from the observed data; a population-scope statement concerns which group the sample can represent. Do not promise that every member of the population follows the fitted line. The conclusion is about a pattern or relationship, not a perfect prediction for each individual.

The sampling method is also separate from model quality. Random selection does not guarantee that a linear model is suitable or that its predictions are accurate. Conversely, a strong-looking pattern in a convenience sample does not establish that the same pattern occurs in a larger population. Keep the scope question distinct from the residual and model-evaluation questions discussed elsewhere in this course.

Key takeaway: Trace a regression conclusion from the observed cases back to the group they were selected from. Random selection from a suitable frame supports generalizing to the represented population; convenience or volunteer selection does not, by itself, support generalizing beyond the observed cases.

Common Mistakes and AP Exam Tips

Population-scope answers earn clarity by naming both the sample and the boundary of the conclusion. Watch for these specific errors:

  • Claiming “random” without saying random what. Identify the population or frame from which cases were selected. Randomly selecting accounts from one city does not create a random sample of all households everywhere.
  • Generalizing because the sample is large. A large convenience sample can still overrepresent the people who were easiest to reach or most willing to respond. Sample size does not replace an appropriate selection method.
  • Confusing random selection with random assignment. Random selection supports generalization to a represented population. Random assignment is a feature of experiments relevant to causal conclusions. Do not use one as evidence for the other.
  • Using an overbroad population name. “All students” may be too broad when the frame contains only ninth graders at one school. State the school, grade, eligibility group, and time period when those details define the frame.
  • Turning a population association into a causal claim. Even if a random sample supports generalizing an association, it does not establish that changing \(X\) causes \(Y\) to change.
  • Confusing representativeness with model reliability. The sampling method tells you who the data may represent. It does not, by itself, show that a regression model is appropriate or predictively accurate.

For a full-credit response, name the observed cases, describe how they were selected, identify the population represented by that process, and limit the regression conclusion to that population. If selection was by convenience or volunteering, say that the relationship describes the observed participants and that the method alone does not justify broader generalization.

Check Your Understanding

For each situation, identify the population scope that the selection process supports, if any, and explain your reasoning.

  1. A random sample of 110 registered library-card holders is selected from a town library’s current roster. Researchers record weekly library visits and minutes spent reading per day. What population does the sampling process represent, and name one group it does not automatically represent.
  2. A student posts a survey link on a personal social-media account and fits a regression using responses from everyone who chooses to answer. Why does the resulting association not automatically describe all students in the region?
  3. A random sample is selected from the roster of tenth graders at one high school. Explain why the selection supports a different scope from a random sample of all tenth graders in the state.
  4. In one or two sentences, distinguish what random selection can support from what it cannot establish about causation.
  5. A large convenience sample shows a clear linear association. Explain why the sample’s size and the pattern’s clarity do not, by themselves, justify generalizing to a broader population.