A Good Fit for Which Cases?
A regression model can fit the data it was built from and still be unreliable for the people or objects someone wants to discuss. In “Population Scope and Generalizing Results,” you traced conclusions back to the population represented by a sample. Here, we apply that idea to regression: how does a convenience or voluntary sample affect conclusions about a fitted line?
A convenience sample includes cases that are easy to reach. A voluntary-response sample includes people who choose to participate, often after seeing an open invitation. Neither method gives every member of a target population a known, fair chance of selection. People who are easiest to contact or most willing to respond may differ from those not included.
This distinction matters because a model can look convincing on its own data. A large \(r^2\) says that the line accounts for a large fraction of the response variation among the observed cases, as discussed in “Interpreting \(r^2\) in Context.” A small \(s\) describes relatively small residuals for those cases, as in “Using \(s\) to Describe Prediction Accuracy.” Neither summary repairs a sample that does not represent the group the researcher wants to describe.
A nonrandom sample does not make its observed association automatically false or useless. The fitted line still summarizes the cases measured. The limitation is that the selection process alone does not support generalizing that relationship to a wider population. Be especially cautious about predictions for people who might not have entered the sample, or whose explanatory-variable values differ from those observed.
How Selection Can Affect a Regression Model
Selection can matter in more than one way. First, the sample may overrepresent certain kinds of individuals. For example, people who volunteer for a study about exercise may be more active than people who do not volunteer. If activity is related to the explanatory variable, the response, or both, the association in the volunteer sample may differ from the association in the intended population.
Second, the selection process may limit which values of \(x\) appear in the data. A study of experienced gardeners, for instance, may include only people who use a narrow range of fertilizer amounts. A fitted line based on that limited range should not be treated as dependable for much smaller or larger amounts. Even when the sample’s line summarizes its observed cases well, it provides little direct information about cases far outside the observed range.
These effects do not have a fixed direction. A convenience sample might produce a stronger, weaker, or otherwise different association from the one that would appear in a broader population. Do not claim that volunteer sampling always raises \(r^2\), changes the slope in a particular direction, or makes every prediction inaccurate. The sound conclusion is that the selection method does not establish how well the sample relationship represents people who were not sampled.
For regression conclusions, check both the selection method and the model’s data range. The regression line is most directly supported for cases like those observed, with explanatory-variable values within the sample’s range. A prediction far beyond that range is an extrapolation, and it is especially difficult to defend when the sample was also chosen by convenience or volunteering. For other aspects of model fit, connect your reasoning to “Fan-Shaped Residual Plots and Changing Spread,” “Residuals and Outliers,” and “What \(r^2\) Does Not Tell You.”
A Sampling-and-Reliability Check
Use the following steps when a prompt gives a regression model but also describes how its cases were collected. The goal is not to discard the model; it is to state clearly what the evidence supports.
Identify who or what was measured and the larger population, if one is proposed.
Decide whether they were selected randomly, reached for convenience, or recruited as volunteers. Note who may have had little or no chance to participate.
Ask whether participation or access could be related to \(x\), \(y\), or a characteristic related to either variable. Also check whether the observed \(x\)-values cover the range where the model will be used.
Use model summaries to describe the observed cases, but do not treat them as proof that the model works equally well for a population the sampling method did not represent.
Describe the association or prediction in context and state which cases it concerns. Explain why the sample method does or does not support extending the conclusion further.
A check for possible connections between selection and \(x\) or \(y\) is a reason to be cautious, not a formula for correcting the fitted line. Without information about cases that were not selected, you generally cannot determine exactly how the population slope, correlation, or prediction errors differ. State the limitation rather than inventing a correction or a definite direction of bias.
Worked Examples: Evaluating Regression From Nonrandom Samples
Worked Example: Screen Time and Sleep Among Student Volunteers
In a fictional study, a school posts an invitation for students to report their average daily recreational screen time, \(x\), in hours, and average nightly sleep, \(y\), in hours. Forty-eight students volunteer. The calculator output for these volunteers gives the line \(\hat{y}=8.40-0.18x\), \(r^2=0.36\), and \(s=0.62\) hours. The observed screen-time values range from 1 to 7 hours.
State. Decide what the model can describe and whether it is reliable for predicting sleep for all students at the school.
Plan. Separate the model’s fit among the volunteers from the sample’s ability to represent all students. Check how participants were recruited and whether a requested prediction is within the observed \(x\)-range.
Do. The 48 students chose to respond to an open invitation, so this is a voluntary-response sample, not a random sample of the school. Students who notice or choose to answer a screen-time survey could differ from nonparticipants in their screen habits or sleep. The line summarizes the volunteer data. For a volunteer reporting \(x=5\) hours, the fitted value is
Five hours is within the observed range of 1 to 7 hours, so this is not an extrapolation relative to the volunteers’ data. Still, that does not make the prediction dependable for every student at the school. The sample’s \(r^2=0.36\) describes the fraction of variation in sleep accounted for by the linear relationship among the volunteers; it does not show that 36% of individual predictions are correct. The residual standard deviation \(s=0.62\) hours describes the typical size of residuals for the fitted model’s cases, not necessarily for students who did not volunteer.
Conclude. “Among the 48 students who volunteered, screen time and nightly sleep were negatively associated, and the fitted line predicts 7.50 hours of sleep for a volunteer reporting 5 hours of screen time. Because participation was voluntary, the results do not, by themselves, establish that the same relationship or prediction accuracy applies to all students at the school.”
Worked Example: Fertilizer Amount and Tomato Yield Among Experienced Gardeners
A fictional gardening club invites members with experience growing tomatoes to submit their fertilizer amounts and harvest weights. The resulting convenience sample contains 32 gardeners. Fertilizer amount, \(x\), ranges from 2 to 4 kilograms per plot. A regression program reports \(\hat{y}=3.2+1.6x\), where \(y\) is tomato yield in kilograms, with \(r^2=0.64\).
State. Evaluate a proposal to use the model to predict yield for a new gardener using 6 kilograms of fertilizer, and decide whether the result can be generalized to all gardeners.
Plan. Consider the convenience sample and its limited \(x\)-range separately. Check whether 6 kilograms lies within the observed range, and avoid treating a fitted value outside that range as established by the data.
Do. The sample includes experienced club members who were easy to recruit, not a random sample of all gardeners. The model’s \(r^2\) indicates that the line accounts for 64% of the variation in yield among these sampled gardeners. It does not show that the model accounts for that same fraction of variation among all gardeners.
The proposed input of 6 kilograms is above the observed maximum of 4 kilograms. Substitution into the line gives
This is a numerical output from the fitted equation, but it is an extrapolation: the sample provides no direct evidence about the line’s behavior at 6 kilograms. The convenience sample and the restricted range both limit the claim. The output should not be presented as a reliable prediction for a new gardener, and the observed association should not automatically be generalized to all gardeners.
Conclude. “For the experienced gardeners in this convenience sample, fertilizer amount and tomato yield had a positive linear association. The equation produces a fitted value of 12.8 kilograms at 6 kilograms of fertilizer, but that input is outside the observed range of 2 to 4 kilograms. Because the sample was nonrandom and the prediction extrapolates beyond its data, this value is not established as reliable for a new gardener or for gardeners generally.”
Worked Example: Cycling Hours and Resting Pulse in a Local Club
A fictional researcher visits one cycling club and asks members who are present to volunteer for a study. The 60 participants report weekly cycling hours, \(x\), and have resting pulse measured in beats per minute, \(y\). Their values range from 3 to 12 cycling hours per week. The reported regression output is \(\hat{y}=78-1.5x\), with \(r^2=0.81\) and \(s=4.0\) beats per minute.
State. Assess whether this apparently strong model can be used to predict resting pulse for all adults in the surrounding community.
Plan. Interpret the model summaries for the participants, then examine who was invited and who chose to participate. Identify whether the sample includes a fair basis for representing adults who are not cycling-club members.
Do. Among these 60 volunteers, the line has a negative slope: for each additional reported hour of weekly cycling, predicted resting pulse decreases by 1.5 beats per minute on average. The \(r^2\) value indicates that 81% of the variation in resting pulse among these participants is accounted for by the linear relationship with cycling hours. The \(s\) value describes the residual standard deviation for this fitted model, in beats per minute.
For a participant reporting 8 weekly cycling hours, the fitted pulse is
Eight hours lies within the sample’s observed range, so this calculation is within the model’s data range. But the participants were recruited at one cycling club and volunteered. Adults who cycle regularly or attend this club may differ from other adults in both cycling time and resting pulse. A high \(r^2\) and a relatively small residual standard deviation do not show that the same line or prediction accuracy applies to all adults in the community.
Conclude. “Among the 60 cycling-club members who volunteered, weekly cycling hours and resting pulse had a strong negative linear association. The model predicts a resting pulse of 66 beats per minute for a participant reporting 8 weekly cycling hours, but the convenience and voluntary selection do not support generalizing that relationship or prediction to all adults in the community.”
Common Mistakes and AP Exam Tips
When evaluating a regression model from a nonrandom sample, make your reasoning about the sampling method explicit. These errors can turn a careful description of the data into an unsupported population claim:
- Treating a high \(r^2\) as proof of broad reliability. State which cases the percentage describes: the observed sample. A high value does not establish that the model represents people who were not included.
- Treating a small \(s\) as guaranteed prediction accuracy elsewhere. Interpret \(s\) in response-variable units for the fitted cases. Do not promise the same typical error for a different population without evidence.
- Assuming a large volunteer sample is representative. A large number of volunteers can still differ systematically from people who did not respond. Sample size does not turn self-selection into random selection.
- Claiming a specific direction of bias without evidence. A nonrandom sample may change the apparent slope, strength, or predictions, but the direction is not automatically known. Explain the concern without asserting an unsupported result.
- Ignoring the predictor range. Check whether a proposed \(x\)-value is inside the range of the observed data. A value outside that range is an extrapolation, and a good fit within the sample does not validate it.
- Confusing generalization with causation. Even a representative sample can show an association without showing that changing \(x\) causes \(y\) to change, as discussed in “Association Versus Causation in Regression.”
A strong AP response usually names the sampled cases, identifies the convenience or voluntary selection, describes what the regression summarizes for those cases, and limits any broader conclusion. If the question asks about a prediction, also identify whether the \(x\)-value falls within the observed range and explain the scope of the prediction.
Check Your Understanding
For each situation, distinguish what the model says about its observed cases from what the sampling process supports beyond them.
- A website invites anyone interested in meal planning to submit weekly planning time and grocery spending. Explain why a regression using respondents may not describe all households.
- A convenience sample produces a regression with \(r^2=0.90\). What does this value say about the sampled cases, and what does it not establish about the target population?
- A model is fitted using predictor values from 10 to 25. Explain why a prediction at 40 deserves extra caution, even if the sample residuals are small.
- Name one way that volunteering might be connected to the explanatory variable or response variable, and explain why that possibility matters for generalization.
- Write a one- or two-sentence conclusion for a strong regression found among volunteers at one neighborhood running club. Limit the conclusion to what the sample method supports.