Two Questions Define a Model’s Scope
A regression model can fit its observed data well and still have a limited scope. Before applying it to a new case, ask two separate questions: Which population do the observed cases represent? And which values of the explanatory variable \(x\) are covered by the data used to fit the model? A model’s usefulness beyond its original data depends on both answers.
In “Population Scope and Generalizing Results,” we traced a conclusion back to the group from which cases were selected. In “Finding the Range of the Explanatory Variable,” we identified the smallest and largest observed \(x\)-values. Here we put those ideas together. A prediction can concern the right population but an unsupported \(x\)-value, or a familiar \(x\)-value but a population the study did not represent.
These limits are different. A representative sample from a well-defined population may support generalizing to that population, but it does not automatically support predictions at \(x\)-values outside the observed range. Conversely, a requested \(x\)-value can fall within the observed range while the model still lacks support for a different population or setting.
Being inside both limits makes applying the model more defensible, not certain. The relationship could differ across groups, measurement procedures, or time periods. A fitted line summarizes the observed association; its scope is not a guarantee that every individual response will match its prediction.
Population Scope: Who Do the Cases Represent?
The population of a regression study is the group of individuals or objects the study aims to describe. The cases are the actual individuals or objects represented by the rows of data. To assess population scope, identify those cases and then ask how they entered the study. As discussed in “Describing the Context of a Regression Study,” a useful account names the cases, variables, and data-collection setting.
A random sample selected from an appropriate frame can support generalizing to the population represented by that frame, subject to concerns such as nonresponse or incomplete coverage. A convenience or voluntary-response sample may describe the sampled cases, but it usually does not establish that they represent the broader population. A large sample does not repair a selection method that systematically misses part of the population.
The exact population matters. If researchers randomly select households from one city, that design does not by itself establish that the households represent all households in the country. Similarly, a model based on one type of device, school, or growing condition may not apply to other types just because the variables have the same names.
Population scope is also distinct from causation. As explained in “Association Versus Causation in Regression,” a regression describes an association and does not by itself show that changing \(x\) causes a change in \(y\). Random sampling can help support generalization to a population; random assignment in an experiment addresses a different question about cause and effect.
x-Scope: Which Values Are Covered?
The model’s \(x\)-scope is anchored to the explanatory-variable values among the cases used to fit the line. The observed range runs from the minimum to the maximum \(x\)-value. A requested value within that range is an interpolation; a value outside it is an extrapolation, as defined in “Interpolation Versus Extrapolation.”
An in-range value is generally better supported by the data than an out-of-range value, but range membership is not enough by itself. There should be relevant observations near the requested value, and the data should come from a setting suitable for the intended application. Also check whether the relationship is reasonably linear for the purpose at hand; “How Unmodeled Curvature Shows Up Outside the Data” explains why an apparently useful line may miss a curved pattern.
The range is specific to the data used to fit that particular model. If one study measures \(x\) from 2 to 20 and another measures it from 10 to 60, their fitted models do not automatically share the same \(x\)-scope. Do not substitute a general idea of what values are possible in the world for the actual range of the model’s data.
Worked Examples: Matching a Model to Its Scope
Worked Example: Stream Measurements in One Watershed
Hypothetical setting. A team randomly selects 48 streams from a complete list of streams in one watershed. During the same month, they measure nitrate concentration \(x\), in milligrams per liter, and algae level \(y\), in a consistent unit. The observed nitrate concentrations range from 0.4 to 3.8 milligrams per liter. The team fits a linear regression model.
State. The relevant population is the streams on the watershed’s sampling list, under the conditions represented by the measurement period and procedures. The model’s observed \(x\)-range is 0.4 to 3.8 milligrams per liter.
Plan. For each proposed use, check population and \(x\)-scope separately. A random sample from the list supports generalizing to that list’s population, if the list is suitable and the sampling process is carried out as intended. Compare each requested nitrate concentration with the observed endpoints. Do not assume the model applies to streams in another watershed or to measurements from a substantially different season without additional support.
Do. A request to estimate the mean algae level for streams in this watershed at 2.6 milligrams per liter passes both basic checks: the target streams are from the represented watershed, and \(2.6\) is between 0.4 and 3.8. A request at 4.5 milligrams per liter is outside the observed range, so applying the line there is extrapolation even if the target is still a stream in the same watershed. A request for streams in a different watershed raises a population-and-setting concern even if their nitrate concentration is 2.6 milligrams per liter.
Conclude in context. The model is most directly supported for streams represented by the watershed sampling list, measured under comparable conditions, and for nitrate concentrations from 0.4 through 3.8 milligrams per liter. A prediction for a different watershed or at 4.5 milligrams per liter is not supported by both parts of this study’s scope.
Worked Example: A Convenient Sample of Cyclists
Hypothetical setting. A repair shop posts an online survey link and receives responses from 120 cyclists who choose to complete it. Respondents report weekly cycling distance \(x\), in miles, and annual repair visits \(y\). Their weekly distances range from 3 to 42 miles. A fitted line has a strong linear pattern among the respondents.
State. The observed cases are the 120 cyclists who responded, not a random sample of all cyclists in the city. The model’s observed \(x\)-range is 3 to 42 miles per week.
Plan. Separate what the model describes in the respondent data from what could be generalized to the city’s cyclists. Then check whether a proposed \(x\)-value lies between 3 and 42. A strong association or a large number of responses does not remove the selection concern.
Do. A model prediction at 18 miles per week is within the respondents’ observed range. It can be used to describe the fitted relationship among the survey respondents at that value, with the usual caution about individual variation. However, cyclists who chose to answer may differ from those who did not—for example, in cycling habits or interest in repairs. The voluntary-response design does not establish that the respondents represent all city cyclists. A prediction at 50 miles per week also falls outside the observed range, adding an extrapolation concern.
Conclude in context. The fitted line summarizes the association between weekly distance and annual repair visits among the 120 responding cyclists, over their observed range of 3 to 42 miles per week. The study does not, by its sampling method alone, justify generalizing the model to all city cyclists. A prediction at 50 miles per week is additionally outside the model’s observed \(x\)-range.
Worked Example: Same Population, Different x-Values
Hypothetical setting. Researchers randomly select farms from a complete register of farms in one county. For the sampled farms, they measure acreage \(x\) and monthly irrigation use \(y\), using the same procedures. The sampled acreage values range from 10 to 200 acres. A county planning group wants to use the fitted line for farms in that county.
State. The random selection supports generalizing to farms represented by the county register, assuming the register and data collection are suitable. The observed \(x\)-range is 10 to 200 acres.
Plan. Check whether each target farm belongs to the represented population and whether its acreage is within the observed range. A supported population does not extend the line’s \(x\)-range.
Do. A prediction for a 75-acre farm in the county passes both basic scope checks: the farm is in the represented county population, and 75 is between 10 and 200. A prediction for a 320-acre farm in the same county passes the population check but not the \(x\)-range check. The line would be extrapolated for that farm. A 75-acre farm in another county has an acreage within the observed range, but its population and growing conditions may differ from those represented by the study.
Conclude in context. The model’s strongest scope is farms represented by the county register with acreages from 10 to 200 acres, measured under comparable conditions. The 320-acre farm is outside the observed \(x\)-range, while a farm in another county raises a population-and-setting question even at an in-range acreage.
Writing a Complete Scope Statement
A clear scope statement does more than say, “The model applies to the data.” It identifies the cases, gives the population or setting supported by the study, and states the observed \(x\)-range. When a proposed use does not fit, say which part of the scope it exceeds. This makes the limitation specific rather than vague.
The statement should not overclaim that every possible case within the limits will follow the line closely. As discussed in “Reliability Within the Range of Data” and “Using Residual Standard Deviation to Gauge Prediction Error,” in-range predictions can still have substantial scatter around the fitted line. Scope answers where the model is supported; it does not promise an exact prediction.
Common Mistakes and AP Exam Tips
- Using “the population” without identifying it. Name the cases and the group from which they were selected. “All farms” is too broad if the study sampled only farms in one county.
- Treating a strong fit as proof of representativeness. A high \(r^2\) describes the linear relationship in the data, as covered in “r-squared and Strength of Association.” It does not show that a voluntary sample represents a larger population.
- Checking only the population or only the \(x\)-value. State both. A target can match the sampled population but be outside the observed \(x\)-range, or be in range but belong to a different population.
- Calling every in-range prediction reliable. In-range status avoids extrapolation, but does not guarantee a good prediction. Consider the data’s scatter, the presence of relevant cases near that \(x\)-value, and whether the setting is comparable.
- Confusing generalization with causation. Sampling can support conclusions about a population; it does not by itself show that changing \(x\) causes a change in \(y\). Keep the study design and the model’s scope distinct.
- Writing only “use with caution.” Explain why. Identify the unsupported population, setting, or \(x\)-value, and name the evidence that sets the limit.
For full-credit communication, be precise about both dimensions: “The model describes the association for the randomly selected streams in this watershed and is supported for nitrate concentrations from 0.4 to 3.8 milligrams per liter. A prediction for a stream in another watershed is not automatically supported, even at an in-range concentration; a prediction at 4.5 milligrams per liter is an extrapolation.”
Check Your Understanding
For each situation, identify the relevant population and \(x\)-scope limits.
- A random sample of apartments in one town has monthly heating costs measured for outdoor temperatures from 5 to 24 degrees. What two details should you report when describing the model’s scope?
- A model is based on a voluntary survey of 80 gardeners, with garden sizes from 20 to 500 square feet. What can the sampling method support, and what can it not establish by itself?
- A randomly selected sample of local buses has ages from 1 to 12 years. A transit agency wants a prediction for a 15-year-old bus in the same fleet. Which scope check does the request pass, and which does it fail?
- Why does an \(x\)-value within the observed range not automatically mean a model should be applied to a different region?
- Write a one-sentence scope statement for a model fitted to randomly selected orchard trees in one county, where trunk diameters range from 4 to 30 centimeters.