Check the Predictor Before Calculating
In “Making Predictions Using the Equation,” you learned to substitute a predictor value into a regression equation \(\hat{y}=a+bx\). Before doing that, check whether the requested \(x\)-value is in the range of predictor values used to fit the line. A calculation can produce a number for almost any \(x\), but that does not mean the prediction is equally trustworthy everywhere.
The observed range runs from the smallest to the largest \(x\)-value in the data used to fit the regression line. Compare the requested predictor value with those endpoints. If it falls between them, using the line is called interpolation. If it falls below the minimum or above the maximum, using the line is called extrapolation.
The range to check is the range of the predictor, not the response. For a line predicting \(y\) from \(x\), find the smallest and largest observed \(x\)-values. Do not use the smallest and largest \(y\)-values to decide whether a requested \(x\) is within range.
Extrapolation is risky because the pattern in the observed data may not continue outside the values studied. The relationship could bend, level off, or change for another reason. Even an equation that returns a plausible-looking number does not provide evidence that the relationship continues that far.
Interpolation is generally better supported by the data, but it is not a promise of an accurate prediction. A value can be within the overall range but fall in a part of the range with few observations. Also, as discussed in “Strong Correlation Does Not Mean a Good Model,” a strong correlation alone does not establish that a linear model is appropriate. The pattern of the data still matters.
A Practical Prediction Check
Use this brief check before interpreting a regression prediction. It keeps the predictor and response roles clear and makes the limits of the estimate explicit.
Determine which variable is \(x\) in the equation and which value the question supplies.
Use the smallest and largest predictor values in the data that fitted the line.
A value between the endpoints is interpolation; a value outside them is extrapolation.
Substitute into \(\hat{y}=a+bx\), report the predicted response with its units, and state whether the value is within or outside the observed range.
If the requested value is outside the range, you can still evaluate the equation when the question asks for the calculation. But label it as an extrapolation and explain that the observed data do not directly support the line’s behavior there. Do not present the numerical result as a dependable forecast just because the arithmetic is straightforward.
Worked Examples
Worked Example: Predict Within the Observed Range
A fictional tutoring program records weekly practice hours \(x\) and assessment scores \(y\), in points, for students. The observed practice hours range from 1 to 8 hours. A fitted regression line is \(\hat{y}=41+5.4x\). Estimate the assessment score for a student who practices 5.5 hours per week.
First compare 5.5 with the observed predictor endpoints:
The requested value is within the observed range, so this is interpolation. Substitute \(x=5.5\) into the equation:
The regression line estimates an assessment score of 70.7 points for students who practice 5.5 hours per week. This is an estimate, not a statement that every student at that practice time will score 70.7 points. The prediction is interpolation because 5.5 hours lies between the smallest and largest practice times in the data.
Worked Example: Recognize an Extrapolation
A fictional delivery service models delivery time \(y\), in minutes, from route distance \(x\), in kilometers. Its training data include routes from 1.5 to 8 kilometers, and the fitted line is \(\hat{y}=6+4.5x\). What does the equation predict for a 12-kilometer route, and how should that result be described?
The requested distance is greater than the largest observed distance:
Thus, this is extrapolation. The equation’s numerical output is:
The line gives a predicted delivery time of 60 minutes for a 12-kilometer route. However, 12 kilometers is outside the observed range of 1.5 to 8 kilometers, so the result is an extrapolation. The data used to fit the line do not show whether the same pattern continues for longer routes. For example, longer routes might use different roads or encounter conditions not represented in the shorter-route data.
A complete response gives both parts: the calculation and the limitation. Saying only “60 minutes” leaves out the important fact that the prediction is beyond the data range. Saying “no prediction can be calculated” is also too strong—the equation can be evaluated, but the result is less supported.
Worked Example: Predict at an Endpoint
A fictional environmental project relates water temperature \(x\), in degrees Celsius, to the measured growth of a type of aquatic plant \(y\), in millimeters over a set period. Temperatures in the data range from 12°C to 30°C. The fitted line is \(\hat{y}=3+1.8x\). Estimate growth at 12°C.
The requested value equals the smallest observed predictor value. It is not outside the range:
The prediction is at the lower endpoint of the observed range. Substitute \(x=12\):
The regression line estimates plant growth of 24.6 millimeters at 12°C. This uses the line at an observed-range endpoint, not beyond the data. It is still an estimate: being within the range does not mean that each plant at 12°C would have exactly this growth. The line summarizes the fitted pattern, and the actual responses can vary.
What the Range Check Can—and Cannot—Tell You
The range check answers a specific question: does the requested predictor value fall among the predictor values represented in the data? It does not, by itself, prove that a linear prediction is accurate. Keep these points in mind:
- Use the observed data that fitted the line. If the regression was fitted using one group of cases, check the predictor range for that group, not a broader range from a different source.
- Check the predictor, not the response. For \(\hat{y}=a+bx\), compare the requested \(x\) with the observed \(x\)-values.
- Do not treat the endpoints as a guarantee. A value exactly at the minimum or maximum is within the observed range, but it is still a prediction and can differ from an actual response.
- Do not treat every in-range value as equally supported. A value between the endpoints may lie in a sparse part of the data. Consider the distribution of the observed \(x\)-values and whether a linear model is appropriate.
- Do not confuse arithmetic with evidence. A calculator can evaluate the line outside the range; that does not show that the linear pattern continues there.
This check also does not turn an association into a cause-and-effect conclusion. As covered in “Why Correlation Does Not Imply Causation,” an observed relationship alone does not show that changing the predictor would cause the response to change. Describe what the regression estimates, without claiming more than the data support.
Common Mistakes and AP Exam Tips
- Checking the wrong variable’s range. A prediction from \(x\) to \(y\) requires checking the observed \(x\)-values. Name the predictor and use its minimum and maximum.
- Calling a value at an endpoint extrapolation. A requested value equal to the observed minimum or maximum is within the range. Show the comparison, such as \(12 \leq 12 \leq 30\).
- Reporting an extrapolated answer without a warning. Include the computed value if requested, then state that it is extrapolation because the predictor is outside the observed range and that the pattern may not continue.
- Describing a prediction as a certainty. Avoid “the response will be 70.7.” A careful statement says the line estimates the response, in context and with units.
- Assuming interpolation guarantees a good model. “Within the range” describes where the predictor value falls; it does not replace checking whether the data support a linear model.
For full-credit communication, identify the observed predictor range, state whether the requested value is within or outside it, show the substitution and arithmetic, and interpret the result as an estimated response in context. If it is outside the range, clearly qualify the prediction as extrapolation.
Check Your Understanding
For each question, check the predictor range before interpreting the regression prediction.
- A line predicting battery life \(y\) from screen brightness \(x\) was fitted using brightness values from 20 to 90 units. Is a prediction at \(x=55\) interpolation or extrapolation? Explain.
- A regression line predicts commute time from distance. The observed distances are 2 to 14 kilometers, and the equation is \(\hat{y}=5+3.6x\). Calculate the prediction for 18 kilometers and explain how it should be qualified.
- A dataset has predictor values from 4 to 16. Is using the line at \(x=4\) interpolation or extrapolation? Explain how the endpoint is treated.
- Why is the observed range to check the range of \(x\), rather than the range of \(y\), when using \(\hat{y}=a+bx\)?
- Write one sentence describing an in-range prediction that makes clear it is an estimate rather than a guaranteed individual outcome.