Classifying a Prediction by Its x-Value
In “Finding the Range of the Explanatory Variable,” we learned to find the smallest and largest observed \(x\)-values among the cases used to fit a regression line. Now we can use those endpoints to classify a prediction. The key question is whether the requested explanatory-variable value falls within that observed range or outside it.
A prediction within the observed \(x\)-range is called interpolation. A prediction below the observed minimum or above the observed maximum is called extrapolation. The classification depends on the requested \(x\)-value—not on the predicted response, how plausible the prediction sounds, or whether the prediction turns out to be close to an actual response.
If the observed \(x\)-range is from \(x_{\min}\) to \(x_{\max}\), a request is within the range when \(x_{\min}\leq x\leq x_{\max}\). Both endpoints count as within the range. A request is outside the range when \(x<x_{\min}\) or \(x>x_{\max}\).
The range belongs to the particular data used to fit the particular line. If a problem gives an observed range, use that range when classifying a request for that model. Keep the units consistent: a requested value in a different unit must be converted before comparing it with the endpoints.
A Practical Comparison
To classify a prediction, identify the model’s explanatory variable and the observed \(x\)-range. Then compare the requested \(x\)-value with both endpoints. Only after making that comparison do you need to calculate the predicted response, if the question asks for it. The fitted equation \(\hat{y}=a+bx\) gives the model’s predicted response at the requested \(x\), whether that request is interpolation or extrapolation.
Check which variable is explanatory and make sure the requested value and range use the same units.
Use the minimum and maximum \(x\)-values among the observations used to fit the model.
A value between or equal to the endpoints is within the range; a value below or above them is outside it.
Substitute the requested \(x\)-value into the fitted line and state what \(\hat{y}\) represents in context.
Interpolation does not require that the exact requested \(x\)-value appeared in the data. The definition uses the range’s endpoints, not a checklist of every recorded value. For instance, a request may fall between the observed minimum and maximum even if there is a gap in the observed \(x\)-values around it. It is still interpolation by this range-based classification.
Conversely, a value may be common, physically possible, or within a variable’s general limits and still be outside the range of the data used to fit the line. The observed range—not the full range of values that could occur—is the relevant comparison. As covered in “What Extrapolation Means,” a fitted line can produce a numerical prediction outside the observed range too; calculating one does not change its classification.
Worked Examples
Worked Example: A Prediction Between the Observed Endpoints
Fictional AP-style scenario. A greenhouse activity records the height of seedlings at different times after transplanting. Let \(x\) be days after transplanting and let \(\hat{y}\) be predicted height in centimeters. The observed \(x\)-range used to fit the model is from 5 to 29 days. The supplied fitted line is \(\hat{y}=12.4+0.85x\).
Classify the requested value. A prediction is requested for \(x=18\) days. The comparison is \(5\leq 18\leq 29\), so 18 days is within the observed range. The requested prediction is interpolation.
Calculate the predicted response. Substitute \(x=18\) into the fitted line:
Conclude in context. For seedlings at 18 days after transplanting, the model predicts a height of 27.7 centimeters. This is interpolation because 18 days lies between the observed endpoints of 5 and 29 days. The classification describes where the requested \(x\)-value falls; it does not claim that each seedling is exactly 27.7 centimeters tall.
Check the endpoint. A request at 29 days is also within the observed range because it equals the maximum. Substitution gives \(\hat{y}=12.4+0.85(29)=12.4+24.65=37.05\) centimeters. The model’s prediction at the endpoint is still interpolation under the inclusive endpoint rule.
Worked Example: Requests Below and Above the Range
Fictional AP-style scenario. A technology club measures the runtime of sample portable batteries after different numbers of charge cycles. Let \(x\) be the number of charge cycles and \(\hat{y}\) be predicted runtime in hours. The observed \(x\)-range used to fit the line is from 100 to 900 cycles. The model is \(\hat{y}=11.8-0.0045x\).
Classify the requests. First consider \(x=500\). Since \(100\leq 500\leq 900\), this is interpolation. A request at \(x=75\) is below the minimum of 100, so it is extrapolation. A request at \(x=1000\) is above the maximum of 900, so it is also extrapolation.
Calculate each prediction. For 500 cycles, \(\hat{y}=11.8-0.0045(500)=11.8-2.25=9.55\) hours. For 75 cycles, \(\hat{y}=11.8-0.0045(75)=11.8-0.3375=11.4625\) hours, or about 11.46 hours rounded to two decimal places. For 1000 cycles, \(\hat{y}=11.8-0.0045(1000)=11.8-4.5=7.3\) hours.
Conclude in context. The model predicts a runtime of 9.55 hours at 500 cycles, an interpolation request because 500 is within the observed range. It predicts about 11.46 hours at 75 cycles and 7.3 hours at 1000 cycles; both requests are extrapolations because they are outside the observed range. These labels come from comparing charge-cycle counts with 100 to 900, not from comparing the predicted runtimes with any range of runtime values.
Worked Example: A Value Within the Range That Was Not Observed
Fictional AP-style scenario. A community group records trip distance and commute time for a set of fictional bicycle trips. Let \(x\) be distance in kilometers and \(\hat{y}\) be predicted commute time in minutes. The observed distances used to fit the line are 1.2, 2.0, 3.1, 6.8, and 8.4 kilometers, so the observed \(x\)-range is from 1.2 to 8.4 kilometers. The supplied model is \(\hat{y}=6.5+3.2x\).
Classify the requests. A request at \(x=5.0\) kilometers is within the range because \(1.2\leq 5.0\leq 8.4\). Although 5.0 kilometers is not one of the listed observed distances, the prediction is interpolation. A request at \(x=9.0\) kilometers is greater than the maximum of 8.4 kilometers, so it is extrapolation. A request at \(x=1.2\) kilometers is exactly at the minimum and is within the range.
Calculate the predictions. At 5.0 kilometers, \(\hat{y}=6.5+3.2(5.0)=6.5+16=22.5\) minutes. At 9.0 kilometers, \(\hat{y}=6.5+3.2(9.0)=6.5+28.8=35.3\) minutes. At 1.2 kilometers, \(\hat{y}=6.5+3.2(1.2)=6.5+3.84=10.34\) minutes.
Conclude in context. The model predicts 22.5 minutes for a 5.0-kilometer trip; this is interpolation even though that exact distance was not observed. It predicts 35.3 minutes for a 9.0-kilometer trip, which is extrapolation because 9.0 exceeds the observed maximum. At 1.2 kilometers, the request is interpolation because the minimum endpoint is included.
Common Mistakes and AP Exam Tips
- Classifying by the predicted response. The label depends on the requested explanatory-variable value. A prediction is not interpolation just because its \(\hat{y}\) seems close to observed response values.
- Leaving out the endpoints. A request equal to \(x_{\min}\) or \(x_{\max}\) is within the observed range. Extrapolation means strictly below the minimum or strictly above the maximum.
- Requiring the requested value to appear in the data. A value between the endpoints is within the range even if it was not recorded. Say “within the observed range,” not “observed,” when the value itself was not in the dataset.
- Using a possible range instead of the observed range. Compare with the \(x\)-values used to fit the model, not with all values that could occur in the setting.
- Comparing values with different units. Convert the requested \(x\)-value or the range endpoints so that the comparison uses the same units.
- Giving a label without showing the comparison. A clear AP response identifies the observed endpoints and shows why the requested \(x\)-value is between them or beyond one of them.
- Calling an extrapolation prediction impossible by definition. The regression equation can produce a prediction outside the range. The correct classification is extrapolation; whether the prediction is useful or reliable is a separate question.
A complete classification might read: “The fitted model used distances from 1.2 km to 8.4 km. A request at 5.0 km is within that observed range, so it is interpolation, even though 5.0 km was not one of the observed distances.” For an outside value, state which endpoint it exceeds or falls below. If asked to calculate \(\hat{y}\), show the substitution and give the predicted response with its units.
Check Your Understanding
Use the stated observed \(x\)-range in each question. Remember that both endpoints are included.
- A model was fitted using values of \(x\) from 12 to 40. Classify a prediction requested at \(x=40\), and explain the comparison.
- The observed \(x\)-range is from 3.5 to 11.0 hours. Is a request at 2.0 hours interpolation or extrapolation? State why.
- A fitted model used distances of 2, 4, 7, and 10 kilometers. Classify a prediction requested at 6 kilometers, even though 6 is not among the listed values.
- A fitted model predicts a response for \(x=75\), and its observed \(x\)-range is from 20 to 90. Is the request interpolation or extrapolation? What determines the classification?
- A model used temperatures from 8°C to 22°C. A prediction is requested for 77°F. Convert the requested value to Celsius, then classify it.