Two Checks, Not One
A regression prediction can pass one important check and fail another. The requested explanatory-variable value might be within the observed range, while the predicted response is impossible in context. Or the prediction might be an extrapolation, while the calculated response is still physically possible. A careful answer checks both.
As in “Finding the Range of the Explanatory Variable” and “Checking Whether a Prediction Is Reasonable,” the checks answer different questions. The range check compares the requested \(x\)-value with the minimum and maximum \(x\)-values in the data used to fit the line. The reasonableness check asks whether the predicted response could make sense in the setting. A reasonable-looking result does not make an extrapolation supported, and an in-range request does not guarantee a sensible prediction.
The observed endpoints are included in the range: a requested \(x\)-value equal to either endpoint is within the observed range. As covered in “Interpolation Versus Extrapolation,” a request inside or at an endpoint is interpolation; a request below the minimum or above the maximum is extrapolation. Neither label, by itself, decides whether the response prediction is sensible.
A useful mixed-practice routine is to identify the requested \(x\), write the observed range, classify the request, calculate the line’s prediction if needed, then check the response’s meaning and limits. Keep the reasoning separate as you draft. That makes it clear whether a concern comes from the predictor being out of range, the response being implausible, or both.
A Four-Part Check for Mixed Questions
Name the explanatory variable and its units. Compare the requested value with the observed minimum and maximum for the cases used to fit the line.
Say whether it is interpolation or extrapolation, and give the range comparison that supports the classification.
Substitute the requested \(x\)-value into the fitted line. State what the result predicts, with response units.
Ask whether the predicted response can occur or makes practical sense. Explain any concern in context, and do not claim that an unsupported prediction is certainly wrong.
The last step is not a demand to reject every prediction that seems surprising. It is a prompt to compare the model’s output with what is possible or reasonable in the stated setting. A percentage cannot be below 0% or above 100%; a distance cannot be negative. Other settings may have practical limits that are less absolute, so explain why a result raises a concern rather than treating a possibility as proven.
Worked Example: An Out-of-Range Request with a Possible Response
Original AP-style question. A greenhouse team models the percentage of seeds that germinate, \(y\), using daily light exposure, \(x\), in hours. The observed exposures ranged from 4 to 10 hours. The fitted line is \(\hat{y}=18+6.5x\). What does the line predict at 12 hours? Check the range and the response.
State. The requested exposure is 12 hours, which is greater than the observed maximum of 10 hours. The prediction is an extrapolation.
Plan. First compare 12 hours with the observed \(x\)-range of 4 to 10 hours. Then substitute 12 into the fitted line. Finally, compare the predicted germination percentage with the possible range of percentages, from 0% to 100%, and consider whether the extrapolation is supported.
Do. Substituting \(x=12\) gives
Conclude. The prediction at 12 hours is an extrapolation because the observed light exposures ranged from 4 to 10 hours. The line predicts that 96% of seeds would germinate at 12 hours. That percentage is within the possible range of 0% to 100%, so it is not impossible on that basis. However, the data do not cover 12 hours, and the relationship might not keep increasing at the same rate—for example, germination could level off. The output passes a basic response-bound check, but that does not remove the concern about extrapolating.
This example shows why “96% is possible” and “12 hours is within the data range” are not interchangeable claims. The first is a response check; the second is false. A full answer gives the range comparison and the response assessment separately.
In-Range Does Not Mean Reasonable
It is tempting to assume that a prediction is dependable whenever its \(x\)-value is between the observed endpoints. Interpolation is generally better supported by the data than extrapolation, as discussed in “Reliability Within the Range of Data,” but that classification is not a guarantee that every fitted value is appropriate. The line may produce an impossible result even at an \(x\)-value represented by the data range.
When a response has a clear mathematical bound, name it. A percentage score, for instance, must be between 0% and 100%. If the line predicts more than 100%, explain that the calculated value is incompatible with the meaning of the response. The arithmetic can be correct even though the model’s output is not reasonable.
Worked Example: An Interpolation with an Impossible Score
Original AP-style question. A school models a practice test score, \(y\), as a percentage, using hours spent on an online review program, \(x\). The observed study times ranged from 5 to 15 hours. The fitted line is \(\hat{y}=30+6x\). A student asks what score the line predicts at 14 hours. Classify the request and assess the prediction.
State. Fourteen hours is between the observed endpoints, 5 and 15 hours. This is interpolation.
Plan. Being in range determines the classification, but not whether the output is sensible. Substitute 14 into the fitted line, then check the result against the possible range for a percentage score.
Do.
Conclude. The request is interpolation because 14 hours is within the observed range of 5 to 15 hours. The fitted line calculates a score of 114%, but a percentage score cannot exceed 100%. Thus, the prediction is not reasonable as a test score, even though the requested study time is in range. The model’s fitted value should not be reported as a plausible score without noting this limitation.
The classification and the reasonableness judgment are both needed. Saying “it is interpolation, so it is reasonable” would confuse the two checks. Saying “114% is extrapolation” would also be incorrect: extrapolation is determined by \(x\), not by the size of \(\hat{y}\).
When Both Checks Raise Concerns
A request can be outside the observed \(x\)-range and produce a response that is impossible or nonsensical. In that case, state both issues. The range comparison explains why the model’s use is extrapolation; the response’s meaning explains why the particular calculated result cannot be accepted literally.
Do not let one concern replace the other. If you report only that a distance is negative, you have not classified the requested \(x\)-value. If you report only that the request is extrapolation, you have left out the more direct problem that a negative distance cannot occur.
Worked Example: An Extrapolation That Predicts a Negative Distance
Original AP-style question. A repair shop models the distance an electric bicycle travels on a full charge, \(y\), in kilometers, using battery age, \(x\), in years. The observed battery ages ranged from 1 to 4 years. The fitted line is \(\hat{y}=42-9x\). What does the line predict for a 6-year-old battery, and how should the result be evaluated?
State. The requested age of 6 years is above the observed maximum of 4 years, so the request is extrapolation.
Plan. Calculate the fitted value at 6 years, then check whether the predicted distance is possible. Since distance traveled cannot be negative, a negative result would conflict with the response’s meaning.
Do.
Conclude. The prediction for a 6-year-old battery is an extrapolation because the observed battery ages ranged only from 1 to 4 years. The fitted line calculates \(-12\) kilometers, which is not a possible travel distance. The output is therefore unreasonable in context, in addition to being an extrapolation. The calculation describes what the line gives, not a sensible claim that the bicycle can travel a negative distance.
Notice the careful wording: the line calculates a negative value; the real bicycle does not travel a negative distance. The result signals that the fitted line should not be used literally for this request. It does not, by itself, establish exactly what distance a 6-year-old battery would achieve.
A Plausible Response Can Still Be Weakly Supported
The reverse combination is also possible: an out-of-range request can produce a response that seems entirely plausible. For example, a model may predict a positive temperature, a reasonable production amount, or a feasible travel distance beyond the observed \(x\)-range. That response check matters, but it does not turn extrapolation into interpolation.
“Why Extrapolation Is Risky” and “Writing an Extrapolation Critique” emphasize that a line’s mathematical pattern may not continue outside the data. A useful justification connects that general concern to the context. Think about whether conditions could change, whether the response might level off, or whether the process behaves differently near a boundary. Use qualified language such as “may” or “could” when the scenario makes a possibility plausible but does not establish it as fact.
Worked Example: A Plausible Calibration Reading Beyond the Data
Original AP-style question. A technician uses a fitted line to relate a sensor’s input voltage, \(x\), in volts, to a temperature reading, \(y\), in degrees Celsius. The voltages used to fit the line ranged from 2 to 5 volts. The line is \(\hat{y}=-8+12x\). What does it predict at 1.5 volts, and what should the technician say about the result?
State. The requested input of 1.5 volts is below the observed minimum of 2 volts. The request is extrapolation.
Plan and do. Substitute 1.5 into the line, then check whether the resulting temperature is at least possible. A plausible temperature does not establish that the sensor’s fitted relationship applies below 2 volts.
Conclude. The line predicts \(10^\circ\text{C}\) at 1.5 volts. This is a possible temperature, but the prediction is an extrapolation because the observed voltages ranged from 2 to 5 volts. Sensor behavior may differ below the calibrated range, so the technician should not treat the plausible-looking result as confirmed by these data.
This example is a reminder to avoid the shortcut “the answer looks reasonable, so the model is reliable.” The output’s feasibility and the evidence supporting its use are separate questions.
Common Mistakes and Full-Credit Communication
- Checking only the predicted response. A plausible value does not establish that the request is within range. State the observed \(x\)-range and compare the requested \(x\)-value with its endpoints.
- Using \(\hat{y}\) to classify extrapolation. The classification depends on \(x\), not on whether the predicted response is large, small, negative, or surprising.
- Calling every interpolation reasonable. In-range requests can still produce outputs that violate response limits. Check the meaning of the response and any stated bounds.
- Calling an extrapolation definitely wrong. Extrapolation means the requested \(x\)-value is outside the observed range; it does not prove the prediction is false. Say why it may be unreliable and connect the concern to the context.
- Leaving out units or context. Report what the fitted line predicts and include the response units. A bare number does not tell the reader what the model’s output represents.
- Mixing up the line’s calculation and reality. If the model calculates an impossible value, say that the line predicts it and explain why it is not a reasonable real-world response. Do not describe the impossible value as an observed fact.
A concise but complete response can follow this pattern: “The request at [requested \(x\)] is [interpolation/extrapolation] because the observed \(x\)-values ranged from [minimum] to [maximum]. The fitted line predicts [\(\hat{y}\) and units]. This result [does/does not] satisfy the response’s practical limits. [Give a context-specific reason for caution if needed.]” Adapt the final sentence to the situation; do not add a concern that the question does not support.
Check Your Understanding
For each situation, distinguish the \(x\)-range classification from the response-reasonableness check. Include a contextual justification where appropriate.
- A model predicts the percentage of batteries passing a test from their charging time. The observed times were 2 to 8 hours. The line is \(\hat{y}=55+4x\). Classify and evaluate the prediction at 7 hours.
- A fitted line predicts a rainfall total at a requested \(x\)-value of 11, although the observed \(x\)-values ranged from 3 to 9. What comparison determines whether this is extrapolation? Does a plausible rainfall prediction change that classification?
- A model predicts a percentage score from practice time. The observed practice times ranged from 4 to 12 hours, and the line is \(\hat{y}=25+7x\). Calculate the prediction at 12 hours and identify any response-limit concern.
- A model predicts distance per charge from battery age. The observed ages ranged from 1 to 3 years, and the fitted line is \(\hat{y}=36-8x\). Calculate the prediction at 5 years. Identify both the range classification and the reasonableness issue.
- Rewrite this statement to separate its two claims: “The predicted value is possible, so it is not extrapolation.”