Tutorials › AP Statistics › What Extrapolation Means

Extrapolation and prediction limits · Tutorial 941 of 1000

What Extrapolation Means

Recognize when a regression prediction extends beyond the observed explanatory-variable values and explain why it may be less reliable.

Intermediate 9 min read

What You'll Learn

  • Define extrapolation using the observed values of the explanatory variable
  • Distinguish extrapolation from prediction within the observed x-range
  • Check whether a requested x-value is inside or outside the data range
  • Calculate and interpret predictions beyond that range without treating them as certain
  • Explain why a linear pattern may not continue beyond the observations
  • Use a domain check to communicate the limits of a regression prediction

When a Regression Prediction Goes Beyond the Data

In “Regression Context Mixed Practice Set,” we interpreted fitted lines in context and checked whether study design supported the claims made from them. Now we add a related question: is the requested prediction being made for an explanatory-variable value represented in the data, or does it go beyond the values used to fit the line?

A regression line can produce a predicted response for many possible values of \(x\). But the fact that a calculation is possible does not mean that the prediction is well supported. A line summarizes the pattern in the observations; beyond the observed \(x\)-values, we have no observations in this data set to show whether that pattern continues.

Definition: Extrapolation is using a fitted model to predict the response for an explanatory-variable value outside the range of explanatory-variable values in the data used to fit the model. A prediction for an \(x\)-value within that observed range is interpolation.

The key comparison is with the observed values of the explanatory variable, not with the range of observed responses. For example, a predicted response may be higher than every response in the data even when the explanatory-variable value is inside the observed range. That fact alone does not make the prediction extrapolation. Conversely, a prediction can be extrapolation even if the predicted response happens to fall among the observed response values.

To check a request, identify the smallest and largest observed \(x\)-values, with their units, and compare the requested \(x\) to those endpoints. If the requested value is below the observed minimum or above the observed maximum, the prediction is extrapolation. If it is between the endpoints, the prediction is interpolation. Knowing that it is interpolation does not guarantee high accuracy; it only means the explanatory-variable value is within the range represented by the data.

Domain check: Before calculating or interpreting a prediction, ask: What \(x\)-values were observed? What \(x\)-value is requested? Is that value inside or outside the observed interval? If it is outside, state that the prediction is an extrapolation and discuss why the relationship might not continue.

Why Extrapolation Calls for Caution

A least-squares line is chosen to summarize the linear pattern among the observed cases. The data can show how the variables were related over the observed \(x\)-range, but they do not directly show what happens beyond it. A linear pattern may bend, level off, or change for reasons that were not visible in the data.

Consider a model relating the number of days a plant has grown to its height. A straight line might describe growth over the first few weeks. Extending that line to several years would be questionable: the plant cannot grow at the same rate indefinitely. The equation will still return a number, but a sensible calculation is not automatically a sensible prediction.

Extrapolation is not automatically forbidden or always useless. A prediction outside the observed range may be considered when there is a strong subject-matter reason to expect the same relationship to continue. Even then, it should be identified as extrapolation, and its extra uncertainty should be acknowledged. New observations in the relevant range would provide more direct evidence than simply extending the old line.

A high \(r^2\) or a small residual standard deviation \(s\), as discussed in earlier tutorials, summarizes how well the model describes the data used to fit it. Neither establishes that the relationship continues outside those data. A model can fit the observed cases closely and still give poor predictions when extended beyond them.

Key distinction: A regression calculation tells you what the fitted line predicts at a chosen \(x\). The observed \(x\)-range tells you whether that prediction is interpolation or extrapolation. The prediction’s credibility also depends on whether the relationship is reasonable beyond the data.

Worked Examples

Worked Example: Seedling Height After the Observation Period

Fictional AP-style scenario. In a classroom greenhouse activity, students record the height of a particular type of seedling on days 6 through 24 after planting. Here \(x\) is the number of days after planting, and \(y\) is height in centimeters. A fitted line is \(\hat{y}=4.5+1.25x\). A student uses it to predict the seedling’s height on day 40.

State. The observed explanatory-variable values run from day 6 to day 24. The requested value, day 40, is greater than the largest observed value, day 24. The requested prediction is therefore an extrapolation.

Plan. Substitute \(x=40\) into the fitted line to find its prediction. Then interpret the result as what the line predicts, not as an observed height or a guaranteed outcome.

Do.

$$ \hat{y}=4.5+1.25(40)=4.5+50=54.5\text{ centimeters}. $$

The fitted line predicts a height of 54.5 centimeters on day 40. However, that is 16 days beyond the latest observation, so the model’s prediction is extrapolation. The seedlings may not keep growing by the same predicted amount per day: growth could slow, or conditions in the greenhouse could change.

For comparison, predicting on day 18 would be interpolation because day 18 is between days 6 and 24. The line would predict \(4.5+1.25(18)=27\) centimeters. This prediction is within the observed \(x\)-range, although it still would not guarantee that a particular seedling is exactly 27 centimeters tall.

Conclude in context. The line gives an extrapolated prediction of 54.5 centimeters for the seedling on day 40. Because the data covered only days 6 through 24, the prediction should be treated cautiously; the observed linear pattern may not continue through day 40.

Worked Example: Extrapolating a Phone Battery Model

Fictional AP-style scenario. A student tests a phone by playing video continuously and recording its battery percentage after different amounts of time. The observed video-play times range from 0.5 to 4.5 hours. Let \(x\) be video-play time in hours and \(y\) be battery percentage. The fitted line is \(\hat{y}=100-12.8x\). Someone asks what the model predicts after 8 hours.

State and plan. The requested \(x\)-value, 8 hours, exceeds the largest observed time, 4.5 hours. This is extrapolation. Substitute 8 into the model, then consider whether the prediction makes sense for battery percentage.

Do.

$$ \hat{y}=100-12.8(8)=100-102.4=-2.4\%. $$

The fitted line predicts \(-2.4\%\) battery after 8 hours. A negative battery percentage is not physically meaningful. This result is a clear warning that extending the fitted line beyond the observed times produces an unreasonable prediction. In actual use, the phone would likely shut down before its battery reached a negative percentage.

A prediction at 3 hours would be interpolation, since 3 is between 0.5 and 4.5 hours. The line predicts \(100-12.8(3)=61.6\%\) battery at that time. Even there, the value is the model’s prediction, not a guarantee for every phone or every video-playing session.

Conclude in context. The model calculates \(-2.4\%\) after 8 hours, but this is an extrapolation and an impossible battery percentage. It should not be reported as a plausible battery level. The example shows why a prediction should be checked against both the observed \(x\)-range and what is possible in context.

Worked Example: Delivery Time Beyond the Recorded Distances

Fictional AP-style scenario. A local delivery service records delivery distance and travel time for a set of deliveries on its usual routes. Distances range from 1 to 12 kilometers. Let \(x\) be distance in kilometers and \(y\) be travel time in minutes. The fitted line is \(\hat{y}=6+3.4x\). A customer asks for a prediction at 20 kilometers.

State and plan. The observed distances extend only to 12 kilometers, while the requested distance is 20 kilometers. Since 20 is outside the observed interval from 1 to 12 kilometers, this is extrapolation. Substitute the requested distance and then consider whether the usual-route pattern is likely to apply.

Do.

$$ \hat{y}=6+3.4(20)=6+68=74\text{ minutes}. $$

The model predicts a travel time of 74 minutes for a 20-kilometer delivery. This is an extrapolated prediction, 8 kilometers beyond the longest distance in the data. Longer deliveries may use different roads or travel conditions, so the average relationship observed for distances up to 12 kilometers may not continue in the same way.

For an 8-kilometer delivery, the line predicts \(6+3.4(8)=33.2\) minutes. Because 8 kilometers is inside the observed distance interval, this is interpolation. The model still gives an average prediction rather than a promise about an individual trip; traffic, route, and other circumstances can affect the actual time.

Conclude in context. The line predicts 74 minutes for a 20-kilometer delivery, but this prediction is extrapolation because the fitted data included distances only from 1 to 12 kilometers. The prediction should be treated cautiously unless there is good reason to believe the same distance-time pattern continues for longer routes.

A Practical Way to Report an Extrapolated Prediction

A clear answer separates the arithmetic from the strength of the prediction. First report what the model calculates. Then compare the requested \(x\)-value with the observed \(x\)-range and name the prediction as extrapolation. Finally, explain a context-specific reason the pattern may not continue, or note that the data do not establish what happens beyond the range.

1
Locate the observed \(x\)-range.
State the lowest and highest explanatory-variable values represented in the data, including units.
2
Compare the requested value.
If the requested \(x\) is below the observed minimum or above the observed maximum, identify the prediction as extrapolation.
3
Calculate and interpret the model prediction.
Use the fitted line and give the predicted response with its context and units.
4
Explain the limit.
State that the observed data do not show whether the relationship continues beyond the range, and give a relevant reason it could change when possible.

For instance, a strong response might say: “The fitted line predicts a travel time of 74 minutes for a 20-kilometer delivery. Because the observed distances ranged only from 1 to 12 kilometers, this is extrapolation; the model does not establish that the same relationship applies to longer routes.” That wording gives the numerical prediction without presenting it as certain.

Common Mistakes and AP Exam Tips

  • Checking the response values instead of the explanatory values. Extrapolation is determined by whether the requested \(x\)-value is outside the observed \(x\)-range. Do not decide based only on whether the predicted \(y\) is outside the observed response values.
  • Calling every prediction beyond the data set extrapolation. The relevant comparison is with the observed range of the explanatory variable, not simply with the number of rows or the last case in a list. A prediction at an \(x\)-value between the observed endpoints is interpolation.
  • Assuming a line must continue because it fits well. A high \(r^2\) describes variation accounted for by the linear relationship in the data. It does not prove that the relationship continues beyond the observed \(x\)-values.
  • Reporting a calculated value as a dependable fact. The fitted line may produce a number even when that number is implausible, as in the negative battery prediction. Check whether the prediction is reasonable in context.
  • Using “extrapolation” without identifying why. A full-credit explanation names the requested \(x\), gives the observed \(x\)-range, and states which endpoint the request exceeds. “It is outside the data” is less precise.
  • Claiming that extrapolation is always invalid. It is a warning about limited evidence, not a proof that the prediction is wrong. A responsible response says the prediction is less supported by these data and explains what could make the relationship change.

In a written answer, keep the model’s prediction separate from the evidence supporting it. The equation supplies a predicted response. The observed range establishes whether the prediction is extrapolation. Context helps you decide whether extending the pattern is reasonable, but it does not turn unobserved values into observed evidence.

Key takeaway: Extrapolation means predicting at an explanatory-variable value outside the range used to fit the regression model. The line still gives a calculation, but the observed data do not show that the relationship continues there, so report the prediction with appropriate caution.

Check Your Understanding

For each question, decide whether the prediction is interpolation or extrapolation, and explain what that classification means in context.

  1. A model relating weekly exercise time \(x\), in hours, to a fitness measure was fitted using \(x\)-values from 1 to 6 hours. Is a prediction at 4 hours interpolation or extrapolation? Explain.
  2. Using the same model, someone requests a prediction at 8 hours. Identify the relevant observed endpoint and explain why the requested value matters.
  3. A regression model for the number of visitors to a park was fitted using daily temperatures from 12°C to 29°C. Is predicting visitors at 5°C extrapolation? What additional caution should accompany the prediction?
  4. A student says, “The model predicts a response larger than every response in the data, so the prediction must be extrapolation.” Explain why that reasoning is not sufficient.
  5. A line fits its observed data closely, but a requested \(x\)-value is beyond the observed maximum. Explain why the close fit alone does not establish that the prediction is reliable.