Tutorials › AP Statistics › Common Errors About Extrapolation

Extrapolation and prediction limits · Tutorial 958 of 1000

Common Errors About Extrapolation

Use the explanatory-variable range—not the size of a prediction or the strength of \(r\)—to identify extrapolation and judge what a fitted line can support.

Intermediate 9 min read

What You'll Learn

  • Classify a prediction by comparing its explanatory-variable value with the observed range.
  • Explain why a large predicted response can still be interpolation.
  • Recognize that a strong correlation does not establish that an out-of-range prediction is reliable.
  • Distinguish “extrapolation” from “definitely wrong.”
  • Write a concise critique that separates classification from concerns about the prediction’s support.

First Identify What the Prediction Is About

Several errors about extrapolation come from looking at the wrong number. A predicted response may be large, surprising, or far from most observed responses. None of those facts, by itself, tells you whether the prediction is extrapolation. The key comparison is between the requested value of the explanatory variable and the explanatory-variable values used to fit the line.

As established in “What Extrapolation Means” and “Finding the Range of the Explanatory Variable,” a prediction is an extrapolation when its \(x\)-value is outside the observed \(x\)-range. If the requested \(x\)-value is within that range, including at either endpoint, the prediction is interpolation. The size of \(\hat{y}\) and the strength of the linear association do not change that classification.

Definition: To check for extrapolation, identify the explanatory variable, find its minimum and maximum observed values for the cases used to fit the model, and compare the requested \(x\)-value with those endpoints. Do not classify the prediction by looking at the predicted response \(\hat{y}\).

Keep two questions separate. First, is the requested \(x\)-value outside the observed range? That determines whether the prediction is extrapolation. Second, how much confidence should we place in the prediction? That calls for considering the model, the context, the distance beyond the observed range, and any practical limits. Confusing these questions leads to claims that sound decisive but are not supported by the information.

Common Error: “The Predicted Response Is Large, So It Must Be Extrapolation”

A large response can be predicted at an \(x\)-value that lies comfortably inside the observed range. That is interpolation, even if the predicted value is bigger than many of the observed responses. Conversely, a prediction can be small and still be extrapolation if its \(x\)-value falls outside the observed range.

This is why it helps to label the variables and their units before deciding. The \(x\)-range belongs to the explanatory variable, not the response variable. Do not compare a requested number of hours with observed dollars, or a requested temperature with observed energy use. The numbers must refer to the same variable and use the same units.

Worked Example: A High Cost Is Not Automatically Extrapolation

Hypothetical setting. A delivery service fits a regression line relating delivery distance \(x\), in kilometers, to delivery cost \(y\), in dollars. The observed distances range from 2 to 10 kilometers. The fitted line is \(\hat{y}=18+27x\). A customer asks for a prediction at 9 kilometers.

Classify using \(x\). The requested distance, 9 kilometers, is between the observed minimum of 2 and maximum of 10 kilometers. Therefore, this is interpolation. The fact that the predicted cost might seem high does not make it extrapolation.

Calculate the prediction.

$$ \hat{y}=18+27(9)=18+243=261\text{ dollars}. $$

The line predicts a delivery cost of 261 dollars at 9 kilometers. That amount might prompt a separate question about whether the model or the situation makes sense, but it does not change the classification: 9 kilometers is within the observed \(x\)-range.

Now suppose someone asks for the prediction at 12 kilometers. The line gives

$$ \hat{y}=18+27(12)=18+324=342\text{ dollars}. $$

This second prediction is an extrapolation because 12 kilometers is greater than the observed maximum of 10 kilometers. The reason is the \(x\)-value, not the predicted cost of 342 dollars. A correct critique would name the distance beyond the observed range and explain why the cost relationship might not continue unchanged there.

Common Error: “A Strong \(r\) Makes Extrapolation Safe”

A strong correlation can describe a clear linear pattern among the observed cases. It cannot supply observations at \(x\)-values that were not represented in those data. As covered in “Writing a Complete \(r\)-squared Interpretation,” a model summary must be interpreted in relation to the response variation and cases being described. It is not a guarantee that the line will keep working outside the observed range.

A large \(|r|\) does not tell us that the relationship must remain linear at more extreme values. The pattern could bend, level off, change rate, or be affected by conditions not represented in the data. As discussed in “How Unmodeled Curvature Shows Up Outside the Data,” a line can summarize a restricted part of a curved relationship quite well while misrepresenting what happens farther along it.

So do not use “\(r\) is close to 1” as a reason to accept an extrapolation. It is relevant information about the observed association, but it does not remove the limitation created by the out-of-range \(x\)-value. Nor does extrapolation alone prove that the prediction is wrong. Describe what the data support and what they leave uncertain.

Worked Example: A Strong Association Still Has a Range Limit

Hypothetical setting. A community energy project records daily sunlight duration \(x\), in hours, and solar energy output \(y\), in kilowatt-hours, for a set of similar test days. The observed sunlight durations range from 2 to 9 hours. A fitted line is \(\hat{y}=1.8+0.72x\), and the reported correlation is \(r=0.98\). A planner wants a prediction for a day with 12 hours of sunlight.

Classify the request. The requested value, 12 hours, exceeds the observed maximum of 9 hours. It is therefore an extrapolation, 3 hours beyond the upper endpoint. The high reported correlation does not alter this comparison.

Calculate what the line gives.

$$ \hat{y}=1.8+0.72(12)=1.8+8.64=10.44\text{ kilowatt-hours}. $$

The line calculates a predicted output of 10.44 kilowatt-hours. That is the equation’s result, not proof that this is a dependable prediction for 12 hours of sunlight.

Critique in context. The prediction at 12 hours is an extrapolation because the observed sunlight durations ranged only from 2 to 9 hours. Although \(r=0.98\) indicates a strong linear association among the observed cases, it does not establish that the same pattern continues beyond 9 hours. Solar output could level off if the equipment reaches a practical production limit, so the 10.44-kilowatt-hour prediction may not be reliable.

Notice what this critique does and does not claim. It does not say the prediction must be wrong, and it does not say a production limit has been demonstrated by the information given. It identifies a plausible reason for caution and explains why the strong observed association cannot settle what happens outside the range.

Common Error: “Any Extrapolation Is Definitely Wrong”

Extrapolation means the requested \(x\)-value is outside the observed range. It does not mean that the fitted line’s calculation is false, or that the real response is known to differ from it. The data simply do not directly show whether the same pattern holds at that \(x\)-value. A careful response says the prediction is less supported or may be unreliable, then gives a reason tied to the situation.

Also avoid treating every extrapolation as equally concerning. A request just beyond an endpoint is still extrapolation, but its distance beyond the data is different from a request far outside the range. “Short-Range Versus Long-Range Extrapolation” explains how that distance helps describe the size of the extension. Distance can sharpen a critique; it does not create a universal cutoff where a prediction switches from safe to unsafe.

Worked Example: A Small Prediction Can Still Be Extrapolation

Hypothetical setting. A repair shop records equipment age \(x\), in years, and remaining battery capacity \(y\), as a percentage, for devices aged from 1 to 5 years. Its fitted line is \(\hat{y}=86-3.2x\). A customer asks for a prediction for a 6-year-old device.

Classify using the observed boundary. Six years is one year above the observed maximum of 5 years. The requested prediction is an extrapolation, even though it is only just beyond the observed range.

Calculate and interpret the response carefully.

$$ \hat{y}=86-3.2(6)=86-19.2=66.8\%. $$

The line calculates 66.8% remaining capacity. That percentage is not obviously impossible, but its plausible size does not make the prediction interpolation. The device age, not the response percentage, determines the classification.

Critique in context. The prediction for a 6-year-old device is an extrapolation because the devices in the data were no older than 5 years. The model calculates 66.8% remaining capacity, but the data do not establish that battery capacity continues to decline at the same rate beyond 5 years. Differences in device use or battery replacement could also affect the capacity of an older device.

This example separates two ideas that are easy to mix up: the calculated response is within the percentage scale, but the requested age is still outside the observed \(x\)-range. As emphasized in “Checking Whether a Prediction Is Reasonable,” a plausible-looking response does not, by itself, make the model’s use appropriate.

A Quick Audit for Extrapolation Claims

Before writing a critique, pause to check which claim the information actually supports. The following sequence catches the most common mix-ups without repeating the full critique structure from “Writing an Extrapolation Critique.”

1
Name the explanatory variable.
Check which variable is \(x\), and note its units. Do not use the response value or its units to classify the prediction.
2
Find the observed \(x\)-range.
Use the minimum and maximum values among the observations used to fit this line, not an unrelated range from another group or model.
3
Compare the requested \(x\).
Outside the range means extrapolation; inside the range or at an endpoint means interpolation. The size of \(\hat{y}\) does not affect this decision.
4
Assess support separately.
Consider how far the request is beyond the range and whether the context suggests a changing pattern or a practical constraint. Neither distance nor a strong \(r\) supplies a guaranteed verdict by itself.

One subtle mistake is comparing a request with a general range that seems reasonable in the real world instead of the range for the cases used to fit the model. A value can be common in the broader population and still be outside the particular study’s observed \(x\)-range. The classification concerns the data for this fitted line.

Another mistake is treating a strong \(r\) as if it measured prediction accuracy at every possible \(x\). It describes linear association in the observed data, not the behavior of the relationship at unobserved values. For that reason, a critique should not simply say “the correlation is high, so the prediction is reliable,” or “the correlation is high, so extrapolation does not apply.” Both statements confuse evidence about the observed cases with evidence beyond them.

Common Mistakes and AP Exam Tips

  • Using the size of \(\hat{y}\) to classify a prediction. Compare the requested \(x\)-value with the observed \(x\)-range. A high predicted response may be interpolation, and a low response may be extrapolation.
  • Comparing numbers with different units. Check that the requested value and the range refer to the same explanatory variable and units.
  • Assuming a strong \(r\) settles the question. State that a strong association describes the observed cases; it does not demonstrate that the linear pattern continues outside their range.
  • Calling extrapolation proof of error. Say the prediction is not directly supported by observations at that \(x\)-value or may be unreliable. Do not say it is definitely wrong without further evidence.
  • Ignoring the model’s data range because a value seems realistic. The relevant range is the one used to fit the model, not an assumed range for the whole population or setting.
  • Writing only “it is extrapolation.” A stronger answer identifies the requested \(x\), the observed endpoint, and a relevant contextual reason for caution, as explained in “Writing an Extrapolation Critique.”

For a clear AP response, make the classification from the \(x\)-range first. Then discuss support separately. For example: “The prediction at 12 hours is an extrapolation because the observed sunlight durations ranged from 2 to 9 hours. The strong correlation describes the observed days but does not establish that output continues to rise linearly beyond 9 hours; output could level off.” This wording identifies the relevant comparison, handles \(r\) appropriately, and gives a qualified contextual concern.

Key takeaway: Extrapolation is determined by the requested explanatory-variable value, not by how large the predicted response is. A strong \(r\) does not remove the range limitation, and an out-of-range prediction is not automatically false. Classify first, then explain what the data do and do not support.

Check Your Understanding

For each situation, identify the common error to avoid and state what comparison or qualification is needed.

  1. A model predicts a very high monthly bill at an \(x\)-value inside the observed range. Is that prediction extrapolation? Explain which values determine the classification.
  2. A model has \(r=0.95\), and a prediction is requested beyond the largest observed \(x\)-value. What does the strong correlation tell you, and what does it not establish?
  3. The observed \(x\)-values range from 4 to 16. A prediction is requested at \(x=16\). Is it extrapolation? Explain how the endpoint is treated.
  4. A line predicts a plausible response at an \(x\)-value well below the observed minimum. Why is the response’s plausibility not enough to call the prediction interpolation?
  5. Rewrite this claim more carefully: “The prediction is definitely wrong because the requested value is outside the data range.”