Tutorials › AP Statistics › Identifying Extrapolation in a Word Problem

Extrapolation and prediction limits · Tutorial 946 of 1000

Identifying Extrapolation in a Word Problem

Practice translating prediction requests into explanatory-variable values and checking whether they fall inside or outside the range used to fit the regression model.

Intermediate 9 min read

What You'll Learn

  • Identify the explanatory variable and the observations used to fit a regression model
  • Translate a word-problem request into a numerical explanatory-variable value with units
  • Compare the requested value with the minimum and maximum observed values
  • Distinguish extrapolation from a prediction that is merely about the future
  • Explain why an out-of-range prediction deserves caution in context
  • Write a complete classification that names the values and range being compared

Turn the Word Problem into a Range Check

In “Interpolation Versus Extrapolation,” we classified a prediction by comparing its explanatory-variable value with the observed range used to fit the model. Word problems can make that comparison less obvious: the requested value may be described in a sentence, expressed in different units, or framed as a future prediction. The first task is to translate the request into the value of \(x\) that the regression model would use.

The date of a prediction does not by itself determine whether it is extrapolation. A future case can have an \(x\)-value inside the observed range, while a past or present case can have an \(x\)-value outside it. What matters is the requested explanatory-variable value and the actual range of explanatory-variable values in the cases used to fit the model.

Definition: To identify extrapolation in a word problem, identify the explanatory variable, find the minimum and maximum values of that variable among the observations used to fit the model, translate the request into an \(x\)-value with matching units, and compare. A requested value below the minimum or above the maximum is extrapolation. A value equal to an endpoint or between the endpoints is within the observed range.

This is a classification, not a calculation of the predicted response. The regression equation may still produce a number for an out-of-range \(x\), but the fact that the arithmetic works does not show that the model is dependable there. As discussed in “Why Extrapolation Is Risky,” the relationship may not continue in the same way beyond the observed data.

A Reliable Word-Problem Checklist

Use the same sequence whether a problem gives a table, a graph, a written description, or a request in a particular unit. Make sure you use the range for the data that fitted this specific model—not a wider range that seems possible in the real world.

1
Name the explanatory variable.
Identify what \(x\) represents and its units, such as training hours, canopy percentage, or calendar year.
2
Locate the model’s observed range.
Use the minimum and maximum \(x\)-values among the cases that were actually used to fit the regression line.
3
Translate the request.
Convert phrases such as “three more hours,” “halfway through the year,” or “a neighborhood with 50 percent canopy” into a numerical \(x\)-value in the same units.
4
Compare and explain.
State whether the requested value is below, within, equal to, or above the observed range. Name the values and say what that means in context.

A range includes every numerical value from its minimum through its maximum for this classification. A requested value does not have to match an \(x\)-value of a particular observed case to be within the range. For example, if observations span from 2 to 18 hours, a request for 11 hours is between the endpoints even if no case had exactly 11 hours.

Be precise about what counts as “the data.” If a model was fitted using only a selected group of observations, use that group’s range—not the range of a larger file, a different group, or all values that could theoretically occur. In “Finding the Range of the Explanatory Variable,” we focused on the observed minimum and maximum for the cases used in the model; that is the range to check here, too.

Worked Examples: Read the Request in Context

Worked Example: A Request for More Training Hours

Hypothetical setting. A career program fits a regression model to predict a certification exam score from the number of hours participants spent in a practice workshop. The 24 participants used to fit the line reported from 2 through 18 workshop hours. An administrator asks the model to predict a score for a participant who completed 22 workshop hours. Is this interpolation or extrapolation?

Identify and compare. The explanatory variable is workshop hours, measured in hours. Its observed range for the 24 participants is 2 to 18 hours. The request is \(x=22\) hours, and

$$ 22>18 $$

The requested value is 4 hours above the largest observed value: \(22-18=4\) hours. The subtraction checks the size of the gap; the comparison itself shows that 22 is outside the fitted model’s range.

Conclusion in context. Predicting an exam score for 22 workshop hours is extrapolation because the participants used to fit the model had between 2 and 18 workshop hours. The line can produce a score at 22 hours, but the data do not show whether the relationship continues beyond 18 hours. The model’s association does not, by itself, establish that adding workshop hours causes a score change.

Contrast. If the request instead concerned a participant with 12 workshop hours, then \(2\leq12\leq18\). That prediction is within the observed range, so it is interpolation by the range-based definition, even if no participant had exactly 12 hours.

Worked Example: A Canopy Percentage Hidden in a Description

Hypothetical setting. A city planning class fits a regression line relating a neighborhood’s tree-canopy coverage to its average summer surface temperature. The model uses 16 neighborhoods with canopy coverage from 8% to 42%. A planner asks for a prediction for a neighborhood where 0.50 of the land area is covered by tree canopy. Is that request in range?

Translate the wording. The explanatory variable is canopy coverage. The observed values are reported as percentages, but the request uses a proportion. Convert the requested proportion to a percentage:

$$ 0.50(100\%)=50\% $$

This conversion checks units: the request is for 50% canopy, not 0.50%. Compare it with the observed range, 8% to 42%:

$$ 50\%>42\% $$

The request is \(50-42=8\) percentage points above the largest observed canopy value. Percentage points are the appropriate units for a difference between two percentages.

Conclusion in context. Predicting the average summer surface temperature for a neighborhood with 50% tree-canopy coverage is extrapolation. The neighborhoods used to fit the line ranged only from 8% to 42% coverage. Even though a 50% value may be possible in the city, it is outside this model’s observed range. A prediction there warrants caution; the model’s fitted relationship may not continue unchanged. This classification describes the model’s data range, not whether increasing canopy would cause a temperature change.

Worked Example: A Future Year That Is Not Extrapolation

Hypothetical setting. A school district fits a regression model relating calendar year to annual electricity use per student. The observations used in the model cover every year from 2010 through 2020, inclusive. In a report written after those data were collected, a staff member asks for a prediction for 2018 and describes it as a “future planning estimate.” Is the request extrapolation?

Check the explanatory-variable value. Here \(x\) is calendar year, measured in years. The observed range is 2010 to 2020. The requested value, 2018, satisfies

$$ 2010\leq2018\leq2020 $$

The request is within the observed range. It is therefore interpolation by the range-based definition, not extrapolation. The staff member’s phrase “future planning estimate” does not change the numerical comparison. In this scenario the report is written after 2020, so its use of observations through 2020 is temporally consistent; 2018 is an earlier year within the observed period.

Contrast with another request. If the staff member asked for 2024 instead, then \(2024>2020\), so that request would be extrapolation. A request for 2008 would also be extrapolation because \(2008<2010\). Thus, being in the future is neither required nor sufficient for extrapolation: a future value can be inside the observed range, and a value from the past can be outside it.

Conclusion in context. The 2018 prediction is within the years used to fit the model. The 2024 prediction is extrapolation because the fitted data end in 2020, and the 2008 prediction is extrapolation because the data begin in 2010. These are classifications based on the \(x\)-range; they do not tell us, by themselves, how accurate any of the predictions will be.

What the Range Check Can—and Cannot—Tell You

The range check answers a specific question: is the requested explanatory-variable value outside the minimum-to-maximum span of the data used to fit the line? It does not assess every aspect of a prediction. For example, it does not tell you whether the relationship is linear, whether an unusual point strongly affected the line, or how much the predicted response might vary around the line.

A value inside the range is not a guarantee of a dependable prediction. It is classified as interpolation, but the model could still fit poorly or fail to represent the pattern in a particular part of the range. Likewise, calling an out-of-range request extrapolation does not prove the prediction is wrong; it signals that the prediction extends beyond the explanatory-variable values supporting the fitted line and needs special caution.

Also keep the model’s cases straight. If a model was fitted to neighborhood averages, its cases are neighborhoods and its response is an average, not an individual resident’s temperature. As covered in “Group Averages Versus Individual Predictions,” the kind of case in the model matters when describing what its prediction represents. For the range check, use the \(x\)-values for those model cases.

Key takeaway: In a word problem, first translate the requested prediction into the explanatory variable’s units. Then compare that value with the minimum and maximum \(x\)-values used to fit the model. The context may explain why the prediction is interesting, but the numerical range determines whether it is extrapolation.

Common Mistakes and AP Exam Tips

  • Classifying by whether the request is in the future. A future date is not automatically extrapolation. Compare its year with the years used to fit the model, as in the 2018 example.
  • Using the possible range instead of the observed range. A neighborhood could have 50% canopy, but that does not place 50% inside a model whose observed maximum is 42%.
  • Comparing numbers with different units or formats. Convert 0.50 to 50% before comparing it with percentage data. For time or distance, put the request and endpoints in the same units.
  • Forgetting which variable is \(x\). Extrapolation is determined by the explanatory variable’s values, not by whether the predicted response seems unusually large or small.
  • Calling a within-range request “extrapolation” because no case had exactly that value. The classification uses the span from the minimum to the maximum, not just a list of repeated values.
  • Claiming an out-of-range prediction must be wrong. A full-credit answer says it is extrapolation and explains that the relationship may not continue beyond the observed range. It does not claim the result is certainly false.

A strong AP response makes the comparison visible: “The model was fitted using values of [explanatory variable] from [minimum] to [maximum]. The requested value is [requested value], which is [inside/below/above] that range, so the prediction is [interpolation/extrapolation].” Then add a contextual sentence explaining what the classification means for the particular cases and why an extrapolated prediction deserves caution.

Check Your Understanding

For each request, identify the explanatory-variable value, compare it with the model’s observed range, and state whether the prediction is interpolation or extrapolation.

  1. A model relating daily screen time to sleep duration uses people with 1 to 7 hours of screen time. Is a prediction for 5.5 hours within the observed range? Explain why an exact matching observation is not required.
  2. A model uses river-flow measurements from 12 to 80 cubic meters per second. A request describes a flow of 0.09 thousand cubic meters per second. Convert the request to cubic meters per second and classify it.
  3. A model uses calendar years from 2005 through 2015. Classify requests for 2003 and 2015. Explain the role of the endpoints.
  4. In one or two sentences, explain why a future prediction is not automatically extrapolation.
  5. A model was fitted using stores with weekly sales from 40 to 250 items. A new store sells 300 items per week. Write a contextual sentence that classifies the prediction and gives an appropriate caution.