Tutorials › AP Statistics › Checking Whether a Prediction Is Reasonable

Extrapolation and prediction limits · Tutorial 950 of 1000

Checking Whether a Prediction Is Reasonable

Use the response’s units and context to decide whether a regression prediction is a meaningful value or a warning about the model’s limits.

Intermediate 9 min read

What You'll Learn

  • Check a predicted value against the response’s possible range.
  • Distinguish hard limits from values that are merely unlikely in context.
  • Evaluate negative counts and percentages outside 0 to 100.
  • Explain why a mathematically correct prediction may not be usable.
  • Communicate a prediction and its limitations without silently changing the result.

A Correct Calculation Can Still Be an Unreasonable Prediction

In “Extrapolating Over Time,” we used a fitted line to calculate a future prediction and considered whether the trend might continue. A further check is whether the predicted response value makes sense at all. A line can produce a number that is mathematically correct but impossible in the situation, such as a negative count or a percentage greater than 100.

This check is about the response, not just the explanatory-variable value. As discussed in “What Extrapolation Means,” a prediction outside the observed range of \(x\) is an extrapolation and deserves caution. But a prediction can be implausible even when \(x\) is within the observed range. Conversely, an extrapolated prediction might still be possible, though its reliability needs careful consideration.

Definition: A plausible prediction is a model prediction that falls within the possible or reasonable values for the response in the stated context. A prediction can be mathematically valid but implausible if it violates a hard limit or conflicts with important contextual information.

Start by identifying exactly what \(\hat{y}\) measures and its units. A negative prediction means different things for different responses: negative temperature can be possible on some scales, while a negative number of service calls cannot. A percentage of a whole must be between 0% and 100%, inclusive. A count cannot be negative, although a model prediction for a count may be fractional because the fitted line is not restricted to whole numbers.

Some limits are absolute. A duration measured in elapsed time cannot be negative, and a percentage of a whole cannot exceed 100%. Other limits depend on the particular setting. A prediction of 38 shuttle trips per day is a possible count in general, but it would not be feasible for a service whose stated maximum capacity is 32 trips per day. A value can also be possible yet surprising or unlikely; deciding that requires relevant context, not just the equation.

A Response-Scale Plausibility Check

Before using a prediction to describe what might happen, make a short response-scale audit. This means checking the output against the meaning, units, and limits of the response. The audit does not replace evaluating the quality of the model. It flags predictions that should not be taken literally and helps explain why.

1
Name the predicted response.
State what \(\hat{y}\) measures and include its units. Do not judge whether a number is reasonable until you know what it represents.
2
Identify possible-value limits.
Check for mathematical bounds such as 0% to 100%, nonnegative counts, or a stated physical or operational maximum.
3
Compare the prediction with those limits.
Calculate the fitted value accurately, then say whether it falls within the possible range and whether it seems reasonable in this particular setting.
4
Explain the consequence.
If the prediction violates a limit, say that the model’s output is not a plausible response in context. Do not present it as a real-world outcome or quietly replace it with a boundary value.

This check is separate from judging whether the fitted line is a good summary of the data. As explained in “What \(r^2\) Does Not Tell You,” a summary of fit does not guarantee that every prediction is sensible. A high \(r^2\) cannot make a negative count possible, nor does it remove a context-specific limit.

Worked Examples: Checking the Predicted Response

Worked Example: A Negative Prediction for Service Calls

Hypothetical setting. A building manager models the monthly number of repair calls received by a maintenance team, \(y\), using the number of years since a new reporting system was introduced, \(x\). A fitted line is \(\hat{y}=42-3.5x\), where \(\hat{y}\) is measured in calls per month. The observed \(x\)-values ran from 0 through 8. The manager asks for the model’s prediction at \(x=15\).

State. We need to calculate and assess the predicted monthly number of repair calls at \(x=15\).

Plan. Substitute \(x=15\) in the line, keeping the response units in mind. A number of calls cannot be negative. Also, \(x=15\) is beyond the observed range of 0 to 8, so this is an extrapolation, as defined in “What Extrapolation Means.” The response-limit check and the extrapolation check are related cautions, but they are not the same check.

Do. Substitute the requested value:

$$ \hat{y}=42-3.5(15)=42-52.5=-10.5\text{ calls per month} $$

The arithmetic gives a prediction of negative 10.5 calls per month. Because \(x=15\) is outside the observed range, we should also be cautious about extending the fitted trend that far. More importantly for plausibility, even a prediction within the observed \(x\)-range could not make negative monthly calls a possible outcome.

Conclude in context. The fitted line calculates \(-10.5\) repair calls per month at \(x=15\), but a negative number of calls is impossible. Therefore, this output is not a plausible prediction of the team’s actual monthly call count. It warns that the line’s declining pattern cannot be used literally at this value; it does not mean the team will receive a negative number of calls. The calculation alone does not tell us what a more suitable future prediction would be.

Worked Example: A Percentage Above 100%

Hypothetical setting. A greenhouse manager uses a fitted line to predict the percentage of seedlings that sprout, \(y\), from the number of days after planting, \(x\). The model is \(\hat{y}=58+4x\), with \(\hat{y}\) measured in percentage points. The manager asks for the prediction at \(x=12\).

Calculate the prediction. Substituting \(x=12\) gives:

$$ \hat{y}=58+4(12)=58+48=106\% $$

The fitted line’s predicted value is 106%. This is not a plausible percentage of seedlings that sprout: the sprouted seedlings are part of the total group, so the proportion cannot exceed the whole. A percentage of this kind must be from 0% through 100%, inclusive. The line’s increase of 4 percentage points per day continues mathematically, but the real percentage cannot keep increasing past 100%.

Interpret the model carefully. We can report that the fitted line calculates 106% at \(x=12\), but we should not say that 106% of the seedlings will sprout. The model output violates the response’s upper bound. This is a warning about using that line for this prediction, not evidence that more than the entire group will sprout. Replacing 106% with 100% would hide the warning and would not be the prediction made by the fitted line.

Check a nearby input. At \(x=10\), the line gives:

$$ \hat{y}=58+4(10)=98\% $$

A value of 98% is within the percentage’s possible range. That alone does not guarantee that it is a dependable prediction: we would still consider whether the model is suitable for that input and whether the context supports the prediction. The comparison shows why checking the response’s bounds is a necessary plausibility check, not a complete evaluation of a regression model.

Worked Example: A Positive Count That Exceeds Capacity

Hypothetical setting. A community shuttle service has four buses. Each bus can complete at most eight scheduled trips per day, so the service’s stated maximum is \(4(8)=32\) trips per day. A model predicts the number of daily trips, \(\hat{y}\), from the number of months since a route change, \(x\). Its fitted line is \(\hat{y}=20+1.5x\). The prediction is requested for \(x=14\).

Calculate the prediction. Substitute \(x=14\):

$$ \hat{y}=20+1.5(14)=20+21=41\text{ trips per day} $$

Forty-one is a positive whole-number count, so it passes the basic nonnegative-count check. But it exceeds the service’s stated maximum of 32 trips per day by \(41-32=9\) trips. The number is therefore not feasible under the operating conditions given in this example.

Conclusion in context. The line calculates 41 daily trips at \(x=14\), but the shuttle service can complete at most 32 trips per day with its current buses and schedule. So 41 is not a reasonable prediction for the service as described. The model’s output may indicate that its upward trend should not be extended this far, or that the operating conditions would need to change. We cannot infer from this calculation alone which explanation applies, and we should not report 41 as a feasible outcome under the stated capacity.

Common Mistakes and AP Exam Tips

  • Checking only whether the arithmetic is correct. Substitution can be flawless and the result still be implausible. A complete response identifies what the predicted response measures and checks its possible range.
  • Treating every negative value as impossible. Whether a negative value is possible depends on the variable and its scale. A negative number of calls is impossible; a negative temperature can be possible on a scale that allows it. Name the response and its context before making the judgment.
  • Forgetting the limits of a percentage. A predicted percentage of a whole below 0% or above 100% is not a possible percentage of that whole. State the relevant bound instead of merely calling the value “too high.”
  • Assuming every count prediction must be a whole number. The fitted line can produce a fractional prediction for a count. That decimal does not by itself show the model is unusable; a prediction can represent an estimated average. A negative count, however, violates the count’s lower limit.
  • Confusing “possible” with “reasonable.” A value may fall within a response’s mathematical bounds but still conflict with a specific capacity or other contextual fact. Explain the stated constraint that makes the prediction questionable.
  • Silently changing the model’s prediction. Do not replace a negative prediction with zero or a prediction above 100% with 100% and present the adjusted number as what the regression predicts. Report the fitted value, explain the violation, and say why it should not be taken literally.
  • Blaming extrapolation for every problem. An out-of-range \(x\)-value is an extrapolation and a reason for caution, but the response-scale check asks a different question: could the predicted \(y\)-value occur in this context? Keep those explanations distinct.

For full-credit communication, include the numerical prediction with units, identify the relevant limit, and connect the comparison to the situation. For example: “The line predicts \(-10.5\) calls per month, but a monthly call count cannot be negative, so this is not a plausible prediction in context.” This is more informative than saying only that the answer “does not make sense.”

Key takeaway: A regression line can calculate a response value that violates a mathematical bound or a stated practical constraint. Check the meaning and units of \(\hat{y}\), compare it with the response’s possible range, and explain why an implausible output should not be treated as a real-world outcome.

Check Your Understanding

For each situation, calculate or assess the prediction and explain the relevant response-scale limit.

  1. A model predicts the number of customer complaints in a day with \(\hat{y}=15-2x\). What does it predict at \(x=9\), and is that a plausible complaint count?
  2. A fitted line predicts that 103% of a group completes a training program. Explain why the predicted value is not possible as a percentage of that group.
  3. A model predicts 2.4 calls per hour for a service line. Does the decimal alone make the prediction impossible? Explain.
  4. A venue’s stated maximum capacity is 250 people, but a model predicts 268 people attending. Name the contextual issue and state how to describe the model output.
  5. Explain why a prediction at an \(x\)-value within the observed range can still fail a plausibility check.