Which Observation Is Predicted Best?
In “Interpreting a Residual in Context,” you learned that a residual is the observed response minus the predicted response for the same case. Its sign tells whether the fitted line overpredicted or underpredicted. To decide which individual’s observed response was closest to the model’s prediction, focus instead on the size of the residual.
A residual of \(2\) and a residual of \(-2\) point in opposite directions, but both put the observation \(2\) response units from its prediction. Taking the absolute value removes the sign while preserving that distance. Comparing absolute residuals lets us rank the predictions for the observations in a data set: the smallest absolute residual is the closest prediction, and the largest is the farthest.
This comparison concerns how close a model’s prediction was to each observed response, not which person or item is “better.” Saying that the model predicted one individual best is shorthand for saying that this individual’s observed response was closest to the value predicted by the model. The largest absolute residual identifies the least close prediction among the cases compared.
Keep the response units attached to the comparison. If the model predicts time in minutes, absolute residuals are in minutes; if it predicts a distance in centimeters, absolute residuals are in centimeters. Absolute residuals for the same response variable can be compared directly. Comparing raw values for different response variables measured in unrelated units would not tell you which prediction was closer in a meaningful common scale.
A Reliable Way to Compare Cases
Start with the signed residuals, whether they are provided or need to be calculated from \(y-\hat{y}\). Then take the absolute value of each one. Do not compare the signed numbers as if a negative residual were automatically larger or smaller in prediction error than a positive residual. For example, \(-6\) is less than \(2\) as a signed number, but its absolute residual, \(6\), is larger than \(2\).
Check that each observed response is paired with its own prediction at the same predictor value.
Calculate \(|y-\hat{y}|\), or take the absolute value of each already calculated signed residual.
The smallest absolute residual identifies the closest prediction; the largest identifies the farthest among the cases listed.
Name the case, give the error size with response units, and avoid claiming more than the comparison shows.
If two cases have equal absolute residuals, their predictions are equally close in response units, even if one residual is positive and the other is negative. If every case has a different absolute residual, you can rank all of them from closest to farthest. In either situation, preserve enough decimal places to make the comparison fair; rounding too early can make unequal values appear tied.
Worked Example: Comparing Flight-Time Predictions
A fictional travel analyst uses the fitted line \(\hat{y}=12+4x\) to predict a flight’s duration, \(y\), in hours, from a route measure, \(x\). The table gives four observed flights. Which flight was predicted best, and which was predicted least well, among these four?
| Flight | \(x\) | Observed time \(y\) (hours) | Prediction \(\hat{y}=12+4x\) (hours) | Residual \(y-\hat{y}\) (hours) | Absolute residual (hours) |
|---|---|---|---|---|---|
| A | 1 | 17.5 | 16 | 1.5 | 1.5 |
| B | 2 | 17 | 20 | −3 | 3 |
| C | 3 | 27 | 24 | 3 | 3 |
| D | 4 | 23 | 28 | −5 | 5 |
State. Identify the flight with the closest prediction and the flight with the farthest prediction, using the absolute sizes of the residuals.
Plan. For each flight, substitute its \(x\)-value into the fitted line to obtain the predicted time. Then calculate \(y-\hat{y}\), take the absolute value, and compare those magnitudes.
Do. For Flight A, \(\hat{y}=12+4(1)=16\) hours, so its residual is \(17.5-16=1.5\) hours and its absolute residual is \(|1.5|=1.5\) hours. For Flight B, \(\hat{y}=12+4(2)=20\), so its residual is \(17-20=-3\) hours and its absolute residual is \(3\) hours. For Flight C, \(\hat{y}=12+4(3)=24\), so its residual is \(27-24=3\) hours and its absolute residual is \(3\) hours. For Flight D, \(\hat{y}=12+4(4)=28\), so its residual is \(23-28=-5\) hours and its absolute residual is \(5\) hours. In order from smallest to largest, the absolute residuals are \(1.5, 3, 3,\) and \(5\) hours.
Conclude. Flight A’s observed duration was closest to its predicted duration, by \(1.5\) hours. Flight D’s was farthest, by \(5\) hours. Flights B and C were equally far from their predictions, by \(3\) hours each, even though B’s residual was negative and C’s was positive.
What the Sign Adds—and What It Does Not
Absolute values are useful for comparing closeness, but they deliberately hide the direction of the error. As explained in “Sign of a Residual: Over- and Underprediction,” a positive residual means the observed response was above the prediction, so the line underpredicted. A negative residual means the observed response was below the prediction, so the line overpredicted. Taking the absolute value does not change either observation; it just answers a different question: how far was the observation from its prediction?
For a clear explanation, give the absolute residual when discussing closeness and use the signed residual if direction matters. For example, “The model’s prediction for case A was \(1.5\) hours below the observed duration” communicates direction, while “The absolute residual was \(1.5\) hours” communicates size. A complete comparison might include both, but do not use the positive or negative sign alone to decide which error was greater.
Also distinguish “best among these cases” from “a good prediction in general.” The smallest absolute residual in a set is simply the smallest in that set; it could still be large enough to matter in context. Likewise, the largest residual does not, by itself, prove that the whole regression model is poor. It identifies the least close observed prediction among the cases you examined.
Worked Example: Comparing Restaurant Wait-Time Predictions
A fictional restaurant manager uses \(\hat{y}=4+2.5x\) to predict a party’s wait time, \(y\), in minutes, based on party size \(x\). The observed wait times for four parties are shown below. Compare the absolute residuals to decide which prediction was closest and which was farthest from the observed wait.
| Party | Size \(x\) | Observed wait \(y\) (minutes) | Prediction \(\hat{y}\) (minutes) | Residual (minutes) | Absolute residual (minutes) |
|---|---|---|---|---|---|
| 1 | 2 | 8 | 9 | −1 | 1 |
| 2 | 3 | 13.5 | 11.5 | 2 | 2 |
| 3 | 5 | 15 | 16.5 | −1.5 | 1.5 |
| 4 | 6 | 22 | 19 | 3 | 3 |
State. Find the party whose wait was predicted most closely and the party whose wait was predicted least closely.
Plan. Calculate the predicted wait from the line for each party, subtract the prediction from the observed wait, and compare the absolute residuals in minutes.
Do. For Party 1, the prediction is \(4+2.5(2)=9\) minutes, giving residual \(8-9=-1\) minute and absolute residual \(|-1|=1\) minute. For Party 2, the prediction is \(4+2.5(3)=11.5\) minutes, giving residual \(13.5-11.5=2\) minutes and absolute residual \(2\) minutes. For Party 3, the prediction is \(4+2.5(5)=16.5\) minutes, giving residual \(15-16.5=-1.5\) minutes and absolute residual \(1.5\) minutes. For Party 4, the prediction is \(4+2.5(6)=19\) minutes, giving residual \(22-19=3\) minutes and absolute residual \(3\) minutes. The absolute residuals rank from smallest to largest as \(1, 1.5, 2,\) and \(3\) minutes.
Conclude. The model predicted Party 1’s wait most closely: the observed wait was \(1\) minute from the prediction. It predicted Party 4’s wait least closely: the observed wait was \(3\) minutes from the prediction. The negative residual for Party 1 means the model overpredicted that wait, while the positive residual for Party 4 means it underpredicted that wait; neither sign changes the magnitude comparison.
Comparisons Are Specific to the Cases and the Response Scale
A ranking of absolute residuals answers a focused question about the observations included. If you compare four individuals, the “closest” and “farthest” labels apply to those four individuals. They do not automatically describe everyone represented by the model, and they do not say which predictor values are generally easier to predict. To make a broader claim about a region of the data, you would need to examine the overall residual pattern there, as in the earlier tutorial on fan-shaped residual plots.
The response scale matters, too. An absolute residual of \(4\) minutes is a difference of \(4\) minutes, regardless of whether the signed residual is \(4\) or \(-4\). That makes absolute residuals directly interpretable in context. But a difference of \(4\) minutes and a difference of \(4\) dollars are not the same kind of error. Do not rank them together as if their numerical values measured a common quantity.
Even within one response variable, absolute residuals focus on distance in the original units. A comparison does not adjust for whether prediction errors tend to vary more in one part of the predictor range than another. A large absolute residual tells you that this case was relatively far from its own prediction in response units; by itself, it does not say whether that distance is unusual compared with other cases in a region. Keep the claim at the level the calculation supports.
Worked Example: Ranking Commute-Time Predictions
A fictional transportation planner uses \(\hat{y}=8+2.5x\) to predict commute time \(y\), in minutes, from travel distance \(x\), in kilometers. Compare the four commuters by absolute residual. Is anyone tied for the closest prediction?
| Commuter | Distance \(x\) (kilometers) | Observed commute \(y\) (minutes) | Predicted commute \(\hat{y}\) (minutes) | Residual (minutes) | Absolute residual (minutes) |
|---|---|---|---|---|---|
| Al | 4 | 16 | 18 | −2 | 2 |
| Bea | 6 | 28 | 23 | 5 | 5 |
| Chen | 8 | 25 | 28 | −3 | 3 |
| Dev | 10 | 31 | 33 | −2 | 2 |
State. Compare all four absolute residuals and identify the closest and farthest predictions, allowing for a tie.
Plan. Use the fitted line to find each predicted commute time, calculate the signed residual as observed minus predicted, and take its absolute value. Then rank the distances from the predictions.
Do. Al’s predicted time is \(8+2.5(4)=18\) minutes, so the residual is \(16-18=-2\) minutes and the absolute residual is \(2\) minutes. Bea’s prediction is \(8+2.5(6)=23\) minutes, so her residual is \(28-23=5\) minutes and her absolute residual is \(5\) minutes. Chen’s prediction is \(8+2.5(8)=28\) minutes, so the residual is \(25-28=-3\) minutes and the absolute residual is \(3\) minutes. Dev’s prediction is \(8+2.5(10)=33\) minutes, so the residual is \(31-33=-2\) minutes and the absolute residual is \(2\) minutes. Thus the absolute residuals are \(2, 5, 3,\) and \(2\) minutes for Al, Bea, Chen, and Dev, respectively.
Conclude. Al and Dev tie for the closest predictions, each with an absolute residual of \(2\) minutes. Bea’s prediction is farthest from the observed commute, by \(5\) minutes. Al and Dev have the same absolute residual even though their predictor values differ; this comparison ranks the observed prediction errors and does not explain why they differ.
Common Mistakes and AP Exam Tips
- Comparing signed residuals instead of their magnitudes. A negative residual is not automatically a larger error than a positive one. Take absolute values before ranking closeness.
- Calling the most positive residual the worst prediction. The worst prediction among the cases is the one with the largest absolute residual, which could be positive or negative.
- Dropping units or context. Report an error in the response variable’s units, such as “\(5\) minutes,” and name the individual or case being discussed.
- Ignoring ties. If two cases have the same absolute residual, say they are equally close to their predictions. Do not force a unique best case.
- Claiming that the smallest error proves the model is good. It establishes only that this case is closest among those compared. Explain the comparison’s scope rather than making a broad claim about model quality.
- Rounding before comparing. Keep the residual calculations precise until the ranking is clear. If values are close, premature rounding can hide which one is smaller or create a false tie.
For full-credit communication, identify the case or cases, compare absolute residuals, and state the error size in context and response units. If useful, add the signed residual’s direction separately: say whether the line overpredicted or underpredicted, but do not confuse that direction with closeness.
Check Your Understanding
Use absolute residuals to compare prediction closeness, and use the response units when explaining your answers.
- Two cases have residuals of \(-4\) centimeters and \(3\) centimeters. Which prediction is closer to its observed response, and by how much does its absolute residual differ from the other one?
- A model’s residuals for three patients’ predicted recovery times are \(2\), \(-5\), and \(1\) days. Which patient’s prediction was least close among the three?
- Two deliveries have residuals of \(1.5\) minutes and \(-1.5\) minutes. What can you conclude about their prediction closeness, and what does the sign difference indicate?
- Why does having the smallest absolute residual among five observations not prove that a regression model gives good predictions for all cases?
- A fitted model predicts a response in liters. One observation has a residual of \(-2.4\) liters. State its absolute residual and explain what that value says about the prediction’s closeness.