Tutorials › AP Statistics › Residuals and Outliers

Residuals · Tutorial 896 of 1000

Residuals and Outliers

Learn how to identify possible outliers by their vertical distance from a regression line—and why an extreme predictor value alone is not enough.

Intermediate 9 min read

What You'll Learn

  • Identify a possible outlier by examining the size of its residual.
  • Use the absolute residual to compare how far observations are from the regression line.
  • Compare a residual with the residual standard deviation \(s\) as a scale check.
  • Distinguish a point that is extreme in the predictor from one with an unusually large residual.
  • Describe a possible outlier cautiously and investigate it before deciding how to handle it.

Two Different Ways a Point Can Be Unusual

A point in a scatterplot can stand out because its predictor value is far from the other \(x\)-values, because its response is far from the linear pattern, or both. These are different features. In regression, the vertical distance between an observed response and the fitted line is measured by the residual, \(y-\hat{y}\). A point with an unusually large residual is a possible outlier relative to the linear pattern. A point with an unusual \(x\)-value is extreme in the predictor; that fact alone does not make its residual large.

In “Constructing a Residual Plot” and “Reading a Residual Plot for Random Scatter,” you learned to plot residuals against \(x\) or the fitted values and look for patterns. Here, the focus is on individual observations that sit unusually far above or below the zero line. The sign tells you the direction of the miss, while the absolute value, \(|y-\hat{y}|\), gives the size of the vertical gap.

Definition: A possible outlier in a regression setting is an observation with a residual that is unusually large in magnitude compared with the residuals for the other observations. An extreme predictor value is an \(x\)-value far from the other observed \(x\)-values; it is not, by itself, evidence of a large residual.

Residuals are measured in the response variable’s units. If a model predicts tree height in meters, a residual of \(-1.2\) means the observed tree is 1.2 meters shorter than predicted. Its absolute residual is 1.2 meters. Comparing that gap with other residuals—and with the typical residual size \(s\)—helps assess whether it is unusually large.

There is no universal cutoff that automatically turns a residual into an outlier. A residual that is large relative to the others may deserve attention, but the judgment depends on the data and context. The residual standard deviation \(s\), described in “Standard Deviation of the Residuals, \(s\),” provides a useful scale for comparison. A residual several times as large as \(s\) is a reason to look more closely, not a formal rule that proves the point is an outlier.

How to Look for a Possible Outlier

Start with the residual plot or a table of residuals. Find observations with large absolute residuals, then compare their sizes with the general spread of residuals and with \(s\), if it is available. In the residual plot, a possible outlier appears far vertically from the zero line. Its horizontal position tells you where its predictor value lies; it does not determine how far the point is from the regression line.

1
Locate the observation.
Match the plotted residual or table entry to its original case and predictor value.
2
Measure the residual’s size.
Use the absolute residual, \(|y-\hat{y}|\), to describe the vertical distance from the line. Keep the signed residual as well if you need to say whether the line overpredicted or underpredicted.
3
Compare with the overall scale.
Compare the residual with other residuals and, when reported, with \(s\). Ask whether it stands out rather than applying a rigid cutoff.
4
Separate vertical and horizontal unusualness.
Check whether the \(x\)-value is extreme, but base the residual-outlier judgment on the vertical gap from the fitted line.

A residual plot is especially useful because it places all residuals on the same vertical scale and includes the zero line. A point far above zero has a large positive residual; a point far below zero has a large negative residual. Either can be a possible outlier because the size of the residual, not its sign, indicates how far the point is from the prediction.

An observation can be extreme in \(x\) and still lie close to the fitted line. Conversely, an observation near the middle of the \(x\)-range can have a very large residual. A point at an extreme \(x\)-value may also affect the fitted line, so it can be important to examine, but being important to investigate is not the same as having an unusually large residual.

Worked Example: Tree Height and Trunk Diameter

A fictional forestry class uses a least-squares line to predict tree height, in meters, from trunk diameter, in centimeters. The fitted model is \(\hat{y}=18+2.5x\), and the residual standard deviation for all observations is \(s=3.2\) meters. The table gives five selected observations from the data.

Trunk diameter \(x\) (cm)Observed height \(y\) (m)Predicted height \(\hat{y}\) (m)
32725.5
528.530.5
744.535.5
93940.5
186463

State. Decide which selected observation has a residual that stands out, and whether the observation with the most extreme trunk diameter is a residual outlier.

Plan. For each observation, calculate \(y-\hat{y}\), keeping the response in meters. Compare the absolute residuals with one another and with \(s=3.2\) meters. The most extreme \(x\)-value is 18 centimeters, but its residual must still be calculated before judging its vertical distance from the line.

Do. The residuals are \(27-25.5=1.5\) meters, \(28.5-30.5=-2\) meters, \(44.5-35.5=9\) meters, \(39-40.5=-1.5\) meters, and \(64-63=1\) meter. The corresponding absolute residuals are 1.5, 2, 9, 1.5, and 1 meter. The largest is 9 meters. Relative to \(s\), that residual is \(9/3.2=2.8125\), or about 2.81 times the residual standard deviation. For the observation at \(x=18\), the absolute residual is only 1 meter, which is \(1/3.2=0.3125\), or about 0.31 times \(s\).

Conclude. The tree with a 7-centimeter trunk has the largest residual, 9 meters, and is a possible outlier relative to the fitted linear pattern because its vertical distance is large compared with the other selected residuals and with \(s\). The tree with the most extreme trunk diameter, 18 centimeters, has a residual of only 1 meter. Its \(x\)-value is extreme among these selected observations, but its residual is not unusually large based on this comparison.

A Large Residual Is a Flag to Investigate

A possible outlier is not automatically a mistake or a reason to remove an observation. First check that the data were recorded and entered correctly. Then consider whether the case came from the same population and was measured under comparable conditions. A correct, unusual observation may reveal that the linear model does not describe every case equally well.

Look at the point in the original scatterplot as well as in the residual plot. The scatterplot shows the observation in context with the fitted line, while the residual plot makes the vertical differences easier to compare. If a point has a large residual, describe how far it is from the prediction and whether the line overpredicted or underpredicted. Avoid saying that it is an outlier merely because it is far to the right or left.

A large residual also does not, on its own, tell you why a prediction missed. The observation could reflect natural variation, a recording issue, a relevant circumstance not represented in the model, or a relationship that is not well described by the line for that case. The residual identifies a discrepancy between an observed response and a prediction; further context is needed to explain it.

Worked Example: Identifying an Unusual Canopy Width

A fictional gardening program predicts a tree’s canopy width, in meters, from its trunk diameter. The residual standard deviation for the fitted model is \(s=0.45\) meter. A residual plot shows these selected observations:

Trunk diameter \(x\) (cm)Residual (m)
8\(-0.20\)
120.15
160.10
201.35
24\(-0.25\)
400.18

State. Decide which observation appears most unusual in its response relative to the fitted line, and assess whether the largest predictor value is also the residual outlier.

Plan. Compare absolute residuals in meters with \(s=0.45\) meter. The residual plot’s horizontal position identifies the trunk diameter; the distance from zero identifies the size of the prediction error.

Do. At \(x=20\), the absolute residual is \(|1.35|=1.35\) meters, and \(1.35/0.45=3\). At the largest listed predictor value, \(x=40\), the absolute residual is \(|0.18|=0.18\) meter, and \(0.18/0.45=0.4\). For comparison, the other absolute residuals are 0.20, 0.15, 0.10, and 0.25 meter, all smaller than 1.35 meters.

Conclude. The observation at \(x=20\) has the largest residual and is a possible outlier relative to the fitted pattern: its residual is 1.35 meters, three times \(s\). The observation at \(x=40\) has the most extreme \(x\)-value among those listed, but its residual is only 0.18 meter. It is extreme in trunk diameter, not in its vertical distance from the regression line.

Why an Extreme \(x\)-Value Is Not Enough

It is tempting to call a point an outlier because it is far from the rest of the data along the horizontal axis. But residuals are vertical differences: they compare the observed \(y\) with the line’s predicted \(\hat{y}\) at that same \(x\). If a point far out in \(x\) follows the linear trend, the prediction may be close and the residual small. If a point near the center of the \(x\)-values is far above or below the trend, its residual may be large.

Worked Example: Comparing Two Delivery Routes

A fictional courier uses a regression line to predict delivery time, in minutes, from route distance, in kilometers. For two routes, the residual standard deviation is \(s=5\) minutes. A route that is much farther than the others has residual \(4\) minutes. Another route near the middle of the observed distance range has residual \(-8\) minutes.

State. Compare the routes’ residual sizes and distinguish the route with the extreme distance from the one with the greater prediction error.

Plan. Calculate the absolute residuals and compare them with \(s\). Use the \(x\)-information to identify which route is extreme in distance, but use the residual magnitudes to compare vertical prediction errors.

Do. The farther route has absolute residual \(|4|=4\) minutes, and \(4/5=0.8\). The route near the middle of the distance range has absolute residual \(|-8|=8\) minutes, and \(8/5=1.6\). The second route’s vertical error is \(8-4=4\) minutes greater in absolute size.

Conclude. The farther route is extreme in its predictor value, but its prediction is 4 minutes from the observed delivery time. The route near the middle of the distance range has the larger residual, 8 minutes, so it is farther from the fitted line. Neither comparison alone establishes a formal outlier cutoff, but the second route is the stronger candidate for further investigation based on residual size.

Common Mistakes and AP Exam Tips

  • Calling a point an outlier because \(x\) is extreme. State whether the point has an unusually large residual. A far-left or far-right position describes the predictor, not the vertical prediction error.
  • Comparing signed residuals without considering direction. A residual of \(-8\) is not smaller in distance than a residual of \(4\). Compare absolute values to assess how far points are from the line; retain the sign when describing overprediction or underprediction.
  • Ignoring units. Residuals and \(s\) are measured in the response variable’s units. If the response is delivery time in minutes, describe residual size in minutes.
  • Treating a rule of thumb as a law. A residual several times \(s\) may be unusual and worth checking, but there is no universal cutoff here that proves a point is an outlier. Compare it with the other residuals and the context.
  • Deleting an unusual point automatically. Verify the data and consider the case’s context before deciding what to do. A genuine observation should not be removed just because it does not fit the model well.
  • Claiming to know the cause of a large residual. A residual shows the difference between observed and predicted response. It does not establish why the difference occurred.

For a clear AP-style explanation, name the observation, give its residual or absolute residual with units, compare its size with other residuals or \(s\), and make a cautious conclusion. For example: “This route’s residual is \(-8\) minutes, so the model overpredicted its delivery time by 8 minutes. Its absolute residual is larger than the other listed residuals and is 1.6 times \(s\), so it is a possible outlier relative to the linear pattern.” Avoid using “extreme in \(x\)” as a substitute for that vertical comparison.

Key takeaway: A possible regression outlier has an unusually large residual—the vertical difference between the observed response and the fitted line. A point that is merely extreme in \(x\) may still have a small residual, so judge unusualness in the response relative to the model, not by horizontal position alone.

Check Your Understanding

Use residual size to assess unusualness, and keep predictor position separate from vertical distance.

  1. A model predicts plant height in centimeters. One observation has residual \(-12\) centimeters, while the other residuals range from \(-4\) to 5 centimeters. What does the sign tell you, and why might this observation deserve attention?
  2. A point has the largest \(x\)-value in a data set but a residual of 0.3 response units. Does its predictor position alone establish that it is an outlier relative to the fitted line? Explain.
  3. Two observations have residuals of \(7\) and \(-9\) minutes. Which has the larger absolute residual, and what is the difference in their distances from the fitted line?
  4. A model has \(s=2\) meters, and an observation’s residual is 6 meters. Calculate the residual-to-\(s\) ratio. What can this comparison suggest, and what does it not prove?
  5. List one check you could make before deciding whether to remove a point with an unusually large residual.