Tutorials › AP Statistics › Outlier With Little Influence

Unusual points and model fit · Tutorial 966 of 1000

Outlier With Little Influence

Learn how a y-outlier near the center of the explanatory-variable values can leave the slope nearly unchanged while shifting the line or changing the correlation.

Intermediate 9 min read

What You'll Learn

  • Identify why a large residual and strong influence are different ideas
  • Explain how a point near the mean of x can have little effect on the regression slope
  • Compare fitted lines with and without a y-outlier
  • Recognize that the intercept and correlation may still change
  • Describe the with-and-without comparison in context

A Large Residual Does Not Always Mean a Large Change to the Line

In “Testing Influence by Removing a Point,” you learned to check influence by comparing regression results with and without a candidate observation. This tutorial looks at a particular case: a point can be far from the line in the vertical direction but have an \(x\)-value near the mean of the other \(x\)-values. Such a point is a y-outlier, but it may change the fitted line only a little.

The reason is that the slope depends on how the \(x\)- and \(y\)-values vary together. A point near the center of the \(x\)-values contributes little to that pattern, even if its \(y\)-value is unusual. By contrast, a point far from the center of the \(x\)-values may have high leverage, as described in “High-Leverage Points in the x-Direction.”

Key idea: A large residual describes a point’s vertical distance from a fitted line. Influence describes how much the fitted line or correlation changes when the point is removed. A y-outlier near the mean of \(x\) can have a large residual and little effect on the slope.

Why the Mean of x Matters

For a fitted line \(\hat{y}=a+bx\), the slope \(b\) describes the predicted change in \(y\) for a one-unit increase in \(x\). In calculating the slope, each observation’s \(x\)-value is compared with the mean of the \(x\)-values. A point whose \(x\)-value is near that mean has a small horizontal deviation from the center.

Consider adding a point to an existing data set, and call the original number of observations \(n\). If the new point’s \(x\)-value is exactly the original \(\bar{x}\), the mean of \(x\) stays the same. The new point adds zero to the sum of squared \(x\)-deviations, \(S_{xx}\). It also adds zero to the sum of cross-product deviations, \(S_{xy}\), because its \(x\)-deviation is zero. Since the slope is \(b=S_{xy}/S_{xx}\), the slope stays exactly the same.

$$ b=\frac{S_{xy}}{S_{xx}} $$

That does not mean every part of the regression output stays the same. A point with an unusual \(y\)-value can move the mean of \(y\), changing the intercept. It can also change \(r\) and \(r^2\), because it changes the variation in the response values. If the candidate’s \(x\)-value is close to, but not exactly at, the original \(\bar{x}\), the slope can change slightly rather than remain exactly fixed.

The useful conclusion is precise: being near \(\bar{x}\) often limits a point’s effect on the slope, but it does not guarantee that the entire fitted line or every measure of association will be unchanged. Use the with-and-without comparison from “Testing Influence by Removing a Point” to see what actually changes.

Worked Examples: Outliers Near the Center of x

Worked Example: A Large y-Outlier at the Mean of x

Original AP-style question. In a fictional equipment test, \(x\) is the number of minutes a device runs and \(y\) is the temperature increase in degrees. Twenty devices have \(x\)-values \(1,2,\ldots,20\), and for each device \(y=10+2x\). One additional device has \(x=10.5\) and \(y=51\). Compare the regression line with and without this candidate point.

State. We will compare the slopes and intercepts from the twenty original devices and from all twenty-one devices. The candidate is a y-outlier if it lies far vertically from the pattern in the original devices.

Plan. Calculate the regression statistics for the original data, then refit with the candidate included. The original mean of \(x\) is \(10.5\), so the candidate is exactly at that mean. We will also compare \(r\) and \(r^2\) to show that an unchanged slope does not mean all fit statistics remain unchanged.

Do: fit the original twenty devices. The original points lie on \(\hat{y}=10+2x\). Their means are \(\bar{x}=10.5\) and \(\bar{y}=31\). The sums of squared deviations and cross-product deviations are \(S_{xx}=665\), \(S_{xy}=1330\), and \(S_{yy}=2660\). Thus,

$$ b=\frac{1330}{665}=2,\qquad a=31-(2)(10.5)=10. $$

The original correlation is \(r=1330/\sqrt{665(2660)}=1\), so \(r^2=1\). At \(x=10.5\), the original line predicts \(10+2(10.5)=31\). The candidate’s observed response is 51, giving a residual of \(51-31=20\) degrees. This is a large vertical departure from the original pattern.

Do: include the candidate. The candidate’s \(x\)-value is the original mean of \(x\), so the new mean of \(x\) remains \(10.5\). Its \(x\)-deviation is zero: \(10.5-10.5=0\). Therefore, \(S_{xx}\) stays 665, and \(S_{xy}\) stays 1330. The slope is still \(1330/665=2\).

The response total for the original data is \(20(31)=620\). After adding the candidate, the mean response is \((620+51)/21=31.9524\), rounded. The new intercept is therefore

$$ a=31.9524-(2)(10.5)=10.9524. $$

The line with the candidate is \(\hat{y}=10.9524+2x\). Its slope is unchanged, and its intercept is about 0.9524 degrees higher than before. For example, at \(x=10.5\), the new line predicts \(31.9524\), compared with the original prediction of 31.

The candidate does change the correlation. With the added point, \(S_{yy}=2660+(20/21)(20^2)=3040.9524\). Therefore,

$$ r=\frac{1330}{\sqrt{665(3040.9524)}}\approx0.9353,\qquad r^2\approx0.8747. $$

Conclude. The candidate is far above the original pattern, with a residual of 20 degrees, but its \(x\)-value is at the center of the original \(x\)-values. The slope remains 2 degrees per minute, and the intercept changes by about 0.9524 degrees. Thus, the point barely changes the line’s slope and shifts the line only modestly in this setting. It does change the correlation and \(r^2\), so it would be inaccurate to say it has no effect on the regression results.

Worked Example: A Centered Outlier Changes r More Than the Slope

Original AP-style question. In a fictional greenhouse exercise, \(x\) is the number of hours of supplemental light and \(y\) is a plant’s measured growth in centimeters. Five plants have observations \((1,12)\), \((2,14)\), \((3,16)\), \((4,18)\), and \((5,20)\). A sixth plant has observation \((3,31)\). Compare the fits with and without the sixth plant.

Do: fit the five original plants. The points follow \(\hat{y}=10+2x\), so the slope is 2 centimeters per hour, the intercept is 10 centimeters, \(r=1\), and \(r^2=1\). The original \(\bar{x}\) is 3. At \(x=3\), the line predicts 16 centimeters, so the candidate’s residual relative to that line is \(31-16=15\) centimeters.

Do: include the candidate. For the original five points, \(\bar{x}=3\), \(\bar{y}=16\), \(S_{xx}=10\), \(S_{xy}=20\), and \(S_{yy}=40\). The candidate’s \(x\)-value is exactly \(\bar{x}\), so the new \(\bar{x}\) stays 3 and \(S_{xx}\) and \(S_{xy}\) do not change. The new response mean is \((5(16)+31)/6=18.5\). Hence,

$$ b=\frac{20}{10}=2,\qquad a=18.5-(2)(3)=12.5. $$

The candidate’s unusual response increases \(S_{yy}\) by \((5/6)(15^2)=187.5\), giving a new \(S_{yy}=227.5\). The correlation and coefficient of determination are

$$ r=\frac{20}{\sqrt{10(227.5)}}\approx0.4193,\qquad r^2\approx0.1758. $$

Conclude. The slope remains 2 centimeters per hour because the candidate lies at the original mean of \(x\). The intercept increases from 10 to 12.5 centimeters, shifting the line upward, and \(r\) falls from 1 to about 0.4193. This example shows why “the slope barely changed” is not the same claim as “the point had no effect on the fit.” The change in \(r^2\) describes the proportion of variation in the response values accounted for by the linear model in each data set; it is not the percentage of predictions that are correct.

Worked Example: A Point Close to, but Not Exactly at, the Mean

Original AP-style question. Return to the equipment-test pattern where twenty devices have \(x\)-values \(1,2,\ldots,20\) and responses \(y=10+2x\). Now add a device with \(x=10.55\) and \(y=51.1\). Its \(x\)-value is close to the original mean of 10.5. Find the new slope and intercept and compare the line with the original one.

The original means are \(\bar{x}=10.5\) and \(\bar{y}=31\), with \(S_{xx}=665\) and \(S_{xy}=1330\). Relative to those means, the candidate has \(x\)-deviation \(10.55-10.5=0.05\) and \(y\)-deviation \(51.1-31=20.1\). Adding this point gives

$$ S_{xx,\mathrm{new}}=665+\frac{20}{21}(0.05^2)=665.0024, $$ $$ S_{xy,\mathrm{new}}=1330+\frac{20}{21}(0.05)(20.1)=1330.9571. $$

The new slope is

$$ b_{\mathrm{new}}=\frac{1330.9571}{665.0024}\approx2.0014. $$

For the new means, \(\bar{x}_{\mathrm{new}}=(210+10.55)/21\approx10.5024\) and \(\bar{y}_{\mathrm{new}}=(620+51.1)/21\approx31.9571\). The new intercept is

$$ a_{\mathrm{new}}=31.9571-(2.0014)(10.5024)\approx10.9373. $$

Conclude. The slope changes from 2 to about 2.0014 degrees per minute, a difference of about 0.0014 degrees per minute. The intercept changes from 10 to about 10.9373 degrees. Across the observed \(x\)-values, the new line is roughly one degree above the original line. The candidate is still far above the pattern: at \(x=10.55\), the new line predicts about 32.0524 degrees, leaving a residual of about \(51.1-32.0524=19.0476\) degrees. Its \(x\)-value is close to the center, so the large vertical departure has very little effect on the slope.

Common Mistakes and AP Exam Tip

  • Calling a point influential just because its residual is large. A large residual makes a point unusual in the \(y\)-direction. To assess influence, compare the fits with and without it.
  • Claiming that a centered point changes nothing. At the original \(\bar{x}\), the slope stays fixed when the point is added, but the intercept, \(r\), and \(r^2\) can change.
  • Using “near the mean” as a guarantee. A point close to \(\bar{x}\) often has little effect on the slope, but the size of any change should be checked by refitting. If it is only near, not exactly at, the mean, the slope may change slightly.
  • Describing only the correlation change. If asked whether the point changes the line, report the slope and intercept comparisons as well as any requested association statistics.
  • Leaving out context and units. State the slope in response units per explanatory-variable unit. Explain an intercept change as a change in the model’s predicted response at \(x=0\), while noting whether that value is meaningful in the setting.

A strong AP-style response separates the observations from the conclusion: “The candidate is a y-outlier, with a residual of about 19 degrees, but its run time is close to the mean run time. Removing it changes the slope only from about 2.0014 to 2 degrees per minute. The point has little effect on the slope, although the intercept and correlation may also change.” This reports evidence and avoids treating an unusual response as automatic proof of influence.

Key takeaway: A y-outlier near the mean of \(x\) can have a large residual but little effect on the slope because it contributes little to the \(x\)-based variation used to determine that slope. Check the intercept and association statistics too: they may change even when the slope barely does.

Check Your Understanding

Use the distinction between a y-outlier and an influential point to answer each question.

  1. Why does adding a point whose \(x\)-value equals the original \(\bar{x}\) leave \(S_{xx}\) and \(S_{xy}\) unchanged?
  2. A centered point leaves the slope unchanged but moves the mean of \(y\). Which part of the fitted line can change as a result?
  3. In the greenhouse example, what were the candidate’s residual relative to the original line and its effect on the slope?
  4. Why is it inaccurate to conclude that a point has no effect just because the slope is unchanged?
  5. Write a contextual sentence distinguishing the candidate’s large residual from its influence on the slope in the equipment-test example.