Tutorials › AP Statistics › Effect of an Outlier on r-squared

Coefficient of determination · Tutorial 915 of 1000

Effect of an Outlier on r-squared

Compare r-squared before and after removing unusual observations, and learn why the change depends on how each point relates to the overall pattern.

Intermediate 9 min read

What You'll Learn

  • Compare r-squared for a data set before and after removing one observation.
  • Distinguish a point with an unusual response from one with an unusual predictor value.
  • Explain how a point can be influential by changing the fitted line and r-squared.
  • Calculate r-squared from sums of squares and paired data.
  • Explain why removing a point is not automatically justified just because r-squared improves.

Why One Point Can Change r-squared

In “Unexplained Variation and Residual Scatter,” you connected \(r^2\) to the fraction of response variation accounted for by a least-squares line. That fraction is calculated from all the observations used to fit the line. If one observation is unusual, including it can change the fitted line, the correlation, and the total variation in the response. As a result, \(r^2\) can change noticeably when that point is included or removed.

A point with an unusually large residual is a possible outlier in the regression setting, as described in “Residuals and Outliers.” A point can also be unusual because its predictor value is far from the other predictor values. Such a point has high leverage: its position in the horizontal direction gives it the potential to pull the fitted line toward itself. A point is influential when removing it substantially changes the fitted line or an important summary such as \(r^2\). A point may have an unusual response, high leverage, or both.

Definition: A point is influential when including it rather than removing it makes a substantial difference to the fitted regression results. To assess its effect descriptively, compare the results for the full data set with the results after removing that observation and refitting the line.

The key idea is that there is no fixed direction for the change in \(r^2\). An unusual point that does not follow the main pattern may weaken the linear association and lower \(r^2\). A distant point that extends the main pattern may strengthen the association and raise \(r^2\). And because removing a point changes the set of response values, it can change \(SST\) as well as \(SSE\). The comparison is not simply a matter of subtracting that point’s residual from \(SSE\).

What to Compare

For simple linear regression, \(r^2\) is the square of the correlation \(r\). It can also be calculated using \(SSE\) and \(SST\). If working from raw paired data, one useful calculation is to first find the centered sums of squares and products:

$$ S_{xx}=\sum x_i^2-\frac{(\sum x_i)^2}{n},\qquad S_{yy}=\sum y_i^2-\frac{(\sum y_i)^2}{n},\qquad S_{xy}=\sum x_i y_i-\frac{(\sum x_i)(\sum y_i)}{n} $$

Here \(S_{yy}=SST\), and for simple linear regression \(r=S_{xy}/\sqrt{S_{xx}S_{yy}}\). Squaring this gives a convenient formula for \(r^2\). Recalculate these quantities after removing the point; do not reuse the original means or sums of squares.

$$ r^2=\frac{S_{xy}^2}{S_{xx}S_{yy}} $$

When comparing two versions of the same data, report both \(r^2\) values and state which observations were included in each. Interpret each value in context: it describes the fraction of variation in the response values included in that particular fit that is accounted for by their linear relationship with the predictor. If the data sets differ, the two percentages refer to different sets of response values.

A useful comparison also considers what happened to the fitted line and the scatterplot. A large change in \(r^2\) tells you that the summary changed; it does not, by itself, explain why. Check whether the point is far out in the predictor direction, far vertically from the main pattern, or both. As you learned in “What r-squared Does Not Tell You,” \(r^2\) alone does not reveal the shape of the data or guarantee that a linear model is appropriate.

Worked Examples: Recalculating r-squared

Worked Example: A High-Leverage Point Weakens the Association

A fictional garden club records the weekly watering time \(x\), in minutes, and the number of new leaves \(y\) on a plant for five plants. The pairs are \((1,2),(2,4),(3,5),(4,4),(5,5)\). A sixth plant has the pair \((10,1)\). Compare \(r^2\) using all six plants with \(r^2\) after removing the sixth plant.

State. The predictor is weekly watering time, and the response is number of new leaves. The sixth plant has a much larger watering time than the others and a low leaf count. We will calculate \(r^2\) for each data set to describe how including that observation affects the linear association in these data.

Plan. Use \(r^2=S_{xy}^2/(S_{xx}S_{yy})\). Calculate the needed sums from the paired observations, first with all six plants and then with only the first five. The point at \(x=10\) is far from the other predictor values, so it has high leverage; the calculation will show whether including it changes the summary.

Do. For all six plants, \(n=6\), \(\sum x=25\), \(\sum y=21\), \(\sum x^2=155\), \(\sum y^2=87\), and \(\sum xy=76\). Therefore,

$$ S_{xx}=155-\frac{25^2}{6}=\frac{305}{6},\qquad S_{yy}=87-\frac{21^2}{6}=13.5,\qquad S_{xy}=76-\frac{25(21)}{6}=-11.5 $$

So \(r^2=\frac{(-11.5)^2}{(305/6)(13.5)}=\frac{132.25}{686.25}\approx0.1927\). As a check, the correlation is negative because \(S_{xy}<0\), and \(r\approx-0.4390\); squaring gives \(r^2\approx0.1927\).

After removing the sixth plant, the five remaining observations have \(n=5\), \(\sum x=15\), \(\sum y=20\), \(\sum x^2=55\), \(\sum y^2=86\), and \(\sum xy=66\). Then \(S_{xx}=55-\frac{15^2}{5}=10\), \(S_{yy}=86-\frac{20^2}{5}=6\), and \(S_{xy}=66-\frac{15(20)}{5}=6\). Thus, \(r^2=\frac{6^2}{(10)(6)}=\frac{36}{60}=0.6000\). Checking through the correlation, \(r=6/\sqrt{60}\approx0.7746\), whose square is \(0.6000\).

Conclude. For these plants, \(r^2\) is about 0.1927 with all six observations and 0.6000 without the sixth. With all six plants, about 19.27% of the variation in their leaf counts is accounted for by the linear relationship with watering time; among the five plants without that observation, about 60.00% is accounted for. The point is influential for this summary: it changes the correlation from positive to negative and substantially changes \(r^2\). This describes the two data sets; it does not establish that the sixth observation should be discarded.

Worked Example: A Distant Point Extends the Pattern

Use the same five fictional plant observations, \((1,2),(2,4),(3,5),(4,4),(5,5)\), but suppose the sixth plant instead has the pair \((10,8)\). Compare \(r^2\) for the five plants with \(r^2\) for all six.

State. The five-plant data set is unchanged, so its \(r^2\) is 0.6000. The new sixth point has a predictor value far from the others, but its response is also high, extending the positive direction of the pattern. We will recalculate \(r^2\) with all six observations.

Plan. Use the centered sums formula for the six paired observations. Then compare the result with the five-plant value. High leverage describes the point’s unusual predictor position, but it does not determine whether \(r^2\) will rise or fall; the point’s response position matters too.

Do. For all six plants, \(n=6\), \(\sum x=25\), \(\sum y=28\), \(\sum x^2=155\), \(\sum y^2=150\), and \(\sum xy=146\). Thus, \(S_{xx}=155-\frac{25^2}{6}=\frac{305}{6}\), \(S_{yy}=150-\frac{28^2}{6}=\frac{58}{3}\), and \(S_{xy}=146-\frac{25(28)}{6}=\frac{88}{3}\).

$$ r^2= \frac{(88/3)^2}{(305/6)(58/3)} =\frac{7744}{8845} \approx 0.8755 $$

As a check, the correlation is positive because \(S_{xy}>0\), and \(r=\frac{88/3}{\sqrt{(305/6)(58/3)}}\approx0.9357\). Squaring \(0.9357\) gives approximately \(0.8755\), allowing for rounding.

Conclude. In the five-plant data, \(r^2=0.6000\); with the sixth plant included, \(r^2\approx0.8755\). For the six plants, about 87.55% of the variation in leaf counts is accounted for by their linear relationship with watering time. Here, the high-leverage point strengthens the linear association because it lies in the same general direction as the existing pattern. This example and the previous one show why high leverage alone does not predict the direction of a change in \(r^2\).

Worked Example: A Vertical Outlier Changes the Response Variation

Return to the five plant observations \((1,2),(2,4),(3,5),(4,4),(5,5)\), but add a sixth plant with the pair \((3,10)\). The sixth plant has a typical predictor value but a much larger response than the others. Compare \(r^2\) before and after removing it.

State. We are comparing the original five-plant fit with a six-plant fit that includes an observation unusual in the response direction. Its predictor value, \(x=3\), is at the center of the other predictor values, so this point does not have an extreme predictor value.

Plan. The five-plant value is \(r^2=0.6000\), calculated earlier. For all six plants, recompute \(S_{xx}\), \(S_{yy}\), and \(S_{xy}\) using the updated sums. In particular, removing the point changes \(S_{yy}\), the total response variation, as well as the fitted relationship.

Do. For all six plants, \(n=6\), \(\sum x=18\), \(\sum y=30\), \(\sum x^2=64\), \(\sum y^2=186\), and \(\sum xy=96\). Therefore, \(S_{xx}=64-\frac{18^2}{6}=10\), \(S_{yy}=186-\frac{30^2}{6}=36\), and \(S_{xy}=96-\frac{18(30)}{6}=6\). So \(r^2=\frac{6^2}{(10)(36)}=\frac{36}{360}=0.1000\). As a check, \(r=6/\sqrt{360}\approx0.3162\), and \(0.3162^2\approx0.1000\).

Conclude. With all six plants, \(r^2=0.1000\); without the plant at \((3,10)\), \(r^2=0.6000\). The six-plant linear relationship with watering time accounts for about 10% of the variation in leaf counts, while the five-plant relationship accounts for about 60%. The unusual response increases the response variation and weakens the linear association in this example. Its effect is not found just by looking at its vertical distance: the summary must be recalculated using all the observations in each version.

Interpreting the Comparison Carefully

When one observation is removed, both the fitted line and the data being summarized can change. The mean of the response may change, so \(SST\) changes. The regression line is refitted, so the residuals and \(SSE\) change too. Recall from “Unexplained Variation and Residual Scatter” that \(r^2=1-SSE/SST\) for a least-squares line with an intercept. After removing a point, calculate the ratio for the new data set rather than assuming only \(SSE\) has changed.

The examples also show why the word “outlier” does not tell the whole story. The point \((10,1)\) had an extreme predictor value and a response that did not follow the positive pattern, so it pulled the association down. The point \((10,8)\) also had an extreme predictor value, but extended the positive pattern and raised \(r^2\). The point \((3,10)\) was unusual vertically while its predictor value was typical, and it reduced \(r^2\). These are different positions relative to the pattern, with different effects.

If an observation seems influential, investigate it rather than removing it automatically. Check for a recording or measurement error. If the value is valid, consider whether the individual belongs to the group the analysis is meant to describe. Report the result with and without the point when that comparison is informative, and explain which data were used for each calculation. A larger \(r^2\) after removal is not, by itself, a sound reason to exclude a valid observation.

Key distinction: A large change in \(r^2\) after removing a point is evidence that the point affects this summary for these data. It is not evidence, by itself, that the observation is erroneous or should be excluded.

Common Mistakes and AP Exam Tips

  • Assuming removing an outlier must increase \(r^2\). A point that follows the main direction can strengthen the association, as in the \((10,8)\) example. Describe the actual comparison instead of predicting the direction from the word “outlier.”
  • Calling every point with an unusual response high leverage. High leverage refers to an unusual predictor value. A point can be unusual vertically without being far out in the predictor direction.
  • Using the original means after removing a point. Recalculate the means or use the updated sums in the centered-sums formulas. The data set, and therefore its centers, have changed.
  • Treating the two \(r^2\) values as if they describe identical response data. State which observations are included. After deletion, the \(r^2\) interpretation applies to the remaining response values.
  • Removing a valid observation just to get a higher \(r^2\). First check whether the point is a data error and whether it belongs in the analysis. A point’s influence is a reason to examine it, not a justification for excluding it.
  • Interpreting \(r^2\) as a percentage of accurate predictions. In context, say what fraction of the variation in the response is accounted for by the linear relationship in that particular data set.

A full-credit comparison names the response and predictor, reports both \(r^2\) values, identifies which observations were included, and interprets the change in context. If the point has a far-out predictor value or an unusual response, describe that position accurately. Avoid claiming that the change proves the point is invalid or that one fit is automatically appropriate.

Key takeaway: An influential observation can raise or lower \(r^2\), depending on how its predictor and response values relate to the overall pattern. Recalculate \(r^2\) for each data set, interpret each value for the observations it includes, and investigate unusual points rather than deleting them solely to improve the fit.

Check Your Understanding

Use the examples and the ideas about leverage and influence to answer these questions.

  1. In the first plant example, compare \(r^2\) with all six observations and with the five observations after removal. Which version has the stronger linear association according to \(r^2\)?
  2. Why did the point \((10,8)\) raise \(r^2\), even though its predictor value was far from the other five values?
  3. A point has a typical predictor value but a response far from the overall pattern. Is it necessarily a high-leverage point? Explain.
  4. Why must you recalculate \(S_{xx}\), \(S_{yy}\), and \(S_{xy}\) after removing an observation?
  5. Give one reason to investigate an influential observation and one reason not to remove it automatically.