Tutorials › AP Statistics › Leverage Without Large Residuals

Unusual points and model fit · Tutorial 967 of 1000

Leverage Without Large Residuals

See why a high-leverage point can have a small residual but still noticeably change the correlation and the proportion of response variation explained by a linear model.

Intermediate 9 min read

What You'll Learn

  • Distinguish high leverage in the x-direction from a large residual in the y-direction
  • Identify when a point far from the center of x lies close to a fitted line
  • Compare r and r-squared with and without a high-leverage point
  • Interpret a change in r-squared in context without treating it as a measure of prediction accuracy
  • Explain why a small residual does not guarantee that a point has little influence

A Point Can Be Far Out in x and Close to the Line

In “Outlier With Little Influence,” you saw that a point can have a large residual without greatly changing the slope. This tutorial considers a different arrangement: a point with an \(x\)-value far from the center of the other \(x\)-values, but a \(y\)-value close to the linear pattern. As explained in “High-Leverage Points in the x-Direction,” its unusual \(x\)-position gives it high leverage.

High leverage and a large residual describe different features. Leverage is about how far a point’s \(x\)-value is from the center of the \(x\)-values. A residual measures the vertical difference between an observed response and the response predicted by a fitted line. So a point can have high leverage and a small residual at the same time.

Key idea: A high-leverage point close to a linear pattern can still affect the regression results. Compare the fits with and without the point, including \(r\) and \(r^2\), rather than judging influence from the residual alone.

Why a Close, High-Leverage Point Can Change r

The correlation \(r\) summarizes the direction and strength of the linear association. In simple linear regression, \(r^2\) is the proportion of variation in the response values accounted for by the linear model using the explanatory variable. As in “Testing Influence by Removing a Point,” a with-and-without comparison shows what changes when a candidate point is removed.

A point far from the center of \(x\) contributes substantially to the variation in \(x\), and it can contribute substantially to the joint pattern of \(x\) and \(y\). If it lies close to the existing line, its distant position may reinforce the linear association. That can make \(r\) stronger in magnitude and raise \(r^2\), even when the point’s residual is small.

There is no rule that every such point must increase \(r\) or \(r^2\). The result depends on the existing data and the point’s position. The essential method is to calculate or compare the statistics for both data sets, then describe the change in context. Remember that \(r^2\) describes variation in the particular data set; it is not the percentage of predictions that are correct.

Comparison to make: Check the candidate’s \(x\)-position relative to the other \(x\)-values, assess its residual from the line, then compare \(r\) and \(r^2\) with and without the point. A small residual does not replace the with-and-without comparison.

Worked Examples: High Leverage and a Small Residual

Worked Example: A Point on the Extended Pattern Raises r-squared

Original AP-style question. In a fictional practice session, \(x\) is minutes spent on a skill drill and \(y\) is the number of successful attempts on a later task. Five participants have observations \((1,7)\), \((2,9)\), \((3,11)\), \((4,12)\), and \((5,11)\). A sixth participant has \((13,21)\). Compare \(r\) and \(r^2\) before and after adding that participant.

State. We will assess whether the sixth observation has high leverage and whether it lies close to the original fitted line. Then we will compare the correlations and coefficients of determination for the two data sets.

Plan. Calculate the original regression summary and predict the response at \(x=13\). Next, use the sums of squared deviations and cross-product deviations to calculate \(r\) and \(r^2\) after adding the sixth point. The point is far beyond the original \(x\)-values, so it has a potentially influential position.

Do: calculate the original fit. For the five participants, \(\bar{x}=3\) and \(\bar{y}=10\). The sums are \(S_{xx}=10\), \(S_{xy}=11\), and \(S_{yy}=16\). Thus the slope is \(11/10=1.1\), and the intercept is \(10-(1.1)(3)=6.7\). The fitted line is \(\hat{y}=6.7+1.1x\). The original correlation and coefficient of determination are

$$ r=\frac{S_{xy}}{\sqrt{S_{xx}S_{yy}}} =\frac{11}{\sqrt{10(16)}}\approx0.8696, \qquad r^2\approx0.7563. $$

At \(x=13\), the original line predicts \(6.7+1.1(13)=21\), exactly the observed response. The residual relative to the original line is \(21-21=0\). The candidate’s \(x\)-value is also well beyond the original range from 1 to 5, so it is high leverage.

Do: include the candidate. Before adding the point, the means are 3 and 10. For a data set of five observations with a new point, the updated sums equal the old sums plus \(5/6\) times the product of the candidate’s deviations from the old means. Here those deviations are \(13-3=10\) and \(21-10=11\). Therefore,

$$ S_{xx,\mathrm{new}}=10+\frac{5}{6}(10^2)=93.3333, \qquad S_{xy,\mathrm{new}}=11+\frac{5}{6}(10)(11)=102.6667, $$ $$ S_{yy,\mathrm{new}}=16+\frac{5}{6}(11^2)=116.8333. $$

The new statistics are

$$ r_{\mathrm{new}} =\frac{102.6667}{\sqrt{93.3333(116.8333)}}\approx0.9832, \qquad r_{\mathrm{new}}^2\approx0.9666. $$

Conclude. The point has high leverage because its drill time is far from the other participants’ times, but it has a residual of zero from the original line. Including it raises \(r\) from about 0.8696 to 0.9832 and \(r^2\) from about 0.7563 to 0.9666. In this fictional data set, the linear association appears stronger with the point included. The increase in \(r^2\) means that the model accounts for a larger proportion of the response variation in the expanded data set; it does not mean that a stated percentage of individual predictions is correct.

Worked Example: A Close High-Leverage Point Changes the Fit Statistics

Original AP-style question. A fictional garden records \(x\), the number of hours of supplemental light, and \(y\), plant growth in centimeters. Five plants have observations \((2,5)\), \((4,8)\), \((6,9)\), \((8,13)\), and \((10,15)\). A sixth plant has \((20,27)\). Compare \(r\) and \(r^2\) with and without the sixth plant.

State and plan. The question is whether the sixth plant’s high-leverage \(x\)-value and small residual affect the reported strength of the linear association. First calculate the original line and residual for the candidate; then recalculate \(r\) and \(r^2\) with the point included.

Do: calculate the original fit. The five original plants have \(\bar{x}=6\) and \(\bar{y}=10\), with \(S_{xx}=40\), \(S_{xy}=50\), and \(S_{yy}=64\). The slope is \(50/40=1.25\) centimeters per hour, and the intercept is \(10-(1.25)(6)=2.5\) centimeters. The original line is \(\hat{y}=2.5+1.25x\). At \(x=20\), it predicts \(2.5+1.25(20)=27.5\) centimeters, so the candidate’s residual from this line is \(27-27.5=-0.5\) centimeter. This is a small vertical departure compared with the candidate’s distant \(x\)-position.

The original correlation and \(r^2\) are

$$ r=\frac{50}{\sqrt{40(64)}}\approx0.9882, \qquad r^2=\frac{50^2}{40(64)}=0.9766. $$

Do: include the candidate. Relative to the original means, the candidate’s deviations are \(20-6=14\) hours and \(27-10=17\) centimeters. Applying the update calculation gives

$$ S_{xx,\mathrm{new}}=40+\frac{5}{6}(14^2)=203.3333, \qquad S_{xy,\mathrm{new}}=50+\frac{5}{6}(14)(17)=248.3333, $$ $$ S_{yy,\mathrm{new}}=64+\frac{5}{6}(17^2)=304.8333. $$

Thus,

$$ r_{\mathrm{new}} =\frac{248.3333}{\sqrt{203.3333(304.8333)}}\approx0.9975, \qquad r_{\mathrm{new}}^2\approx0.9949. $$

Conclude. The point is high leverage because 20 hours is much farther from the center of the other light durations than any of their values. Its residual from the original line is only \(-0.5\) centimeter. Including it raises \(r\) from about 0.9882 to 0.9975 and \(r^2\) from about 0.9766 to 0.9949. The change is modest because the original association was already very strong, but the comparison still shows that a close high-leverage point can affect these statistics.

Worked Example: Interpreting a Change in r-squared

Original AP-style question. A fictional technology club records the number of practice sessions \(x\) and a task score \(y\) for five members: \((1,3)\), \((2,5)\), \((3,4)\), \((4,7)\), and \((5,6)\). A sixth member has \((9,10)\). Find the two values of \(r^2\) and interpret their difference.

Do: calculate the original fit. For the first five members, \(\bar{x}=3\) and \(\bar{y}=5\). The sums are \(S_{xx}=10\), \(S_{xy}=8\), and \(S_{yy}=10\). The original line has slope \(8/10=0.8\) score points per practice session and intercept \(5-(0.8)(3)=2.6\) score points. At \(x=9\), it predicts \(2.6+0.8(9)=9.8\), so the added member’s residual relative to the original line is \(10-9.8=0.2\) score points. The \(x\)-value of 9 is well beyond the other values, indicating high leverage.

For the original data,

$$ r=\frac{8}{\sqrt{10(10)}}=0.8, \qquad r^2=0.8^2=0.64. $$

Do: include the sixth member. The added point’s deviations from the original means are \(9-3=6\) sessions and \(10-5=5\) score points. The updated sums are

$$ S_{xx,\mathrm{new}}=10+\frac{5}{6}(6^2)=40, \qquad S_{xy,\mathrm{new}}=8+\frac{5}{6}(6)(5)=33, $$ $$ S_{yy,\mathrm{new}}=10+\frac{5}{6}(5^2)=30.8333. $$

The new coefficient of determination is

$$ r_{\mathrm{new}}^2 =\frac{33^2}{40(30.8333)} \approx0.8832. $$

Conclude. In this fictional club data, \(r^2\) increases from 0.64 to about 0.8832 when the high-leverage member is included. For the five-member data set, about 64% of the variation in task scores is accounted for by the linear model relating score to practice sessions. For the six-member data set, the corresponding proportion is about 88.32%. These are descriptions of the variation in each data set, not guarantees about future scores or percentages of correct individual predictions.

Common Mistakes and AP Exam Tip

  • Calling a point a y-outlier because its \(x\)-value is extreme. An unusual \(x\)-position indicates high leverage; a y-outlier has an unusually large vertical departure from the pattern. The two descriptions are not interchangeable.
  • Assuming a small residual means no influence. A residual describes vertical distance, not the effect of removing the point. Compare the statistics with and without it.
  • Claiming high leverage always raises \(r\) or \(r^2\). The examples show increases, but the direction and size of a change depend on the data. Report what the comparison actually shows.
  • Interpreting \(r^2\) as accuracy. Say “the proportion of variation in the response accounted for by the linear model,” with the response and explanatory variables identified. Do not call it the percentage of predictions that are correct.
  • Reporting a changed statistic without context. Include which cases are being compared and what the response and explanatory variables measure. Keep the interpretation about association, not causation.

A strong AP-style response might say: “The added plant has high leverage because its light duration of 20 hours is far from the other plants’ values, but its residual from the original line is only \(-0.5\) centimeter. Including it changes \(r^2\) from about 0.9766 to 0.9949, so the linear model accounts for a larger proportion of the variation in growth in the expanded data set.” This distinguishes leverage from residual size and states the change in context.

Key takeaway: High leverage describes an unusual position in \(x\), while a residual describes vertical distance from the line. A point can have high leverage and a small residual yet still change \(r\) and \(r^2\). Use a with-and-without comparison and interpret the change in context.

Check Your Understanding

Use the distinction between leverage and residual size to answer each question.

  1. What feature of an observation makes it high leverage in the \(x\)-direction?
  2. In the garden example, how far was the added point from the original fitted line, and how did \(r^2\) change?
  3. Why does a small residual by itself not show that a point has little influence on \(r\) or \(r^2\)?
  4. Write a contextual interpretation of the technology-club example’s \(r^2\) with the sixth member included.
  5. Why is it inaccurate to describe an \(r^2\) of about 0.8832 as “88.32% of predictions are correct”?