Tutorials › AP Statistics › Testing Influence by Removing a Point

Unusual points and model fit · Tutorial 965 of 1000

Testing Influence by Removing a Point

Refit a regression line after removing one observation and use changes in the slope, intercept, correlation, and \(r^2\) to assess its influence.

Intermediate 10 min read

What You'll Learn

  • Organize a with-and-without comparison that removes only the candidate point.
  • Calculate and compare the slope and intercept from both regression fits.
  • Compare correlation \(r\) and explain how it relates to \(r^2\) in simple linear regression.
  • Interpret changes in each statistic in the context and units of the variables.
  • Explain why high leverage or a large residual alone does not establish influence.
  • Avoid treating a change in \(r^2\) as a measure of accuracy or as evidence of causation.

Test Influence by Fitting the Line Again

In “Influential Points and Their Effect on the Line,” you learned that a point is influential when removing it substantially changes the fitted regression line or correlation. Now we will make that comparison systematic: fit the line with the candidate point included, fit it again with that one point removed, and compare the slope, intercept, \(r\), and \(r^2\).

The goal is not to decide that a point is influential just because it looks unusual. Instead, measure what changes when the point is omitted. The size of a change is a numerical fact; whether it is substantial is a judgment that depends on the context and purpose of the model.

Definition: A with-and-without influence check compares regression results from all observations with results from the same data set after removing one candidate observation. A substantial change in the slope, intercept, \(r\), or \(r^2\) is evidence that the point influences that feature of the fit.

What to Compare

The slope \(b\) tells how much the fitted line predicts the response will change for a one-unit increase in \(x\). Compare the two slopes and retain their units: response units per explanatory-variable unit. The intercept \(a\) is the predicted response when \(x=0\). Compare the intercepts, but remember that its practical meaning depends on whether \(x=0\) makes sense in the setting.

The correlation \(r\) describes the direction and strength of the linear association. In simple linear regression with an intercept, \(r^2\), called the coefficient of determination, is the square of \(r\). It describes the proportion of the variation in the response values that is accounted for by the linear model for those observations. For example, \(r^2=0.64\) means 64% of the variation in those response values is accounted for by the linear model.

Formula: In simple linear regression with an intercept, \(r^2=(r)^2\). Because squaring removes the sign, \(r^2\) does not show whether the association is positive or negative; use the slope or \(r\) for direction.

The comparison of \(r\) and \(r^2\) is not two independent checks: once \(r\) is known, \(r^2\) is determined. Reporting both can still help explain the fit. \(r\) retains direction, while \(r^2\) describes the fraction of response variation accounted for by the linear model.

Interpret changes in \(r^2\) carefully. It is not the percentage of predictions that are correct, and a higher \(r^2\) does not establish causation or guarantee useful predictions. Also, the full-data fit and reduced-data fit use different sets of response values, so their \(r^2\) values summarize variation in different data sets. Treat the change as a description of how the model fit responds to removing the point, not as a standalone score of model quality.

1
Record the full-data results.
Using all observations, record the fitted line \(\hat{y}=a+bx\), the correlation \(r\), and \(r^2\).
2
Remove just the candidate point.
Keep the same variables and measurement units, then refit the line and find the new \(r\) and \(r^2\).
3
Compare the statistics.
State the before-and-after values. Describe changes in the slope in its units, and interpret changes in the other statistics in context.
4
Judge the importance of the changes.
Decide whether a change is substantial for the variables and the intended use of the model. Do not rely on an unusual-looking point alone.

With technology, fit a linear regression using all the observations, note its statistics, omit only the candidate, and run the regression again. On a TI-84, LinReg can provide the regression information when the calculator is set to display \(r\) and \(r^2\). Check that the second list contains exactly the remaining observations and that paired \(x\)- and \(y\)-values still match.

Worked Examples: Refit and Compare

Worked Example: A Long Trial in a Toy-Car Test

Original AP-style question. In a fictional toy-car test, \(x\) is track length in meters and \(y\) is completion time in seconds. Four trials give \((1,2)\), \((2,3)\), \((3,5)\), and \((4,4)\). A fifth trial gives \((9,10)\). Compare the fits with and without the fifth trial. Does it influence the model?

State. We will compare the slope, intercept, \(r\), and \(r^2\) for the five trials with those for the four trials after removing \((9,10)\). Only that observation will be removed.

Plan. Calculate each fit using the same explanatory and response variables, first with all five trials and then with the first four. This is a descriptive comparison of these observations, not an inference about a wider population.

Do: fit the four trials without the candidate. For the four remaining trials, \(\bar{x}=2.5\) and \(\bar{y}=3.5\). The sums of squared and cross-product deviations are \(S_{xx}=5\), \(S_{xy}=4\), and \(S_{yy}=5\). Therefore,

$$ b=\frac{S_{xy}}{S_{xx}}=\frac{4}{5}=0.8, \qquad a=\bar{y}-b\bar{x}=3.5-(0.8)(2.5)=1.5. $$

Without the candidate, the line is \(\hat{y}=1.5+0.8x\). Its correlation and coefficient of determination are

$$ r=\frac{4}{\sqrt{5(5)}}=0.8000, \qquad r^2=(0.8000)^2=0.6400. $$

Do: fit all five trials. For all five observations, \(\sum x=19\), \(\sum y=24\), \(\sum x^2=111\), \(\sum xy=129\), and \(\sum y^2=154\). Thus \(\bar{x}=3.8\) and \(\bar{y}=4.8\), and

$$ S_{xx}=111-5(3.8^2)=38.8,\quad S_{xy}=129-5(3.8)(4.8)=37.8,\quad S_{yy}=154-5(4.8^2)=38.8. $$

The full-data line and association statistics are

$$ b=\frac{37.8}{38.8}\approx0.9742,\qquad a=4.8-(0.9742)(3.8)\approx1.0979, $$ $$ r=\frac{37.8}{\sqrt{38.8(38.8)}}\approx0.9742,\qquad r^2=(0.9742)^2\approx0.9491. $$

Conclude. Removing the 9-meter trial changes the slope from about 0.9742 to 0.8 seconds per meter and the intercept from about 1.0979 to 1.5 seconds. The correlation falls from about 0.9742 to 0.8000, and \(r^2\) falls from about 0.9491 to 0.6400. These are meaningful changes to the description of the association in these trials: the candidate point is influential for this fit. In particular, the change in \(r^2\) describes a less tightly linear fit among the four remaining trials; it does not mean the model became a particular percentage less accurate at prediction.

Worked Example: A Far-Out Trial on the Same Pattern

Original AP-style question. In a fictional light-testing setup, \(x\) is the number of hours a plant receives supplemental light and \(y\) is its measured growth in centimeters. Four plants have observations \((1,4)\), \((2,7)\), \((3,10)\), and \((4,13)\). A fifth plant has \((8,25)\). Compare the fit with and without the fifth plant.

The first four observations lie exactly on \(\hat{y}=1+3x\): substituting \(x=1,2,3,4\) gives predicted growth values 4, 7, 10, and 13. The fifth observation also follows that equation, since \(1+3(8)=25\). Therefore, all five points lie exactly on the same line.

With or without the fifth plant, the slope is 3 centimeters per hour and the intercept is 1 centimeter. Every point lies on the line, so \(r=1\) and \(r^2=1\) in both fits. There is no change in the statistics being compared, so the point is not influential in this example. Its \(x\)-value is far from the other four, but that position alone does not establish influence; its response follows the existing linear pattern.

Worked Example: An Unusual Response at an Ordinary x-Value

Original AP-style question. In a fictional tutoring-program exercise, \(x\) is the number of practice modules completed and \(y\) is the number of assessment points earned. Four groups have observations \((1,2)\), \((2,4)\), \((3,6)\), and \((4,8)\). A fifth group has \((3,1)\). Compare the fit with and without the fifth group.

Do: fit the four groups without the candidate. These four observations lie exactly on \(\hat{y}=2x\). The slope is 2 assessment points per module, the intercept is 0 points, \(r=1\), and \(r^2=1\).

Do: fit all five groups. With the fifth observation included, \(\sum x=13\), \(\sum y=21\), \(\sum x^2=39\), \(\sum xy=63\), and \(\sum y^2=121\). Thus \(\bar{x}=2.6\), \(\bar{y}=4.2\), and

$$ S_{xx}=39-5(2.6^2)=5.2,\quad S_{xy}=63-5(2.6)(4.2)=8.4,\quad S_{yy}=121-5(4.2^2)=32.8. $$

The full-data regression statistics are

$$ b=\frac{8.4}{5.2}\approx1.6154,\qquad a=4.2-(1.6154)(2.6)\approx0, $$ $$ r=\frac{8.4}{\sqrt{5.2(32.8)}}\approx0.6432,\qquad r^2=(0.6432)^2\approx0.4137. $$

Conclude. The candidate has an ordinary \(x\)-value: three modules is between the other groups’ one-to-four-module values. But removing it changes the slope from about 1.6154 to 2 assessment points per module, while \(r\) changes from about 0.6432 to 1 and \(r^2\) from about 0.4137 to 1. The point substantially affects the fitted line and the strength of the linear association in this example. This comparison shows why influence is determined by refitting, not by checking for high leverage alone.

Common Mistakes and AP Exam Tip

  • Removing more than one observation. To test one point’s influence, keep every other observation and remove only the candidate.
  • Reporting \(r^2\) as a percentage of correct predictions. A full-credit explanation says it is the proportion of variation in the observed response values accounted for by the linear model, in the data being analyzed.
  • Forgetting that \(r^2\) loses direction. Since \(r^2\) is squared, it is nonnegative. Use the slope or \(r\) to describe whether the association is positive or negative.
  • Comparing statistics without units or context. State the slope in response units per \(x\)-unit, and identify what the variables represent when explaining the change.
  • Calling any numerical change substantial. Explain why the difference matters for the model or its purpose; do not treat a small change as automatic proof of important influence.
  • Claiming a higher \(r^2\) proves the model is better in every way. \(r^2\) summarizes variation in the particular response values used for that fit. It does not establish causation or guarantee dependable predictions.

A concise AP-style response gives evidence and interpretation: “Removing the fifth group changes the slope from about 1.6154 to 2 assessment points per practice module and increases \(r\) from about 0.6432 to 1. The change is substantial for describing the association in these groups, so this observation is influential for the fit.” This statement reports the comparison and explains what it means.

Key takeaway: Refit the regression after removing only the candidate point. Compare slope, intercept, \(r\), and \(r^2\), and judge the changes in context. An unusual \(x\)-value or response can suggest a candidate, but the with-and-without comparison is the evidence of influence.

Check Your Understanding

Use a with-and-without comparison to answer each question.

  1. A point’s removal changes the slope from 4.1 to 2.0 response units per explanatory-variable unit. What information should you include when explaining whether that change matters?
  2. In simple linear regression with an intercept, \(r\) changes from \(-0.8\) to \(-0.5\) after a point is removed. What are the corresponding \(r^2\) values, and what information about direction does \(r^2\) leave out?
  3. Why must you remove only the candidate point when testing its influence?
  4. A candidate point has an unusually large \(x\)-value but lies on the line formed by the other points. Does its position alone prove it is influential? Explain.
  5. Write one sentence interpreting a change in \(r^2\) without describing it as the percentage of predictions that are correct.