Test Influence by Fitting the Line Again
In “Influential Points and Their Effect on the Line,” you learned that a point is influential when removing it substantially changes the fitted regression line or correlation. Now we will make that comparison systematic: fit the line with the candidate point included, fit it again with that one point removed, and compare the slope, intercept, \(r\), and \(r^2\).
The goal is not to decide that a point is influential just because it looks unusual. Instead, measure what changes when the point is omitted. The size of a change is a numerical fact; whether it is substantial is a judgment that depends on the context and purpose of the model.
What to Compare
The slope \(b\) tells how much the fitted line predicts the response will change for a one-unit increase in \(x\). Compare the two slopes and retain their units: response units per explanatory-variable unit. The intercept \(a\) is the predicted response when \(x=0\). Compare the intercepts, but remember that its practical meaning depends on whether \(x=0\) makes sense in the setting.
The correlation \(r\) describes the direction and strength of the linear association. In simple linear regression with an intercept, \(r^2\), called the coefficient of determination, is the square of \(r\). It describes the proportion of the variation in the response values that is accounted for by the linear model for those observations. For example, \(r^2=0.64\) means 64% of the variation in those response values is accounted for by the linear model.
The comparison of \(r\) and \(r^2\) is not two independent checks: once \(r\) is known, \(r^2\) is determined. Reporting both can still help explain the fit. \(r\) retains direction, while \(r^2\) describes the fraction of response variation accounted for by the linear model.
Interpret changes in \(r^2\) carefully. It is not the percentage of predictions that are correct, and a higher \(r^2\) does not establish causation or guarantee useful predictions. Also, the full-data fit and reduced-data fit use different sets of response values, so their \(r^2\) values summarize variation in different data sets. Treat the change as a description of how the model fit responds to removing the point, not as a standalone score of model quality.
Using all observations, record the fitted line \(\hat{y}=a+bx\), the correlation \(r\), and \(r^2\).
Keep the same variables and measurement units, then refit the line and find the new \(r\) and \(r^2\).
State the before-and-after values. Describe changes in the slope in its units, and interpret changes in the other statistics in context.
Decide whether a change is substantial for the variables and the intended use of the model. Do not rely on an unusual-looking point alone.
With technology, fit a linear regression using all the observations, note its statistics, omit only the candidate, and run the regression again. On a TI-84, LinReg can provide the regression information when the calculator is set to display \(r\) and \(r^2\). Check that the second list contains exactly the remaining observations and that paired \(x\)- and \(y\)-values still match.
Worked Examples: Refit and Compare
Worked Example: A Long Trial in a Toy-Car Test
Original AP-style question. In a fictional toy-car test, \(x\) is track length in meters and \(y\) is completion time in seconds. Four trials give \((1,2)\), \((2,3)\), \((3,5)\), and \((4,4)\). A fifth trial gives \((9,10)\). Compare the fits with and without the fifth trial. Does it influence the model?
State. We will compare the slope, intercept, \(r\), and \(r^2\) for the five trials with those for the four trials after removing \((9,10)\). Only that observation will be removed.
Plan. Calculate each fit using the same explanatory and response variables, first with all five trials and then with the first four. This is a descriptive comparison of these observations, not an inference about a wider population.
Do: fit the four trials without the candidate. For the four remaining trials, \(\bar{x}=2.5\) and \(\bar{y}=3.5\). The sums of squared and cross-product deviations are \(S_{xx}=5\), \(S_{xy}=4\), and \(S_{yy}=5\). Therefore,
Without the candidate, the line is \(\hat{y}=1.5+0.8x\). Its correlation and coefficient of determination are
Do: fit all five trials. For all five observations, \(\sum x=19\), \(\sum y=24\), \(\sum x^2=111\), \(\sum xy=129\), and \(\sum y^2=154\). Thus \(\bar{x}=3.8\) and \(\bar{y}=4.8\), and
The full-data line and association statistics are
Conclude. Removing the 9-meter trial changes the slope from about 0.9742 to 0.8 seconds per meter and the intercept from about 1.0979 to 1.5 seconds. The correlation falls from about 0.9742 to 0.8000, and \(r^2\) falls from about 0.9491 to 0.6400. These are meaningful changes to the description of the association in these trials: the candidate point is influential for this fit. In particular, the change in \(r^2\) describes a less tightly linear fit among the four remaining trials; it does not mean the model became a particular percentage less accurate at prediction.
Worked Example: A Far-Out Trial on the Same Pattern
Original AP-style question. In a fictional light-testing setup, \(x\) is the number of hours a plant receives supplemental light and \(y\) is its measured growth in centimeters. Four plants have observations \((1,4)\), \((2,7)\), \((3,10)\), and \((4,13)\). A fifth plant has \((8,25)\). Compare the fit with and without the fifth plant.
The first four observations lie exactly on \(\hat{y}=1+3x\): substituting \(x=1,2,3,4\) gives predicted growth values 4, 7, 10, and 13. The fifth observation also follows that equation, since \(1+3(8)=25\). Therefore, all five points lie exactly on the same line.
With or without the fifth plant, the slope is 3 centimeters per hour and the intercept is 1 centimeter. Every point lies on the line, so \(r=1\) and \(r^2=1\) in both fits. There is no change in the statistics being compared, so the point is not influential in this example. Its \(x\)-value is far from the other four, but that position alone does not establish influence; its response follows the existing linear pattern.
Worked Example: An Unusual Response at an Ordinary x-Value
Original AP-style question. In a fictional tutoring-program exercise, \(x\) is the number of practice modules completed and \(y\) is the number of assessment points earned. Four groups have observations \((1,2)\), \((2,4)\), \((3,6)\), and \((4,8)\). A fifth group has \((3,1)\). Compare the fit with and without the fifth group.
Do: fit the four groups without the candidate. These four observations lie exactly on \(\hat{y}=2x\). The slope is 2 assessment points per module, the intercept is 0 points, \(r=1\), and \(r^2=1\).
Do: fit all five groups. With the fifth observation included, \(\sum x=13\), \(\sum y=21\), \(\sum x^2=39\), \(\sum xy=63\), and \(\sum y^2=121\). Thus \(\bar{x}=2.6\), \(\bar{y}=4.2\), and
The full-data regression statistics are
Conclude. The candidate has an ordinary \(x\)-value: three modules is between the other groups’ one-to-four-module values. But removing it changes the slope from about 1.6154 to 2 assessment points per module, while \(r\) changes from about 0.6432 to 1 and \(r^2\) from about 0.4137 to 1. The point substantially affects the fitted line and the strength of the linear association in this example. This comparison shows why influence is determined by refitting, not by checking for high leverage alone.
Common Mistakes and AP Exam Tip
- Removing more than one observation. To test one point’s influence, keep every other observation and remove only the candidate.
- Reporting \(r^2\) as a percentage of correct predictions. A full-credit explanation says it is the proportion of variation in the observed response values accounted for by the linear model, in the data being analyzed.
- Forgetting that \(r^2\) loses direction. Since \(r^2\) is squared, it is nonnegative. Use the slope or \(r\) to describe whether the association is positive or negative.
- Comparing statistics without units or context. State the slope in response units per \(x\)-unit, and identify what the variables represent when explaining the change.
- Calling any numerical change substantial. Explain why the difference matters for the model or its purpose; do not treat a small change as automatic proof of important influence.
- Claiming a higher \(r^2\) proves the model is better in every way. \(r^2\) summarizes variation in the particular response values used for that fit. It does not establish causation or guarantee dependable predictions.
A concise AP-style response gives evidence and interpretation: “Removing the fifth group changes the slope from about 1.6154 to 2 assessment points per practice module and increases \(r\) from about 0.6432 to 1. The change is substantial for describing the association in these groups, so this observation is influential for the fit.” This statement reports the comparison and explains what it means.
Check Your Understanding
Use a with-and-without comparison to answer each question.
- A point’s removal changes the slope from 4.1 to 2.0 response units per explanatory-variable unit. What information should you include when explaining whether that change matters?
- In simple linear regression with an intercept, \(r\) changes from \(-0.8\) to \(-0.5\) after a point is removed. What are the corresponding \(r^2\) values, and what information about direction does \(r^2\) leave out?
- Why must you remove only the candidate point when testing its influence?
- A candidate point has an unusually large \(x\)-value but lies on the line formed by the other points. Does its position alone prove it is influential? Explain.
- Write one sentence interpreting a change in \(r^2\) without describing it as the percentage of predictions that are correct.