Influence Is About What Changes
In “High-Leverage Points in the x-Direction,” you learned that a point far from the center of the \(x\)-values has the potential to pull or rotate a fitted line. Potential is not the same as an actual, substantial effect. To assess that effect, compare the regression results with the point included and with it removed.
A point is influential when removing it substantially changes an important feature of the regression analysis. In this tutorial, focus on three features: the slope, the intercept, and the correlation \(r\). The slope describes the predicted change in the response for a one-unit increase in \(x\); the intercept is the predicted response when \(x=0\); and \(r\) describes the direction and strength of the linear association.
“Substantially” is a contextual judgment, not a universal numerical cutoff. A small change in a slope may matter if it reverses the practical conclusion; a larger numerical change may matter less in another setting. Describe the changes and explain why they are important for the variables and units in the situation.
High leverage and a large residual can help draw attention to a candidate point, but neither label automatically means the point is influential. A high-leverage point may lie along the pattern of the other observations and leave the fit nearly unchanged. A point with an ordinary \(x\)-value can also affect the fit if its response is far from the pattern. The decisive evidence is the with-versus-without comparison.
A With-and-Without Comparison
Use the same cases, variables, and regression method in both fits. First fit the line and find \(r\) using all the data. Then omit the candidate observation, refit the line, and find the new \(r\). Compare the results rather than relying on a visual impression of the original scatterplot alone.
Record the fitted line \(\hat{y}=a+bx\) and the correlation \(r\) with all observations included.
Calculate or obtain the new intercept, slope, and correlation from the remaining observations.
State how the slope, intercept, or \(r\) changed. Decide whether the change is substantial for the variables and purpose of the model.
For hand calculations, the slope is \(b=S_{xy}/S_{xx}\), and the intercept is \(a=\bar{y}-b\bar{x}\). The correlation is \(r=S_{xy}/\sqrt{S_{xx}S_{yy}}\). Here, \(S_{xx}\), \(S_{xy}\), and \(S_{yy}\) are the usual sums of squared or cross-product deviations from their means. In a larger data set, technology can fit both regressions; the important part is to make sure the candidate point is excluded only from the second fit.
A line can change in more than one way. Its slope might become steeper or flatter, and its intercept might move. These changes are connected: when a line rotates around part of the data, both coefficients may change. The correlation can also change, but it measures linear association rather than the line’s location. A point can substantially affect the slope or intercept even if \(r\) changes very little.
Worked Examples: Comparing Fits
Worked Example: A Long Study Session and Quiz Score
Original AP-style question. A teacher records study hours, \(x\), and quiz score points, \(y\), for five students. Four students have data \((1,2)\), \((2,4)\), \((3,5)\), and \((4,7)\). A fifth student studied 10 hours and scored 12 points. Does the fifth observation substantially affect the fitted line?
State. We will compare the slope, intercept, and correlation with all five students included and with the 10-hour student removed. This directly assesses whether that observation is influential.
Plan. Fit the least-squares regression line to each data set, then compare the coefficients and \(r\). The comparison assumes the same variables and measurement units in both fits; only the candidate student is omitted from the second fit.
Do: fit the four students without the candidate. For these four observations, \(\bar{x}=2.5\), \(\bar{y}=4.5\), \(S_{xx}=5\), \(S_{xy}=8\), and \(S_{yy}=13\). Thus,
The line without the candidate is \(\hat{y}=0.5+1.6x\). Its correlation is
Do: fit all five students. The totals are \(\sum x=20\), \(\sum y=30\), \(\sum x^2=130\), \(\sum xy=173\), and \(\sum y^2=238\). The means are \(\bar{x}=4\) and \(\bar{y}=6\). Therefore,
So the slope and intercept with all five students are
The full-data line is \(\hat{y}=1.76+1.06x\), and its correlation is
Conclude. Removing the 10-hour student changes the slope from 1.06 to 1.6 quiz-score points per study hour and the intercept from 1.76 to 0.5 points. The correlation changes only from about 0.9842 to 0.9923. The slope change is substantial for describing the predicted score increase per additional study hour, so the student is influential for the fitted line. The small change in \(r\) does not cancel that conclusion: influence can be evident in the slope or intercept even when the correlation changes little.
Worked Example: High Leverage but Little Effect
Original AP-style question. A greenhouse manager records hours of supplemental light, \(x\), and plant height, \(y\), in centimeters. Four plants have observations \((2,5)\), \((3,7)\), \((4,9)\), and \((5,11)\). A fifth plant received 20 hours of light and reached 41 centimeters. Is the fifth plant influential?
The four plants without the candidate lie exactly on \(\hat{y}=1+2x\): for example, \(1+2(2)=5\), and the same equation gives 7, 9, and 11 for the remaining \(x\)-values. The candidate also lies on this line because \(1+2(20)=41\). Thus, including it does not change the fitted line: the slope remains 2 centimeters per hour and the intercept remains 1 centimeter.
All five points lie exactly on a line with positive slope, so \(r=1\) with the candidate included and \(r=1\) without it. Although 20 hours is far from the other plants’ 2-to-5-hour values and therefore has high leverage, removing it changes neither the line nor the correlation. In this example, the point is not influential. Its extreme \(x\)-position gives it the potential to matter, but its response follows the existing pattern.
Worked Example: A Large Residual at an Ordinary x-Value
Original AP-style question. A delivery service compares route length, \(x\), in kilometers, with delivery time, \(y\), in minutes. Four routes have observations \((1,1)\), \((2,2)\), \((3,3)\), and \((4,4)\). A fifth route is 2 kilometers long and takes 8 minutes. Assess the effect of the fifth route.
Without the fifth route, the four observations lie on \(\hat{y}=x\), so the slope is 1 minute per kilometer, the intercept is 0 minutes, and \(r=1\). The candidate’s \(x\)-value of 2 kilometers is within the range of the other route lengths, but the line predicts 2 minutes. Its residual is \(8-2=6\) minutes, a large positive departure from that pattern.
With all five observations, \(\bar{x}=12/5=2.4\) and \(\bar{y}=18/5=3.6\). The sums are \(\sum x^2=34\), \(\sum xy=46\), and \(\sum y^2=94\). Therefore,
The full-data slope, intercept, and correlation are
The slope drops from 1 to about 0.5385 minute per kilometer, and \(r\) drops from 1 to about 0.2272. The fifth route substantially changes the model despite having an ordinary \(x\)-value. It is influential in this data set. This example shows why an influence assessment cannot be limited to high-leverage points: a response far from the pattern can also alter the fit.
Common Mistakes and AP Exam Tip
- Equating high leverage with influence. High leverage means an unusual \(x\)-position and potential to affect the line. A full-credit influence claim points to a substantial change in the fit after removal.
- Equating a large residual with influence. A large residual identifies vertical departure from a line, not the size of the change caused by removing the point. Compare the two fits.
- Looking only at correlation. \(r\) may change very little while the slope or intercept changes substantially. Check all three features named in this tutorial.
- Reporting numbers without comparing them. State the before-and-after values and explain what changed. For a slope, retain the response-units-per-explanatory-unit interpretation.
- Claiming every numerical change is substantial. A fitted value may shift slightly because a point was removed. Explain why a change matters in context instead of treating any difference as proof of important influence.
- Removing more than the candidate point. To assess one observation’s effect, compare the full data set with the same data set minus only that observation.
A strong AP response names the point, reports the relevant with-and-without change, and explains its consequence in context. For example: “Removing the 10-hour student increases the predicted score change from 1.06 to 1.6 points per additional study hour, so this student substantially affects the slope of the quiz-score model.” That states evidence and interpretation rather than relying on the label “influential” alone.
Check Your Understanding
For each question, focus on what changes when the candidate point is removed.
- A high-leverage point lies on the same fitted line as the other observations. Does its high leverage alone show that it is influential? Explain.
- A point’s removal changes the slope from 3.2 to 1.1 response units per explanatory-variable unit but changes \(r\) only from 0.91 to 0.94. Which comparison supports an influence claim?
- Why is a point with an ordinary \(x\)-value not automatically unimportant to the regression fit?
- What must stay the same when you compare a full-data fit with a fit that omits one candidate point?
- Write one sentence explaining influence that includes a before-and-after change and interprets it in context.