Tutorials › AP Statistics › Interpreting Computer Output With Unusual Points

Unusual points and model fit · Tutorial 978 of 1000

Interpreting Computer Output With Unusual Points

Use regression output from fits with and without an unusual point to judge what changed and describe the effect in context.

Intermediate 10 min read

What You'll Learn

  • Identify the regression statistics commonly reported by a calculator or software.
  • Compare intercepts and slopes from fits with and without an observation.
  • Interpret changes in \(r\), \(r^2\), and residual standard deviation.
  • Distinguish an unusual \(x\)-position from an influential observation.
  • Write a conclusion that describes the size and meaning of the output changes in context.

What Regression Output Can Reveal About an Unusual Point

A computer can quickly fit a least-squares regression line, but its output does not decide whether an observation is influential. To assess influence, compare the reported values from a fit using all observations with the values from a fit after removing just the candidate point. The change in the results—not simply the point’s unusual appearance—is the evidence to interpret.

In “Influential Points and Their Effect on the Line” and “Testing Influence by Removing a Point,” we used this with-and-without comparison. Here we focus on reading and explaining the computer output. A point can change the slope or intercept, alter how strong the linear association appears, or change the typical size of the residuals. It may affect some of these more than others.

Definition: An observation is influential when removing it substantially changes a reported feature of the regression, such as the slope, intercept, \(r\), \(r^2\), or residual standard deviation \(s\). Whether a change is substantial depends on its size and meaning in the context.

A typical regression output reports the intercept \(a\), slope \(b\), correlation \(r\), coefficient of determination \(r^2\), residual standard deviation \(s\), and number of observations \(n\). The fitted line is \(\hat{y}=a+bx\). As in earlier tutorials, the slope gives the predicted change in response for a one-unit increase in the explanatory variable, and the intercept is the predicted response when \(x=0\). The units matter: \(b\) has response units per explanatory-variable unit, while \(s\) has response units.

The sign of \(r\) describes the direction of the linear association, and its magnitude summarizes the strength of that association. In simple linear regression with an intercept, \(r^2\) is the proportion of variation in the response accounted for by the linear model. Because \(r^2\) is squared, it does not show the direction of the association. The residual standard deviation \(s\) summarizes the typical size of the residuals. As discussed in “Effect of Unusual Points on \(r\)-squared and \(s\),” these values describe different aspects of fit.

Key comparison: Fit the model with all observations, then refit it after removing only the candidate observation. Compare corresponding output values and describe which changed, by how much, and what that change means in context. Do not infer influence from a single output value in isolation.

A Routine for Reading the Two Fits

1
Locate the two model summaries.
Confirm which output includes all observations and which omits only the candidate point. Check that the same variables and response scale were used in both fits.
2
Compare the fitted lines.
Subtract the slopes and intercepts, or describe their changes directly. Ask whether the fitted relationship changes meaningfully in the context and units.
3
Compare association and scatter.
Check \(r\) and \(r^2\) for changes in direction or strength, and check \(s\) for a change in the typical residual size. Remember that these statistics do not all measure the same feature.
4
Conclude about influence.
State which reported values change substantially and describe the consequence for the model. A high-leverage point may have little influence if the with-and-without results are similar.

A calculator may label the intercept and slope with letters that differ from \(a\) and \(b\). Check the calculator’s convention before interpreting them. For example, some displays list \(a\) as the intercept and \(b\) as the slope, while the equation is still written \(\hat{y}=a+bx\). Do not assume that a displayed \(a\) is the slope.

The number \(n\) will decrease by one in the refit. That change confirms that an observation was removed, but it is not itself evidence that the observation was influential. Also, a large change in \(r^2\) does not necessarily mean the slope changed substantially. Report the changes separately rather than using “the regression changed” without saying what changed.

Worked Examples: Reading and Comparing Output

Worked Example: A High-Leverage Point Changes the Fitted Line

Original AP-style question. An invented exercise records \(x\), the number of practice sessions, and \(y\), a performance score, for six cases: \((1,5),(2,8),(3,11),(4,14),(5,17),(10,40)\). Read the computer summaries below and decide whether the observation at \(x=10\) is influential.

FitIntercept \(a\)Slope \(b\)\(r\)\(r^2\)\(s\)\(n\)
All six observations-0.49183.91800.99330.98671.61966
Without \((10,40)\)231105

State. We will assess whether removing the observation at 10 practice sessions substantially changes the fitted line or other reported regression values.

Plan. Compare the intercept, slope, \(r\), \(r^2\), and \(s\) in the two output rows, keeping their interpretations and units in mind. The observation is far from the other \(x\)-values, so it has high leverage. High leverage alone, however, does not establish influence; the output comparison does.

Do. The five observations without the candidate point follow \(\hat{y}=2+3x\) exactly. For all six observations, the centered sums are \(S_{xx}=305/6\), \(S_{xy}=1195/6\), and \(S_{yy}=4745/6\). Thus the slope is

$$ b=\frac{S_{xy}}{S_{xx}} =\frac{1195/6}{305/6} =\frac{239}{61} \approx 3.9180 $$

The mean values are \(\bar{x}=25/6\) and \(\bar{y}=95/6\), so the intercept is

$$ a=\bar{y}-b\bar{x} =\frac{95}{6}-\frac{239}{61}\left(\frac{25}{6}\right) =-\frac{30}{61} \approx -0.4918 $$

For a check on the remaining output, \(r^2=S_{xy}^2/(S_{xx}S_{yy})=1{,}428{,}025/1{,}447{,}225\approx0.9867\), so \(r\approx0.9933\), positive because the slope is positive. Also, the sum of squared residuals is \(SSE=S_{yy}-S_{xy}^2/S_{xx}=640/61\). With \(n=6\), \(s=\sqrt{SSE/(n-2)}=\sqrt{160/61}\approx1.6196\) score points.

Conclude. Removing the high-leverage observation changes the slope from about 3.918 to 3 score points per practice session and the intercept from about \(-0.492\) to 2 score points. It also changes \(r\), \(r^2\), and \(s\). The observation is influential for this invented data set: including it produces a steeper fitted line and a small amount of residual scatter, whereas the remaining five cases lie exactly on a line.

Worked Example: A Vertical Outlier Changes Fit Statistics More Than Slope

Original AP-style question. In an invented study activity, \(x\) is the number of practice sessions and \(y\) is a score. Five cases are \((1,10),(2,12),(3,14),(4,16),(5,18)\). A sixth case, \((3,24)\), has a notably high score for its number of sessions. Use the output to describe what changes when this case is included.

FitIntercept \(a\)Slope \(b\)\(r\)\(r^2\)\(s\)\(n\)
All six observations9.666720.56950.32434.56446
Without \((3,24)\)821105

State. We want to describe how the added score of 24 affects the fitted relationship and the measures of association and residual scatter.

Plan. Compare the two output rows. Since the candidate case has \(x=3\), which is the center of the \(x\)-values, it is not far out in the explanatory-variable direction. We expect to check especially whether the slope changes, rather than assuming that the unusually high response must change every statistic.

Do. With all six observations, \(\bar{x}=3\) and \(\bar{y}=94/6=47/3\). The centered \(x\)-values are \(-2,-1,0,1,2,0\), so \(S_{xx}=10\). The cross-product sum is \(S_{xy}=20\), giving \(b=20/10=2\). The intercept is \(a=47/3-2(3)=29/3\approx9.6667\). Thus the slope stays at 2, but the intercept increases from 8 to about 9.6667.

The centered response sum of squares is \(S_{yy}=370/3\). Therefore

$$ r^2=\frac{S_{xy}^2}{S_{xx}S_{yy}} =\frac{20^2}{10(370/3)} =\frac{12}{37} \approx0.3243 $$

Since the slope is positive, \(r=\sqrt{12/37}\approx0.5695\). The residual sum of squares is \(SSE=S_{yy}-S_{xy}^2/S_{xx}=370/3-40=250/3\), so \(s=\sqrt{(250/3)/(6-2)}=\sqrt{125/6}\approx4.5644\) score points.

Conclude. The unusual response leaves the slope unchanged at 2 score points per practice session, but it changes the intercept, lowers \(r\) and \(r^2\), and increases the typical residual size from 0 to about 4.5644 score points. In this comparison, it has little effect on the slope but a substantial effect on the reported association and scatter. Describe those effects separately.

Worked Example: High Leverage Does Not Always Mean Influence

Original AP-style question. For an invented practice dataset, five cases are \((1,5),(2,8),(3,11),(4,14),(5,17)\), and a sixth is \((10,32)\). The last \(x\)-value is far from the others. Computer output shows that both the five-case and six-case fits have intercept 2, slope 3, \(r=1\), \(r^2=1\), and \(s=0\). Does the unusual \(x\)-value make the observation influential?

State. We will determine influence by comparing the reported results with and without the high-leverage case, not by judging its \(x\)-position alone.

Plan. Check each reported regression value and explain whether it changes. Then distinguish the case’s high leverage from its influence on the fitted results.

Do. All six cases lie exactly on \(\hat{y}=2+3x\): for the last case, \(2+3(10)=32\). Removing \((10,32)\) leaves the first five observations on the same line. The intercept and slope therefore remain 2 and 3, respectively. Since both fits are exact positive linear patterns, both have \(r=1\), \(r^2=1\), and \(s=0\). Only \(n\) changes, from 6 to 5.

Conclude. The point has high leverage because its explanatory-variable value is far from the others, but it is not influential for these reported regression values: removing it does not change the fitted line, correlation, coefficient of determination, or residual standard deviation. This deliberately exact example shows why high leverage and influence are not interchangeable.

Common Mistakes and AP Exam Tip

  • Calling a point influential just because it looks unusual. An unusual position can signal high leverage or a large residual, but influence is assessed by comparing the fits with and without the point.
  • Reporting only a change in \(r^2\). \(r^2\) summarizes the proportion of response variation accounted for by the linear model; it does not tell you the slope or the direction of the association. State changes in the other requested output values separately.
  • Mixing up \(r\) and \(r^2\). The sign of \(r\) gives the direction of the linear association; \(r^2\) cannot give that direction. In simple linear regression with an intercept, \(r^2=(r)^2\), as explained in “Testing Influence by Removing a Point.”
  • Ignoring units when describing the slope or \(s\). In the examples, the slope is in score points per practice session, while \(s\) is in score points. A careful answer names those units and the context.
  • Claiming every statistic must change in the same way. A point can alter \(r\) and \(s\) while leaving the slope nearly unchanged, or shift the slope substantially. Read each output measure on its own terms.
  • Calling the change meaningful without describing it. A full-credit explanation identifies what changed and gives the direction or approximate size of the change. For example: “Including the observation increases the predicted score change from 3 to about 3.918 points per practice session.”

Computer output is evidence to interpret, not a verdict by itself. Compare corresponding values, explain what a change means in context, and avoid using an output comparison to claim that the explanatory variable causes the response. As in the earlier tutorial “Critiquing a News Claim Based on a Regression,” keep the model’s description of an association separate from any causal claim.

Key takeaway: To interpret the effect of an unusual observation, compare regression output from the full data set with output from a refit that removes only that observation. Explain changes in the slope, intercept, \(r\), \(r^2\), and \(s\) separately and in context; unusual position alone does not establish influence.

Check Your Understanding

Use the ideas in this tutorial to answer each question.

  1. A regression output lists \(a=4\), \(b=1.8\), \(r=0.70\), \(r^2=0.49\), and \(s=6\). Which value gives the typical size of residuals, and what are its units?
  2. After removing a candidate point, the slope and intercept barely change, but \(r\) and \(s\) change substantially. What is a careful way to describe the point’s effect?
  3. Why does a change in \(n\) from 8 to 7 not, by itself, show that the removed observation was influential?
  4. A point has a very unusual \(x\)-value but lies exactly on the fitted line. What comparison would you use to decide whether it is influential?
  5. If \(r^2\) decreases after a point is removed, can you conclude that the direction of the association changed? Explain.