Tutorials › AP Statistics › High-Leverage Points in the x-Direction

Unusual points and model fit · Tutorial 963 of 1000

High-Leverage Points in the x-Direction

Use the distribution of explanatory-variable values to identify high-leverage points and explain why an extreme x-position gives a point potential to pull the regression line.

Intermediate 9 min read

What You'll Learn

  • Identify an x-value that is unusual relative to the other explanatory-variable values.
  • Use the center and spread of the x-values to judge whether a point has high leverage.
  • Explain why a point far from the center of the x-values can have potential to pull the fitted line.
  • Distinguish high leverage from a large residual and from actual influence on the fitted line.
  • Describe a high-leverage point in context without claiming more than the data show.

Look Along the x-Axis

In “Finding Outliers in the y-Direction,” you used residuals to find observations far vertically from the fitted line. This tutorial looks in a different direction: along the horizontal axis, where the explanatory-variable values \(x\) are shown. A point can be unusual because its \(x\)-value is far from the other \(x\)-values, even if it lies close to the line.

The earlier tutorial “Outlier, High-Leverage, and Influential Points Defined” named this feature high leverage. To assess it, compare a point’s explanatory-variable value with the distribution of \(x\)-values for the cases used to fit the line. An \(x\)-value far from the center and the main cluster has high leverage. It may be unusually small or unusually large; the important feature is its position relative to the other \(x\)-values, not whether the number looks large by itself.

Definition: A high-leverage point has an explanatory-variable value that is unusually far from the center of the observed \(x\)-values. Because it is far out along the \(x\)-axis, it has the potential to pull or rotate the fitted regression line toward itself.

“Potential” matters. An unusual \(x\)-value gives a point the opportunity to affect the line, but it does not prove that the fitted line changes substantially. The response value \(y\), the rest of the data, and the actual line all matter too. A high-leverage point is not automatically an outlier in the y-direction, and neither label alone establishes influence.

Judge an x-Value Relative to the Data

Begin by looking at a scatterplot or listing the observed \(x\)-values. Find the general cluster and its center. The mean \(x\)-value, \(\bar{x}\), can help describe that center, but do not use a fixed distance from \(\bar{x}\) as a universal cutoff. High leverage is a relative, visual judgment about whether a value is unusually far from the rest.

Notice both ends of the distribution. A point far to the left can have high leverage just as a point far to the right can. Also consider how widely the ordinary \(x\)-values are spread. A value of 12 might be far from a group ranging from 2 to 8, but not especially far from a group ranging from 0 to 20.

1
Locate the candidate point.
Identify its explanatory-variable value and units. In a scatterplot, use the horizontal coordinate.
2
Compare it with the other x-values.
Look for a cluster, its approximate center, and the overall spread. Consider whether the candidate is separated from the rest on either end.
3
Describe the evidence and the limit.
If the x-value is unusually distant, identify the point as having high leverage. Explain that its position gives it potential to pull the line, not that it necessarily does so.

Why can a point far from the center matter to a line? A regression line represents a pattern across paired \(x\)- and \(y\)-values. A point at an extreme \(x\)-position can help determine the direction and steepness of that pattern because it sits far out along the horizontal axis. If its \(y\)-value does not follow the trend suggested by the central cluster, the line may have to adjust to accommodate it. The actual effect can only be assessed by examining how the fitted line behaves with and without the point, a separate question from identifying high leverage.

Worked Examples: Identifying High Leverage

Worked Example: A Student with an Unusual Practice Time

Original AP-style question. A music teacher records each student’s weekly practice time, \(x\), in hours, and performance score, \(y\). The observed practice times are 1, 2, 2, 3, 3, 4, 4, 5, 6, and 18 hours. Does the student who practiced 18 hours have high leverage?

Compare the x-values. Nine students practiced between 1 and 6 hours, while one practiced 18 hours. The mean practice time is

$$ \bar{x}=\frac{1+2+2+3+3+4+4+5+6+18}{10} =\frac{48}{10}=4.8\text{ hours}. $$

The value 18 hours is 13.2 hours above the mean: \(18-4.8=13.2\). It is also far beyond the next-largest value, 6 hours. Those comparisons support describing the 18-hour student as having high leverage relative to this sample’s practice times.

Interpret in context. The student’s practice time is isolated on the high end of the \(x\)-distribution. That horizontal position gives the observation potential to pull the fitted line toward the student’s practice-time and performance-score pair. We cannot decide from the \(x\)-values alone whether the student’s score actually changes the line substantially. The score and the fitted line matter as well.

Worked Example: The Same x-Value in Two Distributions

Original AP-style question. Two separate studies examine weekly hours spent using an educational app and a response such as a learning measure. In Study A, the observed app-use hours range from 2 to 8, and one student reports 12 hours. In Study B, the observed hours range from 0 to 20, including a student who reports 12 hours. Is 12 hours high leverage in both studies?

Study A. The 12-hour value is 4 hours above the largest value among the other students, 8 hours. It is outside their observed range and separated from the group. Relative to the stated \(x\)-values, the student at 12 hours is a strong candidate for high leverage.

Study B. A value of 12 hours lies within the broader 0-to-20-hour spread. It is not at an extreme of the stated range. Without seeing the individual \(x\)-values, we should not call it high leverage just because 12 is a large number in isolation; it may be fairly ordinary within this study’s distribution.

Conclusion. High leverage depends on the point’s \(x\)-position relative to the other \(x\)-values in the same data set. The number 12 does not carry a fixed high-leverage label. The comparison with the distribution is the evidence.

Worked Example: An Extreme x-Value Close to the Line

Original AP-style question. A technician studies the relationship between daily production hours, \(x\), and electricity use, \(y\), in kilowatt-hours. Most days have production times between 2 and 8 hours. One day has \(x=20\) hours. The fitted line is \(\hat{y}=50+3x\), and the observed electricity use on the 20-hour day is 111 kWh. Is this point high leverage, and is it an outlier in the y-direction?

Assess its x-position. The 20-hour day is far beyond the 2-to-8-hour cluster. Its explanatory-variable value is unusually large relative to the other production times, so it has high leverage.

Calculate its residual. At \(x=20\), the line predicts

$$ \hat{y}=50+3(20)=110\text{ kWh}. $$

The residual is

$$ y-\hat{y}=111-110=1\text{ kWh}. $$

The observed use is only 1 kWh above the predicted use, so this point is close to the fitted line rather than far from it vertically. Its unusual production time supports a high-leverage description; its small residual does not support calling it a y-direction outlier. A point can have high leverage without being far from the line.

Worked Example: An Extreme x-Value Far from the Line

Original AP-style question. In the same kind of production setting, most days have production times between 1 and 6 hours. One day has \(x=16\) hours. For that day, a fitted line predicts \(\hat{y}=44\) kWh, but the observed electricity use is 25 kWh. Describe the point’s x-position and residual, and explain what can and cannot be concluded.

Identify the x-position. Sixteen hours is far beyond the cluster from 1 to 6 hours. This point has high leverage because its explanatory-variable value is unusually distant from the other observed production times.

Calculate the residual.

$$ y-\hat{y}=25-44=-19\text{ kWh}. $$

The negative residual means the observed electricity use is 19 kWh below the fitted prediction. The point is also far from the line in the y-direction, so it may be described as an outlier in that direction if 19 kWh is large compared with the typical residual scatter.

Explain the potential pull. This point is far to the right and below the prediction at its \(x\)-value. Its position gives it potential to pull the right end of the line downward relative to the pattern of the other days. But the high-leverage label and the residual do not, by themselves, quantify how much the fitted line changes. That requires comparing the fitted model with the point included to the fit with it removed.

Keep High Leverage, Outliers, and Influence Separate

The three labels describe different evidence. High leverage concerns an unusual horizontal position. A y-direction outlier concerns an unusually large vertical residual. Influence concerns how much the fitted regression line changes when a point is removed. Keep the relevant comparison in view: compare \(x\)-values to assess leverage, residuals to assess vertical departure, and fitted lines with and without a point to assess influence.

A point may have high leverage and a small residual, as in the 20-hour production example. A point may have a large residual but an ordinary \(x\)-value. Or a point may have both an unusual \(x\)-value and a large residual. These descriptions are compatible, but one does not automatically establish another.

In a scatterplot, high leverage is visible as a point far left or right of the main group of \(x\)-values. In a residual plot, use the horizontal coordinate to look for an unusually distant \(x\)-value; use the vertical coordinate to judge the residual. A point near the zero-residual line can still have high leverage if it is horizontally isolated.

Key takeaway: Identify high leverage by comparing a point’s \(x\)-value with the center and spread of the other explanatory-variable values. An extreme \(x\)-position gives the point potential to pull the line, but only a with-versus-without comparison shows whether it is influential.

Common Mistakes and AP Exam Tip

  • Calling any large number high leverage. A value is unusual only relative to the other \(x\)-values in its data set. State what the rest of the \(x\)-distribution looks like.
  • Checking only the right-hand end. An unusually small \(x\)-value can also be far from the center and have high leverage. Look at both ends of the horizontal distribution.
  • Using the residual to decide leverage. Residuals measure vertical distance. They do not show whether the \(x\)-value is unusual.
  • Calling a high-leverage point influential without checking. High leverage means potential to affect the line, not proof of a substantial change. Compare fits with and without the point before making an influence claim.
  • Claiming a universal cutoff. There is no single distance from \(\bar{x}\) that labels every point as high leverage. Describe the point’s position relative to the observed \(x\)-values and their spread.
  • Leaving out context. Say, for example, “The 18-hour practice time is far above the other students’ 1-to-6-hour times, so this student has high leverage,” rather than merely writing “18 is an outlier.”

A strong AP response identifies the explanatory-variable value, compares it with the rest of the \(x\)-values, and states the limited conclusion: the point has potential to pull the line because it is far from the center of the \(x\)-distribution. Do not claim that it actually changes the fitted line unless that change has been assessed.

Check Your Understanding

Use each point’s position relative to the other explanatory-variable values to assess leverage. Keep horizontal position separate from residual size and influence.

  1. A set of observations has \(x\)-values mostly between 10 and 16. One observation has \(x=31\). What evidence supports describing it as high leverage?
  2. A point has an \(x\)-value near the center of the data but a large negative residual. Which label is supported by the residual, and does that alone show high leverage?
  3. A point is far to the left of the main cluster in a scatterplot but lies close to the fitted line. What can you say about leverage and y-direction outlier status?
  4. Why might the value \(x=12\) be high leverage in one data set but not in another?
  5. What comparison is needed to determine whether a high-leverage point actually has a substantial effect on the fitted line?