Tutorials › AP Statistics › Spotting Outliers in Bivariate Data

Scatterplots and association · Tutorial 807 of 1000

Spotting Outliers in Bivariate Data

Learn to tell whether a point is unusual because it falls far from a scatterplot’s pattern or simply because one of its values is extreme.

Intermediate 9 min read

What You'll Learn

  • Distinguish an outlier in a bivariate pattern from an extreme value in one variable.
  • Check a point against the pattern near its explanatory-variable value.
  • Recognize points that are extreme in the explanatory variable but still follow the trend.
  • Describe a point’s position in context without assuming it is an error or has a particular cause.
  • Explain why unusual points deserve attention but should not be removed automatically.

Unusual Values Are Not Always Outliers

In “Judging Strength of an Association,” you learned to look at how closely points follow a scatterplot’s overall pattern. Now consider what happens when one point seems unusual. It may stand far from the pattern, or it may simply have an unusually large or small value on one axis while still fitting the relationship between the variables.

To spot a bivariate outlier, judge the point in relation to the overall pattern. A point can have an extreme value of the explanatory variable, \(x\), or the response variable, \(y\), without being an outlier in the bivariate data. What matters is whether the point is far from where the pattern would lead you to expect it to fall.

Definition: A bivariate outlier is a point that falls far from the overall pattern in a scatterplot. A point that is unusual in \(x\) or \(y\) alone is not necessarily a bivariate outlier; check how it fits the relationship between the variables.

One helpful question is: “Compared with other points at similar \(x\)-values, is this point far above or below the pattern?” For a roughly linear pattern, imagine the general straight path through the points. For a curved pattern, compare the point with the curve. Do not judge a point only by whether its \(x\)- or \(y\)-coordinate is among the largest or smallest.

This is a visual judgment, not a fixed numerical rule. A point’s apparent distance can depend on the scale and shape of the plot, so read both axes and consider the full cloud of points. If the pattern is weak or unclear, it may be difficult to decide whether one observation stands apart.

A Three-Question Check

When a point catches your eye, separate two questions: “Is one coordinate unusual?” and “Does the point fall far from the pattern?” The first describes its position on an axis; the second describes its fit to the bivariate relationship. These answers can differ.

1
Describe the point’s coordinates.
Is its \(x\)-value near an endpoint of the observed \(x\)-values? Is its \(y\)-value unusually high or low compared with the other response values?
2
Identify the overall pattern.
Describe the general form and direction before deciding whether the point fits. If the pattern is curved, compare the point with the curve rather than an imagined straight line.
3
Judge the point relative to the pattern.
Ask whether the point lies far from the pattern at its \(x\)-value. State whether it appears to be a bivariate outlier, merely extreme in one variable, or both.

A point at an extreme \(x\)-value can be important to inspect because it lies at the edge of the observed explanatory-variable range. But if it continues the pattern, it is not automatically a bivariate outlier. Likewise, a point with a very large \(y\)-value could fit a rising pattern if its \(x\)-value is also large.

Worked Example: A Point Far Above a Rising Pattern

A fictional school garden records weekly sunlight, \(x\), in hours and plant growth, \(y\), in centimeters. The invented observations are shown below. Suppose the scatterplot shows a clear, roughly straight, rising pattern among most of the plants.

PlantSunlight, \(x\) (hours)Growth, \(y\) (cm)
A14
B26
C38
D410
E512
F614
G419

Describe the unusual coordinate. Plant G has a growth value of 19 cm, much higher than the other listed growth values. Its sunlight value, 4 hours, is not at an endpoint; other plants were recorded at nearby sunlight levels.

Compare it with the pattern. At 4 hours of sunlight, the other plant in the table grew 10 cm, and the overall pattern places growth near that level for this \(x\)-value. Plant G, at 19 cm, is far above the rising pattern rather than continuing it.

Conclusion. Plant G appears to be a bivariate outlier: its growth is unusually high relative to the pattern for plants receiving about 4 hours of sunlight. It is also extreme in \(y\), but the key reason to call it a bivariate outlier is that it falls far from the relationship.

The plot alone does not tell us why Plant G differs. The measurement could be correct, or it could reflect a recording problem or some other feature of that plant. The appropriate first step is to check the observation and its context, not to assume a cause.

Extreme in One Variable, but Still on the Pattern

A point can be far to the right or left of the others and still fit the overall trend. In that case, its explanatory-variable value is unusual compared with the rest of the sample, but the point is not far from the bivariate pattern. The same distinction applies to an unusually high or low response value: if the pattern predicts a high or low response at that \(x\)-value, the point may fit well.

This distinction matters because a scatterplot is about paired values. Looking at the \(x\)-values alone or the \(y\)-values alone discards how the values occur together. Keep each observation as a point \((x,y)\), as in “Constructing a Scatterplot by Hand,” and judge the pair against the pattern.

Worked Example: An Extreme Explanatory Value That Fits

A fictional environmental club compares the number of hours a rain barrel is used each week, \(x\), with the number of liters collected, \(y\). These invented observations form a clear rising pattern.

WeekUse, \(x\) (hours)Water collected, \(y\) (liters)
1212
2314
3416
4518
5620
61232

Notice the extreme coordinate. Week 6 has \(x=12\) hours, while the other \(x\)-values range from 2 to 6 hours. This point is far to the right of the rest, so its explanatory-variable value is extreme in this sample.

Check the bivariate pattern. The observations follow a rising straight pattern: as use increases by 1 hour, the collected amount increases by 2 liters. Extending that pattern from the earlier observations places the point at 12 hours near 32 liters. The Week 6 point is consistent with that trend.

Conclusion. Week 6 is extreme in \(x\), but it does not appear to be a bivariate outlier because it lies along the overall rising pattern. Calling it an outlier just because it is far to the right would confuse an unusual coordinate with a point far from the relationship.

Worked Example: Extreme in Both Coordinates and Still Consistent

A fictional recreation center records the number of visitors, \(x\), and the number of towels used, \(y\), on several days. Suppose the plotted points show a steady rising pattern. The values below are invented.

DayVisitors, \(x\)Towels used, \(y\)
11020
22035
33050
44065
55080
690140

Notice both coordinates. Day 6 has the largest number of visitors and the largest number of towels used. It is extreme in both variables compared with the other days.

Judge its position relative to the pattern. The other points show that towel use rises by 15 when visitors increase by 10. From 50 visitors and 80 towels, extending this same pattern to 90 visitors gives 140 towels. Day 6’s point, \((90,140)\), continues the pattern rather than falling far away from it.

Conclusion. Day 6 is extreme in both \(x\) and \(y\), but it is not a bivariate outlier based on this scatterplot because it fits the rising pattern. The coordinate values alone do not determine whether the point is an outlier.

Why Outliers Deserve a Closer Look

A point far from the pattern can affect the visual impression of direction, form, and strength. For example, a single point separated from a cloud may make an association look stronger, weaker, or different in direction than the main group suggests. That is a reason to notice and describe it rather than silently ignore it.

An extreme \(x\)-value that follows the pattern also deserves attention, even though it is not a bivariate outlier. Because it lies far from the other \(x\)-values, it can stretch the range shown in the scatterplot and may have a noticeable effect on how the pattern looks. For this tutorial, the important distinction is descriptive: an unusual \(x\)-value is not the same thing as a point far from the pattern.

Do not remove a point just because it looks unusual. First check whether the data were entered or measured correctly. If the point is valid, retain it when describing the observed data and explain how it differs from the overall pattern. A plot by itself cannot determine whether a point is an error, explain its cause, or justify excluding it.

Common Mistakes and AP Exam Tips

  • Calling the largest \(y\)-value an outlier automatically: Check whether a high response is expected at that point’s \(x\)-value. A high value can fit a rising pattern.
  • Calling a far-left or far-right point an outlier automatically: Describe it as extreme in \(x\), then ask whether it follows the overall pattern.
  • Ignoring the pairing: Do not inspect the two variables separately and lose the fact that each \(x\)-value is paired with its own \(y\)-value. Judge the plotted point \((x,y)\).
  • Comparing a point with the wrong pattern: If the association is curved, compare the point with the curve. A point can look far from an imagined straight line but fit the actual curved pattern.
  • Assuming an outlier is a mistake: A point that falls far from the pattern may be a valid observation. Say what the plot shows; do not invent a reason for the point.
  • Removing a point without explanation: Check unusual observations, but do not discard them solely because they change the appearance of the plot.

For a strong AP response, identify the point’s unusual coordinate, then explain its relationship to the overall pattern. For example: “The observation at 12 hours is extreme in the explanatory variable, but it is not a bivariate outlier because it continues the rising pattern.” Or: “The observation at 4 hours appears to be a bivariate outlier because its growth value lies far above the pattern at nearby sunlight levels.” These statements give a visual reason, not just a label.

Key takeaway: A bivariate outlier is far from the overall pattern, not merely far out on an axis. Describe unusual \(x\)- or \(y\)-values separately, then judge whether the paired point fits the pattern at its location.

Check Your Understanding

For each situation, distinguish an extreme coordinate from a point that falls far from the bivariate pattern.

  1. A scatterplot has a clear rising pattern. A point has the largest \(x\)-value and lies close to the continuation of the pattern. How should you describe it?
  2. At an ordinary \(x\)-value, one point lies far above the pattern and has a much larger \(y\)-value than nearby points. What makes it a possible bivariate outlier?
  3. A point has the largest \(y\)-value, but it is also at the largest \(x\)-value in a clear rising pattern and lies close to the trend. Is its extreme \(y\)-value enough to call it a bivariate outlier? Explain.
  4. Why should you compare a point with a curve rather than an imagined straight line when the scatterplot has a curved pattern?
  5. A point falls far from the overall pattern. What should you check before deciding whether to remove it from the data?