One Point Can Change the Summary
In “Correlation Measures Only Linear Association,” you learned that \(r\) summarizes the direction and strength of the linear association between two quantitative variables. But a single unusual point can have a large effect on that summary. It may pull \(r\) closer to zero, push it farther from zero, or even change its sign.
The effect depends on where the point lies relative to the overall pattern. A point far from the others that continues their linear trend can make the linear association look stronger. A point that lies far from the trend can make it look weaker. A scatterplot helps show what is happening; comparing \(r\) with and without the point makes the numerical change clear.
Why an Outlier Can Affect \(r\)
Correlation is calculated using all the paired observations, including their means and standard deviations. As you saw in “Calculating \(r\) From Standardized Values,” each observation contributes through how its \(x\)- and \(y\)-values differ from their respective means. A point that is far from the rest can affect those means and the overall calculation substantially.
One useful computational form of the correlation formula is:
Here, \(S_{xy}\) measures how the paired deviations in \(x\) and \(y\) vary together, while \(S_{xx}\) and \(S_{yy}\) measure variation in each variable. A point with both a large \(x\)-value and a large \(y\)-value may contribute differently from a point with a large \(x\)-value and a small \(y\)-value. The point also shifts the means used to find the deviations.
For hand calculations, these sums can be found from the raw data:
You do not need to calculate these sums by hand every time you use \(r\). They help explain why the correlation can change when a point is added or removed. In practice, examine the scatterplot first, then compare the calculated correlations.
Worked Examples
Worked Example: An Outlier Lowers a Positive Correlation
A fictional fitness program records weekly practice time, \(x\), in hours and a skills score, \(y\), in points. Four invented observations follow a perfect increasing line: \((1,1)\), \((2,2)\), \((3,3)\), and \((4,4)\). A fifth participant has values \((5,0)\), far below that pattern. Compare \(r\) for the four observations on the line with \(r\) for all five.
With the four observations on the line: Each \(y\)-value equals its \(x\)-value, so the points lie exactly on an increasing straight line. Therefore, \(r=1\). We can also verify this from the sums: \(\sum x=\sum y=10\), \(\sum x^2=\sum y^2=\sum xy=30\), and \(n=4\). Thus \(S_{xx}=30-100/4=5\), \(S_{yy}=5\), and \(S_{xy}=30-100/4=5\). So \(r=5/\sqrt{5(5)}=1\).
With all five observations: The totals are \(\sum x=15\), \(\sum y=10\), \(\sum x^2=55\), \(\sum y^2=30\), and \(\sum xy=30\). Then
Therefore, \(r=0/\sqrt{10(10)}=0\). The point \((5,0)\) lowers \(r\) from \(1\) to \(0\): it contradicts the strong positive linear pattern of the other four observations. For these invented data, the correlation with all five points does not capture the increasing pattern among the four participants without the outlying point.
Worked Example: An Outlier Raises \(r\)
A fictional technology class records the number of practice sessions, \(x\), and a troubleshooting score, \(y\), in points. Four invented observations are \((-2,1)\), \((-1,-1)\), \((1,1)\), and \((2,-1)\). A fifth observation, \((5,5)\), lies well away from the group. Compare the correlations.
Without the fifth observation: The sums are \(\sum x=0\), \(\sum y=0\), \(\sum x^2=10\), \(\sum y^2=4\), and \(\sum xy=-2\). With \(n=4\), this gives \(S_{xx}=10\), \(S_{yy}=4\), and \(S_{xy}=-2\). Thus
With the fifth observation: The totals become \(\sum x=5\), \(\sum y=5\), \(\sum x^2=35\), \(\sum y^2=29\), and \(\sum xy=23\), with \(n=5\). Therefore,
The correlation is
Here the outlier raises \(r\), from about \(-0.316\) to about \(0.671\), and changes its sign. The point’s high \(x\)- and high \(y\)-values contribute to an increasing linear tendency in the combined data. Notice that “raises” refers to the numerical value of \(r\); the conclusion comes from comparing the two values and considering the point’s location.
Worked Example: A Numerical Increase Can Mean Weaker Correlation
A fictional community garden compares weekly watering time, \(x\), in minutes with a plant-stress rating, \(y\), in points. Four invented observations are \((1,4)\), \((2,3)\), \((3,2)\), and \((4,1)\), which lie on a decreasing line. An additional observation, \((5,10)\), is far above that pattern. Compare \(r\) before and after adding it.
Without the additional observation: The sums are \(\sum x=10\), \(\sum y=10\), \(\sum x^2=30\), \(\sum y^2=30\), and \(\sum xy=20\). With \(n=4\),
So \(r=-5/\sqrt{5(5)}=-1\), indicating a perfect negative linear association in these four observations.
With all five observations: The totals are \(\sum x=15\), \(\sum y=20\), \(\sum x^2=55\), \(\sum y^2=130\), and \(\sum xy=70\). Then
Therefore,
The numerical value of \(r\) rises from \(-1\) to about \(0.447\). Yet the strength of the linear association, judged by \(\lvert r\rvert\), falls from \(1\) to about \(0.447\), and the sign changes from negative to positive. This example shows why it is useful to report both the signed values and what changed in the scatterplot. A numerical increase in \(r\) does not always mean a stronger linear association.
A Reliable Way to Describe the Change
When asked how an outlier changes correlation, avoid deciding from the word “outlier” alone. Compare the two values of \(r\), then use the scatterplot to explain the change. Keep the sign and strength distinct: the sign describes the direction of the linear association, while the magnitude \(\lvert r\rvert\) describes how strong that linear association is.
Use the scatterplot to describe where the outlier falls relative to the other observations and their linear trend.
State \(r\) with the point and \(r\) without it. Check whether the numerical value rises or falls and whether the sign changes.
Explain whether the point strengthens or weakens the linear association, using the change in \(\lvert r\rvert\), and relate that change to the point’s position.
This is a descriptive comparison of the observed data, not a rule that unusual observations should be deleted. A point may be unusual and still be a valid observation. If it is a recording or measurement error, investigate and document the issue. Otherwise, report the analysis using all valid observations; a comparison that omits the point can be useful for showing its influence, but does not by itself justify excluding it.
Common Mistakes and AP Exam Tips
- Assuming every outlier lowers \(r\). A point far from the others can raise or lower \(r\), depending on how it sits relative to the linear pattern. State what the actual comparison shows.
- Confusing a numerical increase with greater strength. For example, \(r\) can rise from \(-1\) to \(0.447\), while \(\lvert r\rvert\) decreases. To discuss strength, compare absolute values; to discuss direction, compare signs.
- Describing \(r\) without checking the scatterplot. A change in \(r\) is easier to explain when you identify whether the point follows or contradicts the overall linear pattern. As in “Describing a Scatterplot With DUFS,” include unusual features rather than relying on one summary number.
- Claiming that an outlier should be removed because it changes the answer. Influence is not a reason by itself to discard a valid observation. Investigate whether the point is an error and explain any decision to exclude it.
- Making a causal claim from the change. Comparing \(r\) with and without an observation describes how the observed linear association changes. It does not show that one variable causes the other to change.
A strong response names both correlations and explains the change precisely. For example: “Including the unusually high practice-time and score observation changes \(r\) from negative to positive. The point pulls the combined data toward a positive linear association; the scatterplot shows it lies far from the pattern of the other observations.” Use the context and the actual values in your own answer.
Check Your Understanding
Use the scatterplot and the two correlation values together when answering.
- A data set has \(r=0.82\) with all points and \(r=0.51\) after omitting one outlying point. Did the point raise or lower \(r\)? What happened to the strength of the linear association?
- A correlation changes from \(r=-0.70\) to \(r=0.20\) when an observation is included. Did the numerical value rise or fall? Did the direction change? Did the strength increase or decrease?
- Why is it not enough to know that a point is far from the other observations to predict how it will affect \(r\)?
- What should you examine in the scatterplot to explain why an outlier changes \(r\)?
- A point is a valid observation but substantially changes \(r\). Is that, by itself, a good reason to remove it? Explain.