A High \(r\) Is Not a Check of Model Fit
In “Restricted Range and Correlation,” you saw that the observations included in a data set can affect the correlation. Another important point is that even a high \(r\) does not guarantee that a linear model describes the data well. A curved pattern can rise steadily from left to right, so its points may have a strong positive linear association while still bending away from a straight line.
As in “Correlation Measures Only Linear Association,” \(r\) summarizes the direction and strength of a linear association. It is not a score for how well a straight line captures every feature of the relationship. A scatterplot helps reveal the overall form. A residual plot provides another way to check whether a linear model leaves a systematic pattern behind.
A residual plot displays the residuals vertically against the explanatory variable \(x\) (or, in some settings, against the predicted values \(\hat{y}\)). Include a horizontal reference line at residual \(0\). If the linear model captures the main pattern, residuals should generally scatter around that line without a visible systematic shape. A curved pattern in the residuals indicates that the straight line is missing structure.
Worked Example: A High Correlation Hides a Curve
Suppose a fictional calibration exercise records a device setting \(x\) and a sensor response \(y\), in response units. The invented data follow an upward-curving pattern. Calculate \(r\), find the least-squares regression line, and inspect the residuals.
| Setting \(x\) | Response \(y\) |
|---|---|
| 1 | 1 |
| 2 | 4 |
| 3 | 9 |
| 4 | 16 |
| 5 | 25 |
The means are \(\bar{x}=3\) and \(\bar{y}=11\). The sums of squares and cross-products are \(S_{xx}=10\), \(S_{yy}=374\), and \(S_{xy}=60\). For example, the \(x\)-deviations are \(-2,-1,0,1,2\), and the \(y\)-deviations are \(-10,-7,-2,5,14\). The sum of products of corresponding deviations is \(20+7+0+5+28=60\).
The value \(r\approx0.9811\) indicates a very strong positive linear association. Now find the least-squares regression line. Its slope is \(b=S_{xy}/S_{xx}=60/10=6\), and its intercept is \(a=\bar{y}-b\bar{x}=11-6(3)=-7\). Thus, the predicted response is \(\hat{y}=-7+6x\).
For each setting, subtract the predicted response from the observed response. For example, at \(x=1\), the prediction is \(\hat{y}=-7+6(1)=-1\), so the residual is \(1-(-1)=2\). The full set of predictions and residuals is:
| \(x\) | Observed \(y\) | Predicted \(\hat{y}\) | Residual \(y-\hat{y}\) |
|---|---|---|---|
| 1 | 1 | -1 | 2 |
| 2 | 4 | 5 | -1 |
| 3 | 9 | 11 | -2 |
| 4 | 16 | 17 | -1 |
| 5 | 25 | 23 | 2 |
The residuals are positive at both ends and negative in the middle. A residual plot would show a U-shaped pattern rather than random scatter around zero. The line underpredicts the response at the lowest and highest settings, where residuals are positive, and overpredicts it at the middle settings, where residuals are negative.
Conclusion: The device setting and sensor response have a very strong positive linear association in these observations, but the residual pattern shows that a straight line does not capture the systematic curve. The high correlation alone would not reveal that problem.
Reading Residuals as a Pattern
A residual is not just a numerical error to calculate; the sequence of residuals across \(x\)-values can diagnose how a model misses. If residuals are mostly positive at low and high \(x\)-values and mostly negative in the middle, that U-shape suggests upward curvature. The reverse arrangement—a positive middle and negative ends—suggests a curve bending downward relative to the line.
Residuals also make the direction of prediction errors precise. Since \(e=y-\hat{y}\), a positive residual says the actual response was greater than the model predicted. A negative residual says the actual response was less than predicted. This is why keeping the subtraction order straight matters.
Use the least-squares regression line to calculate a predicted response \(\hat{y}\) for each observed \(x\).
For each case, subtract the prediction from the observed response: \(e=y-\hat{y}\).
Use a horizontal scale for \(x\), a vertical scale for residuals, and mark the reference line at zero.
Random-looking scatter around zero supports a linear model as a useful summary. A curve or another systematic pattern signals that the line misses a feature of the data.
The residual plot is a diagnostic, not a proof. A small data set may not make a pattern easy to see, and a plot without an obvious curve does not establish that the model is correct for every purpose. Consider the scatterplot, the residual plot, the context, and the range of \(x\)-values together.
Worked Example: Residual Signs Reveal Underprediction and Overprediction
A fictional greenhouse trial records a daily light setting \(x\) and plant response \(y\), in response units. The regression line for the observations is \(\hat{y}=-0.6+4.2x\). The residuals below are invented for illustration. Calculate them and decide whether the line captures the pattern well.
| \(x\) | Observed \(y\) | Predicted \(\hat{y}\) |
|---|---|---|
| 0 | 1 | -0.6 |
| 1 | 3 | 3.6 |
| 2 | 6 | 7.8 |
| 3 | 11 | 12.0 |
| 4 | 18 | 16.2 |
Subtract the prediction from the observation in every row. For example, at \(x=0\), \(e=1-(-0.6)=1.6\). At \(x=3\), \(e=11-12.0=-1.0\). The complete residual list, in order of increasing \(x\), is \(1.6,-0.6,-1.8,-1.0,1.8\).
The first and last residuals are positive, while the middle residuals are negative. Thus, the line underpredicts at the low and high ends of the observed settings and overpredicts across the middle. This curved residual pattern shows that the linear model misses a systematic feature, even though the points overall rise as \(x\) increases.
A careful conclusion is about the model and the observed range: “The residuals show an upward-curving pattern, so the straight-line model does not fully describe the relationship between light setting and plant response in these observations.” The residual plot does not by itself establish what model should replace the line.
Worked Example: Residuals Without an Obvious Curve
A fictional delivery service records route length \(x\), in kilometers, and a time measure \(y\), in minutes. The six invented pairs are \((1,4),(2,8),(3,9),(4,12),(5,12),(6,15)\). Find the regression line and residuals, then assess whether the residuals show a clear curve.
Here \(\bar{x}=3.5\) and \(\bar{y}=10\). The sums are \(S_{xx}=17.5\), \(S_{xy}=35\), and \(S_{yy}=74\). Therefore, the slope is \(b=35/17.5=2\), and the intercept is \(a=10-2(3.5)=3\). The least-squares line is \(\hat{y}=3+2x\).
The correlation is very strong and positive. The model predicts \(5,7,9,11,13,15\) minutes for the six route lengths. Subtracting these predictions from the observed times gives residuals \(-1,1,0,1,-1,0\) minutes.
These residuals are small and fluctuate around zero, with no clear U-shape or inverted U-shape. For these observations, the residual plot does not reveal the systematic curvature seen in the earlier examples, so the linear model is a reasonable summary of the pattern. This is not proof that the relationship is exactly linear; it is a conclusion about what these data show.
Common Mistakes and AP Exam Tips
- Treating a high \(r\) as proof of a good linear model. A high correlation describes a strong linear trend, not the absence of curvature. Refer to what the scatterplot or residual plot actually shows.
- Reversing the residual subtraction. Use observed minus predicted, \(e=y-\hat{y}\). A positive residual means the observation is above the prediction; a negative residual means it is below.
- Checking only the size of residuals. Small residuals can still form a systematic curve. Inspect their arrangement across \(x\), not just whether their values seem close to zero.
- Calling every nonzero residual a problem. Observed values commonly differ from predictions. The important warning is a systematic pattern, such as curvature, rather than the mere presence of positive and negative residuals.
- Claiming that a residual plot proves a model is correct. A plot without obvious structure supports using a line as a summary for the observed data; it does not prove the relationship is linear in every setting or outside the observed range.
A full-credit response connects the visual evidence to the model: “Although \(r\) is high, the residuals are positive at the ends and negative in the middle, forming a U-shaped pattern. The linear model underpredicts at the ends and overpredicts in the middle, so it does not adequately capture the curvature in the observed relationship.”
Check Your Understanding
Use correlation, regression predictions, and residual patterns to assess whether a linear model is appropriate.
- What does a positive residual mean when residuals are calculated as \(y-\hat{y}\)?
- A residual plot has negative residuals in the middle and positive residuals at both ends. What pattern does this suggest, and how does the line tend to predict at the ends?
- Why can a data set have a high positive \(r\) and still be poorly described by a straight line?
- What residual pattern would support using a linear model as a summary, and why does it not prove the relationship is exactly linear?
- A student says, “All the residuals are small, so there cannot be a problem with the model.” What feature of the residual plot should the student check before concluding?