Four Pieces of Evidence, One Model-Fit Judgment
A scatterplot, a residual plot, \(r\), and \(r^2\) offer related but different information about a relationship. Combining them helps you judge whether a straight line is a useful description—not just whether the variables move together. As in “Residual Plot Versus Scatterplot Diagnosis,” the scatterplot shows the original data and the residual plot shows how the fitted line’s errors behave. The statistics \(r\) and \(r^2\) add numerical summaries, but neither replaces the plots.
The key is to avoid treating any one clue as a verdict. A strong positive \(r\) describes a strong positive linear association, but a nonlinear relationship can also have a high \(r\). A high \(r^2\) says that the linear model accounts for a large proportion of the observed variation in \(y\), but it does not show whether the remaining errors have a pattern. A residual plot is useful for that diagnostic question.
What the Statistics Add
The correlation \(r\) describes the direction and strength of a linear association between two quantitative variables. Its sign gives the direction, and values closer to \(-1\) or \(1\) indicate a stronger linear association. Correlation does not describe every feature of the data: in particular, it does not reveal curvature or changing residual spread.
In simple linear regression with an intercept, \(r^2\) is the coefficient of determination. It is the proportion of the observed variation in the response variable \(y\) that is accounted for by the least-squares regression of \(y\) on \(x\). As covered in “Effect of Unusual Points on \(r\)-squared and \(s\),” \(r^2\) does not show the direction of the association. It also does not mean that the model predicts that proportion of cases correctly, prove causation, or show that residuals are pattern-free.
For example, \(r=0.90\) and \(r=-0.90\) both give \(r^2=0.81\). The first describes a strong positive linear association; the second describes a strong negative one. In either case, \(r^2=0.81\) means that 81% of the observed variation in \(y\) is accounted for by the linear regression model for these data. It is not a statement about cause or individual prediction accuracy.
Use the statistics as a compact summary of linear association, then return to the plots to check its shape and errors. A high \(r\) or \(r^2\) can coexist with a curved residual pattern, and a single unusual point can affect these summaries. Consider the full evidence and describe any conflicts among it.
A Combined Diagnosis
State the direction, form, and strength of the overall relationship in context. Look for curvature, clusters, and unusual points.
Check for a roughly pattern-free band around zero. A curve, fan, or systematic run of residuals on one side of zero warns that a straight line may not capture the data well.
Report the direction and strength of the linear association using \(r\), and describe the proportion of variation accounted for using \(r^2\), when relevant.
Say whether the combined evidence supports a linear model. If the evidence conflicts, explain which plot or feature raises concern instead of letting a large statistic settle the question.
“Pattern-free” does not mean every residual equals zero or that the points are perfectly on the line. It means there is no obvious systematic structure in the residuals, such as a curve or fan. Even then, the model is a reasonable description rather than a guarantee of accurate predictions for every case. The earlier tutorials on unusual points also matter: investigate a point that may strongly affect the line or correlation, but do not remove a valid observation merely because it changes the result.
Worked Examples: Putting the Evidence Together
Worked Example: Strong Linear Evidence in a Small Data Set
Original AP-style question. In an invented tutoring exercise, \(x\) is weekly practice time in hours and \(y\) is a student’s assessment score in points. The six observed pairs are \((1,14),(2,14),(3,20),(4,23),(5,23),(6,29)\). A scatterplot looks approximately straight, and the residuals from the least-squares line are \(1,-2,1,1,-2,1\) points. Given \(r\approx0.9640\) and \(r^2\approx0.9292\), assess whether a linear model is reasonable.
State. We want to judge whether a straight line reasonably describes the association between weekly practice time and assessment score for these six students.
Plan. Combine the reported scatterplot shape, the residual pattern, and the numerical summaries. A linear model is better supported if the scatterplot is approximately straight and the residuals do not show an obvious systematic bend or changing spread. The values of \(r\) and \(r^2\) summarize the strength of the linear association and variation accounted for.
Do. The given statistics are consistent: \(r^2=(0.9640)^2\approx0.9293\), which is about \(0.9292\) using the unrounded \(r\). Thus \(r^2\approx0.9292\), or about 92.92%. The residuals are relatively small compared with the 15-point range of observed scores. In their order by practice time, their signs switch several times, with two consecutive positive residuals, so they are not literally pattern-free; however, six residuals of this kind do not show a clear sustained curve or a widening fan. The scatterplot’s approximately straight pattern provides additional support.
Conclude. The scatterplot, residuals, and high positive \(r\) and \(r^2\) together support using a linear model as a reasonable description of these data. The residuals do vary in sign, so do not claim that they are exactly pattern-free. The conclusion is about the observed association, not a claim that practice time causes higher scores.
Worked Example: High \(r\) and \(r^2\) Do Not Remove Curvature
Original AP-style question. In an invented study exercise, \(x\) is weekly practice time in hours and \(y\) is an assessment score in points. The five scores at \(x=1,2,3,4,5\) hours are \(51,54,59,66,75\). The scatterplot bends upward. Find the least-squares line, \(r\), and \(r^2\), then decide whether a linear model captures the pattern well.
State. We will compare the numerical summaries with the scatterplot’s form and the residuals from the fitted line.
Plan. Find the least-squares line using the means and sums of squares. Then calculate \(r\) and \(r^2\), and inspect the residuals in order of \(x\). A visible bend in the residuals would warn that a straight line misses systematic structure, even if \(r\) is large.
Do. The means are \(\bar{x}=3\) hours and \(\bar{y}=61\) points. The centered \(x\)-values are \(-2,-1,0,1,2\), so \(S_{xx}=(-2)^2+(-1)^2+0^2+1^2+2^2=10\). The centered \(y\)-values are \(-10,-7,-2,5,14\), giving \(S_{yy}=100+49+4+25+196=374\). The cross-products sum to \(S_{xy}=(-2)(-10)+(-1)(-7)+0(-2)+1(5)+2(14)=60\).
The fitted line is \(\hat{y}=43+6x\), with score points as the response. The correlation and coefficient of determination are:
Thus the linear model accounts for about 96.26% of the observed variation in scores. But the fitted values are \(49,55,61,67,73\), so the residuals are \(51-49=2\), \(54-55=-1\), \(59-61=-2\), \(66-67=-1\), and \(75-73=2\) points. The residuals are positive at both ends and negative in the middle—a curved pattern.
Conclude. The association is strongly positive, and \(r^2\) is high, but the upward bend in the scatterplot and curved residual pattern show that a straight line misses systematic structure. The high statistics summarize strong linear association; they do not establish that a linear model captures the form well.
Worked Example: An Unusual Point Shapes the Summary
Original AP-style question. In an invented equipment test, \(x\) is an operating setting and \(y\) is an output reading. The five pairs are \((1,10),(2,10),(3,10),(4,10),(5,30)\). Calculate \(r\), \(r^2\), and the fitted line, then use the plots and values to assess linear fit.
State. We want to determine whether one positive \(r\) and a moderate \(r^2\) adequately summarize the relationship, or whether the scatterplot and residual plot raise concerns.
Plan. Calculate the means and centered sums, find the least-squares line and residuals, and then interpret the statistics alongside the visible data pattern. Pay particular attention to the point at \(x=5\).
Do. Here \(\bar{x}=3\), \(\bar{y}=14\), \(S_{xx}=(-2)^2+(-1)^2+0^2+1^2+2^2=10\), and \(S_{yy}=4(16)+16^2=320\). The cross-product sum is \(S_{xy}=(-2)(-4)+(-1)(-4)+0(-4)+1(-4)+2(16)=40\). Therefore:
The correlation and its square are:
The fitted values are \(6,10,14,18,22\), so the residuals are \(4,0,-4,-8,8\) output units. The scatterplot shows four points at \(y=10\) followed by one much higher point at \(x=5\), rather than a consistent straight-line cloud. The residuals also show structure, including a large positive residual at the final point.
Conclude. Although \(r\) is positive and \(r^2=0.50\), those summaries do not make the linear model appropriate. The scatterplot and residuals reveal the unusual final point and a poor overall pattern for a straight line. Because an unusual point may affect the fit, investigate it in context; do not discard it without evidence that it is an error.
Common Mistakes and AP Exam Tip
- Using a large \(r\) as proof of linear form. Correlation summarizes linear association, not curvature. A full-credit explanation checks the scatterplot and residual plot for evidence of a bend.
- Treating \(r^2\) as the percentage of correct predictions. It is the proportion of observed variation in \(y\) accounted for by the regression model. State that interpretation in context, with the response variable named.
- Ignoring residuals because they are small. Residuals can be small relative to the response scale and still show a systematic pattern. Describe their arrangement, not just their size.
- Calling residuals pattern-free when signs alternate or form a sequence. Look at their order across \(x\). A systematic run, curve, or changing spread must be acknowledged even when \(r\) and \(r^2\) are high.
- Letting one unusual point decide the model without investigation. A point can affect the line and summaries. Identify what is unusual, assess its influence when appropriate, and retain a valid observation in the primary analysis.
- Claiming causation from a good fit. A linear model describes an association. As emphasized in “Common Errors Linking Regression to Causation,” causal conclusions depend on study design, not on \(r\), \(r^2\), or the quality of the fit alone.
A strong AP response names the evidence specifically: “The scatterplot shows an approximately straight positive association, and the residual plot has no obvious curve or changing spread. The positive \(r\) and large \(r^2\) are consistent with a strong linear association, so a linear model appears reasonable for these data.” If a plot contradicts the statistics, state that conflict and explain why the model-fit conclusion is qualified.
Check Your Understanding
Use the plots and statistics together; explain what each piece of evidence contributes.
- A data set has \(r=0.93\) and \(r^2=0.8649\), but its residual plot forms a clear curve. What does each clue say about the linear model?
- In context, what does \(r^2=0.72\) mean for a regression predicting daily energy use from outdoor temperature?
- Why does \(r^2\) not tell you whether the association is positive or negative?
- A scatterplot looks approximately straight, but residuals spread out more as \(x\) increases. What concern should you include in your fit assessment?
- Why should an unusual point be investigated rather than automatically removed when assessing \(r\), \(r^2\), and the fitted line?