Two Graphs, Two Different Questions
When you analyze two quantitative variables, one graph shows the data themselves and another can show how a fitted line performs. These graphs are related, but they answer different questions. Choosing and describing both clearly is a useful technique for supporting a decision about whether a linear model is appropriate.
As discussed in “Planning a Complete Regression Analysis,” start by identifying the cases, the explanatory variable \(x\), and the response variable \(y\). The scatterplot puts the observed \(x\)- and \(y\)-values together. After fitting a line, a residual plot displays each residual, \(y-\hat{y}\), and helps reveal patterns in the line’s errors. The first graph shows the relationship; the second checks what the line leaves behind.
The scatterplot is the place to describe the original association’s direction, form, and strength. It can also show clusters, gaps, the observed \(x\)-range, or observations that stand apart. A roughly straight pattern is a reason to consider a linear model. A visible bend, separate groups, or an unusual point may change what you conclude or prompt further investigation.
The residual plot has a different job. The horizontal line at residual \(0\) marks where the observed response equals the predicted response. A residual above zero means the observed response is greater than the fitted value; one below zero means it is less. For a linear model to be a reasonable description, residuals should generally form a patternless band around zero with reasonably even vertical spread. A curve or changing spread suggests that the line misses some systematic feature of the data.
Choose and Read the Graphs in Sequence
A scatterplot is usually the first graph because it lets you see whether a straight-line model is sensible to examine. Put \(x\) on the horizontal axis and \(y\) on the vertical axis, label both axes with variable names and units, and choose scales that show the data without hiding important features. Keep each point as one case: do not connect points unless the order of observations itself matters to the question.
If a line is reasonable to examine, fit it and make the residual plot for that fit. The residual plot should use the residuals from the same cases and the same fitted line. Plotting residuals against \(x\) makes it easy to notice whether the errors change across the explanatory-variable range. Plotting them against fitted values can serve a similar purpose. In either version, include or identify the zero reference line.
Put the explanatory variable on the horizontal axis and the response variable on the vertical axis. Label the variables and units.
Look for direction, form, strength, range, clusters, and unusual observations. Decide whether examining a line is reasonable.
After fitting the line, look for residuals scattered around zero without a clear pattern and with reasonably even spread.
Say what the scatterplot shows and what the residual plot adds. Explain whether the evidence supports using a line for these data.
This sequence matters. A residual plot is not a substitute for viewing the original data: by itself, it does not show the original response values or tell the whole story about the relationship. Likewise, a scatterplot that looks roughly straight does not guarantee that a fitted line’s errors are patternless. As you practiced in “Residual Plot Versus Scatterplot Diagnosis,” use each graph for the question it is designed to answer.
Worked Example: A Relationship That Looks Roughly Linear
Worked Example: Battery Charge and Device Run Time
Original AP-style question. In an invented test, each case is one device. The explanatory variable \(x\) is the percentage of battery charge at the start of a test, and the response variable \(y\) is run time in hours. The five observed pairs are \((1,8),(2,10),(3,13),(4,14),(5,17)\). Explain what the scatterplot and residual plot show about using a linear model.
Solution. On the scatterplot, place starting charge percentage on the horizontal axis and run time in hours on the vertical axis. The points rise from left to right and follow a roughly straight pattern, with no obvious isolated point. This supports examining a positive linear model for these five observations; it does not establish that the model is useful for every device or setting.
To make the residual plot, first find the least-squares line. The means are \(\bar{x}=3\) and \(\bar{y}=12.4\). The centered sums are \(S_{xx}=10\) and \(S_{xy}=22\), so the slope is \(b=S_{xy}/S_{xx}\), and the intercept is \(a=\bar{y}-b\bar{x}\):
The fitted values for \(x=1,2,3,4,5\) are \(8,10.2,12.4,14.6,16.8\) hours. Subtract each fitted value from its observed run time. The residuals are \(0,-0.2,0.6,-0.6,0.2\) hours, respectively. For example, at \(x=3\), the residual is \(13-12.4=0.6\) hours.
The residual plot would place those five residuals against starting charge percentage, with a horizontal reference line at zero. They sit fairly close to zero and do not show a clear curve or a strong change in spread. Together, the roughly straight scatterplot and the patternless-looking residuals support using a line as a description of these observations. With only five points, however, the visual check is limited; it does not prove that the relationship is exactly linear.
Worked Example: A Scatterplot’s Bend Appears in the Residuals
Worked Example: Practice Sessions and a Skill Score
Original AP-style question. In an invented exercise, \(x\) is the number of practice sessions and \(y\) is a skill score. The observations are \((1,3),(2,6),(3,11),(4,18),(5,27)\). Describe the evidence in the scatterplot and residual plot, and decide whether a straight line is a good description.
Solution. The scatterplot rises, but the points bend upward: score increases more sharply at higher session counts. To check what a fitted line misses, calculate its equation. Here, \(\bar{x}=3\), \(\bar{y}=13\), \(S_{xx}=10\), and \(S_{xy}=60\). Therefore:
The fitted scores for \(x=1,2,3,4,5\) are \(1,7,13,19,25\). Subtracting these predictions from the observed scores gives residuals \(2,-1,-2,-1,2\) score points. For instance, at \(x=1\), \(3-1=2\); at \(x=5\), \(27-25=2\).
In the residual plot, residuals are positive at both ends of the \(x\)-range and negative in the middle. That systematic bend is not a patternless band around zero. It reflects the curvature visible in the scatterplot: the line tends to underpredict at the ends and overpredict in the middle. As explained in “Curved Residual Pattern Means Nonlinear Relationship,” this pattern is evidence that a straight-line model misses a systematic feature.
Conclusion. The scatterplot’s upward bend and the residual plot’s curved pattern both argue against describing these observations with a straight line. A positive overall trend alone is not enough to establish that a linear model is appropriate.
Worked Example: Look for Changing Residual Spread
Worked Example: Machine Setting and Production Output
Original AP-style question. In an invented production check, \(x\) is a machine setting and \(y\) is the number of units produced during a fixed interval. The eight observed outputs, in order for settings \(x=1,2,\ldots,8\), are \(26,21,24,29,31,30,31,40\). A fitted line is \(\hat{y}=20+2x\). Use the scatterplot and residual plot evidence to describe the fit.
Solution. The fitted outputs at settings 1 through 8 are \(22,24,26,28,30,32,34,36\). The residual is observed output minus fitted output, so the residuals are \(4,-3,-2,1,1,-2,-3,4\) units. For example, at setting 2 the residual is \(21-24=-3\) units. These values put the observed points above or below the fitted line in the original scatterplot.
| Setting \(x\) | Observed output \(y\) | Fitted output \(\hat{y}\) | Residual \(y-\hat{y}\) |
|---|---|---|---|
| 1 | 26 | 22 | 4 |
| 2 | 21 | 24 | -3 |
| 3 | 24 | 26 | -2 |
| 4 | 29 | 28 | 1 |
| 5 | 31 | 30 | 1 |
| 6 | 30 | 32 | -2 |
| 7 | 31 | 34 | -3 |
| 8 | 40 | 36 | 4 |
The scatterplot shows an overall increase in output as the setting increases, but the vertical departures from the line are larger near the ends of the settings and smaller near the middle. The residual plot makes that changing spread easier to see: the residuals are close to zero around settings 4 and 5, while the residuals at settings 1 and 8 are farther away.
The line is consistent with these residuals as a least-squares fit: their sum is zero, and their products with the centered settings sum to zero:
Conclusion. The residual plot suggests that the size of the errors changes across the settings, rather than forming a band with reasonably even spread. As discussed in “Non-Constant Spread in Residuals,” this weakens the case for treating the line’s prediction errors as similarly sized throughout the range. The scatterplot shows the overall relationship; the residual plot makes the changing error spread more apparent.
Common Mistakes and What Full-Credit Communication Includes
- Giving both graphs the same job. Saying that a residual plot shows the original association confuses it with a scatterplot. State that the scatterplot displays observed \(x\) and \(y\); the residual plot displays \(y-\hat{y}\) from a fitted model.
- Describing only direction. “The points go up” identifies a positive direction but leaves out form and possible unusual features. Comment on whether the scatterplot is roughly straight, curved, clustered, or affected by an observation that stands apart.
- Calling a line appropriate because the scatterplot looks straight. Check the residual plot too. A curve or changing spread can expose a feature that is hard to judge from the overall association alone.
- Ignoring the zero line. Residuals are deviations from the fitted values. Explain that positive residuals mean observed responses exceed predictions and negative residuals mean observed responses fall below predictions.
- Calling any residual variation a problem. A fitted line rarely predicts every observation exactly. Look for a systematic pattern, such as curvature or changing spread, not merely residuals that are nonzero.
- Claiming the graphs prove a model will work elsewhere. The graphs provide evidence about the observations used to fit the line. A patternless residual plot supports the fit; it is not proof of a perfect model or a guarantee for other cases.
For a strong AP response, name the graph and describe the visible evidence before stating what it supports. For example: “The scatterplot shows a roughly straight positive association. The residual plot has no clear pattern and a reasonably even spread around zero, so a linear model is a reasonable description of these observations.” If a graph shows a curve or changing spread, identify that pattern and explain why it weakens the case for a line.
Check Your Understanding
For each situation, identify what the graph evidence shows and what conclusion it supports.
- Which variable belongs on the horizontal axis of a scatterplot when the goal is to predict \(y\) from \(x\)?
- A scatterplot of two quantitative variables has a positive, roughly straight pattern. What does that suggest, and what graph should you examine after fitting a line?
- A residual plot has residuals mostly positive at low and high \(x\)-values but negative in the middle. What does this pattern suggest about the line?
- In a residual plot, an observation has residual \(-2.5\). Explain what this says about its observed response compared with its fitted value.
- A residual plot has a noticeably wider vertical spread at higher fitted values. What feature of the errors should you describe?
- Why is it incomplete to say that a scatterplot alone proves a linear model is appropriate?