Tutorials › AP Statistics › Comparing Fit Using Residual Plots

Comparing and communicating regression models · Tutorial 986 of 1000

Comparing Fit Using Residual Plots

Use residual plots to judge which of two regression models leaves less systematic structure in its errors, while accounting for the data and the prediction goal.

Intermediate 9 min read

What You'll Learn

  • Compare residual plots for models fitted to the same response values and cases.
  • Look for systematic curves, trends, changing spread, and residuals centered around zero.
  • Distinguish irregular scatter from a visible zigzag or other ordered pattern.
  • Compare residual-plot structure even when the models use different explanatory variables.
  • Write a qualified conclusion about which model better captures the pattern in the observed data.

Use Residual Plots to Compare What Models Miss

When two regression models describe the same response, their residual plots help show whether either model systematically misses part of the data’s pattern. As in “Comparing Two Linear Models for the Same Data” and “Comparing Models With Different Explanatory Variables,” a fair comparison starts with the same response values and the same cases. This tutorial adds a graphical comparison: look at the structure left behind by each model, not just its numerical summaries.

A residual is the observed response minus the predicted response, \(y-\hat{y}\). A residual plot displays these residuals vertically, usually against the explanatory-variable values or fitted values. A residual above zero means the observed response was higher than predicted; one below zero means it was lower. As “Reading a Residual Plot for Model Fit” explains, a useful linear fit tends to leave residuals scattered around zero without a clear systematic pattern.

Definition: When comparing residual plots, the model that leaves less systematic structure in its residuals is generally better at capturing the pattern shown by the data. Look for residuals centered around zero, no clear curve or trend, and a reasonably even vertical spread. This is a descriptive judgment about the observed data, not proof that the model will be best in every setting.

A plot with a clear curve suggests the line misses a changing relationship across the range. A trend suggests residuals tend to increase or decrease as the explanatory or fitted values change. A fan shape suggests the typical size of errors changes across the plot. These patterns matter even if one model has a larger \(r^2\): a summary number alone does not reveal how the errors are arranged.

The two horizontal axes may represent different explanatory variables when the models use different predictors. Do not compare their horizontal positions as if they had the same units. Instead, compare the overall residual structure: does one plot show a clear pattern that is weaker or absent in the other? Also consider whether both plots use the same response scale, since residuals are measured in the response’s units.

A Practical Comparison Process

Use this sequence to make the comparison specific and defensible. A plot should support your conclusion with a description of what the residuals do, rather than a vague statement that one model “looks better.”

1
Check the comparison is fair.
Confirm that both models use the same response, the same response scale, and the same cases. If cases differ, the plots do not provide a direct comparison of the models on the same observations.
2
Check each plot against zero.
Look for residuals on both sides of zero and whether they are centered around it. Note a sustained run, curve, trend, or other visible order rather than describing it as random scatter.
3
Compare spread and structure.
Identify which model leaves the clearer systematic pattern, if any. Compare whether the vertical spread is reasonably even or changes across a plot.
4
Conclude in context.
Say which model appears to capture the observed pattern better and name the evidence. Qualify the conclusion to these data and the intended prediction task.

A residual plot is not a contest to find the plot with the most perfectly random-looking points. Small samples can look uneven by chance, and a plot can contain an isolated large residual. Focus on systematic features that persist across the range. As in “Outlier, High-Leverage, and Influential Points Defined,” a single unusual point and a broad pattern are different issues; do not let one distract you from the overall arrangement.

Worked Example: One Model Leaves a Curve

Worked Example: Predicting Daily Water Use

Original AP-style question. In an invented data set, two simple linear regression models use the same 18 neighborhoods and predict daily household water use, measured in liters per household. Model A uses household size as its explanatory variable; Model B uses the number of bathrooms. The plots show Model A residuals in a curved pattern: residuals tend to be positive at the low and high ends and negative in the middle. Model B residuals are scattered above and below zero without a clear bend, trend, or change in spread. Which model better captures the pattern in these observations?

State. We will compare the two residual plots for predicting daily water use for the same 18 neighborhoods.

Plan. Both models use the same response, liters per household, and the same neighborhoods, so the plots describe errors for the same observations on the same response scale. We will compare the patterns around zero, including curvature and spread. The horizontal axes differ, so we will not compare particular positions along those axes as though household size and number of bathrooms had the same units.

Do. In Model A, the residuals are mostly positive at both ends and mostly negative in the middle. This systematic bend indicates that the line misses part of the relationship: it tends to underpredict at the ends and overpredict in the middle. In Model B, the residuals are mixed above and below zero without a clear sustained curve or changing spread. That plot leaves less obvious systematic structure. The described evidence favors Model B for these observations.

Conclude. The number-of-bathrooms model appears to capture the observed water-use pattern better than the household-size model, because its residual plot has less systematic curvature and no clear change in spread. This is a conclusion about these 18 neighborhoods; it does not establish that bathrooms cause water use to change or guarantee that Model B will predict better in a different population.

If the regression output also showed \(r^2=0.52\) for Model A and \(r^2=0.49\) for Model B, the residual plots would still be useful. Model A’s larger \(r^2\) does not erase the visible curve. The plots reveal a feature of fit that a single summary does not show, so describe both kinds of evidence rather than letting one number settle the comparison.

Worked Example: A Zigzag Is Still a Pattern

Worked Example: Comparing Residuals for Two Delivery Models

Original AP-style question. An invented delivery service fits two models to delivery time in minutes using the same 12 routes. In the first residual plot, when routes are ordered from low to high fitted time, the residual signs follow a repeated positive, negative, positive, negative sequence across most of the plot. The residuals range from \(-3\) to \(+3\) minutes. In the second plot, the residuals range from \(-4\) to \(+4\) minutes and form a loose, irregular band around zero with no sustained curve, trend, or fan shape. Is it accurate to say that the first plot has no pattern because its residuals alternate around zero?

Solution. No. A repeated left-to-right alternation is itself a visible zigzag pattern. It is not accurate to call the first plot patternless merely because its residuals appear on both sides of zero. The first model leaves a systematic-looking sequence, while the second plot has somewhat larger residual magnitudes but less obvious order.

The ranges alone do not determine which fit is better. The second plot extends one minute farther in each direction, so its residuals may be somewhat larger in magnitude; however, the first plot’s repeated zigzag is evidence of structure. We should describe both observations instead of treating either the range or the sign alternation as the only criterion. A clear zigzag can be a clue that the line does not capture all the structure, though it is not automatically proof of a practically important defect.

Conclusion. The second model appears to leave less systematic structure in these residual plots, although its residuals reach slightly farther from zero. A careful comparison would report the zigzag in the first plot and the larger range in the second, then consider the size and practical importance of each feature. It would be incorrect to say that both plots show no pattern.

Worked Example: Compare Spread as Well as Center

Worked Example: Estimating Plant Growth

Original AP-style question. In an invented greenhouse data set, two models predict plant growth in centimeters from the same 16 plants. Both residual plots are centered around zero, and neither shows a clear curve. In Model A, residuals are mostly between \(-2\) and \(+2\) cm throughout the plot. In Model B, residuals are mostly between \(-1\) and \(+1\) cm at low fitted growth but widen to about \(-5\) to \(+5\) cm at high fitted growth. Which model has the more consistent residual spread?

Solution. Model A has a reasonably similar vertical spread across the plot: its residuals stay mostly within about 2 cm of zero. Model B has a fan shape, with a wider spread at high fitted growth than at low fitted growth. That indicates that prediction errors vary in size across the range. Model A therefore has the more consistent spread in these observations.

Both plots are centered around zero, but centering alone is not enough to call the fits equally satisfactory. The changing spread in Model B is a meaningful residual pattern. If the prediction task focuses on plants with high fitted growth, the wider residuals there are particularly relevant: predictions in that region are less consistent by this graphical measure.

Conclusion. Model A appears to leave more consistent residual variation across the observed range. Model B’s fan shape signals that the size of its errors changes as fitted growth increases. This comparison does not tell us that every Model A prediction is closer, but it identifies a limitation that would be missed by checking only whether residuals lie above and below zero.

Common Mistakes and AP Exam Tips

  • Calling alternation “random.” If the residual signs repeatedly alternate from positive to negative as you move left to right, name the zigzag as a visible pattern. A full-credit answer does not call a repeated sequence patternless just because the points cross zero.
  • Looking only for curvature. A residual plot may have no obvious curve but still show a trend, a zigzag, clusters, or changing spread. Describe the particular feature you see.
  • Choosing the plot with the narrowest vertical range automatically. A smaller range can indicate smaller errors, but a plot’s structure matters too. Compare both typical spread and systematic patterns, and do not claim that a range gives a guaranteed maximum error.
  • Comparing horizontal positions from different predictors. Household size and number of bathrooms, for example, use different units. Compare the residual structure across each plot rather than equating their horizontal locations.
  • Ignoring which observations were used. A graphical comparison is most direct when both models use the same response values and cases. If the cases differ, explain that limitation rather than treating the plots as a clean head-to-head comparison.
  • Making a universal or causal claim. Say which model appears to capture the pattern better in the observed data. A residual plot does not prove causation or establish that the model will work equally well for other cases.

A strong AP response identifies the models and response, confirms that the comparison uses the same cases when that information is available, and names the residual evidence. For example: “For these 18 neighborhoods, Model B appears to fit better because its residuals are scattered around zero without the curved pattern seen for Model A.” If a plot shows regular alternation or a fan, say so explicitly and explain what that pattern indicates.

Key takeaway: Compare residual plots by looking for structure, not just by checking whether points fall above and below zero. A model that leaves residuals centered around zero with less systematic pattern and reasonably even spread generally captures the observed pattern better. Describe visible zigzags, curves, trends, and changing spread accurately, and keep the conclusion tied to the data being compared.

Check Your Understanding

Use the residual-plot comparison ideas from this tutorial to answer each question.

  1. Two models predict the same response for the same cases. One residual plot has a clear curve; the other has residuals scattered around zero without an obvious pattern. Which model appears to capture the pattern better, and what is the evidence?
  2. Residual signs repeatedly alternate as the explanatory values increase. Why is it inaccurate to describe this sequence as patternless?
  3. One plot has residuals mostly within 2 response units of zero throughout; another widens from about 1 unit to 5 units. What does the second plot’s changing spread suggest?
  4. Two models use different explanatory variables with different units. What should you compare in their residual plots, and what should you avoid comparing directly?
  5. Why should a conclusion about a favored model be qualified to the cases and setting represented by the data?