What Does a Zero Mean Residual Tell Us?
In “Finding All Residuals for a Data Set,” you calculated each residual as the observed response minus its prediction, \(y-\hat{y}\), and checked whether the signed residuals added to zero. Here we explain why that check works for a least-squares regression line with an intercept—and, just as importantly, what the check does not tell us about the fit.
A residual’s sign records whether an observation is above or below the fitted line. When we add signed residuals, positive and negative values can cancel. For a least-squares line with an intercept, they cancel exactly when calculations use full precision. Dividing their sum by the number of observations shows that the mean residual is also zero.
Why the Residuals Sum to Zero
For observation \(i\), write its response as \(y_i\), its predictor as \(x_i\), and its residual as \(e_i=y_i-\hat{y}_i\). The fitted line is \(\hat{y}=a+bx\), so the predicted response for observation \(i\) is \(a+bx_i\). Add the residuals across all \(n\) observations:
As established in “Properties of the Least-Squares Line” and “Finding the Line Through the Means,” a least-squares regression line with an intercept passes through \((\bar{x},\bar{y})\). That means the line predicts \(\bar{y}\) when \(x=\bar{x}\), or \(\bar{y}=a+b\bar{x}\). The quantity in parentheses in the final expression is therefore zero, so the sum of residuals is zero. The mean residual, \(\sum e_i/n\), is zero as well.
The intercept matters. Do not assume this zero-sum result for a line fitted without an intercept or for an arbitrary line chosen to describe the data. Also, the result concerns the data used to fit the line. It does not guarantee that residuals for new observations will average to zero.
Worked Example: Deriving the Result From a Fitted Line
Worked Example: Deriving the Result From a Fitted Line
A fictional garden project records the number of watering minutes, \(x\), and the amount of water used, \(y\), in liters. The observations are \((1,5)\), \((2,7)\), \((3,8)\), and \((4,12)\). Find the least-squares line, calculate its residuals, and connect their total to the algebraic result.
State. The predictor is watering time in minutes, and the response is water use in liters. We will calculate the fitted line and then use \(y-\hat{y}\) for each residual.
Plan. First find \(\bar{x}\) and \(\bar{y}\). Use the centered-data slope calculation, \(b=\sum (x_i-\bar{x})(y_i-\bar{y})/\sum (x_i-\bar{x})^2\), then find \(a=\bar{y}-b\bar{x}\). These are the standard least-squares calculations covered earlier in “Computing a Least-Squares Line From Raw Data.” Finally, find the predictions and signed residuals.
Do. The means are \(\bar{x}=(1+2+3+4)/4=2.5\) minutes and \(\bar{y}=(5+7+8+12)/4=8\) liters. The centered predictor values are \(-1.5,-0.5,0.5,1.5\), and the centered response values are \(-3,-1,0,4\). Thus, the sum of the cross-products is \(4.5+0.5+0+6=11\), while the sum of the squared predictor deviations is \(2.25+0.25+0.25+2.25=5\). So \(b=11/5=2.2\) liters per minute, and \(a=8-2.2(2.5)=2.5\) liters. The line is \(\hat{y}=2.5+2.2x\).
For example, at \(x=1\), the prediction is \(2.5+2.2(1)=4.7\) liters, so the residual is \(5-4.7=0.3\) liter. The other rows are calculated the same way.
| Watering time, \(x\) (minutes) | Observed use, \(y\) (liters) | Predicted use, \(\hat{y}\) (liters) | Residual, \(y-\hat{y}\) (liters) |
|---|---|---|---|
| 1 | 5 | 4.7 | +0.3 |
| 2 | 7 | 6.9 | +0.1 |
| 3 | 8 | 9.1 | −1.1 |
| 4 | 12 | 11.3 | +0.7 |
The signed residual total is \(0.3+0.1-1.1+0.7=0\) liters, so the mean residual is \(0/4=0\) liters. The algebra gives the same check: \(4\bigl(8-(2.5+2.2(2.5))\bigr)=4(8-8)=0\) liters.
Conclude. The residuals sum to zero because this least-squares line with an intercept passes through the mean point \((2.5,8)\). The individual residuals are not all zero, even though their signed average is.
Zero Sum Does Not Mean the Line Is Least Squares
It is tempting to treat a zero residual total as proof that a line is the least-squares line. That conclusion is not valid. Any line with an intercept that passes through \((\bar{x},\bar{y})\) has residuals summing to zero, whatever its slope. The zero-sum property alone cannot tell us whether that slope minimizes the sum of squared residuals, or SSE.
Worked Example: Two Lines With Zero-Sum Residuals
A fictional workshop records hours of equipment use, \(x\), and the number of items produced, \(y\): \((1,7)\), \((2,7)\), \((3,9)\), and \((4,9)\). Compare the least-squares line with another line that passes through the mean point.
State. We will check the residual sums for both lines, then compare their SSE values. The second line is a candidate line, not necessarily a least-squares line.
Plan. The means are \(\bar{x}=2.5\) hours and \(\bar{y}=8\) items. Calculate the least-squares line using the centered slope calculation, and calculate residuals for both it and the candidate line \(\hat{y}=3+2x\). For each line, square and add its residuals to find SSE.
Do. The response deviations from 8 are \(-1,-1,1,1\); the predictor deviations from 2.5 are \(-1.5,-0.5,0.5,1.5\). The cross-products sum to \(1.5+0.5+0.5+1.5=4\), and the squared predictor deviations sum to \(5\). Therefore, the least-squares slope is \(4/5=0.8\) items per hour, and its intercept is \(8-0.8(2.5)=6\) items. The least-squares line is \(\hat{y}=6+0.8x\).
The candidate line \(\hat{y}=3+2x\) also passes through the mean point, since \(3+2(2.5)=8\). The table shows the predictions and residuals from both lines.
| \(x\) (hours) | \(y\) (items) | Least-squares prediction | Least-squares residual | Candidate prediction | Candidate residual |
|---|---|---|---|---|---|
| 1 | 7 | 6.8 | +0.2 | 5 | +2 |
| 2 | 7 | 7.6 | −0.6 | 7 | 0 |
| 3 | 9 | 8.4 | +0.6 | 9 | 0 |
| 4 | 9 | 9.2 | −0.2 | 11 | −2 |
Both signed totals are zero: \(0.2-0.6+0.6-0.2=0\) items for the least-squares line, and \(2+0+0-2=0\) items for the candidate line. But their SSE values differ:
Conclude. The candidate line’s residuals also sum to zero, but its SSE is larger. A zero residual sum shows that the signed deviations balance; it does not establish that the line minimizes SSE. That minimizing property is what defines the least-squares line.
Zero Mean Does Not Mean a Good Fit
A mean of zero is a statement about the signed average, not about the typical size of a residual. A few large positive residuals can balance several smaller negative ones. Even a data set with substantial residuals—or a sequence of residuals that shows a pattern—can have a mean residual of zero.
Worked Example: Balanced Residuals With a Pattern
A fictional community energy project records energy use, \(y\), in kilowatt-hours over four numbered observation periods, \(x\). The data are \((1,2)\), \((2,10)\), \((3,2)\), and \((4,10)\). Find the least-squares line and examine what its residuals say about fit.
State. We want to calculate the residuals for the least-squares line and decide whether their zero mean means that the line captures the data pattern well.
Plan. Calculate the means and least-squares slope, obtain the intercept, then find all four predictions and residuals. Inspect the residuals in observation order and calculate SSE to describe their size.
Do. Here, \(\bar{x}=2.5\) periods and \(\bar{y}=6\) kilowatt-hours. The response deviations are \(-4,4,-4,4\), and the predictor deviations are \(-1.5,-0.5,0.5,1.5\). Their cross-products sum to \(6-2-2+6=8\), and the squared predictor deviations sum to \(5\). Thus, \(b=8/5=1.6\) kilowatt-hours per period, and \(a=6-1.6(2.5)=2\) kilowatt-hours. The fitted line is \(\hat{y}=2+1.6x\).
| Period, \(x\) | Observed use, \(y\) (kilowatt-hours) | Predicted use, \(\hat{y}\) (kilowatt-hours) | Residual, \(y-\hat{y}\) (kilowatt-hours) |
|---|---|---|---|
| 1 | 2 | 3.6 | −1.6 |
| 2 | 10 | 5.2 | +4.8 |
| 3 | 2 | 6.8 | −4.8 |
| 4 | 10 | 8.4 | +1.6 |
The signed sum is \(-1.6+4.8-4.8+1.6=0\) kilowatt-hours, so the mean residual is zero. Yet the residuals alternate negative, positive, negative, positive, and some are large compared with the eight-kilowatt-hour range of the observed responses. Their SSE is \(1.6^2+4.8^2+(-4.8)^2+1.6^2=2.56+23.04+23.04+2.56=51.2\) kilowatt-hours squared.
Conclude. The residuals have a mean of zero, but that fact does not show that the line fits well. Their substantial sizes and alternating pattern indicate that the line misses an important feature of these observations.
This example illustrates two separate questions. The zero-sum property is an algebraic fact about how a least-squares line with an intercept is fitted. Assessing fit requires looking at residual sizes and patterns, not just their average. In later work with residual plots, patterns will be a useful clue that a straight-line model may not describe the data adequately.
Common Mistakes and AP Exam Tips
- Claiming all residuals equal zero. A zero mean says their signed total is zero, not that every observation lies on the line. State that positive and negative residuals can cancel.
- Using the result for any line. The guaranteed result is for a least-squares regression line with an intercept fitted to those observations. An arbitrary line need not have a zero residual sum. A different line through the mean point does have a zero sum, but that fact alone does not make it least squares.
- Calling zero mean a measure of fit. The mean residual does not describe typical error size. To compare how far points are from a line under the least-squares criterion, consider SSE and inspect the residuals.
- Confusing a signed sum with a sum of magnitudes. Positive and negative residuals cancel in \(\sum e_i\). Absolute values do not cancel, and neither do squared residuals.
- Ignoring rounding. A displayed equation may have rounded coefficients, and a residual table may show rounded predictions. Those rounded values can produce a small nonzero total. Keep precision during calculations and describe a slight discrepancy as approximate when appropriate.
- Leaving out the condition. For a precise explanation, identify that the line is a least-squares regression line with an intercept and that the residuals are for the observations used to fit it.
A strong AP response gives the reason, not only the result: the least-squares line with an intercept passes through \((\bar{x},\bar{y})\), so the average prediction at \(\bar{x}\) equals \(\bar{y}\), which makes the signed residual sum—and therefore the mean residual—zero. Avoid claiming that this proves a good fit.
Check Your Understanding
Use the zero-mean result carefully. Explain what it establishes and what it does not establish.
- In one sentence, explain why the residuals from a least-squares line with an intercept sum to zero.
- A fitted data set has six residuals whose signed sum is zero. Must all six residuals equal zero? Explain.
- A line passes through \((\bar{x},\bar{y})\) but is not the least-squares line. What can you conclude about its residual sum, and what can you not conclude from that sum alone?
- Why does a mean residual of zero not establish that a linear model fits well?
- What conditions should you name when explaining the zero-sum residual property in an AP response?