Tutorials › AP Statistics › Why Squared Residuals Not Absolute Values

Least-squares regression · Tutorial 862 of 1000

Why Squared Residuals Not Absolute Values

See how squaring residuals changes which candidate line is preferred, and compare the least-squares fit with a line that minimizes absolute errors.

Intermediate 9 min read

What You'll Learn

  • Calculate the sum of absolute residuals as well as the sum of squared residuals for a candidate line.
  • Compare how the two criteria weight small and large prediction errors.
  • Find a least-squares line for a tiny data set and verify its squared-error total.
  • Show that a different line can minimize absolute errors for the same observations.
  • State precisely what a comparison of candidate lines does—and does not—establish.

Two Ways to Measure Prediction Error

In “The Idea of the Least-Squares Criterion,” you learned that the least-squares regression line is the line \(\hat{y}=a+bx\) that minimizes the sum of squared residuals. That choice raises a natural question: why square the residuals instead of adding their absolute values? Both calculations measure prediction errors, but they can favor different lines.

For each observation, the residual is \(y-\hat{y}\): the observed response minus the line’s predicted response. Its sign indicates whether the point is above or below the line. Squaring the residual removes its sign and makes larger errors count more heavily. Taking its absolute value also removes the sign, but the error contributes in direct proportion to its size.

Definition: The sum of absolute residuals for a candidate line is the sum of \(|y-\hat{y}|\) across the observations. The sum of squared residuals, or SSE, is the sum of \((y-\hat{y})^2\). The least-squares regression line minimizes SSE, not the sum of absolute residuals.

These criteria have different units, too. If the response is measured in seconds, the sum of absolute residuals is measured in seconds, while SSE is measured in seconds squared. Their numerical totals should not be compared directly with each other. Instead, compare candidate lines using the same criterion and the same data.

Neither criterion makes every prediction perfect. The important difference is how each combines the errors: absolute residuals give an error twice the size twice the contribution, whereas squared residuals give it four times the contribution. A line with a relatively large miss may therefore be less attractive under squared errors than under absolute errors.

Compare Both Criteria on Three Points

Consider three invented measurements of a device’s operating time, \(x\), and temperature, \(y\), in degrees Celsius. The observed points are \((1,0)\), \((2,0)\), and \((3,10)\). We will compare a least-squares line with a line that minimizes the sum of absolute residuals.

Worked Example: The Squared-Error and Absolute-Error Lines Differ

First find the least-squares line. The means are \(\bar{x}=2\) and \(\bar{y}=\frac{10}{3}\). Using the slope calculation from earlier in the regression sequence, the numerator is:

$$ (1-2)\left(0-\frac{10}{3}\right) +(2-2)\left(0-\frac{10}{3}\right) +(3-2)\left(10-\frac{10}{3}\right) =\frac{10}{3}+0+\frac{20}{3}=10 $$

The sum of squared deviations in \(x\) is \((1-2)^2+(2-2)^2+(3-2)^2=2\), so the slope is \(b=10/2=5\). The intercept is:

$$ a=\bar{y}-b\bar{x} =\frac{10}{3}-5(2) =-\frac{20}{3} $$

Thus, the least-squares line is \(\hat{y}=-\frac{20}{3}+5x\). Its predictions and residuals are:

\(x\)Observed \(y\)Least-squares prediction \(\hat{y}\)Residual \(y-\hat{y}\)Squared residualAbsolute residual
10\(-\frac{5}{3}\)\(\frac{5}{3}\)\(\frac{25}{9}\)\(\frac{5}{3}\)
20\(\frac{10}{3}\)\(-\frac{10}{3}\)\(\frac{100}{9}\)\(\frac{10}{3}\)
310\(\frac{25}{3}\)\(\frac{5}{3}\)\(\frac{25}{9}\)\(\frac{5}{3}\)

The totals for this line are:

$$ \text{SSE} =\frac{25}{9}+\frac{100}{9}+\frac{25}{9} =\frac{150}{9} =\frac{50}{3} \approx 16.6667 $$
$$ \text{Sum of absolute residuals} =\frac{5}{3}+\frac{10}{3}+\frac{5}{3} =\frac{20}{3} \approx 6.6667 $$

Now consider the line through the first and third points, \(\hat{y}=-5+5x\). It predicts \(0,5,10\), so its residuals are \(0,-5,0\). Its SSE is \(0^2+(-5)^2+0^2=25\), and its sum of absolute residuals is \(|0|+|-5|+|0|=5\). The least-squares line has the smaller SSE, but this endpoint line has the smaller sum of absolute residuals.

We can verify that the endpoint line really does minimize the sum of absolute residuals, rather than merely beating the least-squares line. For any line, its predicted values at \(x=1,2,3\) form an arithmetic sequence. If the residuals are \(e_1,e_2,e_3\), this means:

$$ e_1-2e_2+e_3 =y_1-2y_2+y_3 =0-2(0)+10 =10 $$

By the triangle inequality, \(10=|e_1-2e_2+e_3|\leq |e_1|+2|e_2|+|e_3|\), which is at most \(2(|e_1|+|e_2|+|e_3|)\). Therefore every line has a sum of absolute residuals of at least \(5\). The endpoint line attains \(5\), so it minimizes that total. The two criteria select different lines for these data.

Why the Two Totals Can Prefer Different Lines

For the three-point example, the least-squares line distributes its residuals across all three cases: \(\frac{5}{3},-\frac{10}{3},\frac{5}{3}\). The endpoint line has zero residuals at two observations and a residual of \(-5\) at the middle observation. The second line’s single large miss makes its SSE larger, even though its total absolute error is smaller.

This is the central tradeoff. Squared errors put extra weight on large residuals; absolute errors do not increase that quickly. In some settings, a method that is less affected by an unusually large error may be useful. But that is a different fitting criterion. In AP Statistics, “least-squares line” has a specific meaning: the line that minimizes SSE.

A lower sum of absolute residuals does not prove that a line is better by the least-squares criterion. Conversely, a lower SSE does not mean a line minimizes absolute errors. Be explicit about which total you calculated and which criterion the question asks you to use.

Key distinction: For a fixed data set, the least-squares line minimizes the sum of squared residuals. A line that minimizes the sum of absolute residuals may be different. To compare candidates, calculate the same type of total for each line.

Worked Comparison With a Second Data Set

A second small example makes the ranking reversal visible with another set of values. A fictional sensor records an input setting \(x\) and a response \(y\). The points are \((0,2)\), \((1,2)\), and \((2,8)\). We will compare the least-squares line, the line through the endpoints, and a horizontal candidate line.

Worked Example: Rank Three Candidate Lines Two Ways

The means are \(\bar{x}=1\) and \(\bar{y}=4\). The centered-value slope numerator is:

$$ (0-1)(2-4)+(1-1)(2-4)+(2-1)(8-4) =2+0+4=6 $$

The sum of squared deviations in \(x\) is \(1+0+1=2\), so the least-squares slope is \(b=6/2=3\). Its intercept is \(a=4-3(1)=1\), giving \(\hat{y}=1+3x\). The predictions are \(1,4,7\), and the residuals are \(1,-2,1\). Thus:

$$ \text{SSE}=1^2+(-2)^2+1^2=6, \qquad \text{Sum of absolute residuals}=|1|+|-2|+|1|=4 $$

The line through the endpoints is \(\hat{y}=2+3x\). Its predictions are \(2,5,8\), with residuals \(0,-3,0\). Its SSE is \(0^2+(-3)^2+0^2=9\), and its sum of absolute residuals is \(3\).

For the horizontal line \(\hat{y}=4\), the residuals are \(-2,-2,4\). Its SSE is \((-2)^2+(-2)^2+4^2=4+4+16=24\), and its sum of absolute residuals is \(2+2+4=8\). The results are:

Candidate lineSSESum of absolute residuals
Least-squares line, \(\hat{y}=1+3x\)64
Endpoint line, \(\hat{y}=2+3x\)93
Horizontal line, \(\hat{y}=4\)248

Among these three candidates, the least-squares line has the smallest SSE, while the endpoint line has the smallest sum of absolute residuals. As in the first example, the endpoint line minimizes absolute residuals over all lines: for any line, its residuals satisfy \(e_1-2e_2+e_3=2-2(2)+8=6\), so \(6\leq 2(|e_1|+|e_2|+|e_3|)\). Every line therefore has a sum of absolute residuals of at least \(3\), and the endpoint line reaches that value.

A Reliable Way to Compare Error Criteria

When a task asks you to compare squared and absolute errors, keep the calculations separate. Use observed response minus predicted response for each residual, as established in “The Idea of the Least-Squares Criterion.” Then apply the requested operation to every residual.

1
Write each candidate line.
Keep the explanatory variable \(x\) and response variable \(y\) in their stated roles.
2
Calculate predictions and residuals.
Substitute each observed \(x\) into the line, then calculate \(y-\hat{y}\).
3
Calculate the requested total.
For SSE, square each residual and add. For the sum of absolute residuals, take each residual’s absolute value and add.
4
Compare like with like.
Use SSE totals to compare squared-error fit, or absolute-error totals to compare absolute-error fit. Do not compare a number from one criterion directly with a number from the other.
5
State what the result establishes.
A candidate with the smaller total beats the other candidate by that criterion. Claim it is the overall minimizing line only when the minimizing property has been established.

Common Mistakes and AP Exam Tips

  • Calling absolute errors “least squares.” The least-squares criterion squares residuals. If you take absolute values instead, name the sum of absolute residuals.
  • Comparing unlike totals. Do not say an SSE of \(6\) is larger or smaller in a meaningful model-selection sense than an absolute-error total of \(4\). They use different criteria and have different units.
  • Forgetting the effect of a large error. A residual of \(5\) contributes \(25\) to SSE but \(5\) to the absolute-error total. Explain how the criterion changes the contribution rather than saying only that one line “looks better.”
  • Assuming a candidate comparison proves a minimum. If you compare only two lines, you know which of those two has the lower total. You have not necessarily proved that either one minimizes over all possible lines.
  • Omitting signs before applying the criterion. The residual convention is observed minus predicted. For SSE and absolute values the final contribution is nonnegative, but showing the signed residual first helps prevent arithmetic and interpretation errors.
  • Claiming absolute-error fitting is always superior. The criteria make different tradeoffs. Least squares is the defined method for the AP Statistics regression line; absolute errors are an alternative way to score prediction misses, not a replacement definition.

A strong AP response names the criterion, shows the residuals and resulting total, and limits its conclusion to that criterion. For example: “The endpoint line has a smaller sum of absolute residuals, but the least-squares line has a smaller SSE for these observations.” That statement distinguishes the two objectives without confusing which line is least squares.

Key takeaway: Squared and absolute residuals measure errors differently. Squaring makes large residuals count disproportionately more, so the least-squares line can differ from the line that minimizes absolute errors. Calculate and compare the same criterion across candidate lines.

Check Your Understanding

For each question, distinguish the sum of squared residuals from the sum of absolute residuals.

  1. A candidate line has residuals \(3,-1,-2\). Calculate its SSE and its sum of absolute residuals.
  2. Why does a residual of \(4\) contribute more to SSE than a residual of \(-2\), but only twice as much to the sum of absolute residuals?
  3. For the points \((1,0),(2,0),(3,10)\), the endpoint line has SSE \(25\) and absolute-error total \(5\). The least-squares line has SSE \(\frac{50}{3}\) and absolute-error total \(\frac{20}{3}\). Which line is preferred by each criterion?
  4. If two lines are compared and line A has the smaller SSE, what conclusion is justified? What broader claim is not automatically justified?
  5. If the response is measured in meters, what are the units of the sum of absolute residuals and of SSE?