What Makes a Regression Line “Least Squares”?
In “Linear Model Interpretation Mixed Review,” you practiced reading and using a regression line. This tutorial asks a different question: among all possible lines that could predict \(y\) from \(x\), how does the least-squares method choose one? The answer is in its name: it chooses the line with the smallest sum of squared residuals.
For a case with observed response \(y\), the residual is \(y-\hat{y}\), where \(\hat{y}\) is the response predicted by a candidate line at that case’s \(x\)-value. A positive residual means the observed response is above the line; a negative residual means it is below. Geometrically, a residual is a vertical difference between an observed point and the line, measured at that point’s \(x\)-value. It is not the shortest, perpendicular distance to the line.
Each candidate line gives a predicted value for every observed \(x\), and therefore a residual for every observed response. To assess the line by this criterion, square each residual and add the results:
The least-squares line is selected by comparing this total across possible lines. It minimizes the total squared error for the observations; it does not necessarily minimize every individual residual, make every prediction close, or guarantee that the line is a useful model for every purpose. As you learned in “Strong Correlation Does Not Mean a Good Model,” you still need to consider the pattern of the data and the context.
Why Square the Residuals?
Adding the signed residuals without squaring would not measure the overall size of the prediction errors. Positive and negative residuals could cancel. For example, residuals of \(+2\) and \(-2\) add to zero, even though both predictions miss their observed responses by 2 units. Squaring makes both contributions positive: \(2^2=4\) and \((-2)^2=4\). The two errors then contribute 8 to the total rather than disappearing through cancellation.
Squaring also makes the size of an error matter more as the error grows. A residual of 3 contributes \(3^2=9\); a residual of 1 contributes only \(1^2=1\). Thus, a large miss can add substantially to a line’s SSE. This is one reason an unusual observation can have considerable influence on a least-squares fit.
There is also a useful mathematical feature: for a line \(\hat{y}=a+bx\), the SSE is a quadratic expression in the intercept \(a\) and slope \(b\). That structure makes it possible to find the values that minimize the total. You do not need calculus to use the criterion in AP Statistics: you can calculate and compare SSE values, or use regression technology to find the least-squares line.
Calculate and Compare SSE Values
The following example uses a small set of points to distinguish a line that is merely a possible candidate from the least-squares line. A candidate line’s SSE can be calculated directly, but a low-looking or simple equation should not be called the least-squares line unless it actually has the minimum SSE.
Worked Example: Check Two Lines for Three Data Points
A fictional workshop records the number of hours a device is tested, \(x\), and its temperature reading, \(y\), in degrees Celsius. The three observed points are \((1,2)\), \((2,4)\), and \((3,4)\). Compare the SSE for the candidate line \(\hat{y}=2+\frac{2}{3}x\) with the SSE for the least-squares line.
First evaluate the candidate line at each observed \(x\). The residual is observed \(y\) minus predicted \(\hat{y}\):
| \(x\) | Observed \(y\) | Candidate prediction \(\hat{y}\) | Residual \(y-\hat{y}\) | Squared residual |
|---|---|---|---|---|
| 1 | 2 | \(2+\frac{2}{3}(1)=\frac{8}{3}\) | \(-\frac{2}{3}\) | \(\frac{4}{9}\) |
| 2 | 4 | \(2+\frac{2}{3}(2)=\frac{10}{3}\) | \(\frac{2}{3}\) | \(\frac{4}{9}\) |
| 3 | 4 | \(2+\frac{2}{3}(3)=4\) | 0 | 0 |
Therefore, the candidate line has:
That calculation establishes the SSE for this candidate, but it does not establish that the candidate is the least-squares line. To identify the least-squares line, calculate the regression slope and intercept. The means are \(\bar{x}=2\) and \(\bar{y}=\frac{10}{3}\). Using the slope calculation based on centered values, the numerator is:
The sum of squared deviations in \(x\) is \((1-2)^2+(2-2)^2+(3-2)^2=2\), so the least-squares slope is \(b=2/2=1\). The intercept is:
Thus, the least-squares line is \(\hat{y}=\frac{4}{3}+x\), not the earlier candidate. Its predictions at \(x=1,2,3\) are \(\frac{7}{3},\frac{10}{3},\frac{13}{3}\). The corresponding residuals are \(-\frac{1}{3},\frac{2}{3},-\frac{1}{3}\), so its SSE is:
Since \(\frac{2}{3}<\frac{8}{9}\), the least-squares line has a smaller SSE than the candidate line, as it should. The key distinction is that \(2+\frac{2}{3}x\) was one possible line to evaluate; \(\frac{4}{3}+x\) is the least-squares line for these data.
Use Squared Contributions to Understand the Criterion
SSE is built from individual contributions. Each squared residual adds a nonnegative amount, and a larger residual adds more than a smaller one. The next example follows a candidate line’s predictions for several plants. The purpose is to see the contribution of each residual, not to claim that this candidate is the least-squares line.
Worked Example: See How a Large Residual Counts
A fictional greenhouse uses a candidate line to predict plant height, \(y\), in centimeters from time \(x\), in weeks after planting:
For four plants observed at \(x=1,2,3,4\) weeks, the measured heights are 15, 13, 15, and 17 centimeters. Find the residuals and SSE for this candidate line.
The predictions are \(12,14,16,18\) centimeters. Subtract each prediction from its observed height, then square each result:
| Weeks \(x\) | Observed height \(y\), cm | Predicted height \(\hat{y}\), cm | Residual \(y-\hat{y}\), cm | Squared residual, \(\text{cm}^2\) |
|---|---|---|---|---|
| 1 | 15 | 12 | 3 | 9 |
| 2 | 13 | 14 | -1 | 1 |
| 3 | 15 | 16 | -1 | 1 |
| 4 | 17 | 18 | -1 | 1 |
The sum is:
The residual of 3 centimeters contributes \(9\text{ cm}^2\), while each residual of \(-1\) centimeter contributes \(1\text{ cm}^2\). Its contribution is larger because the residual is larger in magnitude, not because it is positive. The sign tells whether the observed height is above or below the candidate line; after squaring, either sign gives a positive contribution.
This candidate’s SSE is 12 \(\text{cm}^2\). That number alone does not say whether the line is least squares: another line might have a smaller SSE for the same observations. To make that claim, its SSE must be at least as small as that of every other line of the form \(\hat{y}=a+bx\).
Compare Candidate Lines Carefully
For a fixed data set, the same observed \(y\)-values are used to evaluate every candidate line. Only the predictions—and therefore the residuals and SSE—change. A lower SSE means a better fit according to this particular criterion for these observations. It does not automatically mean the line is more sensible in context or that it captures every feature of the data.
Worked Example: Choose the Better of Two Candidates
A fictional cycling club records a rider’s training time, \(x\), in hours and a practice score, \(y\), in points. The observations are \((1,8),(2,11),(3,11),(4,14)\). Compare candidate lines A, \(\hat{y}=6+2x\), and B, \(\hat{y}=7+1.8x\), by their SSE values.
For line A, the predictions are \(8,10,12,14\). The observed-minus-predicted residuals are \(0,1,-1,0\), so:
For line B, the predictions are \(8.8,10.6,12.4,14.2\). Its residuals are \(-0.8,0.4,-1.4,-0.2\), giving:
Because \(2<2.80\), line A has the lower SSE of these two candidates for these observations. That comparison does not prove line A is the least-squares line: there may be another line with an even smaller SSE. The least-squares regression line is the one that minimizes SSE over all possible intercepts and slopes.
A Reliable SSE Workflow
When a question asks for the least-squares criterion or asks you to compare lines, keep the procedure organized. The residual sign matters when you describe whether a point lies above or below the line, but SSE uses each residual squared.
Write the line being evaluated and keep the explanatory variable \(x\) and response variable \(y\) in their stated roles.
Substitute each observed \(x\)-value into the candidate line to calculate \(\hat{y}\).
For each observation, use observed response minus predicted response: \(y-\hat{y}\).
Square every residual separately, then sum those nonnegative contributions to get SSE.
A smaller SSE identifies the better of the candidates compared. Call a line the least-squares line only when it is the line that minimizes SSE over all possible lines.
Common Mistakes and AP Exam Tips
- Adding signed residuals instead of squared residuals. Positive and negative errors can cancel. Show the squares explicitly before adding to demonstrate that you are using SSE.
- Reversing the residual subtraction. The standard residual is observed \(y\) minus predicted \(\hat{y}\). Reversing it changes the sign and would alter an interpretation of whether a point is above or below the line, even though its square is unchanged.
- Calling any evaluated line the least-squares line. A line with a calculated SSE is only a candidate unless it is known to minimize SSE. In the three-point example, the candidate \(2+\frac{2}{3}x\) has SSE \(\frac{8}{9}\), while the actual least-squares line is \(\frac{4}{3}+x\) with SSE \(\frac{2}{3}\).
- Using perpendicular distances. Regression residuals are vertical differences in the response direction at the observed \(x\)-values. Do not substitute shortest distances from points to the line.
- Forgetting that large residuals have greater influence. A residual twice as large contributes four times as much when squared. Do not say that squaring treats all errors equally.
- Giving SSE in the wrong units. If the response is measured in centimeters, residuals are in centimeters and SSE is in square centimeters. State this when units are relevant.
- Claiming the minimum makes every prediction close. Least squares minimizes the total for the observed data. A line can still have a substantial residual for one case, and the criterion alone does not establish that a linear model is appropriate.
A strong AP response identifies the candidate line, shows predicted values, calculates residuals as \(y-\hat{y}\), squares them, and adds the contributions. When comparing lines, state which has the smaller SSE and limit the conclusion to that comparison unless the least-squares property has been established.
Check Your Understanding
Use observed response minus predicted response when calculating each residual. Show enough work to distinguish a candidate line from the least-squares line.
- An observation has \(y=18\) and a candidate line predicts \(\hat{y}=16\). Find the residual and its contribution to SSE. What does the residual’s sign indicate?
- A candidate line has residuals \(2,-2,1\). Calculate its SSE and explain why adding the signed residuals would not give the criterion.
- For one data set, candidate A has SSE \(7.2\) and candidate B has SSE \(6.9\). Which candidate is better by the least-squares criterion? What additional claim cannot be made from this comparison alone?
- Explain why a residual of 4 contributes more to SSE than a residual of \(-2\), even though the latter is negative.
- If a response is measured in liters, what are the units of a residual and of SSE?