Two Measures of Regression Fit
In “Effect of Unusual Points on Correlation,” you compared \(r\) and its magnitude before and after adding a point. Two related measures describe other aspects of regression fit: \(r^2\) summarizes the proportion of variation in the response accounted for by the least-squares regression on the explanatory variable, and \(s\) summarizes the typical size of the residuals.
These measures answer different questions. \(r^2\) is a proportion, so it has no units. The residual standard deviation \(s\) is in the response variable’s units. Neither measure alone tells you whether a point is unusual or influential; to judge its effect, compare the values with and without that point.
A larger \(r^2\) means a larger proportion of the response variation is accounted for by the linear model. A smaller \(s\) means the residuals are typically closer to the fitted line in the response’s units. These statements are about the data and model fit; they do not establish that changes in \(x\) cause changes in \(y\).
How \(r^2\) and \(s\) Are Calculated
Use \(S_{xx}\) for the sum of squared deviations of the \(x\)-values from \(\bar{x}\), \(S_{yy}\) for the sum of squared deviations of the \(y\)-values from \(\bar{y}\), and \(S_{xy}\) for the sum of the paired products of those deviations. In this setting, the total sum of squares, \(\mathrm{SST}\), is \(S_{yy}\). The regression sum of squares, \(\mathrm{SSR}\), is the part accounted for by the fitted line. The error sum of squares, \(\mathrm{SSE}\), is the sum of squared residuals.
For a simple linear regression with an intercept, the residual standard deviation is calculated using the number of observations \(n\) and \(n-2\) degrees of freedom. The two fitted quantities are the intercept and slope, which is why the denominator is \(n-2\), not \(n\).
When one point is added, the fitted line can change, so the residuals for the original points can change too. The new \(\mathrm{SSE}\) is calculated from the refitted line, not by simply adding the squared residual from the old line. The new \(r^2\) also uses the new total response variation, and the new \(s\) uses the new sample size.
As in “Effect of Unusual Points on Correlation,” the sums can be updated using the point’s displacements from the original means. For an added point \((x_0,y_0)\), let \(d_x=x_0-\bar{x}\), \(d_y=y_0-\bar{y}\), and suppose the original data contain \(n\) points.
Use those updated sums to find the new \(\mathrm{SSR}\), \(\mathrm{SSE}\), \(r^2\), and \(s\). The calculations make an important point clear: an unusual point can increase one measure of fit while making the other seem less favorable. \(r^2\) is relative to total response variation, while \(s\) measures residual size in response units.
Worked Examples: Comparing Fit With and Without a Point
Worked Example: \(r^2\) Rises While \(s\) Rises
Original AP-style question. In a fictional greenhouse test, \(x\) is the number of hours a lamp is used and \(y\) is a seedling’s height in centimeters. Four observations are \((1,2)\), \((2,3)\), \((3,5)\), and \((4,6)\). A fifth observation, \((6,8)\), is added. Compare \(r^2\) and \(s\) before and after.
State. We will compare the proportion of response variation accounted for by the fitted line and the typical residual size in centimeters. This is a descriptive comparison of the two data sets, not an inference procedure; randomization and distribution conditions are not needed to calculate these summaries.
Plan. Calculate the original sums and fit measures. Then update the sums for the added point, refit through the updated sums, and calculate the new \(r^2\) and \(s\), using \(n-2\) degrees of freedom for each data set.
Do. For the four original observations, \(\bar{x}=2.5\) and \(\bar{y}=4\). The deviations give \(S_{xx}=5\), \(S_{yy}=10\), and \(S_{xy}=7\). Thus,
There are four observations, so \(n-2=2\). The original fit measures are
For the added point, \(d_x=6-2.5=3.5\) and \(d_y=8-4=4\). The update factor is \(4/5=0.8\), so
For the five-point data set, \(\mathrm{SSR}=18.2^2/14.8=331.24/14.8\approx22.3811\). Therefore, \(\mathrm{SSE}=22.8-22.3811\approx0.4189\). With \(n-2=3\),
Conclude. After adding the fifth seedling, \(r^2\) rises from \(0.9800\) to about \(0.9816\), while \(s\) rises from about \(0.3162\) cm to \(0.3737\) cm. The line accounts for a slightly larger proportion of the response variation, but the typical residual size in centimeters is larger. The two changes are not contradictory: \(r^2\) is relative to the total response variation, whereas \(s\) is an absolute measure in response units.
Worked Example: A Point at the Mean of \(x\) Lowers \(r^2\)
Original AP-style question. A fictional class project records \(x\), weekly study hours, and \(y\), quiz points. Four students have \((1,2)\), \((2,4)\), \((3,5)\), and \((4,7)\). A student with \((2.5,10)\) is added. Compare the two fit measures.
State and plan. The new point has \(x\) equal to the original mean, but its quiz score is high compared with the others. We will calculate \(r^2\) and \(s\) for both data sets, using updated sums after adding the point. These are descriptive summaries, so no inference conditions are required.
Do. For the original four observations, \(\bar{x}=2.5\) and \(\bar{y}=4.5\). The deviation sums are \(S_{xx}=5\), \(S_{yy}=13\), and \(S_{xy}=8\). Thus, \(\mathrm{SSR}=8^2/5=12.8\), and \(\mathrm{SSE}=13-12.8=0.2\). With two degrees of freedom,
For the added point, \(d_x=2.5-2.5=0\) and \(d_y=10-4.5=5.5\). The update factor is \(0.8\). Therefore,
The new regression sum of squares is still \(8^2/5=12.8\), so the new \(\mathrm{SSE}\) is \(37.2-12.8=24.4\). With five observations, there are three degrees of freedom:
Conclude. Adding the student lowers \(r^2\) from about \(0.9846\) to \(0.3441\), and raises \(s\) from about \(0.3162\) to \(2.8519\) quiz points. Because the point is at the original mean \(x\), it adds no \(x\)-variation or cross-product contribution to the updated sums. Its high quiz score adds substantial response variation without adding corresponding linear co-movement, so a much smaller proportion of that variation is accounted for by the line.
Worked Example: A High-Leverage Point Lowers \(s\)
Original AP-style question. In a fictional greenhouse test, \(x\) is supplemental lamp hours per day and \(y\) is seedling growth in centimeters. Four observations are \((1,3)\), \((2,4)\), \((3,6)\), and \((4,7)\). A fifth observation, \((10,16.5)\), is added. What happens to \(r^2\) and \(s\)?
State and plan. The new observation has an \(x\)-value far above the original values, so it has high leverage. We will calculate both fit measures before and after adding it. As in the earlier examples, this is a descriptive comparison, not an inference procedure.
Do. For the original observations, \(\bar{x}=2.5\) and \(\bar{y}=5\). The deviation sums are \(S_{xx}=5\), \(S_{yy}=10\), and \(S_{xy}=7\). Then \(\mathrm{SSR}=7^2/5=9.8\) and \(\mathrm{SSE}=10-9.8=0.2\). The original fit measures are
The original regression line has slope \(S_{xy}/S_{xx}=7/5=1.4\) and intercept \(\bar{y}-1.4\bar{x}=1.5\). At \(x=10\), it predicts \(1.5+1.4(10)=15.5\) cm, so the added observation is 1 cm above that original line. To calculate the fit after adding it, use \(d_x=10-2.5=7.5\), \(d_y=16.5-5=11.5\), and the update factor \(0.8\):
Thus, \(\mathrm{SSR}=76^2/50=115.52\) and \(\mathrm{SSE}=115.8-115.52=0.28\). With five observations and three degrees of freedom,
Conclude. Adding this high-leverage observation raises \(r^2\) from \(0.9800\) to about \(0.9976\) and lowers \(s\) from about \(0.3162\) cm to \(0.3055\) cm. The point extends the positive pattern and is close to the original fitted line, so it contributes substantial variation that the new line accounts for. The values describe this fictional data set; they do not by themselves show that lamp hours cause seedling growth.
Interpreting the Comparison
The examples show why you should calculate and compare both measures rather than treating one as a substitute for the other. In the first example, \(r^2\) rose while \(s\) also rose. In the second, \(r^2\) fell while \(s\) rose. In the third, \(r^2\) rose while \(s\) fell. The point’s position and the response’s scale both matter.
A useful way to organize the comparison is to ask two separate questions. First, what fraction of the response variation is accounted for by the line? That is the \(r^2\) question. Second, how large are the residuals, measured in response units? That is the \(s\) question. When adding a point, keep track of changes in \(\mathrm{SSE}\), total response variation, and the degrees of freedom used to calculate \(s\).
If you use technology to compare regressions, make sure the point is included for one fit and omitted for the other. Record \(r^2\) and the residual standard deviation for each fit, then explain the changes in context. A point’s unusual appearance alone does not establish how much either measure changes. The with-and-without comparison does.
Common Mistakes and AP Exam Tip
- Claiming that a higher \(r^2\) always means smaller residuals. \(r^2\) is a proportion of total response variation, while \(s\) is in response units. Compare each measure separately.
- Comparing \(s\) without accounting for the number of observations. The formula uses \(n-2\) degrees of freedom. Recalculate this denominator after adding or removing a point.
- Using the old line’s residual as the new \(\mathrm{SSE}\). Adding a point changes the fitted line and can change every residual. Find \(\mathrm{SSE}\) for the refitted data.
- Describing \(r^2\) as the percent of points on the line. It is the proportion of variation in the response accounted for by the linear regression, not a count of points close to the line.
- Giving only numbers, without interpretation. State which fit has the larger \(r^2\), whether \(s\) rises or falls, and what the \(s\) change means in response units.
- Making a causal claim from a change in fit. A point can alter descriptive summaries, but that alone does not show that changing the explanatory variable causes a response change.
A strong AP response identifies the comparison and uses both units and context: “After the high-study-hours observation is added, \(r^2\) decreases, so a smaller proportion of the quiz-score variation is accounted for by the fitted line. The residual standard deviation increases, meaning the typical residual is larger in quiz points.” Use that form only when the calculated results support it.
Check Your Understanding
Use the meanings of \(r^2\) and \(s\) to answer each question.
- What does \(r^2\) measure, and why does it have no units?
- A point is added and \(r^2\) rises, but \(s\) also rises. Is that possible? Explain the difference between the measures.
- Why must you use the refitted line when finding the \(\mathrm{SSE}\) after adding a point?
- For five observations in a simple linear regression, what denominator is used to calculate \(s\) from \(\mathrm{SSE}\)?
- A new observation has \(x=\bar{x}\) for the original data. What does that imply about its contribution to the updated \(S_{xx}\) and \(S_{xy}\)?