Tutorials › AP Statistics › r-squared and Residual Standard Deviation Together

Coefficient of determination · Tutorial 913 of 1000

r-squared and Residual Standard Deviation Together

Use \(r^2\) to describe variation accounted for and \(s\) to describe the typical size of residuals in the response’s units.

Intermediate 9 min read

What You'll Learn

  • Distinguish the proportion of response variation accounted for from the typical size of residuals.
  • Interpret \(r^2\) and residual standard deviation \(s\) in context, using the right units.
  • Connect \(r^2\), \(SST\), \(SSE\), and \(s\) for a least-squares regression line.
  • Calculate \(s\) from \(r^2\), sample size, and the sample standard deviation of the response.
  • Compare model fit summaries without treating either statistic as a complete measure of prediction quality.

Two Summaries, Two Kinds of Information

In “Comparing Models Using \(r^2\),” you used the coefficient of determination to compare the proportion of response variation accounted for by a linear model. Earlier, in “Standard Deviation of the Residuals, \(s\),” you learned that \(s\) summarizes the typical size of residuals around a least-squares regression line. These summaries describe different aspects of fit: \(r^2\) is a proportion with no units, while \(s\) is measured in the response variable’s units.

That difference matters when interpreting a model. An \(r^2\) of 0.64 means that about 64% of the variation in the response is accounted for by its linear relationship with the explanatory variable. It does not tell you, by itself, how far observed responses typically are from their predicted values. The residual standard deviation \(s\) gives a scale for those vertical prediction errors, but it does not tell you what proportion of the response’s variation the line accounts for.

The two statistics are related for a least-squares regression line with an intercept. The connection uses the quantities you met in “Total Variation in the Response Variable” and “What Variability Explained Means”: \(SST\), the total sum of squares, and \(SSE\), the sum of squared residuals.

Key distinction: \(r^2\) describes the fraction of variation in the response accounted for by its linear relationship with the explanatory variable. The residual standard deviation \(s\) describes the typical size of residuals in the response’s units. Use both when you want to describe both relative variation accounted for and the scale of prediction errors.

How the Statistics Are Connected

For a least-squares regression line with an intercept, \(r^2=1-\frac{SSE}{SST}\). Rearranging gives \(SSE=(1-r^2)SST\). The residual standard deviation is \(s=\sqrt{\frac{SSE}{n-2}}\) for simple linear regression, where \(n\) is the number of paired observations. The denominator \(n-2\) accounts for estimating the intercept and slope.

The sample standard deviation of the response, \(s_y\), is related to total variation by \(SST=(n-1)s_y^2\). Combining these relationships gives a useful connection between \(r^2\) and \(s\):

$$ s=\sqrt{\frac{(1-r^2)SST}{n-2}} =s_y\sqrt{\frac{(1-r^2)(n-1)}{n-2}} $$

This formula helps explain why the two statistics convey distinct information. For the same response observations, \(n\) and \(SST\) are fixed. A larger \(r^2\) then means a smaller \(SSE\), which means a smaller \(s\). But when the response values or their spread differ, the same \(r^2\) can go with very different values of \(s\). The formula should not lead you to treat \(s\) as a percentage or \(r^2\) as a measure in response units.

Formula: For simple linear regression with an intercept, \(s=\sqrt{\frac{SSE}{n-2}}\), and \(SSE=(1-r^2)SST\). If \(s_y\) is known, then \(s=s_y\sqrt{\frac{(1-r^2)(n-1)}{n-2}}\). Keep the response’s units when reporting \(s\).

Worked Examples: Reading Both Summaries

Worked Example: Same \(r^2\), Different Residual Scales

Two fictional groups of 25 students are used to fit separate linear models that predict a response measured in points. Each model has \(r^2=0.64\). In Group A, the sample standard deviation of the response is \(s_y=5\) points; in Group B, it is \(s_y=20\) points. Find and interpret \(s\) for each model.

State. Both models account for the same proportion of response variation, but the response values have different spreads in the two groups. We will calculate each residual standard deviation to describe the typical size of residuals in points.

Plan. For each group, use \(SST=(n-1)s_y^2\), then \(SSE=(1-r^2)SST\), and finally \(s=\sqrt{\frac{SSE}{n-2}}\). Here, \(n=25\), so the denominator for \(s\) is \(25-2=23\).

Do. For Group A, \(SST=(25-1)(5^2)=24(25)=600\) points squared. Then \(SSE=(1-0.64)(600)=0.36(600)=216\) points squared, so \(s=\sqrt{\frac{216}{23}}\approx3.0645\), or about 3.06 points. Equivalently, the multiplier from the formula using \(s_y\) is \(\sqrt{\frac{(0.36)(24)}{23}}=\sqrt{\frac{8.64}{23}}\approx0.6129\), and \(s=5(0.6129)\approx3.06\) points.

For Group B, \(SST=(25-1)(20^2)=24(400)=9600\) points squared. Then \(SSE=(1-0.64)(9600)=0.36(9600)=3456\) points squared, so \(s=\sqrt{\frac{3456}{23}}\approx12.2581\), or about 12.26 points. Using the same multiplier, \(s=20(0.6129)\approx12.26\) points.

Conclude. Both models account for about 64% of the variation in their group’s response values. However, residuals typically differ from zero by about 3.06 points in Group A and about 12.26 points in Group B. Thus, equal \(r^2\) values do not imply equal prediction-error scales; the response values in Group B are more spread out.

Worked Example: Calculating \(s\) From \(r^2\) and Response Spread

A fictional horticulture class models the number of days until a seed sprouts using a measurement of soil moisture. The model uses \(n=16\) seed observations, has \(r^2=0.75\), and the sample standard deviation of sprouting time is \(s_y=8\) days. Find the residual standard deviation and explain what the two statistics say.

State. The response is sprouting time, measured in days. We want to describe both the proportion of its variation accounted for by the linear relationship with soil moisture and the typical residual size.

Plan. Use \(SST=(n-1)s_y^2\), \(SSE=(1-r^2)SST\), and \(s=\sqrt{\frac{SSE}{n-2}}\). These formulas apply to the fitted simple linear regression with an intercept. No inference procedure is being used, so this is a descriptive calculation for the observed data.

Do. First, \(SST=(16-1)(8^2)=15(64)=960\) days squared. The unexplained sum of squared residuals is \(SSE=(1-0.75)(960)=0.25(960)=240\) days squared. Therefore, \(s=\sqrt{\frac{240}{16-2}}=\sqrt{\frac{240}{14}}\approx4.1404\), or about 4.14 days. Checking with the formula in terms of \(s_y\), \(s=8\sqrt{\frac{(0.25)(15)}{14}}=8\sqrt{\frac{3.75}{14}}\approx8(0.5175)\approx4.14\) days.

Conclude. About 75% of the variation in sprouting time among these seeds is accounted for by its linear relationship with soil moisture. Residuals typically have a size of about 4.14 days. The percentage and the number of days answer different questions; neither says that every prediction is within 4.14 days of the observed time.

Worked Example: Comparing Both Measures for Two Models

A fictional recreation program uses the same 20 participants and response values to fit two linear models predicting a fitness score in points. The response has sample standard deviation \(s_y=10\) points. A model using weekly activity hours has \(r^2=0.81\); a model using weekly practice sessions has \(r^2=0.64\). Calculate each model’s \(s\) and compare the summaries.

State. The response and observations are the same for both models, so their \(r^2\) values are directly comparable. Since both models use the same response values, they also share the same \(SST\).

Plan. Calculate \(SST=(n-1)s_y^2\), use each \(r^2\) to obtain \(SSE=(1-r^2)SST\), then divide by \(n-2\) and take the square root to find \(s\). Report \(r^2\) as a proportion and \(s\) in fitness-score points.

Do. The shared total variation is \(SST=(20-1)(10^2)=19(100)=1900\) points squared. For the activity-hours model, \(SSE=(1-0.81)(1900)=0.19(1900)=361\) points squared, so \(s=\sqrt{\frac{361}{20-2}}=\sqrt{\frac{361}{18}}\approx4.4783\), or about 4.48 points. For the practice-sessions model, \(SSE=(1-0.64)(1900)=0.36(1900)=684\) points squared, so \(s=\sqrt{\frac{684}{18}}=\sqrt{38}\approx6.1644\), or about 6.16 points.

Conclude. The activity-hours model accounts for about 81% of the variation in fitness scores, compared with about 64% for the practice-sessions model, a difference of 17 percentage points. Its residual standard deviation is also smaller: residuals typically have a size of about 4.48 points rather than 6.16 points. For these same response observations, the higher \(r^2\) and lower \(s\) agree: the activity-hours model leaves less unexplained variation and has a smaller residual scale. These summaries alone do not establish that the model is appropriate for every purpose or that activity hours cause fitness scores to change.

Interpret Both Measures Carefully

Use units to keep the interpretations separate. An \(r^2\) of 0.81 is interpreted as about 81% of the variation in the response accounted for by its linear relationship with the explanatory variable. An \(s\) of 4.48 points is interpreted as a typical size of residuals around the regression line of about 4.48 points. Do not say that \(r^2\) is “81% accurate,” or that \(s\) means every prediction is off by exactly 4.48 points.

A value of \(s\) is a summary of residual size, not a guarantee about any individual case. Some residuals may be much smaller than \(s\), and some may be much larger. To discuss a particular observation, use its residual as in “Interpreting a Residual in Context.” To assess whether the line misses in an organized way, inspect the residual plot, as in “Reading a Residual Plot for Random Scatter,” “Curved Patterns in a Residual Plot,” and “Fan-Shaped Residual Plots and Changing Spread.”

It is also important to notice what happens when the response’s units change. Converting a response from points to a scale that is twice as large doubles the numerical residuals and \(s\), while \(r^2\) remains the same. The coefficient of determination describes a proportion; the residual standard deviation describes a distance on the response scale. That is one reason a proportion alone cannot communicate the practical size of prediction errors.

When two models use the same response observations, \(r^2\) and \(s\) are linked: the one with a larger \(r^2\) has a smaller \(SSE\) and a smaller \(s\). Still, state what each measure means rather than treating one as a substitute for the other. When comparing models fitted to different response data or different sets of observations, a direct comparison may not be fair; describe the data behind each value and avoid implying that the numbers alone settle which model is better overall.

Common Mistakes and AP Exam Tips

  • Giving \(s\) without response units. State that the residual standard deviation is in the response variable’s units, such as points, days, or centimeters.
  • Describing \(r^2\) as prediction accuracy. Full-credit wording says the percentage of variation in the response accounted for by its linear relationship with the explanatory variable.
  • Treating \(s\) as a maximum error. Say that \(s\) summarizes the typical size of residuals. Do not claim that every observed response is within \(s\) of its prediction.
  • Using the wrong denominator. For a simple linear regression line, calculate \(s=\sqrt{\frac{SSE}{n-2}}\), not \(\sqrt{\frac{SSE}{n}}\) or \(\sqrt{\frac{SSE}{n-1}}\).
  • Confusing sums of squares with response units. \(SST\) and \(SSE\) are in squared response units; taking the square root in the formula for \(s\) returns to the response’s units.
  • Claiming a model is best in every way. Explain what the two summaries show, then remember that model suitability and residual patterns matter too. Neither statistic establishes causation.

A strong response names the response and explanatory variable, interprets \(r^2\) as a proportion of response variation, and interprets \(s\) as a typical residual size with units. If you compare models, check that the response observations match and explain what each statistic adds to the comparison.

Key takeaway: \(r^2\) tells how much of the response’s variation is accounted for by a linear relationship; \(s\) tells the typical residual size in response units. Together, they describe relative variation accounted for and the scale of prediction errors, not every feature of model quality.

Check Your Understanding

Use both statistics and their units to answer each question.

  1. A model has \(r^2=0.70\) and \(s=6\) centimeters for a response measuring plant height. Interpret each value in context.
  2. Two models use the same response observations. Model A has \(r^2=0.52\), and Model B has \(r^2=0.68\). Which model must have the smaller residual standard deviation, and why?
  3. Two separate groups each have \(n=25\) and \(r^2=0.64\). Their response standard deviations are 5 points and 20 points. Why should their residual standard deviations differ?
  4. Explain why it is incorrect to say that \(s=3\) days guarantees every prediction is within 3 days of the observed response.
  5. What happens to the numerical value of \(s\) if a response is converted to a scale twice as large? What happens to \(r^2\)?