Tutorials › AP Statistics › Standard Deviation of the Residuals, s

Residuals · Tutorial 894 of 1000

Standard Deviation of the Residuals, s

Learn to find and interpret residual standard deviation in regression output, including what it says—and does not say—about prediction errors.

Intermediate 9 min read

What You'll Learn

  • Identify residual standard deviation \(s\) in computer output.
  • Interpret \(s\) as the typical vertical size of prediction errors, in response units.
  • Connect residual standard deviation to the sum of squared residuals and degrees of freedom.
  • Distinguish \(s\) from an individual residual and from the standard deviation of the response.
  • Compare residual standard deviations for models predicting the same response on the same data.
  • Avoid treating \(s\) as a guarantee about every prediction or as an average absolute error.

A Typical Size for Regression Prediction Errors

In “Comparing Residuals Across Observations,” you used absolute residuals to decide which observed cases were closest to their predictions. That compares individual prediction errors. To describe the overall size of the errors for a fitted regression line, we use the standard deviation of the residuals, denoted by \(s\). Computer output may call it the “residual standard error” or “residual standard deviation.”

The value of \(s\) gives a typical size for the vertical distances between observed responses and the responses predicted by the regression line. It is expressed in the response variable’s units. For example, if the model predicts plant height in centimeters and \(s=4\), the residuals typically have a size of about 4 centimeters. This is a useful summary of prediction error, not a claim that every error is exactly 4 centimeters.

Definition: The standard deviation of the residuals, \(s\), summarizes the typical size of the residuals around a least-squares regression line. It is measured in the response variable’s units and is interpreted as a typical vertical prediction error for the cases used to fit the line.

The calculation is based on the sum of squared residuals, abbreviated SSE. As in the earlier tutorials on the least-squares criterion and calculating residuals, each residual is \(y-\hat{y}\). For a least-squares line with an intercept, the residual standard deviation is calculated using \(n-2\) degrees of freedom.

$$ s=\sqrt{\frac{\text{SSE}}{n-2}} =\sqrt{\frac{\sum (y-\hat{y})^2}{n-2}} $$

The subtraction \(n-2\) reflects that the line was fitted by estimating two coefficients: the intercept and the slope. Squaring the residuals makes their contributions nonnegative, and dividing by \(n-2\) gives a variance-like quantity in squared response units. Taking the square root returns \(s\) to the response’s original units. Thus, if the response is measured in centimeters, \(s\) is in centimeters—not centimeters squared.

Regression software often reports \(s\) directly along with its degrees of freedom. If the output says “Residual standard error: 4.00 on 16 degrees of freedom,” the reported typical prediction-error size is 4.00 response units. The degrees of freedom also let you recover the sample size for a simple linear regression with an intercept: \(n-2=16\), so \(n=18\).

Reading and Interpreting \(s\) from Output

When you read regression output, first identify the response variable and its units. Then find the residual standard error or residual standard deviation. State what that value means for predictions of that response. Do not interpret it as a slope, as a percentage, or as the standard deviation of the predictor.

1
Locate the residual standard deviation.
Look for \(s\), “residual standard error,” or “residual standard deviation” in the regression output.
2
Identify the response and its units.
The units for \(s\) match the units of the response variable \(y\), not those of the predictor \(x\).
3
Describe the typical error size.
Say that the observed responses typically differ from the line’s predictions by about \(s\) response units for the fitted cases.
4
Keep the claim appropriately limited.
Do not say that every residual equals \(s\), or that \(s\) alone guarantees accurate predictions for new cases.

An individual residual retains its sign and describes one case’s direction and error size. By contrast, \(s\) is a single summary for the set of residuals. It does not tell whether a particular case was overpredicted or underpredicted. The sign of an individual residual still has the meaning explained in “Sign of a Residual: Over- and Underprediction.”

Also, \(s\) is not the standard deviation \(s_y\) of all observed response values. The standard deviation \(s_y\) describes how responses vary around their mean. Residual standard deviation \(s\) describes the typical vertical differences between observed responses and the regression line’s predictions. A response can vary widely overall while still having a smaller residual spread if the line accounts for much of that variation.

Worked Example: Plant-Height Regression Output

A fictional garden project uses a regression line to predict tomato plant height, in centimeters, from daily sunlight, in hours. The software reports the following summary.

Output itemValue
Number of plants, \(n\)18
Sum of squared residuals, SSE256 cm²
Residual standard error4.00 cm on 16 degrees of freedom

State. Interpret \(s\) in context and verify how the reported value relates to SSE and the degrees of freedom.

Plan. The response is plant height, so \(s\) must be interpreted in centimeters. Use \(s=\sqrt{\text{SSE}/(n-2)}\), and describe the result as a typical prediction-error size for the plants used to fit the line.

Do. The degrees of freedom are \(18-2=16\), consistent with the output. Substituting the reported SSE gives \(s=\sqrt{256/16}=\sqrt{16}=4.00\) centimeters. As a check, \(4.00^2(16)=16(16)=256\) cm², the stated SSE.

Conclude. For these plants, the observed heights typically differ from the heights predicted by the fitted line by about 4 centimeters. This describes the general size of the residuals; it does not say that every plant’s prediction is off by exactly 4 centimeters.

How \(s\) Relates to Individual Residuals

Because \(s\) summarizes a whole set of residuals, it should not be assigned to one observation as though it were that observation’s error. For a particular case, calculate or read its residual and interpret that signed value in context. Use \(s\) to describe the typical scale of prediction errors across the fitted data.

A residual can be smaller or larger in magnitude than \(s\). A residual larger than \(s\) is not automatically an error in the calculation, and \(s\) is not a maximum. The value alone does not state what proportion of residuals fall within a particular range. Such claims would need additional information and assumptions; they do not follow just from reading the residual standard deviation.

Check the residual plot as well. In “Reading a Residual Plot for Random Scatter,” you learned to look for residuals scattered above and below zero without a systematic pattern. A small \(s\) does not erase a curved pattern or changing spread. Conversely, a value of \(s\) has little practical meaning without considering the response’s scale and the context in which prediction errors matter.

Worked Example: Tablet Battery-Time Predictions

A fictional technology club fits a regression line to predict how many hours a tablet takes to fall from a full charge to 20 percent, based on its screen-brightness setting. The response is battery time in hours. The output reports “Residual standard error: 1.8 on 30 degrees of freedom.” One tablet has a residual of \(-4.5\) hours.

State. Explain what the reported \(s\) says about prediction errors, and interpret the individual tablet’s residual in relation to its prediction.

Plan. Interpret the residual standard error in hours because battery time is the response. Interpret the negative residual separately: observed minus predicted is negative, so that tablet’s observed time was below the fitted prediction.

Do. The output gives \(s=1.8\) hours, so residuals typically have a size of about 1.8 hours for the tablets used to fit the line. The individual residual has magnitude \(|-4.5|=4.5\) hours. Its magnitude is \(4.5/1.8=2.5\) times the reported \(s\). The negative sign means the model’s predicted battery time for this tablet was 4.5 hours greater than its observed battery time.

Conclude. The model’s typical prediction error for the fitted tablets is about 1.8 hours, while this tablet’s observed time was 4.5 hours below its prediction. The residual is larger in magnitude than \(s\), but \(s\) alone does not establish whether this individual error is unusual or explain why it occurred.

Comparing Residual Standard Deviations

A smaller \(s\) means a smaller typical vertical prediction error in the response’s units, for the data and model being considered. Comparisons are clearest when the models use the same response variable, the same units, and the same observations. If the response variables or their scales differ, comparing the numerical \(s\) values directly may be misleading.

Even a fair comparison has limits. A lower residual standard deviation describes a smaller residual scale for the fitted data; by itself, it does not prove that predictions will be better for every case or for future observations. Consider the residual plots and the purpose of the model, too. The residual plots may reveal structure that a single summary number cannot show.

Keep the units visible when comparing values. If monthly water use is measured in cubic meters, a residual standard deviation of 8.6 means a typical error size of about 8.6 cubic meters per month. It does not mean 8.6 percent, and it is not a statement about how much water use changes when the predictor increases.

Worked Example: Comparing Two Water-Use Models

A fictional environmental group fits two simple regression lines to the same 20 neighborhoods, using different predictors to predict each neighborhood’s monthly household water use. The response is measured in cubic meters per month. Model A uses household size; Model B uses lawn area.

ModelPredictorResidual standard errorDegrees of freedom
AHousehold size8.6 m³/month18
BLawn area7.4 m³/month18

State. Compare the typical prediction-error sizes for these two fitted models, keeping the conclusion limited to the data and response described.

Plan. Both models predict the same response for the same neighborhoods, so their residual standard deviations can be compared directly. The smaller \(s\) indicates the smaller typical residual size in cubic meters per month.

Do. Model A’s typical residual size is about 8.6 m³/month. Model B’s is about 7.4 m³/month. The difference is \(8.6-7.4=1.2\) m³/month, so Model B’s reported typical residual size is 1.2 m³/month smaller. The shared degrees of freedom, \(20-2=18\), are consistent with each model being a simple regression fitted to the same 20 neighborhoods.

Conclude. For these neighborhoods, Model B has the smaller residual standard deviation, so its fitted predictions of monthly water use typically differ from observed water use by less than Model A’s do. This comparison does not establish that Model B will predict every neighborhood better or that lawn area causes water use to change.

Common Mistakes and AP Exam Tips

  • Using predictor units instead of response units. The residual standard deviation uses the units of \(y\). If the model predicts hours, \(s\) is in hours, even if the predictor is measured in dollars or centimeters.
  • Calling \(s\) an average absolute residual. It is calculated from squared residuals and a degrees-of-freedom adjustment. Do not describe it as the mean of \(|y-\hat{y}|\).
  • Saying every prediction is wrong by exactly \(s\). Use “typically differs by about \(s\) response units.” Individual residuals vary and can be either smaller or larger.
  • Confusing \(s\) with the standard deviation of \(y\). The response’s standard deviation describes spread around its mean; residual standard deviation summarizes spread around the fitted line.
  • Interpreting \(s\) as direction. Since residual standard deviation is nonnegative, it does not tell whether the line tends to overpredict or underpredict a particular case. Use that case’s signed residual for direction.
  • Claiming that a lower \(s\) proves a model is best in every way. State that its typical residual size is smaller for the specified response, cases, and units. Do not claim universal or causal superiority from \(s\) alone.

For full-credit communication, name the response, give the units, and state that observed responses typically differ from the regression line’s predictions by about \(s\) units. If comparing two values, identify the data and response being compared and say which has the smaller typical residual size.

Key takeaway: The residual standard deviation \(s\) summarizes the typical size of prediction errors around a least-squares line. Interpret it in the response variable’s units, and remember that it describes a general scale—not the direction or exact size of every individual residual.

Check Your Understanding

Use response units and describe \(s\) as a typical error size, not a guarantee for each individual case.

  1. A regression model predicts the length of a bird’s wing in millimeters. Its residual standard error is 5.2. Interpret this value in context.
  2. A simple regression uses 25 observations. How many degrees of freedom are reported for its residual standard deviation?
  3. A model predicts delivery time in minutes and has \(s=3\) minutes. One delivery has residual \(-6\) minutes. What does the sign say, and how does the absolute residual compare with \(s\)?
  4. Why is residual standard deviation not the same as the standard deviation of the observed response values?
  5. Two models predict the same response, using the same cases and units. Their residual standard deviations are 6.3 and 5.8. Which has the smaller typical residual size, and by how much?