Tutorials › AP Statistics › Using Residual Standard Deviation to Gauge Prediction Error

Extrapolation and prediction limits · Tutorial 954 of 1000

Using Residual Standard Deviation to Gauge Prediction Error

Use residual standard deviation to describe the typical response-unit error around a prediction made within the observed range.

Intermediate 9 min read

What You'll Learn

  • Describe what residual standard deviation says about the typical size of prediction errors.
  • Interpret \(s\) in the response variable’s units.
  • Calculate \(s\) from residuals for a fitted regression line.
  • Use \(s\) to put an interpolated prediction’s error scale in context.
  • Explain why \(s\) is not a maximum error or a guarantee for an individual prediction.

A Prediction Comes with an Error Scale

In “Reliability Within the Range of Data,” we saw that interpolation is generally better supported than extrapolation, but that being inside the observed \(x\)-range does not guarantee an accurate prediction. One useful way to describe the scatter around a regression line is the residual standard deviation, \(s\). It tells us the typical size of the vertical differences between observed responses and the responses predicted by the line.

For an interpolated prediction, \(s\) helps put the prediction in context. If a line predicts a response of 40 units and \(s\) is about 3 units, the model’s typical residual size is about 3 response units. That is more informative about error size than a unitless summary such as \(r^2\). But \(s\) does not say exactly how far a particular response will be from its prediction.

Definition: The residual standard deviation, \(s\), summarizes the typical size of the residuals—the observed responses minus the responses predicted by the least-squares line. It is measured in the response variable’s units. In simple linear regression, \(s=\sqrt{\frac{SSE}{n-2}}\), where \(SSE\) is the sum of squared residuals and \(n\) is the number of observed cases.

As in “r-squared and Residual Standard Deviation Together,” keep the summaries distinct: \(r^2\) describes the fraction of response variation accounted for by the linear relationship, while \(s\) describes residual scatter in response units. Here our goal is to use \(s\) to describe a typical error scale, not to claim certainty about one new response.

How to Use \(s\) for an Interpolated Prediction

Start by identifying the explanatory variable and its observed range. Compare the requested \(x\)-value with that range: a value between the endpoints is interpolation. Then find the line’s predicted response at that \(x\)-value. The prediction is the line’s estimate, not the response that must occur.

Next, report \(s\) with its units and connect it to the response. For example, “The residual standard deviation is about 2.3 minutes, so observed times typically differ from the fitted line’s predictions by about 2.3 minutes.” This statement describes the overall scatter of the observed cases around the line. It does not say that the prediction is off by exactly 2.3 minutes, or that every error is no more than 2.3 minutes.

Finally, consider whether the overall scatter is meaningful for the particular prediction. A value within the \(x\)-range can still be in a sparsely observed region, and the scatter may not be equally wide across the range. A residual plot or the data’s pattern can offer warnings about a linear model or changing scatter. The single summary \(s\) does not reveal every feature of the residuals.

How to phrase the error scale: “At [explanatory-variable value], the fitted line predicts [response value and units]. This is interpolation because [value] is within the observed \(x\)-range. The residual standard deviation is about [\(s\)] [response units], so residuals typically differ from the line’s predictions by about that amount. This is a typical scale, not a guaranteed error bound for this case.”

Worked Examples: Describing Typical Prediction Error

Worked Example: Prediction Between Observed Test Settings

Hypothetical setting. A technician records a machine setting \(x\), in setting units, and the machine’s output \(y\), in output units, for five invented test runs: \((1,12),(2,14),(3,17),(4,18),(5,24)\). Describe the prediction at \(x=3.5\) using \(s\).

State. The observed \(x\)-range is 1 to 5 setting units. Since 3.5 is between those endpoints, this is interpolation. We will find the fitted line and its residual standard deviation, then describe the prediction and its typical error scale.

Plan. Calculate the least-squares line, use it to find the five fitted responses and residuals, then calculate \(SSE\) and \(s\). Interpret \(s\) in output units; do not treat it as the error for any one test run.

Do. The means are \(\bar{x}=3\) setting units and \(\bar{y}=17\) output units. The sum of products of deviations is:

$$ \sum (x-\bar{x})(y-\bar{y}) =(-2)(-5)+(-1)(-3)+0(0)+1(1)+2(7)=28 $$

The sum of squared \(x\)-deviations is \((-2)^2+(-1)^2+0^2+1^2+2^2=10\). Therefore the slope is \(28/10=2.8\) output units per setting unit, and the intercept is \(17-2.8(3)=8.6\) output units. The fitted line is:

$$ \hat{y}=8.6+2.8x $$

At \(x=3.5\), the predicted output is \(8.6+2.8(3.5)=18.4\) output units. For the observed settings 1 through 5, the line predicts 11.4, 14.2, 17.0, 19.8, and 22.6 output units. Subtracting each prediction from its observed response gives residuals \(0.6,-0.2,0,-1.8,1.4\) output units. Thus:

$$ SSE=0.6^2+(-0.2)^2+0^2+(-1.8)^2+1.4^2 =0.36+0.04+0+3.24+1.96=5.60 $$

With \(n=5\), the residual standard deviation is:

$$ s=\sqrt{\frac{SSE}{n-2}} =\sqrt{\frac{5.60}{3}} \approx 1.366\text{ output units} $$

Conclude in context. The fitted line predicts an output of 18.4 units at a setting of 3.5 units, an interpolation. For these five invented test runs, the typical residual size is about 1.366 output units. This describes the overall scatter around the line; it does not guarantee that a new run at 3.5 will be within 1.366 units of 18.4.

Worked Example: A Larger Typical Error in a Service-Time Model

Hypothetical setting. A library tracks the number of minutes \(x\) a new checkout system has been operating during a practice session and the time \(y\), in seconds, needed to complete a test checkout. Five invented pairs are \((2,21),(4,25),(6,24),(8,31),(10,29)\). Describe the prediction at \(x=7\).

State. The observed operating-time range is 2 to 10 minutes. The requested value of 7 minutes lies inside the range, so this is interpolation. The size of \(s\), measured in seconds, will describe typical residual scatter for these test checkouts.

Plan. Find the regression line and residuals, calculate \(SSE\), and then use \(s=\sqrt{SSE/(n-2)}\). Interpret the prediction in seconds and relate the typical error scale to the checkout-time context.

Do. The means are \(\bar{x}=6\) minutes and \(\bar{y}=26\) seconds. The sum of products of deviations is:

$$ \sum (x-\bar{x})(y-\bar{y}) =(-4)(-5)+(-2)(-1)+0(-2)+2(5)+4(3)=44 $$

The sum of squared \(x\)-deviations is \((-4)^2+(-2)^2+0^2+2^2+4^2=40\). The slope is \(44/40=1.1\) seconds per minute, and the intercept is \(26-1.1(6)=19.4\) seconds. The fitted line is:

$$ \hat{y}=19.4+1.1x $$

At 7 minutes, the predicted checkout time is \(19.4+1.1(7)=27.1\) seconds. At the five observed operating times, the fitted responses are 21.6, 23.8, 26.0, 28.2, and 30.4 seconds. The residuals are \(-0.6,1.2,-2,2.8,-1.4\) seconds, so:

$$ SSE=(-0.6)^2+1.2^2+(-2)^2+2.8^2+(-1.4)^2 =0.36+1.44+4+7.84+1.96=15.60 $$

The residual standard deviation is:

$$ s=\sqrt{\frac{15.60}{5-2}} =\sqrt{5.20} \approx 2.280\text{ seconds} $$

Conclude in context. At 7 minutes of operation, the line predicts a checkout time of 27.1 seconds. The observed test checkouts typically differed from the line’s predictions by about 2.280 seconds. That is a description of the model’s typical residual size, not a claim that the particular checkout will take between 24.820 and 29.380 seconds or that it must be within 2.280 seconds of the prediction.

Worked Example: Comparing Typical Error with a Practical Scale

Hypothetical setting. A group tests a sensor at five calibration settings \(x\), measured in setting units, and records output \(y\), in output units. The invented observations are \((0,100),(2,104),(4,109),(6,111),(8,116)\). The sensor is to be checked at setting 5, and the team regards a difference of 2 output units as practically noticeable. Use \(s\) to describe the model’s typical error scale without making a guarantee.

State. Setting 5 is within the observed range from 0 to 8, so the requested prediction is interpolation. We will calculate the fitted prediction and \(s\), then compare the typical residual size with the stated 2-unit practical scale.

Plan. Find the regression equation and residuals from the five invented observations. Calculate \(s\) in output units. A comparison with 2 units can give context for the size of typical scatter, but it cannot tell us what fraction of individual errors will be below 2 units.

Do. The means are \(\bar{x}=4\) setting units and \(\bar{y}=108\) output units. The sum of products of deviations is:

$$ \sum (x-\bar{x})(y-\bar{y}) =(-4)(-8)+(-2)(-4)+0(1)+2(3)+4(8)=78 $$

The sum of squared \(x\)-deviations is \((-4)^2+(-2)^2+0^2+2^2+4^2=40\). The slope is \(78/40=1.95\) output units per setting unit, and the intercept is \(108-1.95(4)=100.2\) output units. Thus:

$$ \hat{y}=100.2+1.95x $$

At setting 5, the predicted output is \(100.2+1.95(5)=109.95\) output units. The fitted outputs for settings 0, 2, 4, 6, and 8 are 100.2, 104.1, 108.0, 111.9, and 115.8. The residuals are \(-0.2,-0.1,1,-0.9,0.2\) output units. Therefore:

$$ SSE=(-0.2)^2+(-0.1)^2+1^2+(-0.9)^2+0.2^2 =0.04+0.01+1+0.81+0.04=1.90 $$

With five observations, the residual standard deviation is:

$$ s=\sqrt{\frac{1.90}{5-2}} =\sqrt{0.6333\ldots} \approx 0.796\text{ output units} $$

Conclude in context. The line predicts an output of 109.95 units at setting 5, which is interpolation. The typical residual size, about 0.796 output units, is smaller than the team’s 2-unit practical scale. This suggests that the model’s overall residual scatter is modest relative to that scale. It does not establish that every sensor output will be within 2 units of its prediction, or that a particular future sensor reading will meet a requirement.

Common Mistakes and AP Exam Tips

  • Calling \(s\) the error for the requested case. \(s\) summarizes residual scatter across the observed data. A careful answer says the typical residual size is about \(s\) response units, not that this case’s error equals \(s\).
  • Treating \(s\) as a maximum error. Some residuals can be larger than \(s\). Do not say all observations or future responses must lie within \(s\) of the line.
  • Leaving off units. If the response is measured in seconds, \(s\) is in seconds. The explanatory variable’s units do not belong on \(s\).
  • Calling a prediction “reliable” only because it is interpolation. State that it is within the observed \(x\)-range, then describe the residual scatter and any relevant pattern or gaps in the data.
  • Turning a typical size into a probability. \(s\) alone does not tell us the probability that an individual response will be within a given distance of its prediction. Avoid claims about percentages or guaranteed ranges unless a separate method justifies them.
  • Comparing \(s\)-values without considering context. The same numerical value can be small for one response scale and large for another. Interpret \(s\) in the response’s units and, when possible, compare it with a meaningful scale in that setting.

A strong AP-style interpretation names the prediction, confirms interpolation, and explains \(s\) in context: “At a sensor setting of 5 units, the fitted line predicts an output of 109.95 units. This is interpolation because 5 lies within the observed range of 0 to 8 units. The residual standard deviation is about 0.796 output units, so observed outputs typically differ from the fitted line’s predictions by about 0.796 units. This describes typical scatter, not a guaranteed error for this sensor.”

Key takeaway: For an interpolated prediction, \(s\) gives a response-unit scale for the typical residual error around the fitted line. State the prediction and its context, interpret \(s\) in the response units, and make clear that typical error is not a maximum, a probability, or a guarantee for an individual response.

Check Your Understanding

Use the meaning and limitations of residual standard deviation to answer each question.

  1. A model predicts a delivery time of 18 minutes at an \(x\)-value inside the observed range. Its residual standard deviation is 2.5 minutes. Write a careful sentence interpreting \(s\) in context.
  2. A student says, “Since \(s=1.8\) centimeters, every observed height is within 1.8 centimeters of its predicted height.” What is wrong with this statement?
  3. A fitted model has \(s=4\) seconds, and an interpolated prediction is 35 seconds. What does \(s\) add to the prediction, and what does it not tell you?
  4. Two regression models have residual standard deviations of 3 units and 5 units. What additional context is needed before deciding which model has smaller typical errors in a practically meaningful sense?
  5. Why is it not justified to use \(s\) alone to claim that 95% of individual responses are within a specified distance of their predicted values?