Tutorials › AP Statistics › Using s to Describe Prediction Accuracy

Residuals · Tutorial 895 of 1000

Using s to Describe Prediction Accuracy

Use the residual standard deviation and a practical tolerance together to assess the scale of prediction errors without promising accuracy for every case.

Intermediate 9 min read

What You'll Learn

  • Define a practical tolerance for a prediction in the response variable’s units.
  • Compare the residual standard deviation with that tolerance to judge the scale of typical errors.
  • Calculate and interpret the ratio of residual standard deviation to tolerance.
  • Explain why a favorable comparison does not guarantee that each prediction meets the tolerance.
  • Check the data range and residual pattern before relying on the comparison.
  • Compare models’ typical error scales when they predict the same response for the same cases.

When Is a Prediction Accurate Enough?

In “Standard Deviation of the Residuals, \(s\),” you learned that \(s\) summarizes the typical size of the residuals around a least-squares regression line. A typical error size is useful, but whether it is acceptable depends on the purpose of the prediction. Being off by 3 seconds might be fine when estimating a long delivery time, but not when timing a brief process that must finish within a few seconds.

A practical tolerance is the largest prediction error that is acceptable for a particular use. It is stated in the response variable’s units. If the tolerance is 5 minutes, an observed response is within tolerance when its absolute residual is at most 5 minutes. Since \(s\) describes a typical error scale, comparing \(s\) with the tolerance helps judge whether errors are typically small enough for that use.

Definition: A practical tolerance \(T\) is a context-based limit on the acceptable size of a prediction error, expressed in the response variable’s units. Comparing \(s\) with \(T\) indicates whether the model’s typical residual size is small or large relative to that limit.

A convenient comparison is the scale ratio \(s/T\). This ratio has no units. If it is well below 1, the typical residual size is smaller than the tolerance, which supports the judgment that predictions are typically accurate enough for that purpose. If it is near 1, the typical error scale is about as large as the tolerance. If it is greater than 1, the typical residual size exceeds the tolerance, which is a warning that the model may not meet the practical need.

$$ \text{scale ratio}=\frac{s}{T} $$

The ratio is a comparison tool, not a probability. For example, \(s/T=0.5\) does not mean that a particular percentage of predictions will be within tolerance. The residual standard deviation alone does not tell us the proportion of errors that fall inside any chosen range. It describes a typical scale, not a guarantee or a count of successful predictions.

The comparison also makes sense only when \(s\) and \(T\) refer to the same response and units. If \(s=2\) hours but the tolerance is stated in minutes, convert one of the values before comparing. And a tolerance is not a statistical constant: it comes from the context and must be chosen based on what counts as acceptable.

A Practical Way to Use \(s\)

Before deciding whether the comparison supports using a model, check what the model is being asked to do. As in “Predicting Within the Data Range,” a prediction for an \(x\)-value outside the observed range is extrapolation; \(s\) from the fitted data does not make that extrapolation reliable. Also consult the residual plot. The earlier tutorials “Reading a Residual Plot for Random Scatter,” “Curved Patterns in a Residual Plot,” and “Fan-Shaped Residual Plots and Changing Spread” explain how patterns can indicate that the line misses systematically or that error spread changes across the range.

1
Specify the practical tolerance.
State how far a prediction may be from the observed response while still being useful, in response units.
2
Identify \(s\) and its units.
Use the residual standard deviation for the fitted model. Its units are the response variable’s units.
3
Compare the two scales.
Compare \(s\) directly with \(T\), or calculate \(s/T\). Describe whether the typical error scale is well below, near, or above the tolerance.
4
Qualify the judgment.
Use the data range and residual pattern to assess whether the comparison is relevant. Do not claim that every prediction—or a stated percentage of predictions—will meet the tolerance.

This is a practical assessment, not a new inference procedure. It does not produce a confidence level or a guaranteed error bound. For predictions about future cases, the conditions represented by the fitted data should also be relevant to those future cases. If prediction errors are especially costly, the tolerance comparison can flag that \(s\) is too large, but a value of \(s\) by itself cannot certify that every case will be safe or accurate enough.

Worked Example: Predicting Appointment Wait Times

A fictional clinic uses a least-squares line to predict a patient’s wait time, in minutes, from the number of scheduled appointments in the hour. The observed predictor values range from 8 to 20 appointments. The residual plot for the fitted data shows scatter above and below zero without a clear pattern. The software reports \(s=2.4\) minutes. For scheduling, staff decide that an error of up to 5 minutes is practically acceptable.

State. Judge whether the model’s typical prediction-error scale appears small enough for the 5-minute tolerance.

Plan. Both \(s\) and the tolerance \(T\) are measured in minutes. Compare \(s=2.4\) with \(T=5\), and calculate \(s/T\). The requested judgment concerns predictor values within the stated observed range, and the residual plot does not show an obvious systematic pattern.

Do. The scale ratio is \(s/T=2.4/5=0.48\). Thus the reported typical residual size is less than half the 5-minute tolerance. In direct units, \(2.4\) minutes is \(5-2.4=2.6\) minutes below the tolerance.

Conclude. For predictions in the observed range, the residual standard deviation is substantially smaller than the clinic’s 5-minute tolerance, supporting the judgment that the model’s typical error scale is acceptable for this scheduling purpose. This does not establish that every patient’s wait time will be predicted within 5 minutes, nor does \(0.48\) give the proportion of patients whose errors meet the tolerance.

When \(s\) Is Near or Above the Tolerance

A comparison can also show that a model’s typical error scale is not comfortably below a practical limit. When \(s\) is close to \(T\), the typical residual size is comparable to the largest acceptable error. When \(s>T\), the typical residual size is larger than the tolerance. In either case, the model may be inadequate for a use that requires errors to stay within that limit. The appropriate conclusion is about the typical scale—not that every prediction will fail.

It can help to say exactly what the comparison does and does not establish. If \(s\) exceeds the tolerance, that is evidence that the model’s overall prediction-error scale is too large for the stated goal. It does not prove that every observed residual exceeds the tolerance. Likewise, \(s<T\) does not prove that all residuals fall within it. Individual residuals can be larger or smaller than \(s\), as discussed in “Comparing Residuals Across Observations.”

Worked Example: Estimating a 3D-Printer’s Job Time

A fictional technology class fits a regression line to predict the time, in minutes, for a 3D printer to complete a job, using the job’s estimated material volume as the predictor. For the fitted data, there are 14 jobs and the sum of squared residuals is 24. The class needs predictions to be within 1 minute to plan a tightly scheduled demonstration. Assume the jobs being considered are within the observed predictor range, and the residual plot shows no clear curved or fan-shaped pattern.

State. Calculate \(s\), then judge whether the typical residual size is small relative to the 1-minute tolerance.

Plan. For a simple linear regression with an intercept, use \(s=\sqrt{\text{SSE}/(n-2)}\). The response is job time, so \(s\) is in minutes. Compare the result with \(T=1\) minute after checking the relevant range and residual-pattern information given in the situation.

Do. The degrees of freedom are \(n-2=14-2=12\). Therefore, \(s=\sqrt{24/12}=\sqrt{2}\approx1.414\) minutes, rounded to 1.41 minutes. As a check, \(1.414^2(12)\approx2(12)=24\), consistent with the stated SSE. The scale ratio is \(s/T=1.414/1\approx1.41\).

Conclude. The typical residual size, about 1.41 minutes, is larger than the class’s 1-minute tolerance. This suggests that the model’s typical error scale may be too large for the demonstration’s timing requirement. It does not mean that all job-time predictions will miss by more than 1 minute, but the comparison is a reason to be cautious about relying on this model for that tightly timed purpose.

Comparing Models Against the Same Tolerance

If two models predict the same response for the same cases and use the same response units, their values of \(s\) can be compared directly. Against a practical tolerance, the model with the smaller \(s/T\) has the smaller typical residual scale relative to that requirement. This does not make it better for every purpose, but it helps describe which model has more room between its typical error scale and the chosen tolerance.

The tolerance itself must be the same for a fair comparison. If models predict different responses, or if their response values use different scales, first clarify what is being predicted and convert units when appropriate. A smaller numerical \(s\) is not automatically better if it refers to a different response scale or different cases.

Worked Example: Choosing a Model for Delivery-Time Estimates

A fictional delivery service fits two regression models to the same set of routes to predict delivery time in minutes. Model A uses route distance as its predictor, and Model B uses the number of planned stops. For the same 24 routes, Model A has \(s=6.0\) minutes and Model B has \(s=3.6\) minutes. Dispatchers consider a prediction acceptable if it is within 10 minutes of the observed delivery time. Both requested predictions are within the observed predictor ranges, and both residual plots show no obvious systematic pattern.

State. Compare the models’ typical residual sizes relative to the 10-minute tolerance.

Plan. Both models predict the same response, in minutes, for the same routes, so their residual standard deviations are directly comparable. Calculate each scale ratio using \(T=10\) minutes, then compare the results.

Do. For Model A, \(s/T=6.0/10=0.60\). For Model B, \(s/T=3.6/10=0.36\). Both typical residual sizes are below the tolerance: 6.0 minutes for A and 3.6 minutes for B. The difference between their residual standard deviations is \(6.0-3.6=2.4\) minutes.

Conclude. For these routes, Model B has the smaller typical prediction-error scale relative to the 10-minute tolerance. Both values of \(s\) are below the tolerance, while Model B’s \(s/T\) is lower. These comparisons support Model B as having more margin on this measure, but they do not guarantee that every delivery prediction from either model will be within 10 minutes.

Common Mistakes and AP Exam Tips

  • Treating \(s\) as a maximum error. \(s\) is a typical scale, not the largest possible residual. A prediction error can be greater than \(s\), and this comparison alone does not establish how often that occurs.
  • Turning \(s/T\) into a percentage of predictions. A ratio of 0.4 does not mean 40% of predictions are within tolerance. State only how the typical residual scale compares with the tolerance.
  • Forgetting the response units. The tolerance and \(s\) must be in the same units. If a model predicts minutes, report and compare both quantities in minutes.
  • Declaring a model adequate from \(s\) alone. Check whether the requested prediction is within the data range and whether the residual plot suggests a systematic pattern. A favorable ratio cannot repair a curved or changing-spread pattern.
  • Making a claim about every case. Say that the typical error scale is smaller or larger than the practical tolerance. Do not say that every prediction meets—or fails to meet—the limit unless the individual errors have actually been examined.
  • Choosing a tolerance without context. The tolerance represents what is practically acceptable for the intended use. Explain the stated limit rather than presenting it as a property of the regression model.

For full-credit communication, identify the response and its units, state the tolerance, compare it with \(s\), and make a limited conclusion about the typical error scale. If useful, include \(s/T\), but explain it as a scale comparison—not a probability or percentage.

Key takeaway: Compare the residual standard deviation \(s\) with a context-based tolerance \(T\) to judge whether the model’s typical prediction-error scale is practically small enough. A favorable comparison supports—but does not guarantee—that predictions will meet the tolerance.

Check Your Understanding

For each situation, compare the typical error scale with the practical tolerance and keep the conclusion appropriately limited.

  1. A model predicts the time, in seconds, for a school’s 3D printer to finish a small job. It has \(s=4\) seconds, and the user’s tolerance is 12 seconds. Calculate \(s/T\) and describe what it indicates.
  2. A regression model predicts monthly rainfall in millimeters. Its residual standard deviation is 18 millimeters, while a planning task requires predictions within 10 millimeters. What does comparing these values suggest?
  3. A response is measured in hours, but \(s=0.5\) hour and the tolerance is stated as 20 minutes. Convert to matching units and compare.
  4. Why does \(s/T=0.3\) not mean that 30% of predictions will be within the practical tolerance?
  5. Two models predict the same response for the same cases. Model A has \(s=4.8\) and Model B has \(s=3.9\), with a shared tolerance of 8 response units. Which has the smaller typical error scale relative to the tolerance, and what can you conclude?