Finding Residuals and \(s\) in Regression Output
In “Residuals and Outliers,” you used residuals to identify observations that may stand out from a fitted linear pattern. Computer output can provide those residuals directly, but the values may appear in a different place from the regression equation or summary statistics. This tutorial focuses on locating each value, matching it to the correct observation, and interpreting the residual standard deviation \(s\).
An individual residual belongs to one observed case. It is the observed response minus the predicted response for that same case, \(y-\hat{y}\). The residual standard deviation \(s\), by contrast, summarizes the overall scale of residuals around the fitted line. As described in “Standard Deviation of the Residuals, \(s\),” it is measured in the response variable’s units.
Software does not use one universal layout. Individual residuals may appear in a table beside the original observations, in a calculator list named RESID, or in a separate diagnostic output. A compact regression summary may report \(s\) without listing every residual. Look for labels such as “Residual,” “Residuals,” “RESID,” “Residual standard error,” or “S.” Some software uses another label, such as “Root MSE,” for a closely related summary of residual spread. Check the output’s description and units rather than relying on a label alone.
The residual list must be matched to the observations in the same order used for the regression. In a calculator, for example, the first value in RESID belongs to the first paired entries in the predictor and response lists. If the data table has been sorted or rearranged since the regression was run, confirm that the rows still correspond before interpreting a residual.
A Reliable Way to Read the Output
Look for a residual column or list. A regression equation and a value of \(s\) alone do not identify the residual for a particular observation.
Use the row label, case name, or original data order. Do not match by position after the data have been reordered unless you verify the correspondence.
A positive residual means the observed response is above the prediction; a negative residual means it is below. State its size in the response variable’s units.
Look for \(s\) or a software label such as “Residual standard error.” Interpret it as the typical size of prediction errors around the fitted line, not as the residual for any one case.
The sign matters when describing direction, but the absolute value \(|y-\hat{y}|\) describes the distance from the fitted line. The \(s\) value helps describe the general scale of that distance across the fitted data. As in “Using \(s\) to Describe Prediction Accuracy,” a comparison with \(s\) can add context, but it does not create a universal rule that classifies a point as an outlier.
Worked Example: Locating a Residual in a Software Table
A fictional school garden records the number of hours of sunlight \(x\) and the weekly growth of seedlings \(y\), in centimeters. Software reports this fitted line and case-level output:
| Case | Sunlight \(x\) (hours) | Observed growth \(y\) (cm) | Fitted growth \(\hat{y}\) (cm) | Residual (cm) |
|---|---|---|---|---|
| A | 1 | 13 | 12 | 1 |
| B | 2 | 13 | 14 | \(-1\) |
| C | 3 | 15 | 16 | \(-1\) |
| D | 4 | 19 | 18 | 1 |
The regression output gives \(\hat{y}=10+2x\) and “Residual standard error: 1.414 cm.” The task is to interpret case B’s residual and the reported \(s\).
State. Identify the residual for case B, explain what its sign means, and interpret the residual standard error in context.
Plan. Read case B’s residual from the row labeled B, checking that it agrees with observed growth minus fitted growth. Then interpret the reported residual standard error as the typical size of the prediction errors, in centimeters.
Do. Case B has observed growth 13 cm and fitted growth 14 cm. The residual calculation is \(y-\hat{y}=13-14=-1\) cm, matching the software’s residual entry. The output reports \(s=1.414\) cm, rounded to three decimal places. For a check, the four displayed residuals are \(1,-1,-1,1\) cm; their squared values sum to \(1+1+1+1=4\). With four observations and two fitted line coefficients, the residual standard deviation is \(\sqrt{4/(4-2)}=\sqrt{2}\approx1.414\) cm.
Conclude. Case B’s residual is \(-1\) cm, so its observed weekly growth was 1 cm less than the model predicted. The model’s residual standard deviation is about 1.414 cm, meaning that prediction errors in this fitted data set are typically about 1.414 cm in size. Case B’s residual is one individual error, not the value of \(s\).
Reading Residual Lists and Matching the Data Order
A calculator may show the residuals in a list rather than next to the original cases. The list position is the link: residual number 1 corresponds to the first \(x\)- and \(y\)-values used in the regression, residual number 2 to the second pair, and so on. A residual list without case names is useful only if you preserve or reconstruct that order.
Before describing a particular residual, verify the case in two ways when possible: check its position in the original data lists and check the sign by comparing the observed response with the predicted response. The residual should equal \(y-\hat{y}\). This is a practical way to catch a shifted row, reversed subtraction, or a residual list left over from an earlier regression.
Worked Example: Matching TI-84 Residuals to Observations
A fictional recreation program predicts a participant’s trail-walk time, in minutes, from the route distance, in kilometers. The predictor values are in \(L1\), the response values are in \(L2\), and the TI-84’s RESID list shows the residuals from the regression.
| List position | Distance \(L1\) (km) | Observed time \(L2\) (min) | RESID (min) |
|---|---|---|---|
| 1 | 2 | 29 | 1 |
| 2 | 3 | 32 | \(-2\) |
| 3 | 4 | 39 | 1 |
| 4 | 5 | 42 | 0 |
The fitted equation is \(\hat{y}=20+4x\). Identify the residual for the 3-kilometer route and interpret it. The output also reports \(s=1.581\) minutes.
State. The 3-kilometer route is in list position 2. Its residual is the second entry of RESID, but the match should be checked against the observed and predicted times.
Plan. Substitute the route distance into the fitted equation to find its predicted time. Subtract that prediction from the observed time using \(y-\hat{y}\), then interpret the sign and units. Interpret \(s\) separately as the typical prediction-error size.
Do. For \(x=3\) km, the predicted time is \(\hat{y}=20+4(3)=32\) minutes. The observed time is 32 minutes, so the residual is \(32-32=0\) minutes. This agrees with the second residual-list entry, which is \(-2\), only if the data table is read correctly? It does not agree, so check the output and case order before making an interpretation. In this example, the displayed values reveal a mismatch: the residual list’s second entry cannot belong to the listed second case if the fitted equation and response are correct. The residuals must be reconciled with the regression output rather than copied uncritically.
Conclude. The table, fitted equation, and stated residual list are inconsistent for the 3-kilometer case. From the observed value and fitted line, its residual is 0 minutes; the provided list instead shows \(-2\). This is a signal to verify that the residual list came from the same regression and unchanged data. The separately reported \(s=1.581\) minutes describes the overall residual scale only if it also belongs to this same fitted model.
Interpreting the Reported \(s\)
When output reports a residual standard error or \(s\), connect the number to both the model and the response variable. For example, “The residual standard deviation is about 2.4 minutes, so prediction errors from this model are typically about 2.4 minutes in size.” That statement gives the units and the meaning. It does not say that every prediction is wrong by exactly 2.4 minutes, or that every error is within 2.4 minutes of zero.
A reported \(s\) is not an individual residual, the slope, or the standard error of the slope. A residual refers to one case; \(s\) summarizes residual spread for the fitted model. The slope describes the predicted change in response associated with a one-unit increase in the predictor, as covered in “Regression Does Not Mean Causation in Slope Statements.” These quantities answer different questions even if they appear in the same software output.
Output may round residuals and \(s\). Use the displayed precision unless a question asks for more, and do not imply greater accuracy than the output supports. If you calculate a residual independently from rounded coefficients, it may differ slightly from software’s residual calculated using unrounded coefficients. A small difference from rounding is not necessarily a data-order mistake; check the values and precision before deciding.
Worked Example: Interpreting Residual Standard Error
A fictional community program predicts the time, in minutes, needed to assemble a set of donated books from the number of boxes in the set. A software summary reports “Residual standard error: 2.4 on 18 degrees of freedom.” A separate diagnostic table shows that one set has a residual of \(-3.1\) minutes.
State. Interpret the reported residual standard error and the individual residual for that set, keeping their meanings separate.
Plan. Use the response units, minutes, for both values. Interpret \(s\) as the typical size of prediction errors around the line, and interpret the negative residual as an observed time below the prediction. Do not treat \(s\) as a limit on every residual.
Do. The output gives \(s=2.4\) minutes. The individual residual has absolute size \(|-3.1|=3.1\) minutes, which is \(3.1/2.4\approx1.29\) times \(s\). Since the residual is negative, the observed assembly time was less than the time predicted by the model.
Conclude. The model’s prediction errors are typically about 2.4 minutes in size. For this particular book set, the model overpredicted the assembly time by 3.1 minutes. That error is about 1.29 times the reported \(s\); this comparison describes its size relative to the general scale but does not, by itself, establish that the case is an outlier.
Common Mistakes and AP Exam Tips
- Reading \(s\) as a case’s residual. A summary value such as “Residual standard error” describes the model’s overall residual scale. Find the case-level row or list entry to report one observation’s residual.
- Using the wrong row or list position. Match the residual to the same case and data order used for the regression. If the data were sorted, edited, or re-entered, verify the correspondence.
- Reversing the subtraction. Residuals use observed minus predicted, \(y-\hat{y}\). A negative value means the observed response was below the prediction; a positive value means it was above.
- Leaving out context or units. A residual of \(-3.1\) is incomplete if the response and units are known. State, for example, that the observed assembly time was 3.1 minutes less than predicted.
- Claiming \(s\) is a maximum error. The residual standard deviation describes a typical scale, not a guarantee that every residual is within \(s\) of zero.
- Confusing residual standard error with slope uncertainty. Use the label and context in the output. \(s\) describes residual spread in response units; it is not the regression slope or its standard error.
For a full-credit interpretation, name the case, report its signed residual with response units, and explain whether the observed response was above or below the prediction. When interpreting \(s\), identify it as the model’s typical prediction-error size and give the response units. If output labels or values appear inconsistent, state what you checked and avoid assigning a value to a case until the mismatch is resolved.
Check Your Understanding
Use the output labels, case order, and response units to distinguish individual residuals from the overall residual scale.
- A software table lists a residual of \(-2.5\) seconds for a particular device. What does the sign say about its observed response compared with its predicted response?
- A TI-84 residual list contains six values. What information do you need to match the fourth residual to the correct observation?
- A regression summary reports “Residual standard error: 3.2 kg.” Interpret this value in context, without treating it as a maximum error.
- A case-level output gives an observed response of 41 minutes and a fitted response of 44 minutes. Calculate its residual and interpret it.
- A residual calculated from rounded coefficients differs slightly from the software’s displayed residual. Give one reason this can happen and one check to make before concluding the output is wrong.