A Prediction Needs More Than a Number
A fitted value such as “19.8 kilograms” is hard to judge on its own. A reader also needs to know what input produced the prediction, how much residual scatter is typical, and whether that input falls within the explanatory-variable values used to fit the model.
In “Explaining Regression to a Nontechnical Audience,” you practiced translating regression results into clear language. Here, we focus on a compact but informative report of a prediction. The tutorials “Using Residual Standard Deviation to Gauge Prediction Error” and “Reliability Within the Range of Data” established what \(s\) and the observed \(x\)-range tell us. The task now is to bring those pieces together in a sentence without implying more certainty than the model supports.
A useful report usually identifies the case or setting, gives the explanatory-variable value and the fitted prediction, describes \(s\) in response units, and states the observed range of \(x\). If the prediction is outside that range, say so plainly. These details let readers distinguish the model’s output from how dependable its use may be in context.
A Clear Reporting Pattern
A sentence can be concise and still answer four questions: What is being predicted? At what \(x\)-value? What is the typical size of residuals? Is that \(x\)-value inside the range used to fit the line? Keep the units visible: \(x\) and \(y\) generally have different units, and \(s\) has the same units as \(y\).
Give the predicted response for a specific explanatory-variable value, and include the response units. If useful, show the substitution into \(\hat{y}=a+bx\).
Say that observed responses typically differed from the line’s predictions by about \(s\) units. Do not turn this typical scatter into a guarantee about an individual case.
State the minimum and maximum explanatory-variable values among the cases used to fit the model, with units. Then compare the requested \(x\)-value with those endpoints.
Say whether the input is within the observed range or outside it. An in-range prediction is interpolation; an out-of-range prediction is extrapolation. Neither label guarantees accuracy.
The range comparison concerns the explanatory variable, not the size of the predicted response. The tutorial “Common Errors About Extrapolation” emphasizes that distinction. Also remember that the observed range is not a boundary that certifies every prediction inside it as reliable. As discussed in “Reading a Residual Plot for Model Fit,” examine whether residual behavior and the broader model context support using the line.
A complete sentence is often more helpful than listing separate statistics. For example: “For a garden bed of 7.5 square meters, the line predicts a harvest of 19.8 kilograms; residuals typically differ from the line’s predictions by about 3.1 kilograms, and the bed area is within the 2-to-12-square-meter range of the beds used to fit the model.” The sentence keeps prediction, scatter, and range distinct.
Worked Example: Predict a Garden Bed’s Harvest
Worked Example: Report an In-Range Prediction
Situation. In an invented study, a community garden group recorded bed area and harvest mass for garden beds. The least-squares line predicting harvest mass \(y\), in kilograms, from bed area \(x\), in square meters, is \(\hat{y}=1.8+2.4x\). The observed bed areas used to fit the line ranged from 2 to 12 square meters, and the residual standard deviation was \(s=3.1\) kilograms. Report the prediction for a bed with area 7.5 square meters.
State. We want the fitted harvest mass for a bed with an area of 7.5 square meters. This is a prediction for one bed, not a claim about the exact harvest it will produce.
Plan. Substitute \(x=7.5\) into the fitted line. Then report \(s\) in kilograms and compare 7.5 square meters with the observed area range of 2 to 12 square meters. This comparison identifies whether the prediction is interpolation or extrapolation; it does not, by itself, establish how close the prediction will be.
Do. The fitted prediction is
The requested bed area, 7.5 square meters, is between 2 and 12 square meters, so this prediction is an interpolation. The residual standard deviation is 3.1 kilograms, which describes the typical size of the residuals for this fitted model.
Conclude. “For a garden bed with an area of 7.5 square meters, the line predicts a harvest of 19.8 kilograms. In these data, observed harvests typically differed from the line’s predictions by about 3.1 kilograms; the bed area is within the 2-to-12-square-meter range used to fit the line.” This gives the prediction, its input, typical scatter, and range without promising that this particular bed will produce exactly 19.8 kilograms or will be within 3.1 kilograms of it.
Keep Prediction, Scatter, and Range Separate
These three reported quantities answer different questions. The fitted value answers, “What response does the line predict at this input?” The residual standard deviation answers, “How large are residuals typically, in response units?” The observed range answers, “Was this input among the explanatory-variable values used to fit the line?” Including all three does not turn the fitted value into a guaranteed outcome.
| Report element | What it communicates | Units |
|---|---|---|
| Fitted prediction \(\hat{y}\) | The response value given by the line at the stated \(x\) | Response units |
| Residual standard deviation \(s\) | The typical size of residuals around the fitted line | Response units |
| Observed \(x\)-range | The smallest and largest explanatory-variable values in the data used to fit the line | Explanatory-variable units |
Do not confuse the units of \(s\) with those of the explanatory variable. If the model predicts liters of water from household size, then \(s\) is in liters, while the observed range is stated in people per household. The range provides context for the input, not a measure of prediction error.
Likewise, \(s\) is not an interval around an individual prediction. Saying “the prediction is 105 liters, plus or minus 25 liters” can sound like a guaranteed band, which \(s\) does not provide. The tutorial “Using Residual Standard Deviation to Gauge Prediction Error” explains \(s\) as a typical residual scale. It does not say that every residual has magnitude at most \(s\), or that a specified percentage of individual outcomes falls within one \(s\) of the fitted value.
Worked Example: Report a Household Water-Use Prediction
Worked Example: Put the Units in the Right Places
Situation. An invented local planning exercise used household size to predict daily water use. The fitted line is \(\hat{y}=42+18x\), where \(x\) is the number of people in a household and \(y\) is daily water use in liters. The households used to fit the line had 1 to 6 people, and \(s=25\) liters. Report the prediction for a household of 3.5 people as a model calculation.
Plan. Evaluate the line at \(x=3.5\), keep the prediction and \(s\) in liters, and give the observed range in people per household. Although 3.5 people is not a possible size for an actual household, it can be a specified value for evaluating a fitted line; here the requested task is to report that model calculation, not to describe a particular real household.
Do. Substitution gives
The input, 3.5 people, is between the observed values of 1 and 6 people, so it lies within the range used to fit the line. The residual standard deviation is 25 liters per day.
Conclude. “At an input of 3.5 people per household, the fitted line predicts daily water use of 105 liters. In the data, observed daily use typically differed from the line’s predictions by about 25 liters; the input is within the 1-to-6-person range used to fit the model.” The sentence assigns liters per day to the response and \(s\), and people per household to the explanatory-variable range. It does not describe 25 liters as a maximum error.
Worked Example: Flag a Prediction Beyond the Data Range
Worked Example: Report an Extrapolation Honestly
Situation. An invented makerspace recorded the volume of small 3D-printed objects and the time required to print each one. A fitted line predicts printing time \(y\), in minutes, from object volume \(x\), in cubic centimeters: \(\hat{y}=14+0.9x\). The observed volumes ranged from 10 to 80 cubic centimeters, and \(s=8\) minutes. A user asks for a prediction at 95 cubic centimeters.
Plan. Calculate the fitted value at 95 cubic centimeters, report \(s\) in minutes, and compare 95 with the observed volume range. Because this input exceeds the maximum observed volume, describe the result as an extrapolation and avoid presenting the fitted number as dependable just because \(s\) is known.
Do. The model gives
The requested volume, 95 cubic centimeters, is greater than the maximum observed volume of 80 cubic centimeters. Thus the prediction is an extrapolation. The residual standard deviation of 8 minutes describes typical residual size for the fitted model’s data; it does not establish that an extrapolated prediction is accurate to within 8 minutes.
Conclude. “For an object with a volume of 95 cubic centimeters, the fitted line predicts a printing time of 99.5 minutes. Residuals for the data used to fit the line typically differed from predictions by about 8 minutes, but the observed object volumes ranged only from 10 to 80 cubic centimeters, so the 95-cubic-centimeter prediction is an extrapolation and may not be dependable.” The sentence reports the numerical result while making its limitation unmistakable. As in “Writing an Extrapolation Critique,” the explanation connects the requested value to the observed endpoint instead of relying on a vague warning.
Common Mistakes and AP Exam Tips
- Reporting only the fitted value. “The prediction is 19.8” leaves out what is being predicted, the input, and the units. A complete report names the response and states the input that produced the fitted value.
- Putting \(s\) in \(x\)-units. If \(y\) is measured in kilograms, \(s\) is measured in kilograms, even when \(x\) is measured in square meters. State each unit with the quantity it describes.
- Making \(s\) sound like a guarantee. “Every prediction is within 3.1 kilograms” is not what \(s\) means. Say that observed responses typically differed from fitted predictions by about 3.1 kilograms.
- Giving \(s\) without the data range. A typical residual size does not show whether a requested \(x\)-value is within the values used to fit the line. Report the explanatory-variable range separately and compare the requested input with it.
- Treating an in-range input as automatically reliable. Being within the observed range makes the prediction interpolation, but it does not erase residual scatter or a poor fit. Use the model-fit evidence and limitations relevant to the situation.
- Reporting an extrapolation without naming it. For an out-of-range input, identify which endpoint it exceeds and state that the prediction is extrapolation. A small \(s\) from the fitted data does not guarantee that the relationship continues beyond the observed range.
For full credit, make the prediction and its context explicit, describe \(s\) as typical residual scatter in response units, and state the observed \(x\)-range with explanatory-variable units. Then identify whether the requested input is inside or outside that range and qualify the prediction accordingly. A clear report distinguishes what the fitted line calculates from how much confidence the data support in applying it.
Check Your Understanding
Answer using complete sentences. Keep the prediction, typical scatter, and observed range distinct.
- A model predicts package-delivery time in minutes from distance in kilometers. The distance range used to fit the line is 3 to 18 kilometers, and \(s=6\) minutes. What units should be used for \(s\), and what units should be used for the range?
- A fitted line predicts 42 liters for a stated input. Write one careful sentence explaining what \(s=9\) liters means without making it a guarantee.
- A requested input is 20 kilometers when the observed range is 3 to 18 kilometers. What should the report say about the prediction’s status?
- Why does knowing \(s\) alone not establish that an out-of-range prediction is dependable?
- List the main pieces of information you would include in a clear report of an in-range individual prediction.