Tutorials › AP Statistics › Reliability Within the Range of Data

Extrapolation and prediction limits · Tutorial 953 of 1000

Reliability Within the Range of Data

Learn why predictions within the observed range are generally better supported than extrapolations, yet still depend on how closely the data follow a linear pattern.

Intermediate 10 min read

What You'll Learn

  • Distinguish being within the observed x-range from having a dependable prediction.
  • Use residual standard deviation s to describe typical residual size in response units.
  • Use r-squared to describe the fraction of response variation accounted for by a linear relationship.
  • Explain why a high r-squared does not necessarily mean small prediction errors.
  • Assess how scatter, data coverage, and the linear pattern affect an interpolation.
  • Communicate what a regression prediction can and cannot establish.

Within the Range Is Not the Same as Certain

In “Interpolation Versus Extrapolation,” we classified a prediction by comparing its explanatory-variable value with the observed \(x\)-range. Interpolation uses an \(x\)-value within that range; extrapolation uses one outside it. Interpolation is generally safer because the prediction is made among the \(x\)-values used to fit the model, rather than by extending the line beyond them. But “within the range” does not mean “guaranteed accurate.”

A line might pass close to the observed points, or the points might be widely scattered around it. The fit also might be summarized by a high or low \(r^2\). As in “r-squared and Residual Standard Deviation Together,” these summaries answer different questions: \(r^2\) describes the fraction of response variation accounted for by the linear relationship, while \(s\) describes the typical size of residuals in response units. Both help assess a prediction, but neither tells us exactly how far an individual response will be from its predicted value.

Key idea: Interpolation is usually better supported by the observed data than extrapolation, but a prediction within the observed \(x\)-range can still be unreliable if the data are scattered, the linear pattern is weak, or the requested \(x\)-value has little nearby data.

Three Questions to Ask About an Interpolation

First, check the range. A prediction within the observed \(x\)-range is interpolation, including a prediction at either endpoint. This classification tells you whether the requested \(x\)-value lies among the values used to fit the line; it does not by itself measure the quality of the fit.

Second, consider how closely the observed points follow a straight-line pattern. A scatterplot and residual plot can reveal whether the linear model is reasonable for the data. A pattern of residuals that bends or otherwise changes systematically is a warning that a straight line may not represent the relationship well, even within the observed range. As discussed in “How Unmodeled Curvature Shows Up Outside the Data,” a linear model can miss a bend.

Third, put \(s\) and \(r^2\) to work without treating either as a guarantee. A smaller \(s\), compared with meaningful changes in the response, indicates less typical residual scatter in response units. A larger \(r^2\) indicates that a greater fraction of the response variation in the observed cases is accounted for by the linear relationship. A high \(r^2\) does not automatically mean small errors in the response’s units; a low \(r^2\) does not tell you the typical error in those units. That is one reason to examine both summaries.

Keep the summaries distinct: \(s\) is measured in the response variable’s units and describes the typical distance of observed responses from the fitted line. \(r^2\) is unitless and describes the fraction of response variation accounted for by the linear relationship. Neither is a guarantee for an individual prediction.

Also check whether the requested \(x\)-value is in a part of the range with observations nearby. A value can be between the minimum and maximum while still lying in a gap between clusters of observed \(x\)-values. Such a prediction is technically interpolation, but the data may offer less direct support there than they do where observations are concentrated. This does not change the definition of interpolation; it adds useful context when judging its reliability.

Worked Examples: What Fit Summaries Tell Us

Worked Example: A Prediction Within a Close Linear Pattern

Hypothetical setting. A technician records operating time \(x\), in hours, and a device’s response \(y\), in response units, for five test runs. The invented observations are \((1,12),(2,15),(3,17),(4,20),(5,21)\). We will fit the least-squares line, calculate \(s\) and \(r^2\), and predict at 3.5 hours.

State. The observed \(x\)-range is 1 to 5 hours. Since 3.5 is within this range, predicting at 3.5 hours is interpolation. We will assess the prediction using the fitted line and the scatter summaries.

Plan. Calculate the regression line from the observations, then find the residuals to calculate \(s\) and \(r^2\). The residuals and summaries describe fit for these observed test runs; they do not guarantee the response in a particular new run.

Do. The means are \(\bar{x}=3\) hours and \(\bar{y}=17\) response units. The sum of products of deviations is 23, and the sum of squared \(x\)-deviations is 10, so the slope is \(23/10=2.3\) response units per hour. The intercept is \(17-2.3(3)=10.1\) response units. Thus:

$$ \hat{y}=10.1+2.3x $$

At 3.5 hours, the predicted response is \(10.1+2.3(3.5)=18.15\) response units. The predicted response is a value from the fitted line, not a promise about an individual test run.

For the observed \(x\)-values 1 through 5, the line predicts 12.4, 14.7, 17.0, 19.3, and 21.6. Observed minus predicted gives residuals \(-0.4,0.3,0,0.7,-0.6\) response units. Squaring and adding them gives:

$$ SSE=(-0.4)^2+(0.3)^2+0^2+(0.7)^2+(-0.6)^2 =0.16+0.09+0+0.49+0.36=1.10 $$

With \(n=5\), the residual standard deviation is:

$$ s=\sqrt{\frac{SSE}{n-2}} =\sqrt{\frac{1.10}{3}} \approx 0.606\text{ response units} $$

The observed \(y\)-values have mean 17, so their total sum of squares is \(SST=(-5)^2+(-2)^2+0^2+3^2+4^2=54\). Therefore:

$$ r^2=1-\frac{SSE}{SST} =1-\frac{1.10}{54} \approx 0.9796 $$

As a check, \(1-r^2\approx 0.0204\), and \(0.0204(54)\approx 1.10\), matching the calculated \(SSE\) after rounding.

Conclude in context. The prediction at 3.5 hours is an interpolation, and these five invented test runs lie close to a linear pattern: \(r^2\) is about 0.9796 and \(s\) is about 0.606 response units. Those summaries support the line as a description of these data, but they do not tell us the exact response for a new run at 3.5 hours.

Worked Example: Interpolation with Substantial Scatter

Hypothetical setting. A community garden records daily sunlight exposure \(x\), in hours, and a plant’s measured height \(y\), in centimeters, for five plants. The invented observations are \((1,10),(2,18),(3,11),(4,17),(5,14)\). Predict the height at 3 hours and assess whether being within the observed range is enough to trust that prediction.

State. The observed \(x\)-range is 1 to 5 hours, so 3 hours is within the range and the requested prediction is interpolation. We also need to examine the scatter around the fitted line before judging how informative it is.

Plan. Find the least-squares line, then use its residuals to calculate \(s\) and \(r^2\). Interpret \(s\) in centimeters and \(r^2\) as a fraction of the response variation accounted for by the linear relationship.

Do. Here \(\bar{x}=3\) hours and \(\bar{y}=14\) centimeters. The sum of products of deviations is 7, and the sum of squared \(x\)-deviations is 10. The slope is \(7/10=0.7\) centimeters per hour, and the intercept is \(14-0.7(3)=11.9\) centimeters. The fitted line is:

$$ \hat{y}=11.9+0.7x $$

At 3 hours, the line predicts \(11.9+0.7(3)=14.0\) centimeters. The predictions at \(x=1,2,3,4,5\) are 12.6, 13.3, 14.0, 14.7, and 15.4 centimeters. The residuals are \(-2.6,4.7,-3,2.3,-1.4\) centimeters, giving:

$$ SSE=(-2.6)^2+4.7^2+(-3)^2+2.3^2+(-1.4)^2 =6.76+22.09+9+5.29+1.96=45.10 $$

With \(n=5\), the residual standard deviation is:

$$ s=\sqrt{\frac{45.10}{5-2}} =\sqrt{15.0333\ldots} \approx 3.877\text{ centimeters} $$

The \(y\)-values’ deviations from their mean of 14 are \(-4,4,-3,3,0\), so \(SST=16+16+9+9+0=50\). Thus:

$$ r^2=1-\frac{45.10}{50}=0.098 $$

As a check, \(1-r^2=0.902\), and \(0.902(50)=45.10\), matching \(SSE\).

Conclude in context. Although 3 hours is within the observed range, the prediction of 14.0 centimeters comes from a line with substantial residual scatter: \(s\) is about 3.877 centimeters, and the linear relationship accounts for only about 9.8% of the observed variation in plant height. Interpolation alone does not make this prediction dependable. The summaries describe these five invented plants, not every plant grown under those conditions.

Worked Example: A High r-squared with Noticeable Response-Unit Scatter

Hypothetical setting. An engineer records a setting \(x\), in adjustment units, and an instrument’s output \(y\), in output units, for five invented trials: \((0,100),(1,140),(2,170),(3,220),(4,250)\). Assess the line’s prediction at 2.5 adjustment units, taking care not to treat a high \(r^2\) as proof of a tiny prediction error.

State. The observed \(x\)-range is 0 to 4 adjustment units. Since 2.5 is inside that range, the request is interpolation. We will use both \(r^2\) and \(s\) to describe the fit, because they express different aspects of scatter.

Plan. Calculate the line, residuals, \(SSE\), \(SST\), \(s\), and \(r^2\). Then interpret \(r^2\) as a fraction of output variation and \(s\) in output units.

Do. The means are \(\bar{x}=2\) adjustment units and \(\bar{y}=176\) output units. The sum of products of deviations is 380, and the sum of squared \(x\)-deviations is 10. Thus the slope is \(380/10=38\) output units per adjustment unit, and the intercept is \(176-38(2)=100\) output units. The fitted line is:

$$ \hat{y}=100+38x $$

At 2.5 adjustment units, the line predicts \(100+38(2.5)=195\) output units. At the five observed settings, the predictions are 100, 138, 176, 214, and 252 output units. The residuals are \(0,2,-6,6,-2\) output units, so:

$$ SSE=0^2+2^2+(-6)^2+6^2+(-2)^2 =0+4+36+36+4=80 $$

The residual standard deviation is:

$$ s=\sqrt{\frac{80}{5-2}} =\sqrt{26.6667\ldots} \approx 5.164\text{ output units} $$

The observed outputs have mean 176. Their deviations are \(-76,-36,-6,44,74\), so \(SST=5776+1296+36+1936+5476=14520\). Therefore:

$$ r^2=1-\frac{80}{14520} \approx 0.9945 $$

For a check, \(1-r^2\approx 0.00551\), and \(0.00551(14520)\approx 80\), with the small difference due to rounding.

Conclude in context. The line predicts an output of 195 units at 2.5 adjustment units, an interpolation. The \(r^2\) value of about 0.9945 means that about 99.45% of the observed output variation is accounted for by the linear relationship in these trials. However, \(s\) is about 5.164 output units. A high \(r^2\) does not mean every output will be close to its prediction; \(s\) provides a response-unit description of typical residual size.

Common Mistakes and AP Exam Tips

  • Writing “interpolation, so the prediction is reliable.” Interpolation identifies the prediction as being within the observed \(x\)-range. To discuss reliability, also consider the linear pattern, residual scatter, and whether observations support that part of the range.
  • Treating \(r^2\) as an error size. \(r^2\) is a unitless fraction of response variation accounted for by the linear relationship. It is not a number of response units and is not the percentage of predictions that are correct.
  • Calling \(s\) a maximum error. \(s\) describes typical residual size; it is not a boundary that every observed or future response must stay within.
  • Using only one fit summary. \(r^2\) and \(s\) provide different information. Report \(r^2\) in terms of response variation and \(s\) in response units; do not substitute one for the other.
  • Ignoring where observations occur within the range. A request can be inside the minimum and maximum \(x\)-values but far from many observed \(x\)-values. Mention that limited local data may make the prediction less directly supported.
  • Claiming certainty about a new case. A regression line gives a predicted response, not the exact value that must occur. Keep the conclusion tied to the cases and data described.

A careful AP-style statement might say: “At 2.5 adjustment units, the fitted line predicts an instrument output of 195 units. This is interpolation because 2.5 lies within the observed range of 0 to 4 adjustment units. The model has \(r^2\approx0.9945\), so about 99.45% of the output variation in these trials is accounted for by the linear relationship, while \(s\approx5.164\) output units describes the typical residual size. These summaries do not guarantee the output for an individual trial.”

Key takeaway: Being within the observed \(x\)-range makes a prediction interpolation, which is generally better supported than extrapolation. Still assess the scatter and linear pattern. Use \(r^2\) to describe variation accounted for and \(s\) to describe typical residual size in response units; neither guarantees an individual prediction.

Check Your Understanding

Use the distinction between interpolation, residual scatter, and \(r^2\) to answer each question.

  1. A fitted model uses observations with \(x\)-values from 4 to 18. Is a prediction at \(x=12\) interpolation or extrapolation? What additional information would help assess its reliability?
  2. In a regression model, \(r^2=0.81\). Explain what this means in context and identify one conclusion it does not justify.
  3. A model has residual standard deviation \(s=2.4\) centimeters. What does this describe, and why is it incorrect to call 2.4 centimeters the maximum prediction error?
  4. Two models have similarly high \(r^2\), but their response variables use very different scales. Why should you compare their \(s\)-values in context rather than assuming their prediction errors are equally small?
  5. A requested \(x\)-value is within the observed range but lies in a large gap between clusters of observations. How can you describe its classification and the limitation in support from the data?