When a Straight Line Misses a Bend
In “Why Extrapolation Is Risky,” we saw that a fitted line continues its mathematical pattern beyond the observed data, even though the real relationship may not. One reason for that mismatch is curvature: the response may change at different rates across values of the explanatory variable, rather than following one straight-line trend.
A linear model uses a constant slope. For each additional unit of \(x\), it predicts the same change in \(y\). A curved relationship does not have that feature: the size, or sometimes the direction, of its changes varies as \(x\) changes. A line fitted to a limited part of the curve can summarize the general trend in that part, but extending its fixed slope can send the predictions away from the curved pattern.
The examples below use invented, deliberately simple curved patterns so that we can compare a line with a specified curve. In real data, we usually do not know the true relationship outside the observed range. The examples show how curvature can make extrapolation unreliable; they do not tell us exactly what an unknown real-world response will do.
Why the Mismatch Can Grow Outside the Data
Suppose a curved pattern increases slowly at first and then more quickly. A line fitted to the observed portion averages the changes there into one slope. Beyond the largest observed \(x\), that line keeps predicting changes at the same average rate. If the curve continues to bend upward, the line can increasingly underpredict it. For a curve that bends downward, a line may instead overpredict beyond the observed range.
The same reasoning applies when extrapolating below the smallest observed \(x\). A line continues its slope in that direction too, while a curved pattern may change at a different rate or bend the other way. The direction of the mismatch depends on the shape and direction of the curve; extrapolation does not always produce an overprediction or always an underprediction.
Residuals can offer a warning before making a prediction. If residuals tend to be positive at both ends of the observed range and negative near the middle, or show another systematic pattern, a straight line may be missing curvature. But even residuals that look small within the observed range do not guarantee that the line will work beyond it. That is why a close fit to the data is not, by itself, a reason to trust a distant extrapolation.
Worked Examples: Lines and Curved Patterns
Worked Example: A Process That Speeds Up
Hypothetical setting. A demonstration machine starts with 100 completed units. Let \(x\) be operating time in hours and \(y\) be the cumulative number of units completed. For illustration, the machine’s curved pattern over the time period of interest is \(y=100+4x+0.5x^2\). Measurements are taken at \(x=0,2,4,6\), giving \(y=100,110,124,142\). We will fit a line to these four observations, then compare predictions beyond 6 hours with values from the illustrative curve.
State. Find the least-squares line for the observed data, and compare its predictions at 8 and 10 hours with the values given by the curved pattern.
Plan. Calculate the slope using the sum of products of deviations divided by the sum of squared \(x\)-deviations. Then find the intercept from \(\bar{y}-b\bar{x}\). Both requested times are beyond the observed \(x\)-range of 0 to 6 hours, so they are extrapolations.
Do. The means are \(\bar{x}=3\) hours and \(\bar{y}=119\) units. The slope calculation is:
The fitted slope is 7 units per hour. The intercept is \(119-7(3)=98\) units, so the line is:
At 8 hours, the line predicts \(98+7(8)=154\) units. The curved pattern gives \(100+4(8)+0.5(8^2)=100+32+32=164\) units, 10 units more than the line predicts. At 10 hours, the line predicts \(98+7(10)=168\) units, while the curve gives \(100+40+50=190\) units, a difference of 22 units.
As a check on the fitted line, its predictions at the observed times are 98, 112, 126, and 140 units. Subtracting those predictions from the observed values gives residuals of 2, \(-2\), \(-2\), and 2 units. These residuals are modest, but the sign pattern—positive at the ends and negative in the middle—hints that the line misses a bend.
Conclude in context. The line summarizes the machine’s observations from 0 to 6 hours, but it underpredicts the specified curved pattern at 8 and 10 hours. The mismatch is larger at 10 hours. Both requested times are extrapolations, and the curved equation is part of this hypothetical illustration; the comparison does not establish how a real machine would perform beyond its measurements.
Worked Example: A Line Continues Up After a Curve Turns Down
Hypothetical setting. A sensor produces a response score \(y\) at a chosen setting \(x\). For this illustration, the response follows the curve \(y=80+12x-x^2\) over the settings being considered. Observations at settings 0, 2, 4, and 6 have scores 80, 100, 112, and 116. Find the least-squares line and compare its prediction at setting 8 with the curved value.
State. We need to see whether the line fitted to settings 0 through 6 represents the curved pattern at setting 8.
Plan. Compute the slope and intercept from the four observations. Since 8 is greater than the largest observed setting, the requested prediction is an extrapolation. Compare the line’s result with the value calculated from the illustrative curve.
Do. The means are \(\bar{x}=3\) and \(\bar{y}=102\). The sum of products of deviations is \(66+2+10+42=120\), and the sum of squared \(x\)-deviations is 20. Thus:
The least-squares line is \(\hat{y}=84+6x\). At setting 8, it predicts \(84+6(8)=132\). But the curved pattern gives \(80+12(8)-8^2=80+96-64=112\). The line is 20 score units higher than the curved value.
The observed predictions from the line are 84, 96, 108, and 120, so the residuals are \(-4,4,4,-4\). The negative residuals at the ends and positive residuals in the middle also suggest a bend. In this example the curve has reached its maximum at setting 6 and is lower at setting 8, while the line continues increasing at a constant rate.
Conclude in context. The line predicts a score of 132 at setting 8, whereas the specified curved pattern gives 112. The line overpredicts by 20 units in this illustration. Setting 8 is outside the observed range, and the comparison shows why a line can be especially misleading if the relationship turns after the data end. It does not establish that an actual sensor’s response will turn in this way.
Worked Example: A Narrower Data Range, a Larger Miss
Hypothetical setting. In a separate demonstration, an output \(y\) follows the illustrative curve \(y=10+x^2\), where \(x\) is time in minutes. Only the observations at \(x=0,1,2\), with outputs 10, 11, and 14, are used to fit a line. Compare its predictions at 3 and 5 minutes with the values from the curve.
State. Find the fitted line from the three observed points, then assess the two predictions beyond the observed range of 0 to 2 minutes.
Plan. Use the least-squares slope and intercept formulas. Compare each prediction with the illustrative curve, while identifying both requested times as extrapolations.
Do. The means are \(\bar{x}=1\) and \(\bar{y}=35/3\). The sum of products of deviations is \(4\), and the sum of squared \(x\)-deviations is \(2\). Therefore:
The fitted line is \(\hat{y}=\frac{29}{3}+2x\). At 3 minutes, it predicts \(\frac{29}{3}+6=\frac{47}{3}\), or about 15.667 units. The curve gives \(10+3^2=19\), which is \(\frac{10}{3}\), or about 3.333 units, higher. At 5 minutes, the line predicts \(\frac{29}{3}+10=\frac{59}{3}\), or about 19.667 units. The curve gives \(10+5^2=35\), which is \(\frac{46}{3}\), or about 15.333 units, higher.
Conclude in context. The fitted line underpredicts the specified curve at both times, and the difference is greater at 5 minutes than at 3 minutes. Both times lie outside the narrow range used to fit the line. This example illustrates how extending a fixed-slope line can drift farther from a continuing upward curve; it does not show that a real output must follow \(10+x^2\) or keep increasing indefinitely.
Common Mistakes and AP Exam Tips
- Assuming the line follows the curve. A line’s slope is constant. If the relationship bends, do not describe the line as if it captures every change in the response.
- Claiming the direction of the miss without checking the curve. An upward bend can lead to underprediction beyond the data, while a downward bend can lead to overprediction. Identify the particular pattern before stating which occurs.
- Treating an illustrative curve as known real-world truth. A constructed equation lets us compare the line with a specified pattern. In an actual study, the future or out-of-range response is unknown; describe the risk, not an unverified outcome.
- Thinking small residuals guarantee safe extrapolation. Residuals describe differences for observed cases. They do not establish that the linear pattern will continue outside the observed \(x\)-range.
- Forgetting to identify extrapolation. State the observed \(x\)-range and compare the requested value with its endpoints. Then explain separately how curvature could make the continued line inaccurate.
- Reporting a calculation without context. Give the predicted response with its units, identify whether it is above or below the comparison pattern when that pattern is specified, and say that the input is outside the data range.
A careful AP-style conclusion for the first example would be: “At 10 hours, the fitted line predicts 168 completed units, while the specified illustrative curve gives 190 units. The line underpredicts by 22 units there. Since 10 hours is beyond the observed range of 0 to 6 hours, this is an extrapolation; the example illustrates how unmodeled upward curvature can make the line miss.” In a real data setting without a known curve, avoid claiming a precise actual miss.
Check Your Understanding
Use the stated data or illustrative curves to analyze each line and explain what the comparison means.
- A curved pattern is \(y=20+2x^2\). Measurements are taken at \(x=0,1,2\). Explain why a line fitted to those observations may not give a reliable prediction at \(x=5\), even if it summarizes the observed points reasonably.
- A line underpredicts a specified curve at an \(x\)-value beyond the data. What does “underpredicts” mean, and why is the prediction still an extrapolation?
- Residuals are negative at both ends of the observed range and positive near the middle. What might this pattern suggest about a fitted line?
- A downward-bending curve is compared with a line beyond the largest observed \(x\)-value. What should you check before saying whether the line overpredicts or underpredicts?
- Why can a worked example with a known curved equation demonstrate extrapolation risk but not prove what an unknown real-world response will be outside its observed range?