What Does “Unexplained” Variation Look Like?
In “r-squared and Residual Standard Deviation Together,” you connected \(r^2\), \(SSE\), and the residual standard deviation \(s\). This tutorial focuses on the part of response variation not accounted for by the fitted line: the fraction \(1-r^2\). It gives a relative summary of the squared residual scatter compared with the response’s total variation.
A residual is the vertical difference between an observed response and its predicted value, \(y-\hat{y}\). A residual plot displays these differences around the zero line. The residuals’ squared values contribute to \(SSE\), so \(SSE\) summarizes their combined squared size. The value \(1-r^2\) compares that sum with \(SST\), the total squared variation of the response around its mean.
For example, if \(1-r^2=0.20\), then \(SSE\) is 20% of \(SST\). Equivalently, the linear relationship accounts for 80% of the response’s variation. This is a statement about a ratio of sums of squared deviations—not a claim that 20% of observations are predicted incorrectly or that each residual is 20% of its observed response.
From a Fraction to Residual Scatter
The connection to the residual plot is through the vertical residuals. When residuals are generally larger in magnitude, their squares tend to contribute more to \(SSE\). If the response observations are fixed, then \(SST\) is fixed too. In that setting, a larger \(SSE\) means a larger \(SSE/SST\), and therefore a larger \(1-r^2\): more of the response’s variation remains in the residuals.
This comparison depends on having the same response observations. If two analyses use different response values, their \(SST\) values may differ. Then the same \(SSE\) need not produce the same unexplained fraction. In general, compare \(SSE\) with \(SST\), not in isolation.
The fraction summarizes magnitude relative to total response variation, but it does not preserve the order or arrangement of the residuals in a plot. A residual plot may show random scatter, a curve, or changing spread. Those patterns matter when judging whether a linear model is appropriate; they cannot be read from \(1-r^2\) alone. This builds on “Reading a Residual Plot for Random Scatter,” “Curved Patterns in a Residual Plot,” and “Fan-Shaped Residual Plots and Changing Spread.”
There is also a useful link to \(s\), the residual standard deviation. For simple linear regression, \(s=\sqrt{\frac{SSE}{n-2}}\). Thus, with \(n\) fixed, a larger \(SSE\) also gives a larger \(s\). But \(1-r^2\) has no units, while \(s\) is measured in response units. One describes relative unexplained variation; the other gives a typical scale for residuals.
Worked Examples: Interpreting Unexplained Variation
Worked Example: Finding the Unexplained Fraction
A fictional engineering class fits a linear model to predict the operating time, in hours, of 12 small devices from a charging measurement. For this fitted line, \(SST=800\) hours squared and \(SSE=200\) hours squared. Find \(r^2\) and \(1-r^2\), and describe what the unexplained fraction says about residual scatter.
State. The response is operating time in hours. We want to describe how much of its total variation remains in the residuals relative to the total variation in the observed operating times.
Plan. Use \(1-r^2=SSE/SST\), then use \(r^2=1-SSE/SST\). The ratio has no units because both sums of squares use squared hours.
Do. The unexplained fraction is \(1-r^2=\frac{200}{800}=0.25\). Therefore, \(r^2=1-0.25=0.75\). As a check, \(SSE/SST=200/800=0.25\), and \(1-0.25=0.75\).
Conclude. About 25% of the total variation in operating time remains in the residuals, and about 75% is accounted for by the linear relationship with the charging measurement. The value 0.25 summarizes residual scatter relative to total response variation; it does not say that 25% of devices have inaccurate predictions or that all residuals have the same size.
Worked Example: Same Unexplained Fraction, Different Residual Patterns
Two fictional linear models are fitted to the same 20 response observations, using different explanatory variables. The shared response values have \(SST=500\) points squared. Each model has \(SSE=125\) points squared. One residual plot shows residuals scattered above and below zero without a clear pattern; the other shows a curved pattern. Compare their unexplained fractions and explain what the comparison does—and does not—tell you.
State. Because both models use the same response observations, their \(SST\) values are the same. We will compare the residual sums of squares relative to this shared total variation.
Plan. For each model, calculate \(1-r^2=SSE/SST\). Then calculate \(r^2\) as its complement. The residual plots must also be considered because the ratio itself does not describe their patterns.
Do. For either model, \(1-r^2=\frac{125}{500}=0.25\), so \(r^2=1-0.25=0.75\). Both models therefore have the same unexplained fraction, 25%, and the same coefficient of determination, 0.75. As an additional check on the residual scale, \(s=\sqrt{\frac{SSE}{n-2}}=\sqrt{\frac{125}{20-2}}=\sqrt{\frac{125}{18}}\approx2.6352\), or about 2.64 points for each model.
Conclude. Both models leave the same fraction of the response’s total variation in their residuals, and both have the same residual standard deviation for these data. Yet their residual plots differ: one has unstructured scatter, while the other has a curve. Equal \(SSE\) values imply equal \(1-r^2\) here because the fits use the same response observations and thus have the same \(SST\). The curved pattern still signals that the corresponding line misses systematically, so the equal summaries do not make the models equivalent in every respect.
Worked Example: Same Fraction, Different Absolute Scatter
Two fictional environmental monitoring programs fit separate models to predict a response measured in milligrams per liter. Program A has \(n=16\), \(SST=960\) squared response units, and \(SSE=240\) squared response units. Program B also has \(n=16\), but has \(SST=240\) and \(SSE=60\) in the same squared units. Compare their unexplained fractions and residual standard deviations.
State. These are different sets of response observations. We will compare the proportion of variation left in the residuals and also calculate \(s\), which describes residual size in the response’s units.
Plan. For each program, divide \(SSE\) by \(SST\) to find \(1-r^2\). Then find \(s=\sqrt{\frac{SSE}{n-2}}\). Since \(n=16\), the denominator for both residual standard deviations is \(16-2=14\).
Do. For Program A, \(1-r^2=\frac{240}{960}=0.25\), so \(r^2=0.75\). Its residual standard deviation is \(s=\sqrt{\frac{240}{14}}=\sqrt{17.1429}\approx4.1404\), or about 4.14 milligrams per liter. For Program B, \(1-r^2=\frac{60}{240}=0.25\), so \(r^2=0.75\) as well. Its residual standard deviation is \(s=\sqrt{\frac{60}{14}}=\sqrt{4.2857}\approx2.0702\), or about 2.07 milligrams per liter.
Conclude. Each model leaves 25% of its response variation in the residuals. But Program A has the larger residual standard deviation, about 4.14 rather than 2.07 milligrams per liter. The equal fractions do not imply equal absolute residual scatter: the total response variation differs between the programs. This is why \(1-r^2\) and \(s\) answer different questions.
How to Read the Residual Plot Alongside \(1-r^2\)
Think of \(1-r^2\) as a measure of the amount of residual variation relative to the response’s total variation. Think of a residual plot as a display of where residuals occur and how their signs and sizes change across the horizontal range. The two views are complementary. The fraction provides a compact numerical summary; the plot can reveal structure that the summary hides.
For example, a small \(1-r^2\) indicates that \(SSE\) is small compared with \(SST\). It does not guarantee that the residuals are patternless. A few influentially large residuals or a systematic curve may still deserve attention. Conversely, a larger unexplained fraction tells you that the residual sum of squares is large relative to total response variation, but the fraction alone does not identify where the line misses or whether the pattern is curved, fan-shaped, or otherwise organized.
When comparing two fitted lines for the same response observations, the common \(SST\) makes the comparison direct: the fit with smaller \(SSE\) has smaller \(1-r^2\) and larger \(r^2\). Even then, inspect residual plots before deciding what the fit says about the relationship. When the response observations differ, the denominator may differ too; a raw comparison of \(SSE\) values alone can be misleading.
Common Mistakes and AP Exam Tips
- Calling \(1-r^2\) the percent of wrong predictions. It is the fraction of the response’s total variation that remains in the residuals, not the percentage of observations predicted incorrectly.
- Comparing \(SSE\) without checking \(SST\). Equal \(SSE\) values give equal \(1-r^2\) only when the total response variation is also equal. Fits to the same response observations have the same \(SST\); fits to different response data may not.
- Assuming equal \(1-r^2\) means identical residual plots. The ratio summarizes the total squared residual size, not the arrangement of residuals. Describe any visible curve, fan shape, or other pattern separately.
- Mixing up relative and absolute scatter. \(1-r^2\) is unitless. The residual standard deviation \(s\) is in the response’s units. Equal unexplained fractions can occur with different values of \(s\).
- Interpreting the fraction as a statement about every case. A fraction summarizes the whole data set. It does not guarantee that each prediction is close or that each residual has a particular size.
A careful AP response identifies the response variable and states that \(1-r^2\) is the fraction of its total variation remaining in the residuals. If comparing fits, explain whether they share the same response observations. Then use the residual plot to discuss pattern and \(s\) to describe typical residual size with units.
Check Your Understanding
Use \(1-r^2=SSE/SST\) and distinguish relative variation from residual-plot patterns.
- A model has \(SST=600\) and \(SSE=150\). Find \(1-r^2\) and \(r^2\), then interpret the unexplained fraction in context if the response is plant height.
- Two models use the same response observations and have equal \(SSE\). What can you conclude about their \(1-r^2\) values? What can you not conclude about their residual plots?
- Why might two models fitted to different response data have equal \(SSE\) but different \(1-r^2\) values?
- Two programs have the same \(1-r^2\), but one has a larger \(s\). Explain how both facts can be true.
- Give one reason a residual plot can be useful even after you know \(r^2\) and \(1-r^2\).