Total Variation Starts With a Baseline
In “Residuals Mixed Practice Set,” residuals described how far individual observed responses were from the values predicted by a fitted line. Before asking how well a regression line predicts, it helps to describe how much the response values vary in the first place. A natural reference point is the response mean, \(\bar{y}\).
For each observation, \(y-\bar{y}\) is its deviation from the mean. Some deviations are positive and some are negative; their signed sum is zero. To measure the overall size of these deviations without allowing positive and negative values to cancel, square them and add them. This sum is called the total variation in the response, or the total sum of squares.
Here \(y_i\) is the response for observation \(i\), \(\bar{y}\) is the mean of all \(n\) responses, and \(SST\) is a single nonnegative number. It is zero only when every response equals the mean. Because deviations are squared, \(SST\) is measured in squared response units: for example, minutes squared or score-points squared.
The mean provides a useful baseline because it is the constant prediction that minimizes the sum of squared prediction errors when no predictor is used. Predicting \(\bar{y}\) for every case gives prediction errors \(y_i-\bar{y}\). Their squared sum is exactly \(SST\). A regression line uses \(x\) to make different predictions for different cases; it aims to account for some of the response’s departures from this mean baseline. The next tutorial develops what it means for a regression to account for variability.
Calculating the Total Sum of Squares
To calculate \(SST\), find the response mean, subtract it from each response, square each deviation, and add the squares. Keep the deviations’ signs during subtraction, even though squaring makes each final contribution nonnegative.
Calculate \(\bar{y}\) from the response values.
For each response, calculate \(y_i-\bar{y}\).
Square each deviation and add the results to obtain \(SST\).
Describe the result as a sum of squared departures from the mean, in squared response units.
This calculation is different from finding the range or adding absolute deviations. The range uses only the largest and smallest values, while \(SST\) uses every observation. Squaring also makes larger deviations contribute disproportionately more. A response far from the mean can therefore contribute much more to \(SST\) than a response only a little above or below the mean.
Worked Example: Calculate Total Variation in Practice Times
A fictional music club records the number of minutes five students practiced on a particular day: 12, 15, 15, 18, and 20. Find the mean practice time and the total variation in practice time.
State. Calculate \(SST\), the sum of the squared deviations of the five practice times from their mean.
Plan. Find the mean, calculate each value’s deviation from it, and square and add those deviations. The result will be in minutes squared.
Do. The mean is \(\bar{y}=(12+15+15+18+20)/5=80/5=16\) minutes. The deviations from 16 are \(-4,-1,-1,2,4\) minutes. Therefore,
Conclude. The total variation in these practice times is \(38\) minutes squared. This describes the combined squared departures of the five times from their mean of 16 minutes; it is not itself a typical practice time.
As a check, the signed deviations add to \(-4-1-1+2+4=0\), as deviations from the mean should. The squared deviations add to 38, so the positive and negative departures do not cancel in \(SST\).
How \(SST\) Relates to the Response Standard Deviation
The total sum of squares is connected to the sample standard deviation \(s_y\) of the response. The sample variance \(s_y^2\) is the sum of squared deviations divided by \(n-1\). Thus, if \(s_y\) is known, \(SST\) can be recovered by multiplying \(s_y^2\) by \(n-1\).
The denominator is \(n-1\) for the sample variance; it is not part of the definition of \(SST\). The standard deviation \(s_y\) is in the original response units, whereas \(SST\) is in squared response units. For the practice-time data, \(n=5\) and \(SST=38\text{ minutes}^2\), so \(s_y=\sqrt{38/4}=\sqrt{9.5}\approx3.08\) minutes. Reversing the calculation gives \((5-1)(3.08)^2\approx38\text{ minutes}^2\), allowing for rounding.
This relationship offers a convenient check on a calculation, but the two quantities answer different questions. \(s_y\) gives a summary of the typical distance of responses from their mean in response units. \(SST\) is the total of all squared distances and depends on the number of observations as well as their spread.
Worked Example: Use the Mean as a Prediction Baseline
A fictional community garden records the time, in minutes, that five volunteers take to complete a routine task: 6, 8, 8, 10, and 13. Consider a baseline prediction that gives every volunteer the mean task time. Calculate the baseline’s sum of squared prediction errors and connect it to \(SST\).
State. Find the mean-only prediction and its total squared error for these five observed times.
Plan. The mean-only prediction is the response mean. Subtract that predicted value from each observed time, square each error, and add. This is the same calculation as finding total variation around the mean.
Do. The mean is \(\bar{y}=(6+8+8+10+13)/5=45/5=9\) minutes. Predicting 9 minutes for each volunteer gives errors \(6-9=-3\), \(8-9=-1\), \(8-9=-1\), \(10-9=1\), and \(13-9=4\) minutes. The squared errors add to
The deviations from the mean are exactly those same five errors. Therefore, the baseline’s sum of squared prediction errors is \(28\text{ minutes}^2\), which is also \(SST\) for these task times.
Conclude. Predicting the mean for every volunteer produces a total squared error equal to the total variation in task time around the mean: \(SST=28\text{ minutes}^2\). A regression line can use a predictor to make case-specific predictions rather than giving everyone the same baseline prediction.
Total Variation Depends on the Response Units
Changing the response’s measurement units changes the numerical size of its deviations from the mean, and therefore changes \(SST\). The underlying observations have not become more or less spread out in a practical sense; only their numerical scale has changed. Since \(SST\) squares deviations, a change in the size of the unit has a squared effect on its numerical value.
Worked Example: Convert Total Variation to Larger Units
Three fictional seedlings have measured lengths of 12, 14, and 16 centimeters. Find \(SST\) in centimeters squared, then express the same measurements in decimeters and find \(SST\) in decimeters squared. One decimeter is 10 centimeters.
State. Compare the numerical total variation in the two measurement units.
Plan. Calculate the mean and squared deviations in centimeters. Then convert each length and its mean to decimeters and repeat. Since a decimeter is a larger unit, the numerical lengths and deviations should be one tenth as large.
Do. In centimeters, the mean is \((12+14+16)/3=42/3=14\) cm. The deviations are \(-2,0,2\) cm, so \(SST=(-2)^2+0^2+2^2=4+0+4=8\text{ cm}^2\).
In decimeters, the lengths are 1.2, 1.4, and 1.6 dm, with mean \((1.2+1.4+1.6)/3=4.2/3=1.4\) dm. The deviations are \(-0.2,0,0.2\) dm, so \(SST=(-0.2)^2+0^2+(0.2)^2=0.04+0+0.04=0.08\text{ dm}^2\).
Each numerical deviation in decimeters is one tenth its value in centimeters, so the sum of squares is one hundredth as large numerically: \(8/100=0.08\). The conversion is consistent.
Conclude. The total variation is \(8\text{ cm}^2\), or \(0.08\text{ dm}^2\). The different numerical values reflect the unit conversion, not a change in how much the seedlings’ lengths vary.
How the Baseline Connects to a Regression Line
The mean-only prediction is a useful reference for regression. In a least-squares regression line with an intercept, the fitted predictions vary with \(x\), while their residuals measure the differences between observed and predicted responses. As established in “The Idea of the Least-Squares Criterion,” the least-squares line is chosen to minimize the sum of squared residuals among candidate lines.
A horizontal line at \(\bar{y}\) is one possible line with an intercept. Its squared errors add to \(SST\). Because the least-squares regression line chooses a line with the smallest sum of squared residuals, its SSE cannot be greater than the mean-only line’s \(SST\). This comparison gives a baseline for judging the fitted line. It does not, by itself, tell us what proportion of variability the line accounts for; that interpretation belongs in “What Variability Explained Means.”
Worked Example: Compare a Fitted Line With the Mean Baseline
For four fictional observations, a least-squares line relates study sessions \(x\) to a quiz score \(y\), in points. The observed pairs are \((1,3),(2,5),(3,6),(4,10)\), and the fitted line is \(\hat{y}=0.5+2.2x\). Calculate \(SST\) and SSE, then compare them.
State. Find the response’s total variation around its mean and the sum of squared residuals from the fitted line.
Plan. Calculate \(\bar{y}\) and square the observed responses’ deviations from it to find \(SST\). Then predict each score from the line, calculate \(y-\hat{y}\), and square and add the residuals to find SSE. Both sums use squared score points.
Do. The mean score is \(\bar{y}=(3+5+6+10)/4=24/4=6\) points. The deviations from the mean are \(-3,-1,0,4\), so \(SST=(-3)^2+(-1)^2+0^2+4^2=9+1+0+16=26\text{ points}^2\).
The line predicts \(0.5+2.2(1)=2.7\), \(0.5+2.2(2)=4.9\), \(0.5+2.2(3)=7.1\), and \(0.5+2.2(4)=9.3\) points. The residuals are \(3-2.7=0.3\), \(5-4.9=0.1\), \(6-7.1=-1.1\), and \(10-9.3=0.7\) points. Thus,
Conclude. The mean-only baseline has a squared error total of \(26\text{ points}^2\), while the fitted line’s SSE is \(1.80\text{ points}^2\). For these observations, the fitted line’s predictions leave a smaller total squared error than predicting the mean score for everyone. The comparison is a way to assess the line against a baseline; it is not yet an interpretation of variability explained.
Common Mistakes and AP Exam Tips
- Subtracting the wrong center. \(SST\) uses deviations from \(\bar{y}\), the response mean, not deviations from a particular predicted value or from zero.
- Adding signed deviations instead of squared deviations. The signed deviations from the mean sum to zero. For total variation, square each deviation before adding.
- Calling \(SST\) the variance or standard deviation. \(SST\) is a sum of squares. The sample variance is \(SST/(n-1)\), and the sample standard deviation is its square root.
- Reporting original response units. If \(y\) is measured in minutes, \(SST\) is in minutes squared. The sum is not a prediction or a typical number of minutes.
- Reversing the effect of a unit conversion. A larger measurement unit gives smaller numerical values and deviations. If the unit is twice as large, deviations are halved and \(SST\) is one quarter as large—not four times as large.
- Calling the baseline comparison “variability explained.” \(SST\) measures total variation around the mean. Comparing it with SSE sets up a later interpretation, but the comparison alone is not a complete explanation of the proportion or meaning of variability explained.
For full-credit communication, identify the response mean as the reference point, show the squared deviations or a clear equivalent calculation, and report \(SST\) in squared response units. If comparing the result with a regression line, state what each sum measures: \(SST\) is variation around the mean, while SSE is the sum of squared residuals around the fitted line.
Check Your Understanding
For each item, show the requested calculation or explanation and include response units where appropriate.
- For response values 4, 7, 7, and 10, calculate the mean and \(SST\).
- In words, explain why predicting the response mean for every observation produces a squared-error total equal to \(SST\).
- A response is measured in meters and has \(SST=18\text{ m}^2\). If the same values are expressed in centimeters, what is the numerical value of \(SST\) in \(\text{cm}^2\)? Explain the conversion.
- A data set has \(n=9\) and response standard deviation \(s_y=2.5\) points. Use the relationship between \(SST\) and \(s_y\) to find \(SST\).
- Why should a comparison between \(SST\) and a regression’s SSE not be described as the full interpretation of variability explained?