A Mixed-Practice Routine for \(r^2\)
In “Free-Response Practice With the Coefficient of Determination,” you practiced interpreting \(r^2\) and combining it with residual evidence. This tutorial brings together several common tasks: calculating \(r^2\) from a correlation, calculating it from sums of squares, checking whether reported values agree, and deciding what a result does and does not support.
The calculation depends on the information provided. If you know the correlation \(r\), square it. If you know the total variation \(SST\) and the unexplained variation \(SSE\), compare those sums of squares. If software gives \(R\)-sq, that is the reported coefficient of determination in simple linear regression. Whichever route you use, the interpretation still refers to variation in the response—not the percentage of predictions that are correct.
The relationship between sums of squares gives one useful calculation route. As covered in “Unexplained Variation and Residual Scatter,” \(1-r^2=SSE/SST\) for a least-squares line with an intercept. Rearranging gives \(r^2=1-SSE/SST\). The ratio \(SSE/SST\) is the fraction of total response variation left unexplained by the line; subtracting it from 1 gives the fraction accounted for.
As you work, keep two checks in mind. A calculated \(r^2\) must be between 0 and 1, and it must agree with \(r\) if both are reported: squaring \(r\) should produce \(r^2\). A negative correlation still gives a nonnegative \(r^2\); the sign describes direction, while \(r^2\) does not. These checks can help catch a copied value, a calculator-entry error, or a mistaken interpretation.
Worked Examples: Calculate, Interpret, and Critique
Worked Example: Calculate from a Negative Correlation
A fictional outdoor-equipment shop records the price, in dollars, and weekly units sold for 50 products. Weekly units sold is the response, and price is the explanatory variable. The correlation is \(r=-0.68\). Calculate and interpret \(r^2\). Does the result show that lowering a product’s price causes sales to increase?
State. The task is to find the coefficient of determination for the linear relationship between product price and weekly units sold in these 50 observations, then interpret its scope.
Plan. Use \(r^2=r\times r\), as in “Computing r-squared From r.” Square the negative correlation; the result will be nonnegative. Interpret the percentage as variation in weekly units sold, and do not treat an observational association as proof of cause and effect.
Do. The calculation is
A check using the magnitude gives \(0.68^2=68^2/10000=4624/10000=0.4624\), the same result. As a percentage, \(100(0.4624)=46.24\%\).
Conclude. About 46.24% of the variation in weekly units sold among these 50 products is accounted for by the linear relationship between weekly units sold and price. The negative \(r\) indicates a negative direction of association, but \(r^2\) itself does not give a direction. These data do not establish that changing a product’s price causes a change in sales; other factors could also be related to sales.
Worked Example: Calculate from Sums of Squares
A fictional environmental team models the dissolved-oxygen level, in milligrams per liter, at 30 stream locations using distance downstream as the explanatory variable. For the fitted least-squares line, \(SST=1250\) and \(SSE=450\). The residual plot shows residuals scattered above and below zero without an obvious pattern. Find and interpret \(r^2\), then comment on the evidence for using a line to describe these observations.
State. The response is dissolved-oxygen level at the 30 observed stream locations. The task asks for the fraction of its total variation accounted for by the line and what the residual plot adds.
Plan. Use \(r^2=1-SSE/SST\). Check that the unexplained sum of squares does not exceed the total sum of squares. Then convert \(r^2\) to a percentage and assess the residual plot separately from the calculation.
Do. Substitute the reported sums of squares:
To check, \(450/1250=45/125=9/25=0.36\), and \(1-0.36=0.64\). The percentage is \(100(0.64)=64\%\). Also, \(SSE=450\) is less than \(SST=1250\), as expected for these sums of squares.
The residuals show no obvious systematic pattern, which supports using a linear model to describe the relationship in these observations. That evidence is separate from the \(r^2\) calculation: a percentage alone does not reveal whether residuals follow a curve or another pattern.
Conclude. About 64% of the variation in dissolved-oxygen level among the 30 stream locations is accounted for by its linear relationship with distance downstream. The residual plot supports a linear description for the observed locations, though it does not establish that distance causes changes in dissolved oxygen or guarantee the model will work beyond the observed range.
Worked Example: Audit Conflicting Regression Output
A fictional sports program studies the relationship between weekly training hours and 100-meter sprint time for 40 athletes. A summary reports \(r=0.40\) and \(R\)-sq \(=0.64\). Sprint time is the response. Check whether the two reported values are consistent, and state what the correlation would imply for the coefficient of determination if \(r=0.40\) is correct.
State. Both reported values are intended to describe the same simple linear regression, so the reported \(R\)-sq should equal the square of the reported correlation.
Plan. Calculate \(r^2\) from \(r=0.40\), then compare it with the reported \(R\)-sq. If they disagree, identify the inconsistency rather than choosing whichever value seems more plausible. Check the calculation a second way.
Do. Squaring the reported correlation gives
A second check gives \(40^2/100^2=1600/10000=0.16\). This is not the reported \(R\)-sq of \(0.64\). In fact, a coefficient of determination of \(0.64\) would require a correlation with magnitude \(\sqrt{0.64}=0.80\), not \(0.40\).
Conclude. The reported \(r\) and \(R\)-sq values are inconsistent for the same simple linear regression. If \(r=0.40\) is correct, then \(r^2=0.16\), meaning about 16% of the variation in sprint time among these 40 athletes is accounted for by its linear relationship with weekly training hours. The output should be checked for a copying, labeling, or transcription error before making a model evaluation.
Worked Example: A Large \(r^2\) Does Not Cancel a Residual Pattern
A fictional school project models the number of minutes students spend on a reading app using their weekly reading time, in hours. The data include 55 students. The regression output reports \(R\)-sq \(=0.81\), but the residual plot shows a clear curve: residuals are mostly negative in the middle of the horizontal range and mostly positive near both ends. Interpret the coefficient of determination and critique the linear model.
State. The response is app-use time in minutes, and the explanatory variable is weekly reading time in hours. The question asks both what \(r^2\) describes and whether the residual evidence supports a straight-line model.
Plan. Convert \(0.81\) to a percentage and name the response, explanatory variable, and observed students. Then use the residual pattern to assess the linear form. Do not let the large \(r^2\) override evidence of systematic misses.
Do. The percentage conversion is
The result checks directly: \(0.81\) is 81 hundredths, or 81%. The negative residuals in the middle mean observed app-use times there tend to be below the line’s predictions; the positive residuals near both ends mean observed times tend to be above the predictions. This organized curve is not random scatter around zero.
Conclude. About 81% of the variation in app-use time among the 55 students is accounted for by its linear relationship with weekly reading time. However, the curved residual pattern suggests that the line misses the relationship systematically, so the large \(r^2\) alone is not enough to say that a straight-line model is adequate. This conclusion is limited to the observed students and range of reading times.
Common Mistakes and a Reliable Final Check
Mixed questions can switch quickly from arithmetic to interpretation. A useful final check is to ask whether the calculation, percentage, and claim all refer to the same model, response, and observations. Then check whether any extra evidence—such as a residual plot—changes how you should describe the model.
- Keeping a negative sign after squaring. For example, \((-0.68)^2=0.4624\), not \(-0.4624\). The negative sign indicates direction in \(r\); \(r^2\) is not negative.
- Using the wrong sums of squares. When \(SST\) and \(SSE\) are given, calculate \(1-SSE/SST\), not \(1-SST/SSE\). The unexplained fraction is \(SSE/SST\).
- Reading \(r^2\) as prediction accuracy. Do not say that a percentage of cases were predicted correctly. Say that the percentage of variation in the response is accounted for by its linear relationship with the explanatory variable.
- Ignoring a disagreement in output. In the same simple linear regression, \(R\)-sq should equal \(r^2\). If the values conflict, show the calculation and say that the report needs checking.
- Judging a model from \(r^2\) alone. A high value does not rule out a curved residual pattern, and a low value is not an automatic verdict that a model is useless. Use the task and other evidence provided.
- Claiming cause and effect. A strong linear relationship in observational data does not, by itself, show that changing the explanatory variable causes a change in the response.
For full-credit communication, make the calculation visible, convert to a percentage correctly, and use a sentence that names the response and the explanatory variable. For critique, point to the evidence: for instance, a curved residual pattern suggests systematic misses by a line. Avoid replacing that explanation with an unsupported label such as “good fit.”
Check Your Understanding
Show the calculation where appropriate, then keep the interpretation and critique tied to the stated context.
- A fictional study relates daily outdoor temperature to iced-drink sales at 45 kiosks. The correlation is \(r=0.75\). Calculate \(r^2\) and interpret it in context.
- For a line relating travel time to distance for 24 fictional shuttle trips, \(SST=800\) and \(SSE=280\). Calculate \(r^2\) and interpret the result as a percentage.
- A computer report for one simple linear regression gives \(r=-0.50\) and \(R\)-sq \(=0.25\). Are the values consistent? What does \(r^2\) describe in context?
- A model relating weekly practice hours to a fictional team’s free-throw percentage has \(R\)-sq \(=0.92\), but its residual plot shows a clear curve. What does \(r^2\) say, and what does the residual pattern add?
- A student writes, “The line predicts 64% of the stream locations correctly.” Explain why this is not a correct interpretation of \(r^2\), and state what the 64% should describe instead.