Tutorials › AP Statistics › Defining the Coefficient of Determination r-squared

Coefficient of determination · Tutorial 903 of 1000

Defining the Coefficient of Determination r-squared

Learn what the coefficient of determination measures, how it connects to correlation and sums of squares, and how to interpret it without confusing it with direction or causation.

Intermediate 9 min read

What You'll Learn

  • State the relationship between the coefficient of determination and correlation.
  • Interpret \(r^2\) as a fraction or percentage of variation in the response accounted for by a linear model.
  • Connect \(r^2\) to the variation accounted for and total variation.
  • Describe what the remaining fraction represents for a least-squares line.
  • Distinguish the information in \(r^2\) from the direction of an association and from a claim of causation.

From Sums of Squares to a Fraction

In “What Variability Explained Means,” you saw how a least-squares regression line with an intercept divides the total variation in the response into variation represented by fitted values and variation left in the residuals. Those amounts, \(SST\) and \(SSE\), are measured in squared response units. The coefficient of determination, written \(r^2\), puts the amount accounted for by the line on a relative scale: it describes the fraction of the response’s total variation accounted for by the linear model.

In simple linear regression, \(r\) is the correlation between the explanatory variable \(x\) and the response variable \(y\). The coefficient of determination is the square of that correlation. For a least-squares regression line with an intercept, it is also the fraction of total variation accounted for by the line.

Definition: The coefficient of determination, \(r^2\), is the square of the correlation \(r\). In simple linear regression, it is the fraction of variation in the response variable \(y\) accounted for by the least-squares linear model using \(x\).
$$ r^2=(r)^2 =\frac{\text{variation accounted for by the line}}{SST} =\frac{\sum_{i=1}^{n}(\hat{y}_i-\bar{y})^2} {\sum_{i=1}^{n}(y_i-\bar{y})^2} =1-\frac{SSE}{SST} $$

The fraction has no units, even though its numerator and denominator are sums of squares with squared response units. Multiplying \(r^2\) by 100% expresses the fraction as a percentage. For example, a coefficient of determination of \(0.72\) corresponds to 72% of the response’s variation accounted for by the fitted line.

For a simple linear regression, \(r^2\) is between 0 and 1, inclusive. A value closer to 1 means the line accounts for a larger fraction of the response’s total variation in the data. A value closer to 0 means it accounts for a smaller fraction. This describes how the line fits the observed data; it does not say that the model predicts every individual response closely.

What the Fraction Means—and Does Not Mean

Suppose \(r^2=0.64\) for a least-squares regression of response \(y\) on explanatory variable \(x\). The appropriate interpretation is that 64% of the variation in \(y\), among the observations in the data set, is accounted for by the linear regression of \(y\) on \(x\). The remaining 36% is variation left in the residuals relative to this model.

The word “variation” matters. The statement is about how responses differ across the data set, compared with the response mean as a baseline. It is not a statement that 64% of the response values are predicted correctly, that predictions are 64% accurate, or that the model explains 64% of each person’s response. It is a summary of the overall variation, not a score for individual observations.

Also, “accounted for” does not mean “caused by.” A large \(r^2\) alone does not show that changing \(x\) causes changes in \(y\). Other variables, the way data were collected, or other features of the situation may be involved. The wording describes the association and the fit of the linear model.

Interpretation pattern: “About \(100r^2\)% of the variation in [response], among [the observations or population described], is accounted for by the linear regression of [response] on [explanatory variable].” Use the response variable in the variation statement, and do not imply causation unless the study design justifies a causal conclusion.

The remaining fraction \(1-r^2\) corresponds to the fraction of total variation left in the residuals for a least-squares line with an intercept. It does not mean the model is useless: whether the remaining variation matters depends on the context and the purpose of the model. A small leftover fraction can still matter when precise predictions are important, while a larger one might be acceptable for a rough description.

Worked Example: Interpret a Fraction From Sums of Squares

A fictional coastal science class records water temperature \(x\), in degrees Celsius, and the number of small fish observed during a fixed survey \(y\), for several sampling locations. For the least-squares regression of fish count on temperature, the total variation is \(SST=125\) fish-count-squared units, and the variation accounted for by the line is \(100\) fish-count-squared units. Interpret the coefficient of determination.

State. Find the fraction and percentage of variation in fish count accounted for by its linear regression on water temperature.

Plan. Divide the variation accounted for by the line by \(SST\). This ratio is \(r^2\). Check the complementary fraction by comparing it with the residual variation, then state the result in context.

Do. The accounted-for fraction is

$$ r^2=\frac{100}{125}=0.80 $$

The residual variation is \(SSE=125-100=25\) fish-count-squared units. Its share of the total is \(25/125=0.20\), and \(0.80+0.20=1\). As a percentage, the accounted-for fraction is \(0.80(100\%)=80\%\).

Conclude. About 80% of the variation in fish counts across the sampled locations is accounted for by the least-squares linear regression of fish count on water temperature. The remaining 20% is variation left in the residuals. This describes the fit of the line; it does not establish that water temperature caused differences in fish counts.

Keep \(r^2\) Separate From the Sign of \(r\)

Correlation \(r\) carries information about the direction of a linear association: its sign indicates whether the association is positive or negative. Squaring \(r\) removes that sign. Therefore, \(r^2\) describes the fraction accounted for, but it does not tell you whether the association slopes upward or downward. To describe direction, look at \(r\), the slope, or the scatterplot as appropriate.

For example, \(r=0.6\) and \(r=-0.6\) both produce \(r^2=0.36\). These correlations have opposite directions, but the same coefficient of determination. A statement that “36% of the variation is accounted for” does not tell the reader whether larger \(x\)-values tend to go with larger or smaller \(y\)-values.

A value \(r^2=0\) means the fitted linear model accounts for none of the response’s total variation beyond the mean baseline. In simple linear regression, that corresponds to \(r=0\). It does not rule out every possible relationship between \(x\) and \(y\): a nonlinear pattern, for example, may exist even when the linear correlation is zero. The coefficient of determination summarizes linear fit, not every kind of association.

Worked Example: Same \(r^2\), Different Directions

Two fictional teams study the relationship between weekly training time \(x\) and a performance measure \(y\). In one data set, the correlation is \(r=0.6\); in another, it is \(r=-0.6\). Compare what the coefficient of determination tells you about the two linear relationships.

State. Compare the coefficients of determination and identify which feature of the relationships is not captured by \(r^2\).

Plan. Square each correlation to find its coefficient of determination. Then compare the signs of the original correlations, because squaring does not retain direction.

Do. For the first relationship, \(r^2=(0.6)^2=0.36\). For the second, \(r^2=(-0.6)^2=0.36\). Both coefficients of determination are 0.36, or 36%. The correlations have different signs: the first is positive and the second is negative.

Conclude. In each data set, 36% of the variation in the performance measure is accounted for by its linear regression on weekly training time. The first relationship is positive and the second is negative; \(r^2\) alone cannot distinguish those directions. Neither result, by itself, demonstrates that training time causes a change in performance.

Interpreting a Reported Coefficient of Determination

In practice, a question may provide \(r^2\) directly in a calculator or computer report. Your job is then to identify which variable is the response, convert the fraction to a percentage if useful, and write a complete interpretation in context. Do not describe \(r^2\) as the percentage of variation in the explanatory variable. In a regression of \(y\) on \(x\), the coefficient of determination concerns variation in \(y\).

If \(r^2=0.36\), the remaining fraction is \(1-0.36=0.64\), or 64%. For a least-squares line with an intercept, this is the fraction of the total response variation left in the residuals. The two percentages refer to the division of variation, not to two groups of observations.

Worked Example: Interpret a Reported \(r^2\)

A fictional school technology team studies the relationship between the number of hours students use a review app each week \(x\) and their score on a short quiz \(y\), in points. A computer report gives \(r^2=0.36\) for the least-squares regression of quiz score on app-use hours. Explain the value and the remaining fraction.

State. Interpret the reported coefficient of determination for quiz scores and find the fraction of response variation not accounted for by the line.

Plan. Convert \(r^2\) to a percentage for the contextual interpretation. Subtract it from 1 to obtain the complementary fraction, which corresponds to the residual share for a least-squares line with an intercept.

Do. The accounted-for percentage is \(0.36(100\%)=36\%\). The remaining fraction is \(1-0.36=0.64\), or \(0.64(100\%)=64\%\).

Conclude. About 36% of the variation in quiz scores among the students in the data is accounted for by the linear regression of quiz score on weekly app-use hours. The remaining 64% is variation left in the residuals. The result does not say that 36% of students’ scores were predicted correctly, and it does not show that using the app caused higher or lower scores.

Common Mistakes and AP Exam Tips

  • Interpreting \(r^2\) as a direction. The square is nonnegative, so it does not show whether the relationship is positive or negative. Use the sign of \(r\) or the slope to describe direction.
  • Putting the variation in the wrong variable. In a regression of \(y\) on \(x\), \(r^2\) describes variation in the response \(y\), not variation in \(x\).
  • Saying “percent of values predicted correctly.” The coefficient of determination is a fraction of variation, not a count or percentage of observations.
  • Leaving out the context. A full interpretation names the response, the explanatory variable, and the setting or observations being described.
  • Claiming that \(x\) caused \(y\). “Accounted for by the linear regression” describes a statistical fit. It does not by itself establish a cause-and-effect relationship.
  • Calling the remaining fraction another percentage of observations. For a least-squares line with an intercept, \(1-r^2\) is the fraction of response variation left in residuals, not the percentage of cases the model failed to predict.

For full credit, write the interpretation as a statement about the percentage of variation in the named response accounted for by its linear regression on the named explanatory variable, in context. If direction is requested, discuss it separately using the sign of \(r\) or the slope. If causation is requested, consider the study design rather than treating \(r^2\) as causal evidence.

Key takeaway: In simple linear regression, \(r^2\) is the square of \(r\) and the fraction of variation in the response accounted for by the least-squares linear model. It summarizes the amount accounted for, not direction, individual prediction accuracy, or causation.

Check Your Understanding

Answer in context, and distinguish the response’s variation from the direction of the association.

  1. A least-squares line for predicting garden yield from weekly watering time has \(r^2=0.49\). Write an interpretation in context without implying causation.
  2. For a least-squares line with an intercept, what does the fraction \(1-r^2\) represent?
  3. Two data sets have correlations \(r=0.8\) and \(r=-0.8\). What do their coefficients of determination have in common, and what differs?
  4. Why is “49% of the garden yields were predicted correctly” not an appropriate interpretation of \(r^2=0.49\)?
  5. Can \(r^2=0\) rule out every possible relationship between the explanatory and response variables? Explain briefly.