Tutorials › AP Statistics › Low and High r-squared Values

Coefficient of determination · Tutorial 910 of 1000

Low and High r-squared Values

Interpret low and high \(r^2\) values by describing how much variation in the response the linear relationship accounts for.

Intermediate 9 min read

What You'll Learn

  • Interpret \(r^2=0.05\) as 5% of response variation accounted for by the linear relationship.
  • Interpret \(r^2=0.95\) as 95% of response variation accounted for by the linear relationship.
  • Describe what the unaccounted-for percentage means without assigning it to particular causes.
  • Compare \(r^2\) values in context and identify which line accounts for more response variation.
  • Avoid treating a high \(r^2\) as proof of perfect predictions or a low \(r^2\) as proof of no relationship.

What “Low” and “High” Mean for \(r^2\)

In “Interpreting \(r\)-squared in Context,” you learned to describe \(r^2\) as the proportion of variation in the response accounted for by its linear relationship with the explanatory variable. This tutorial applies that interpretation to values near the low and high ends of the scale. The central idea is that \(r^2\) summarizes how much of the response’s variation is accounted for by the fitted linear model.

A low \(r^2\), such as \(0.05\), means the linear relationship accounts for a small proportion of the variation in the response. A high \(r^2\), such as \(0.95\), means it accounts for a large proportion. These numbers are proportions, so converting them to percentages makes their meaning easier to express: \(0.05\) is \(5\%\), and \(0.95\) is \(95\%\).

Key idea: A low \(r^2\) means a small percentage of the response’s variation is accounted for by the linear relationship. A high \(r^2\) means a large percentage is accounted for. Always name the response, the explanatory variable, and the context when interpreting the value.

“Fit” here refers to how much variation in the observed response is accounted for by the line, not to whether a prediction for one individual is exactly right. A larger \(r^2\) means the line accounts for more of the response’s variation in the data being described. A smaller \(r^2\) means it accounts for less. Neither value, by itself, identifies why the remaining variation exists.

The complement of \(r^2\) is the proportion of response variation not accounted for by the linear model. For \(r^2=0.05\), the complement is \(1-0.05=0.95\), or \(95\%\). For \(r^2=0.95\), the complement is \(1-0.95=0.05\), or \(5\%\). This complement describes variation not accounted for by the line; it does not say that a particular other variable or cause accounts for all of it.

Interpreting a Low \(r^2\)

Suppose a regression of a response on an explanatory variable has \(r^2=0.05\). The linear relationship accounts for about \(5\%\) of the variation in the response among the observations. The other \(95\%\) is not accounted for by this linear model. That is a careful description of what the statistic says.

It would be too strong to conclude from \(r^2=0.05\) alone that the variables are unrelated. The value describes the strength of the linear model’s accounting for response variation; it does not summarize every possible feature of a relationship. It would also be incorrect to say that the line predicts \(5\%\) of cases correctly. As explained in “\(r\)-squared Is Not Percent Correct Predictions,” \(r^2\) is about variation, not a count of accurate predictions.

A low value often means the fitted line leaves substantial response variation unaccounted for. Whether that makes the line useful for a particular task depends on the purpose and context. For example, a modest account of variation might still be useful in an exploratory setting, while a prediction task requiring high precision may need more. The number alone does not settle that practical decision.

Worked Example: Interpreting \(r^2=0.05\)

A fictional community garden records the number of hours of sunlight received by each garden bed and the kilograms of tomatoes harvested from that bed. A simple linear regression uses sunlight hours to predict harvest, and its coefficient of determination is \(r^2=0.05\). Interpret this value and its complement.

State. The response is tomato harvest in kilograms, and the explanatory variable is sunlight hours. The reported coefficient of determination is \(0.05\).

Plan. Convert \(r^2\) to a percentage to describe the variation in the response accounted for by the linear relationship. Then calculate the complementary percentage not accounted for by the line.

Do. The percentage accounted for is \[ 0.05\times 100\%=5\%. \] The complementary proportion is \[ 1-0.05=0.95, \] which is \[ 0.95\times 100\%=95\%. \]

Conclude. In these garden beds, about \(5\%\) of the variation in tomato harvest is accounted for by its linear relationship with sunlight hours. About \(95\%\) of the variation in harvest is not accounted for by this linear model. The \(r^2\) value does not identify what accounts for that remaining variation.

Interpreting a High \(r^2\)

Now suppose \(r^2=0.95\). The fitted line accounts for about \(95\%\) of the variation in the response through its linear relationship with the explanatory variable. This indicates that the linear model accounts for a large proportion of the response variation in the data. The complement is \(5\%\), the proportion not accounted for by the line.

A high \(r^2\) is evidence that the line accounts for a large share of variation in the response; it does not mean the line passes through every observed point. Individual observations can still differ from their predicted values. In “Standard Deviation of the Residuals, \(s\)” and “Using \(s\) to Describe Prediction Accuracy,” you learned about another summary that describes the typical size of prediction errors in the response’s units. The \(r^2\) interpretation is not a replacement for that information.

A high \(r^2\) also does not establish that changes in the explanatory variable cause changes in the response. Describe the relationship in terms of the variables and the variation accounted for, without making a causal claim unless the study design supports one. A high value is about the line’s account of variation in the observed data, not proof of a cause-and-effect mechanism.

Worked Example: Interpreting \(r^2=0.95\)

A fictional school district examines the relationship between the number of practice problems students complete and their scores on a particular mathematics quiz. A simple linear regression uses practice problems completed to predict quiz score, and the output reports \(r^2=0.95\). Interpret the value.

State. The response is quiz score, and the explanatory variable is the number of practice problems completed. The reported coefficient of determination is \(0.95\).

Plan. Convert \(r^2\) to a percentage to describe the share of variation in quiz scores accounted for by the linear relationship. Find the complement to describe the share not accounted for.

Do. The accounted-for percentage is \[ 0.95\times 100\%=95\%. \] The proportion not accounted for is \[ 1-0.95=0.05, \] or \[ 0.05\times 100\%=5\%. \]

Conclude. For the students in this data set, about \(95\%\) of the variation in quiz scores is accounted for by the linear relationship between quiz score and practice problems completed. About \(5\%\) of the variation is not accounted for by this linear model. This interpretation does not say that \(95\%\) of students receive accurate predictions or that practice problems alone cause the scores.

Comparing Low and High Values

When interpreting several coefficients of determination, compare the percentages of response variation accounted for by their linear models. State which model accounts for more variation, and keep the comparison tied to the named response. The size of \(r^2\) does not depend on the units used to measure the response, but the interpretation still needs the actual response and explanatory variable to be meaningful.

For a straightforward comparison, the models should be describing the same response in a relevantly comparable setting. If one model predicts a different response, saying that its \(r^2\) is larger does not automatically make it the more useful model overall. Consider the study’s purpose and what counts as a useful prediction. This comparison is about the proportion of response variation accounted for, not every feature of the models.

Worked Example: Comparing Two Values of \(r^2\)

A fictional transit team fits two simple linear models to describe variation in the daily number of bicycle rentals at the same set of stations. Model A uses the number of hours of daylight and has \(r^2=0.40\). Model B uses the daily maximum temperature and has \(r^2=0.80\). Compare how much variation each line accounts for.

State. Both models use the number of bicycle rentals as the response, but they use different explanatory variables. Model A has \(r^2=0.40\); Model B has \(r^2=0.80\).

Plan. Convert each coefficient of determination to a percentage. Because both models describe variation in the same response, compare those percentages in context.

Do. For Model A, \[ 0.40\times 100\%=40\%. \] For Model B, \[ 0.80\times 100\%=80\%. \] The difference is \[ 80\%-40\%=40\text{ percentage points}. \]

Conclude. In these data, the linear relationship between bicycle rentals and maximum temperature accounts for about \(80\%\) of the variation in bicycle rentals, while the linear relationship between rentals and daylight hours accounts for about \(40\%\). Model B accounts for \(40\) percentage points more of the response variation. This comparison does not establish that temperature causes rentals to change.

Common Mistakes and AP Exam Tips

  • Leaving out the response variable. “The model explains \(95\%\)” is incomplete. A full-credit interpretation says, for example, “About \(95\%\) of the variation in quiz scores is accounted for by the linear relationship between quiz score and practice problems completed.”
  • Calling \(r^2\) a percentage of accurate predictions. An \(r^2\) of \(0.95\) does not mean that \(95\%\) of predictions are correct. It describes the proportion of variation in the response accounted for by the linear relationship.
  • Assigning the unexplained percentage to a specific cause. If \(r^2=0.05\), the line does not account for \(95\%\) of response variation. That statement does not show that a particular unmeasured variable accounts for the remaining variation.
  • Describing a high value as a perfect fit. A high \(r^2\) means the line accounts for a large proportion of the response variation, not that every observed response equals its prediction.
  • Using “no relationship” to describe a low value. A low \(r^2\) says that the linear model accounts for a small proportion of response variation. It does not, by itself, justify a broad claim that the variables have no relationship at all.
  • Mixing up percentage and percentage points. When comparing \(40\%\) with \(80\%\), the difference is \(40\) percentage points. State the individual percentages first so the comparison is clear.
  • Making a causal claim from \(r^2\). The coefficient of determination describes a linear association in the data. It does not by itself show that the explanatory variable causes changes in the response.

For an AP response, use “about” when the reported value is rounded, name the response, and use “accounted for by its linear relationship with” to make the interpretation precise. For example: “About \(5\%\) of the variation in harvest is accounted for by its linear relationship with sunlight hours.” If discussing the complement, say that variation is “not accounted for by the linear model,” rather than claiming to know its causes.

The next tutorial, “What \(r\)-squared Does Not Tell You,” develops an important caution: \(r^2\) is useful, but it does not summarize everything one might want to know about a regression relationship.

Key takeaway: A low \(r^2\), such as \(0.05\), means the line accounts for a small percentage of response variation; a high \(r^2\), such as \(0.95\), means it accounts for a large percentage. Interpret the value in context, and do not turn it into a claim about prediction accuracy, specific causes, or causation.

Check Your Understanding

For each question, identify the response and interpret the coefficient of determination in context.

  1. A fictional environmental project uses stream-flow rate to predict the number of insects counted at a sampling site. The regression has \(r^2=0.12\). What percentage of variation in the response is accounted for, and what percentage is not accounted for by the line?
  2. A model uses weekly hours of instrument practice to predict a music-performance score and has \(r^2=0.91\). Write a careful interpretation of the value.
  3. A student says, “An \(r^2\) of \(0.08\) means the model predicts \(8\%\) of cases correctly.” Explain the error and give a correct interpretation using a response of your choice.
  4. Two models predict the same response and have \(r^2=0.35\) and \(r^2=0.70\). Which accounts for more response variation, and by how many percentage points?
  5. Why is it not justified to say that the unaccounted-for percentage of variation is caused by one particular factor just from knowing \(r^2\)?