Tutorials › AP Statistics › Writing a Complete r-squared Interpretation

Coefficient of determination · Tutorial 918 of 1000

Writing a Complete r-squared Interpretation

Practice writing complete r-squared interpretations that identify the response’s variation, name both variables, and describe their linear relationship without overclaiming.

Intermediate 10 min read

What You'll Learn

  • Use a sentence structure that identifies the response, explanatory variable, and linear relationship.
  • Convert a coefficient of determination to a percentage accurately.
  • Make clear which observed cases the interpretation describes.
  • Improve vague or incomplete r-squared statements for AP-style communication.
  • Avoid claims about individual predictions or causation that r-squared alone cannot support.

From an r-squared Value to a Complete Sentence

In “Common Errors Interpreting r-squared,” you reviewed what the coefficient of determination can and cannot say. Now the task is to communicate its meaning completely. A strong AP-style interpretation does more than translate a decimal into a percentage: it names the response, identifies the explanatory variable, and states that the percentage refers to variation accounted for by a linear relationship.

The sentence should also fit the study. If the data describe a sample of homes, for example, say “among the homes in this study” rather than making a broader claim about all homes. That small detail keeps the interpretation attached to the cases the data actually describe.

Interpretation frame: About [percentage]% of the variation in [response variable] among [the observed individuals or items] is accounted for by the linear relationship between [response variable] and [explanatory variable].

The frame is a guide, not a sentence that must be copied word for word. “Accounted for by the linear relationship between” and “accounted for by the linear relationship with” can both work if the variables are unmistakable. The important points are to describe variation in the response, name both variables, and use the word linear.

As covered in “Interpreting r-squared in Context,” \(r^2\) is a fraction. To state it as a percentage, multiply by 100. For example, if \(r^2=0.72\), then \(100(0.72)=72\%\). Use “about” when the reported value is rounded or when the percentage is a convenient summary. Do not attach response units to the percentage: \(r^2\) is unitless.

$$ \text{percentage of response variation accounted for}=100r^2\% $$

Before writing, identify the roles of the variables. The response variable is the outcome the regression predicts; its variation is the subject of the interpretation. The explanatory variable is the variable used to explain or predict the response. Naming both variables in context prevents an answer from quietly reversing their roles.

A Reliable Writing Process

Use this short process whenever a question asks you to interpret \(r^2\). It moves from the numerical value to a sentence that is specific enough to earn credit and cautious enough to stay within what the statistic says.

1
Identify the response.
Ask what variable is being predicted. That is the variable whose variation you will describe.
2
Convert to a percentage.
Multiply the reported \(r^2\) by 100. Keep the word “about” if the value is rounded.
3
Name the variables and cases.
Use the variables’ real context names, not just \(x\) and \(y\). State which group of observed cases the data represent when that is known.
4
Describe the linear relationship.
Say the response’s variation is accounted for by its linear relationship with the explanatory variable. Do not turn that description into a claim about causes or individual predictions.

The word “linear” matters because \(r^2\) in this setting summarizes variation accounted for by a least-squares linear model. Leaving it out makes the interpretation less precise. Likewise, naming only one variable is not enough: “72% of electricity use is explained” does not tell the reader whether you mean variation in electricity use, what it is related to, or what cases are being described.

Worked Examples: Writing the Interpretation

Worked Example: Name the Response and the Cases

A fictional study records each of 40 homes’ average daily electricity use, in kilowatt-hours, and the average outdoor temperature, in degrees Celsius, over the same month. The response is daily electricity use, the explanatory variable is outdoor temperature, and the regression output reports \(r^2=0.72\). Write a complete interpretation.

State. The requested interpretation concerns variation in daily electricity use among the 40 homes, not variation in temperature.

Plan. Convert \(r^2\) to a percentage. Then name the response, the observed cases, and the explanatory variable, and describe the association as linear. The units help identify the variables but do not become units for \(r^2\).

Do. The conversion is

$$ 100(0.72)=72\% $$

Daily electricity use is the response because it is being predicted from outdoor temperature. The cases are the 40 homes in the study. Thus the percentage refers to variation in their electricity use and to the linear relationship between electricity use and outdoor temperature.

Conclude. About 72% of the variation in average daily electricity use among the 40 homes in this study is accounted for by the linear relationship between average daily electricity use and average outdoor temperature. This statement describes an association in the observed homes; it does not say that temperature caused 72% of electricity use or that the model predicts each home’s use with 72% accuracy.

Worked Example: Repair an Incomplete Statement

A fictional environmental project records the distance from a trailhead to each of 30 monitoring sites, in kilometers, and the number of pieces of litter found in a fixed survey area at each site. The number of pieces of litter is the response, and distance is the explanatory variable. The reported coefficient of determination is \(r^2=0.37\). A student writes, “The model explains 37% of the results.” Revise the statement so it is complete.

State. The student has converted the statistic correctly in words, but “the results” is too vague to identify the response. The statement also omits the explanatory variable and the linear relationship.

Plan. Calculate the percentage, then replace “results” with the response name. Specify the observed sites and connect litter counts with distance using the phrase “linear relationship.”

Do. The percentage is

$$ 100(0.37)=37\% $$

The response is the number of pieces of litter found in the survey area. The observations are the 30 monitoring sites. Distance from the trailhead is the explanatory variable. A complete statement must say that the percentage is about variation in litter counts, not the percentage of sites, pieces of litter, or predictions.

Conclude. About 37% of the variation in the number of pieces of litter found at the 30 monitoring sites is accounted for by the linear relationship between litter count and distance from the trailhead. This identifies what varies, which cases are summarized, and the two variables in the relationship.

Worked Example: Interpret a Rounded Computer Output

A fictional school project examines the relationship between students’ weekly minutes of independent reading and their scores on a vocabulary assessment, measured in points. Vocabulary score is the response, reading minutes are the explanatory variable, and the computer output reports \(R\)-sq \(=0.806\). The data include 52 students. Write an AP-style interpretation using an appropriate level of precision.

State. \(R\)-sq in this simple linear regression output is the reported coefficient of determination. The response is vocabulary-assessment score, and the observed cases are the 52 students.

Plan. Convert the reported value to a percentage and round sensibly. Then use the context to describe variation in assessment scores accounted for by the linear relationship with weekly reading time. Do not interpret the percentage as a score increase or as a prediction success rate.

Do. Convert the displayed decimal:

$$ 100(0.806)=80.6\% $$

The output is rounded to three decimal places, so reporting “about 81%” is a clear summary. The outcome being described is variation in vocabulary-assessment scores among these students. Weekly reading minutes is the explanatory variable, so include it as the other variable in the linear relationship.

Conclude. About 81% of the variation in vocabulary-assessment scores among the 52 students in this project is accounted for by the linear relationship between vocabulary-assessment score and weekly minutes of independent reading. This does not mean that 81% of students received accurate predictions, nor does it establish that reading time caused the scores to differ.

Precision, Scope, and Common Mistakes

A complete interpretation is not necessarily a long one. It is a sentence in which each phrase does a clear job: the percentage gives the size of the fraction, the response identifies the variation, the explanatory variable completes the relationship, and the case description sets the scope. The examples show that a concise answer can still include all of these details.

  • Writing only “\(r^2\) is 72%.” This gives a percentage but no interpretation. State what variation the percentage describes and name the linear relationship.
  • Describing the wrong variable’s variation. The response is the focus of an \(r^2\) interpretation. Use the explanatory variable as the other variable in the relationship, not as the variable whose variation is being accounted for.
  • Leaving out “linear.” The word signals that the percentage refers to variation accounted for by the linear relationship. Include it in the final statement.
  • Using broad or undefined labels. “Results,” “data,” and “the outcome” may be unclear on their own. Substitute the actual response name, such as “vocabulary-assessment scores.”
  • Turning the percentage into a count or prediction rate. About 37% of variation does not mean 37% of sites, observations, or predictions. As discussed in “r-squared Is Not Percent Correct Predictions,” \(r^2\) is not a tally of successful predictions.
  • Overstating who the result describes. If the data are from a particular group of observed cases, keep the sentence anchored to that group unless the study’s sampling design supports broader generalization.
  • Using causal wording without support. “Accounted for by” describes the relationship in the data. It is not a substitute for evidence from a design that supports cause-and-effect conclusions.

For full-credit communication, an AP response should not force a reader to guess what “variation” refers to or which two variables are related. If the study names a particular group, include that group naturally: “among the homes in this study,” “for these monitoring sites,” or “among the students observed.” Avoid adding claims about the slope, the direction of the association, or causation unless the question asks for them and the available information supports them. As covered in “r-squared and Strength of Association,” \(r^2\) alone does not give the direction of the relationship.

A final check is to read the interpretation without looking at the regression output. Could someone identify what is varying, the cases being described, and the two variables in the linear relationship? If not, add the missing context. This check is especially useful when an exam prompt uses generic labels such as \(x\) and \(y\): translate those labels into the variables’ names before writing your final sentence.

Key takeaway: A complete interpretation turns \(r^2\) into a percentage of variation in the named response, identifies the observed cases, and connects that response to the named explanatory variable through a linear relationship.

Check Your Understanding

For each item, focus on the response’s variation and write or assess the interpretation in context.

  1. A study of 24 urban parks relates the number of visitors per day to the number of available parking spaces. Visitors per day is the response and \(r^2=0.58\). Write a complete interpretation.
  2. A fictional analysis relates weekly time spent practicing a musical instrument to a performance-rating score. The score is the response and \(r^2=0.43\). A student says, “43% of practice time is explained.” Identify the error and state which variable’s variation the interpretation should describe.
  3. A data set contains 35 delivery orders. Delivery time in minutes is the response, distance in kilometers is the explanatory variable, and \(R\)-sq is 0.913. Write an interpretation using sensible rounding.
  4. Why is “About 60% of the predictions were correct” not an appropriate interpretation of an \(r^2\) value of 0.60?
  5. A report gives an \(r^2\) value for a regression of monthly water use on household size. Name the response and explanatory variable, and identify which variable’s variation belongs in the interpretation.