What \(r\) Tells You—and What It Does Not
In “Comparing Strength Using Correlations,” you used the absolute value of \(r\) to compare the strength of linear associations and the sign to compare their directions. This tutorial focuses on a related skill: keeping the meaning of \(r\) within its limits. A correlation is easy to overstate because its value is a number, but it is not a percentage, a slope, or a measure of cause and effect.
As established in “What the Correlation Coefficient Measures,” \(r\) summarizes the direction and strength of the linear association between two quantitative variables. It is unitless and ranges from \(-1\) to \(1\). A positive value indicates a positive linear association; a negative value indicates a negative linear association. The magnitude describes how closely the data follow a straight-line pattern. Those features do not make \(r\) a measure of how much one variable changes when the other changes.
A careful interpretation names both variables and describes their linear association in context. For example, if \(r=-0.68\) for daily outdoor temperature and building heating demand, the data show a moderately strong negative linear association: days with higher temperatures tend to have lower heating demand. The correlation alone does not say that heating demand decreases by 68%, or by any fixed amount, for a one-unit increase in temperature.
Misinterpretation 1: Reading \(r\) as a Percent
The decimal form of a correlation can look like a percentage. But \(r=0.80\) does not mean “80%,” “80% of observations,” or “an 80% increase.” Likewise, \(r=-0.80\) does not mean an 80% decrease. The sign and magnitude summarize the direction and strength of a linear association; they do not describe a percentage change.
A different statistic, \(r^2\), is often expressed as a percentage when interpreted as the proportion of variation in the response variable accounted for by the linear model. For example, if \(r=-0.80\), then \(r^2=0.64\), or 64%. This does not mean that 64% of the response values changed, that 64% of individuals follow the pattern, or that one variable caused 64% of an outcome. As in “Finding \(r\) With LinReg on the Calculator,” keep \(r\) and \(r^2\) distinct.
Worked Example: Is \(r=-0.68\) a 68% Decrease?
A fictional environmental data set records outdoor temperature and the heating demand of several buildings. The correlation between temperature and heating demand is \(r=-0.68\). A student writes, “A one-degree increase in temperature causes heating demand to decrease by 68%.” Is this interpretation correct?
No. The value \(-0.68\) is the correlation, not a percentage change and not a rate of change. Its negative sign indicates a negative linear association: in these observations, higher outdoor temperatures tend to go with lower heating demand. Its magnitude indicates the strength of that linear association.
If a student wants to report the coefficient of determination, then \(r^2=(-0.68)^2=0.4624\). In this example, about 46.2% of the variation in heating demand among the buildings in the data is accounted for by the linear relationship with outdoor temperature. That statement concerns variation accounted for by the linear model; it is not a statement that heating demand falls by 46.2%, and it does not establish that temperature caused the differences.
A more accurate interpretation of \(r\) is: “There is a moderately strong negative linear association between outdoor temperature and building heating demand in these data; buildings observed at higher temperatures tend to have lower heating demand.” This describes the correlation without turning it into a percentage change or causal claim.
Misinterpretation 2: Reading \(r\) as a Slope
The correlation and the slope of a regression line are related, but they are not the same quantity. Correlation is unitless. A slope has units: response-variable units per explanatory-variable unit. For instance, a slope might describe the predicted change in heating demand, measured in kilowatt-hours, for each one-degree increase in temperature.
In a least-squares regression line with an intercept, the slope \(b\) can be calculated from \(r\) and the sample standard deviations:
Here, \(s_x\) is the sample standard deviation of the explanatory-variable values and \(s_y\) is the sample standard deviation of the response-variable values. Because the ratio \(s_y/s_x\) depends on the data’s scales, a correlation alone does not determine the numerical slope. Two data sets can have the same \(r\) but different slopes.
Worked Example: Same Correlation, Different Slopes
Two fictional data sets relate outdoor temperature \(x\), in degrees Celsius, to building heating demand \(y\), in kilowatt-hours. Both have \(r=-0.60\) and \(s_y=120\) kilowatt-hours. In Set A, \(s_x=5\) degrees Celsius; in Set B, \(s_x=15\) degrees Celsius. Find the slope for each set and explain why \(r\) is not the slope.
For Set A, substitute the given values into \(b=r(s_y/s_x)\):
For Set B:
Both sets have the same correlation, so their linear associations have the same signed correlation summary. But their regression slopes differ: \(-14.4\) and \(-4.8\) kilowatt-hours per degree Celsius. In the fitted lines, the slope describes the predicted change in heating demand for a one-degree Celsius increase in temperature. The correlation does not; it describes direction and strength and has no units.
A common error would be to say that \(r=-0.60\) means heating demand decreases by 0.60 kilowatt-hours per degree Celsius. That gives \(r\) the wrong meaning and the wrong units. A correct response keeps the summaries separate: “The correlation is \(-0.60\), indicating a negative linear association. The regression slope, calculated using the standard deviations, is \(-14.4\) kilowatt-hours per degree Celsius in Set A.”
Misinterpretation 3: Reading \(r\) as a Causal Effect
A correlation describes how two variables vary together in observed data. A causal effect is a stronger claim: it says that changing one variable would produce a change in another. As emphasized in “Why Correlation Does Not Imply Causation,” an association—even a strong one—does not, by itself, establish cause and effect.
One reason is that a third variable may be related to both variables being studied. In “Lurking Variables and Confounding,” you learned to consider how such a variable could contribute to an observed association. Another reason is that the relationship may partly run in the opposite direction from the one claimed. A correlation does not tell you which variable, if either, is causing the other to change.
Worked Example: Does a Correlation Show That Screen Time Reduces Sleep?
A fictional observational data set records evening screen time and hours of sleep for a group of teenagers. The correlation is \(r=-0.74\). A report claims, “Increasing screen time causes teenagers to sleep less.” Does the correlation alone support that conclusion?
No. The negative value indicates that higher evening screen time tends to be associated with fewer hours of sleep in these observations, and the magnitude indicates a fairly strong linear association. But these are observed data, not evidence from an experiment that randomly assigned different amounts of screen time. The correlation alone cannot establish that increasing screen time caused the decrease in sleep.
For example, a student’s workload or schedule might be related to both evening screen use and sleep duration. That would be a possible lurking variable to consider, though the correlation does not tell us whether it actually explains the pattern. Other explanations might also be possible. To claim a causal effect, we would need an appropriate study design and evidence beyond the correlation itself.
A careful conclusion stays with the evidence: “In this group, evening screen time and hours of sleep have a fairly strong negative linear association: teenagers with more evening screen time tend to report fewer hours of sleep. These data alone do not show that screen time causes less sleep.” This accurately describes the observed pattern and makes its causal limit explicit.
A Quick Audit Before You Interpret \(r\)
When a correlation appears in a question, check what the number can support before writing a conclusion. The following sequence helps catch the three misinterpretations in this tutorial.
Identify the two quantitative variables and use the sign of \(r\) to describe a positive or negative linear association.
Use the magnitude of \(r\) to describe how closely the observations follow a straight-line pattern. Do not turn the value into a percent change.
If a claim describes how much \(y\) changes for a one-unit increase in \(x\), it is about a slope, which has units—not about \(r\).
A correlation describes association. It does not establish that changing one variable causes a change in the other.
Common Mistakes and AP Exam Tips
- Converting \(r\) into a percent change. “\(r=0.70\) means a 70% increase” is incorrect. State that the value indicates a positive linear association and use its magnitude to describe linear strength.
- Calling \(r\) the slope. A slope has response units per explanatory-variable unit; \(r\) has no units. If the question asks for predicted change per unit of \(x\), use the slope, not the correlation.
- Confusing \(r\) with \(r^2\). The correlation summarizes signed direction and linear strength. The squared correlation is nonnegative and can be used to describe the proportion of response variation accounted for by the linear model. Neither is the percentage of individuals affected.
- Using causal language for an observed association. “\(x\) causes \(y\)” goes beyond what a correlation alone establishes. Say that the variables are associated, and refer to the study design when evaluating a causal claim.
- Leaving out the context. “There is a negative correlation” is less informative than naming both variables and explaining what a negative linear association means in the situation.
For full-credit communication, describe the direction and strength of the linear association between the named variables, in context. Do not state a percentage change, give \(r\) units, or claim a causal effect unless the study design and evidence support that claim. Also remember, as in “Correlation Measures Only Linear Association,” that a correlation summarizes linear pattern; it does not describe every possible feature of a scatterplot.
Check Your Understanding
For each item, distinguish what the correlation supports from what it does not establish.
- A study reports \(r=0.62\) between weekly outdoor exercise time and a fitness score. What does the sign indicate, and why is it incorrect to call this a 62% increase?
- A student says \(r=-0.50\) means that the response decreases by 0.50 units for each one-unit increase in the explanatory variable. What statistic and units would be needed to describe that rate of change?
- If \(r=-0.70\), calculate \(r^2\). What kind of variation can that squared value describe, and what does it not mean?
- A fictional observational study finds a strong positive correlation between a neighborhood’s number of trees and residents’ reported outdoor activity. Why does this alone not prove that adding trees causes activity to increase?
- Write one careful, context-based sentence interpreting \(r=-0.55\) for two quantitative variables of your choice, without using percentage, slope, or causal language.