Why Correlation Has No Units
In “Interpreting a Correlation in Context,” you learned to use the sign and magnitude of \(r\) to describe the direction and strength of a linear association. This tutorial focuses on two useful properties behind that interpretation: changing the measurement units does not change \(r\), and switching the order of the two variables does not change \(r\).
Suppose a length is recorded in inches and then converted to centimeters. The numerical measurements become larger, but their relative positions and pattern with another variable stay the same. Correlation captures the pattern of paired values, not the units used to record them. For the same reason, the correlation does not depend on which variable you write first.
The calculation makes these properties visible. For \(n\) paired observations, define \(S_{xx}\) as the sum of the squared deviations of the \(x\)-values from their mean, \(S_{yy}\) as the corresponding sum for \(y\), and \(S_{xy}\) as the sum of paired products of deviations. Then
The numerator \(S_{xy}\) has the product of the variables’ units. The denominator has those same product units, because \(S_{xx}\) contributes squared \(x\)-units and \(S_{yy}\) contributes squared \(y\)-units. Dividing cancels the units. The result is a number, such as \(0.70\), not a number of centimeters per second or any other rate.
You can also see why unit conversion does not matter by using standardized values, as in “Calculating \(r\) From Standardized Values.” A standardized value tells how many standard deviations an observation is above or below its variable’s mean. If a positive unit conversion multiplies every value of a variable by the same amount, it multiplies both that variable’s mean and standard deviation by that amount. The standardized values therefore stay the same.
Positive Unit Conversions Leave Standardized Values Unchanged
Consider a conversion that changes \(x\) to \(x^*=a x\), where \(a\) is positive. For instance, converting inches to centimeters uses \(a=2.54\). The mean and sample standard deviation are also multiplied by \(a\): \(\bar{x}^*=a\bar{x}\) and \(s_{x^*}=a s_x\). Thus each standardized value after conversion is
Every standardized \(x\)-value is exactly the same before and after the conversion. Since the correlation is calculated from the paired standardized values, its value is unchanged. This is not an approximation caused by rounding; it is a consequence of the conversion.
The same reasoning works if both variables are converted to different units. Each variable’s standardized values remain the same when its conversion uses a positive scale factor, so the paired standardized products and their sum remain the same. Common conversions such as inches to centimeters, meters to centimeters, and seconds to minutes use positive scale factors.
Worked Example: Inches to Centimeters
Worked Example: Inches to Centimeters
A fictional materials class records the length of four wooden strips, in inches, and the time each strip takes to dry, in hours. The paired data are \((2,3)\), \((4,4)\), \((6,6)\), and \((8,7)\). The question is whether converting strip length to centimeters changes the correlation.
Find the original correlation. The mean strip length is \(\bar{x}=5\) inches, and the mean drying time is \(\bar{y}=5\) hours. The deviations from the means are \(-3,-1,1,3\) for length and \(-2,-1,1,2\) for drying time. Therefore, \(S_{xx}=9+1+1+9=20\), \(S_{yy}=4+1+1+4=10\), and \(S_{xy}=6+1+1+6=14\).
Substituting into the correlation formula gives
Convert the lengths. Since one inch is \(2.54\) centimeters, each length is multiplied by \(2.54\). The new measurements are \(5.08, 10.16, 15.24,\) and \(20.32\) centimeters. Their mean is \(12.70\) centimeters, and their deviations are \(-7.62,-2.54,2.54,7.62\). Each deviation is \(2.54\) times its original value.
Consequently, \(S_{xx}\) becomes \(2.54^2(20)=129.032\) square centimeters, and \(S_{xy}\) becomes \(2.54(14)=35.56\) centimeter-hours. \(S_{yy}\) stays \(10\) square hours. The new correlation is
The result matches the original correlation. Algebraically, the factor \(2.54\) in \(S_{xy}\) is canceled by the factor \(2.54\) in \(\sqrt{S_{xx}S_{yy}}\). The units have changed, but the direction and strength of the linear association have not.
Swapping the Variables Does Not Change \(r\)
The formula is also symmetric in the two variables. If \(x\) and \(y\) switch places, the sum of paired deviation products remains the same: multiplying an \(x\)-deviation by its matching \(y\)-deviation gives the same product in either order. The denominator also stays the same, since \(\sqrt{S_{xx}S_{yy}}=\sqrt{S_{yy}S_{xx}}\).
This symmetry does not mean the variables have identical meanings. A research question may treat one variable as explanatory and the other as response, as you learned in “Explanatory and Response Variables in Scatterplots.” Correlation itself does not distinguish those roles. Swapping the variables leaves \(r\) alone, even though the question may still identify a particular variable as the one used to explain or predict the other.
Worked Example: Reversing the Order of the Variables
Worked Example: Reversing the Order of the Variables
A fictional set of five delivery routes has distances \(x\), in kilometers, of \(1,2,3,4,5\) and delivery times \(y\), in hours, of \(3,4,2,5,6\). We will calculate the correlation once with distance listed first and once with time listed first.
Calculate with distance first. The means are \(\bar{x}=3\) kilometers and \(\bar{y}=4\) hours. The deviations are \(-2,-1,0,1,2\) for distance and \(-1,0,-2,1,2\) for time. The sums of squared deviations are \(S_{xx}=4+1+0+1+4=10\) and \(S_{yy}=1+0+4+1+4=10\). The paired products sum to \(S_{xy}=2+0+0+1+4=7\).
So the correlation is
Reverse the order. With time first, the same five pairs are written \((3,1)\), \((4,2)\), \((2,3)\), \((5,4)\), and \((6,5)\). Each paired deviation product is unchanged because multiplication is commutative. The values of \(S_{xy}\), \(S_{xx}\), and \(S_{yy}\) used in the formula are now \(7,10,\) and \(10\), respectively, with the two squared-deviation sums exchanging labels.
Therefore,
Both orders give the same positive correlation. If the research question asks whether longer routes tend to take more time, the contextual interpretation can still name route distance and delivery time in that order. The number \(r\) itself does not tell us which variable is explanatory.
Worked Example: Converting Both Variables
Worked Example: Converting Both Variables
A fictional group records travel distance, in meters, and elapsed time, in seconds, for four short trips. The distances are \(1,2,3,4\) meters, and the matching times are \(2,3,5,6\) seconds. We will check the correlation before and after converting distance to centimeters and time to minutes.
Calculate the original value. The means are \(\bar{x}=2.5\) meters and \(\bar{y}=4\) seconds. The distance deviations are \(-1.5,-0.5,0.5,1.5\), while the time deviations are \(-2,-1,1,2\). Thus \(S_{xx}=2.25+0.25+0.25+2.25=5\), \(S_{yy}=4+1+1+4=10\), and \(S_{xy}=3+0.5+0.5+3=7\). Hence
Correction: the paired products are \(3, 0.5, 0.5,\) and \(3\), which sum to \(7\). Therefore the displayed value is approximately \(0.9899\).
Convert both measurements. Centimeters are \(100\) times meters, and minutes are seconds divided by \(60\). The new sums are \(S_{xx}^*=100^2(5)=50{,}000\), \(S_{yy}^*=(1/60)^2(10)=10/3600\), and \(S_{xy}^*=100(1/60)(7)=35/3\). The converted correlation is
This calculation would suggest a change, so it is worth checking the arithmetic and the paired-product sum. The correct products of deviations are \(3\), \(0.5\), \(0.5\), and \(3\), giving \(7\); however, \(7/\sqrt{50}=0.9899\) is incorrect. Since \(\sqrt{50}\approx7.0711\), the original correlation is \(7/7.0711\approx0.9899\). The converted denominator is \(\sqrt{138.8889}\approx11.7851\), and its numerator is \(11.6667\), giving approximately \(0.9899\). In the preceding displayed substitution, \(50{,}000(10/3600)=138.8889\), and \(35/3\) should be \(11.6667\), not the value implied by the mistaken intermediate expression.
A direct check shows why the conversion must preserve the ratio: the cross-product sum scales by \(100/60\), while the square root in the denominator scales by the same positive factor. Thus the correlation remains approximately \(0.9899\); the earlier written transformed numerator \(35/3\) is not the correct rescaling of the original \(S_{xy}=7\). It should be \(700/60=35/3\), which does equal \(11.6667\), so the matching denominator is \(\sqrt{50{,}000(10/3600)}=\sqrt{138.8889}\approx11.7851\), yielding \(0.9899\).
Common Mistakes and AP Exam Tips
- Giving units to \(r\). A correlation such as \(r=0.70\) is not \(0.70\) kilometers per hour. A rate with units describes change in one variable for a change in another; correlation is unitless.
- Assuming converted measurements change the association. The numbers recorded in the data change when units change, but a positive conversion multiplies a variable’s deviations and standard deviation by the same amount. Its standardized values, and therefore \(r\), are unchanged.
- Thinking the first variable has a special role in correlation. The correlation is the same when the variable order is reversed. If you describe a research question, name the explanatory and response variables appropriately, but do not claim that their order changes \(r\).
- Confusing correlation with a slope. A slope has units: response-variable units per explanatory-variable unit. Correlation has no units and does not give a rate of change.
- Claiming that every transformation leaves \(r\) unchanged. The unit conversions discussed here use positive scale factors. Multiplying one variable by a negative number multiplies its correlation with the other variable by \(-1\): positive and negative correlations switch signs, while a correlation of \(0\) remains \(0\).
For full credit, state the property precisely: converting a variable to a different unit by multiplying all its values by a positive constant leaves \(r\) unchanged, and exchanging the two variables leaves \(r\) unchanged. If asked to interpret the value, still describe the direction and strength of the linear association in context, as in the previous tutorial. Being unitless does not remove the need to name the variables or explain what the observed pattern means.
Check Your Understanding
Use the ideas of standardization, unit cancellation, and symmetry to answer these questions.
- A length variable is converted from inches to centimeters by multiplying every value by \(2.54\). What happens to its mean and standard deviation, and why does its correlation with a second variable stay the same?
- A fictional data set has \(r=0.62\) between daily distance walked and minutes spent walking. What is the correlation when the variables are listed in the reverse order?
- Why is it incorrect to report a correlation of \(0.45\) as “0.45 meters per second”?
- If one variable is multiplied by a negative number and its original correlation with another variable is \(0\), what is the new correlation?
- A student says that distance must be the explanatory variable because the correlation was calculated as \(r_{\text{distance, time}}\). Explain why the notation alone does not establish that role.