From Paired Values to \(r\)
In “What the Correlation Coefficient Measures” and “Properties of \(r\): Range and Sign,” you learned that \(r\) summarizes the direction and strength of the linear association between two quantitative variables. In “Matching Correlations to Scatterplots,” you used the sign and magnitude of \(r\) to describe a pattern. Now calculate \(r\) directly from a small set of paired observations.
The calculation uses standardized values, or z-scores. A z-score tells how many sample standard deviations an observation is above or below its variable’s mean. For each individual, multiply the x-variable z-score by the matching y-variable z-score. The pattern of these products helps determine the correlation: pairs that are high together or low together contribute positive products, while pairs high on one variable and low on the other contribute negative products.
The denominator is \(n-1\), not \(n\). The z-scores use sample standard deviations, which are calculated with \(n-1\) in the denominator. Accordingly, the correlation formula uses \(n-1\) when combining those sample z-scores. In this tutorial, “average product” means this correlation calculation: sum the paired products and divide by \(n-1\). Do not divide the sum by the number of observations.
A Step-by-Step Calculation
Keep the original pairs together throughout the work. First find \(\bar{x}\), \(s_x\), \(\bar{y}\), and \(s_y\). Then calculate \(z_x=(x-\bar{x})/s_x\) and \(z_y=(y-\bar{y})/s_y\) for every pair. Finally, multiply the two z-scores from each individual, add those products, and divide by \(n-1\).
Calculate \(\bar{x}\) and \(s_x\) for the explanatory-variable values, and \(\bar{y}\) and \(s_y\) for the response-variable values.
For each observation, subtract its variable’s mean and divide by that variable’s sample standard deviation.
For each individual, calculate \(z_xz_y\). Do not pair an x-value with a y-value from a different individual.
Sum the products and divide by \(n-1\). With five pairs, divide by \(4\).
An equivalent calculation can help check the work. If you already have the deviations from the means, add their paired products, then divide by \((n-1)s_xs_y\). This gives the same correlation as the z-score method. The worked examples use both forms as a check, while keeping the z-score products at the center of the calculation.
Worked Example: A Positive Linear Association
Worked Example: A Positive Linear Association
A class invents five paired observations comparing minutes spent reviewing a topic with a quiz score. The values are illustrative, not results from a real study.
| Observation | Review time \(x\), minutes | Quiz score \(y\), points |
|---|---|---|
| 1 | 1 | 8 |
| 2 | 2 | 9 |
| 3 | 3 | 11 |
| 4 | 4 | 10 |
| 5 | 5 | 12 |
Step 1: Find the means and sample standard deviations. The means are \(\bar{x}=15/5=3\) minutes and \(\bar{y}=50/5=10\) points. The x-deviations are \(-2,-1,0,1,2\), so their squared deviations sum to \(10\), and \(s_x=\sqrt{10/4}=\sqrt{2.5}\approx1.5811\) minutes. The y-deviations are \(-2,-1,1,0,2\), whose squares also sum to \(10\), so \(s_y=\sqrt{10/4}=\sqrt{2.5}\approx1.5811\) points.
Step 2: Standardize and multiply each pair. Both standard deviations equal \(\sqrt{2.5}\). The table shows z-scores and their paired products, rounded to four decimal places.
| Observation | \(z_x\) | \(z_y\) | \(z_xz_y\) |
|---|---|---|---|
| 1 | \(-1.2649\) | \(-1.2649\) | \(1.6000\) |
| 2 | \(-0.6325\) | \(-0.6325\) | \(0.4000\) |
| 3 | \(0\) | \(0.6325\) | \(0\) |
| 4 | \(0.6325\) | \(0\) | \(0\) |
| 5 | \(1.2649\) | \(1.2649\) | \(1.6000\) |
Step 3: Sum the products and divide by \(n-1\). Using the unrounded z-scores, the products sum to \(3.6\). There are five observations, so \(n-1=4\).
The equivalent deviation calculation checks the result. The paired deviation products sum to \(4+1+0+0+4=9\), and \((n-1)s_xs_y=4(\sqrt{2.5})(\sqrt{2.5})=10\). Thus \(r=9/10=0.9\), matching the z-score method. The positive value indicates a positive linear association in these five paired observations.
Worked Example: A Negative Linear Association
Worked Example: A Negative Linear Association
A fictional workshop records five participants’ practice scores \(x\) and time to finish a task \(y\), in minutes. The values are chosen to make the z-score products easy to inspect.
| Participant | Practice score \(x\) | Task time \(y\), minutes |
|---|---|---|
| 1 | 1 | 22 |
| 2 | 2 | 20 |
| 3 | 3 | 19 |
| 4 | 4 | 21 |
| 5 | 5 | 18 |
Find the means and standard deviations. Here \(\bar{x}=3\) and \(\bar{y}=20\). The x-deviations \((-2,-1,0,1,2)\) have squared sum \(10\), so \(s_x=\sqrt{10/4}=\sqrt{2.5}\approx1.5811\). The y-deviations are \((2,0,-1,1,-2)\), with squared sum \(4+0+1+1+4=10\), so \(s_y=\sqrt{10/4}=\sqrt{2.5}\approx1.5811\) minutes.
Standardize, multiply, and calculate. The x z-scores are approximately \((-1.2649,-0.6325,0,0.6325,1.2649)\). The y z-scores are approximately \((1.2649,0,-0.6325,0.6325,-1.2649)\). Their corresponding products, using unrounded z-scores, are \(-1.6,0,0,0.4,-1.6\).
Check with the deviations: their paired products sum to \((-2)(2)+(-1)(0)+(0)(-1)+(1)(1)+(2)(-2)=-7\). The denominator is \(4(\sqrt{2.5})(\sqrt{2.5})=10\), so \(r=-7/10=-0.7\). This confirms the standardized calculation. The negative correlation indicates that higher practice scores tend to occur with shorter task times in this small set of observations.
Worked Example: Different Spreads, Same Calculation
Worked Example: Different Spreads, Same Calculation
A fictional environmental class records five days of a sensor’s operating time \(x\), in hours, and a related measurement \(y\), in units. The two variables have different spreads, so each must be standardized with its own sample standard deviation.
| Day | Operating time \(x\), hours | Measurement \(y\), units |
|---|---|---|
| 1 | 2 | 3 |
| 2 | 4 | 5 |
| 3 | 6 | 4 |
| 4 | 8 | 8 |
| 5 | 10 | 10 |
Find the summaries. The means are \(\bar{x}=6\) hours and \(\bar{y}=6\) units. The x-deviations \((-4,-2,0,2,4)\) have squared sum \(40\), giving \(s_x=\sqrt{40/4}=\sqrt{10}\approx3.1623\) hours. The y-deviations \((-3,-1,-2,2,4)\) have squared sum \(34\), giving \(s_y=\sqrt{34/4}=\sqrt{8.5}\approx2.9155\) units.
Calculate the paired z-score products. The x z-scores are approximately \((-1.2649,-0.6325,0,0.6325,1.2649)\), and the y z-scores are approximately \((-1.0290,-0.3430,-0.6860,0.6860,1.3720)\). Using unrounded values, the paired products are approximately \(1.3014,0.2170,0,0.4340,1.7354\). Their sum, calculated before rounding, is \(34/\sqrt{85}\approx3.6878\).
For a second check, the paired deviation products sum to \((-4)(-3)+(-2)(-1)+(0)(-2)+(2)(2)+(4)(4)=34\). The denominator is \(4(\sqrt{10})(\sqrt{8.5})=4\sqrt{85}\). Therefore \(r=34/(4\sqrt{85})\approx0.9220\), the same result. Because the variables have different spreads, their z-scores are not identical; standardizing each with its own \(s\) accounts for that difference.
What the Products Tell You
A positive product occurs when the two z-scores have the same sign: both observations are above their respective means, or both are below. A negative product occurs when one z-score is positive and the other is negative. A product of zero occurs when at least one observation equals its variable’s mean. These contributions combine to give the overall direction and strength of the linear association.
The size of an individual product matters too. A pair that is far from both means in the same direction can make a relatively large positive contribution; a pair far from the means in opposite directions can make a relatively large negative contribution. This is one reason to keep the paired observations visible and to inspect a scatterplot as well. As discussed in “Spotting Outliers in Bivariate Data,” unusual points can affect a correlation.
The calculation does not make \(r\) a measure of slope, and it does not make correlation an appropriate summary of a curved pattern. As in “Matching Correlations to Scatterplots,” check the form of the relationship and interpret \(r\) as a summary of linear association. For a valid calculation, the values of each variable must have nonzero sample standard deviations; if one variable has no variation, its z-scores and the correlation are undefined.
Common Mistakes and AP Exam Tips
- Dividing by \(n\). With five pairs, dividing by \(5\) instead of \(4\) gives the wrong result. Use \(n-1\) in the z-score correlation formula.
- Pairing values from different observations. Each \(z_x\) must be multiplied by the \(z_y\) belonging to the same individual or case. Reordering only one variable changes the pairings and can change \(r\).
- Using one standard deviation for both variables. Calculate \(s_x\) from the x-values and \(s_y\) from the y-values. The two variables can have different units and different spreads.
- Rounding too early. Keep full calculator precision for means, standard deviations, z-scores, and products until the final value. Rounded table entries may not add exactly to the displayed total.
- Forgetting the sign. Negative z-score products reduce the sum. Preserve their signs rather than adding their absolute values.
- Reporting more than \(r\) says. A correlation describes direction and strength of linear association. It is not a slope, a causal effect, or a description of a curved relationship.
For full-credit work, show how each variable was standardized, keep the matched products identifiable, state that the five products are divided by \(4\), and report \(r\) with suitable rounding. A contextual interpretation should name the variables and describe the direction and strength of their linear association without claiming causation.
Check Your Understanding
Use \(r=\frac{\sum z_xz_y}{n-1}\) and keep each pair matched.
- For five paired observations, what number divides the sum of the z-score products?
- If a paired observation has \(z_x=1.2\) and \(z_y=0.5\), what is its product? What does its positive sign indicate about the two values relative to their means?
- If a paired observation has \(z_x=-0.8\) and \(z_y=1.1\), what is its product, and how does it contribute to the sum?
- Why must \(x\)-values and \(y\)-values be standardized using their own sample standard deviations?
- A student calculates a correlation of \(1.14\). What should the student check about the formula or arithmetic?