From a Scatterplot to One Number
In “Scatterplot Description Synthesis Review,” you practiced describing direction, form, strength, and unusual features by examining a scatterplot. The correlation coefficient, written \(r\), provides a numerical summary of two of those features when the overall pattern is roughly linear: its direction and strength.
A number is convenient, but it cannot replace the scatterplot. The same value of \(r\) can accompany different patterns, and a curved pattern can have a correlation near zero even when the variables are clearly associated. Use \(r\) to summarize a linear association—not as a complete description of every feature in the data.
How \(r\) Summarizes a Linear Pattern
The sign of \(r\) corresponds to the direction of a linear association. A positive \(r\) goes with an upward trend from left to right: larger values of the explanatory variable tend to occur with larger values of the response variable. A negative \(r\) goes with a downward trend: larger explanatory-variable values tend to occur with smaller response-variable values.
The size of \(r\), considered by its distance from zero, describes how closely the points follow a straight-line pattern. Values nearer to \(1\) or \(-1\) indicate a closer linear pattern; values nearer to \(0\) indicate a weaker linear pattern. There is no single cutoff that divides every context into “weak,” “moderate,” and “strong.” Look at the scatterplot and consider the setting when describing strength.
Correlation is unitless. It is not measured in the units of either variable, and it is not a percentage. In particular, \(r=0.8\) does not mean that the response increases by 0.8 units for every one-unit increase in the explanatory variable. That kind of change-per-unit description belongs to the slope of a regression line, not to correlation.
Where the Formula Comes From
The formula makes the idea behind \(r\) more precise. For each observation, compare its \(x\)-value with the mean of the \(x\)-values and its \(y\)-value with the mean of the \(y\)-values. If the two deviations tend to have the same sign, the points tend to rise together. If they tend to have opposite signs, the points tend to move in opposite directions. The formula combines these comparisons and adjusts for the spread of each variable.
Here, \(n\) is the number of paired observations, \(x_i\) and \(y_i\) are the two values for observation \(i\), \(\bar{x}\) and \(\bar{y}\) are the sample means, and \(s_x\) and \(s_y\) are the sample standard deviations. Each fraction is a standardized deviation: it tells how far a value is from its mean in standard-deviation units. The product is positive when both values are above their means or both are below; it is negative when one is above its mean and the other is below.
For hand calculations, an equivalent formula is often simpler. It uses the sums of the cross-products of deviations and the squared deviations. The worked examples use this version; both formulas give the same \(r\).
In practice, technology can calculate \(r\) from the paired data. On a TI-84, enter the \(x\)-values and matching \(y\)-values in two lists, then use LinReg to display the correlation coefficient if diagnostics are enabled. Before interpreting the result, make a scatterplot and check that the data show a roughly linear pattern without an influential outlier.
Worked Example: Practice Sessions and Quiz Scores
Worked Example: Practice Sessions and Quiz Scores
A fictional class records the number of practice sessions completed by five students and their scores, in points, on a short quiz. Each row represents one student.
| Practice sessions, \(x\) | Quiz score, \(y\) |
|---|---|
| 1 | 2 |
| 2 | 3 |
| 3 | 5 |
| 4 | 4 |
| 5 | 6 |
Question: Calculate \(r\), then interpret it in context.
First, the means are \(\bar{x}=(1+2+3+4+5)/5=3\) sessions and \(\bar{y}=(2+3+5+4+6)/5=4\) points. The deviations and products are:
| \(x_i-\bar{x}\) | \(y_i-\bar{y}\) | Product | Squared \(x\)-deviation | Squared \(y\)-deviation |
|---|---|---|---|---|
| \(-2\) | \(-2\) | 4 | 4 | 4 |
| \(-1\) | \(-1\) | 1 | 1 | 1 |
| 0 | 1 | 0 | 0 | 1 |
| 1 | 0 | 0 | 1 | 0 |
| 2 | 2 | 4 | 4 | 4 |
The cross-products sum to \(9\); the squared \(x\)-deviations sum to \(10\); and the squared \(y\)-deviations sum to \(10\). Substituting:
The value \(r=0.900\) indicates a strong positive linear association in these five students: students with more practice sessions tended to have higher quiz scores, and the points follow a fairly close upward linear pattern. The coefficient is not a prediction that each additional session raises a score by 0.900 points. Nor does it show that practice sessions caused the scores; it summarizes the association in these paired observations.
Worked Example: Evening Device Use and Sleep
Worked Example: Evening Device Use and Sleep
For a fictional classroom data activity, students record evening device-use hours and sleep hours for five observations. The values are invented for practice.
| Device use, \(x\) (hours) | Sleep, \(y\) (hours) |
|---|---|
| 1 | 9 |
| 2 | 7 |
| 3 | 8 |
| 4 | 4 |
| 5 | 2 |
Question: Use the deviation formula to find \(r\) and describe what it says.
The means are \(\bar{x}=3\) hours of device use and \(\bar{y}=6\) hours of sleep. The \(x\)-deviations are \(-2,-1,0,1,2\); the \(y\)-deviations are \(3,1,2,-2,-4\). Their cross-products are \(-6,-1,0,-2,-8\), which sum to \(-17\). The squared \(x\)-deviations sum to \(10\), while the squared \(y\)-deviations sum to \(34\).
The calculation says there is a strong negative linear association among these five observations: greater evening device use tends to go with fewer hours of sleep. The negative sign describes direction, while the magnitude, about \(0.922\), indicates that the data follow a fairly close straight-line pattern. This is a description of the invented paired data, not evidence by itself that device use causes less sleep.
Worked Example: A Clear Curve with \(r=0\)
Worked Example: A Clear Curve with \(r=0\)
A fictional workshop tests five settings, recorded as degrees above or below a recommended setting. For each setting, the team records the number of defects on a sample item.
| Setting offset, \(x\) (degrees) | Defects, \(y\) |
|---|---|
| \(-2\) | 4 |
| \(-1\) | 1 |
| 0 | 0 |
| 1 | 1 |
| 2 | 4 |
Question: What does \(r\) summarize here, and what important feature would it miss?
The means are \(\bar{x}=0\) degrees and \(\bar{y}=2\) defects. The \(x\)-deviations are \(-2,-1,0,1,2\), and the \(y\)-deviations are \(2,-1,-2,-1,2\). The cross-products are \(-4,1,0,-1,4\), which sum to \(0\). The squared deviations sum to \(10\) for \(x\) and \(14\) for \(y\).
Here \(r=0\) means there is no linear association summarized by the correlation coefficient. It does not mean there is no association. The table shows a clear U-shaped pattern: settings closer to the recommended setting have fewer defects, while settings farther above or below it have more. The upward and downward parts of the curve balance in the linear calculation. This example shows why you should inspect the scatterplot before treating a small \(r\) as evidence that the variables have no relationship.
Common Mistakes and AP Exam Tips
- Calling \(r\) a slope. A slope describes the predicted change in response for a one-unit increase in the explanatory variable, with units. Correlation has no units and summarizes linear direction and strength.
- Treating \(r\) as a percentage. A value such as \(0.900\) is not “90% of scores explained” and does not mean the response rises by 90%. State what the sign and magnitude say about the linear association.
- Claiming that a small \(r\) proves there is no relationship. A curved pattern may be strong even when \(r\) is near zero. Check form in the scatterplot and distinguish “little linear association” from “no association.”
- Interpreting strength without checking form or unusual points. Correlation is designed to summarize a linear pattern, and an outlier can strongly affect its value. Inspect the plot and mention notable departures instead of letting one number stand in for the whole description.
- Using causation language. A correlation between two variables does not, by itself, show that changing one causes a change in the other. Describe the observed association unless the study design justifies a causal conclusion.
- Leaving out context. On an exam, name the variables and state what tends to happen as one increases. For instance, say “greater evening device use tends to be associated with fewer hours of sleep” rather than writing only “strong negative correlation.”
A careful AP response connects the number to the scatterplot and the situation: identify the direction, describe the strength of the linear pattern, and name the two quantitative variables. Avoid translating \(r\) into a slope, percentage, causal claim, or summary of a curve. “Describing a Scatterplot With DUFS” remains useful: \(r\) can support your description of direction and strength, but you must still consider form and unusual features.
Check Your Understanding
Answer each question using the examples and the distinction between a linear summary and a complete description of paired data.
- In the practice-session example, what do the positive sign and the value \(0.900\) each tell you?
- Why would it be incorrect to say that each extra practice session raises a quiz score by 0.900 points?
- In the device-use example, interpret the sign of \(r\) in context without making a causal claim.
- The setting-offset data have \(r=0\). What pattern does the table show, and why does \(r\) fail to summarize it well?
- Before interpreting a calculated correlation, name two features of the scatterplot you should check.