Why the Range of \(x\) Matters
In “Correlation of Averages Versus Individuals,” you saw that the cases represented by the points matter when interpreting \(r\). Another feature of the data can also affect the observed correlation: how much the explanatory variable \(x\) varies. If a data set includes only cases from a narrow range of \(x\)-values, the linear pattern can look weaker than it does across a wider range.
The reason is that correlation measures how closely the points follow a linear pattern relative to their variation. When \(x\) varies widely, changes associated with the linear pattern may stand out against the scatter. When the \(x\)-values are packed into a narrow interval, that pattern has less horizontal variation to distinguish it from the remaining scatter in \(y\). The result can be a smaller \(|r|\).
This effect is a tendency, not a rule that applies to every subset. Correlation also depends on the particular points selected, including how their vertical scatter lines up with \(x\). As in “How Outliers Change the Correlation,” changing which observations are included can change \(r\); here, the key issue is that the selected observations have less variation in \(x\).
A Comparison Using the Same Data
The first example isolates the effect of looking at a narrower \(x\)-range. We will calculate the correlation for all the cases and then recalculate it using only cases whose \(x\)-values fall in the middle of the full range.
Worked Example: Restricting the Range Weakens the Correlation
Suppose a fictional environmental data set records an explanatory measurement \(x\) and a response measurement \(y\) at 11 sites. The values are invented for illustration. First, use every site; then use only the five sites with \(x\)-values from 3 through 7.
| \(x\) | \(y\) |
|---|---|
| 0 | 0 |
| 1 | 1 |
| 2 | 2 |
| 3 | 5 |
| 4 | 2 |
| 5 | 5 |
| 6 | 4 |
| 7 | 9 |
| 8 | 8 |
| 9 | 9 |
| 10 | 10 |
All 11 sites: The means are \(\bar{x}=5\) and \(\bar{y}=5\). Using deviations from those means, the sum of squared deviations for \(x\) is \(S_{xx}=110\), the sum of squared deviations for \(y\) is \(S_{yy}=126\), and the sum of cross-products is \(S_{xy}=110\). For example, the \(x\)-deviations are \(-5,-4,\ldots,5\), whose squares sum to 110. The \(y\)-deviations are \(-5,-4,-3,0,-3,0,-1,4,3,4,5\), whose squares sum to 126. Multiplying corresponding \(x\)- and \(y\)-deviations and adding gives \(S_{xy}=110\).
Only the five sites with \(3\leq x\leq 7\): Their points are \((3,5),(4,2),(5,5),(6,4),(7,9)\). Their means are \(\bar{x}=5\) and \(\bar{y}=5\). The \(x\)-deviations are \(-2,-1,0,1,2\), so \(S_{xx}=10\). The \(y\)-deviations are \(0,-3,0,-1,4\), so \(S_{yy}=0^2+(-3)^2+0^2+(-1)^2+4^2=26\). The cross-products sum to \(0+3+0-1+8=10\), so \(S_{xy}=10\).
Interpret in context: Across all 11 sites, \(x\) and \(y\) have a very strong positive linear association. Among just the five sites with \(x\) between 3 and 7, the association is still positive but is less strong. The restricted subset has much less variation in \(x\), while the response values still show scatter around a straight-line pattern. In this example, that makes the linear association less pronounced relative to the variation in the selected data.
The comparison can also be understood by looking at the construction of the values. In both the full data and the restricted subset, \(y=x+e\), where \(e\) is a departure from the simple linear pattern. For the full data, the variation associated with \(x\) is relatively large compared with the departures. Within the restricted subset, the \(x\)-values cover only five consecutive integers, while the departures remain noticeable. They account for a larger share of the pattern’s remaining variation.
This explanation is not a formula that predicts the exact correlation for every selected subset. Which points remain matters. A narrow-range subset might happen to contain points that lie unusually close to a line, or points whose vertical departures follow their own pattern. The calculation must use the actual paired observations in the subset.
Comparing Two Ranges in a Second Data Set
Range restriction can be considered directly by comparing a broad set of observations with a subset from its middle. The next example uses another invented data set and repeats the comparison with a different linear trend and different departures from it.
Worked Example: Comparing Broad and Narrow \(x\)-Ranges
A fictional technology test records \(x\), an input setting, and \(y\), a measured output. The five paired observations are \((0,0),(1,5),(2,2),(3,11),(4,12)\). Compare the correlation for all five settings with the correlation for the middle three settings, \(x=1,2,3\).
All five observations: The means are \(\bar{x}=2\) and \(\bar{y}=6\). The \(x\)-deviations are \(-2,-1,0,1,2\), giving \(S_{xx}=10\). The \(y\)-deviations are \(-6,-1,-4,5,6\), giving \(S_{yy}=36+1+16+25+36=114\). The cross-products sum to \(12+1+0+5+12=30\), so \(S_{xy}=30\).
Middle three observations: The selected points are \((1,5),(2,2),(3,11)\). Their means are \(\bar{x}=2\) and \(\bar{y}=6\). The \(x\)-deviations \(-1,0,1\) give \(S_{xx}=2\). The \(y\)-deviations \(-1,-4,5\) give \(S_{yy}=1+16+25=42\). The cross-products sum to \(1+0+5=6\), so \(S_{xy}=6\).
Interpret in context: For all five settings, the input setting and measured output have a strong positive linear association. Among the three middle settings, they have a moderately strong positive linear association. The narrower range corresponds to a smaller observed correlation in this particular data set. The values do not show that the relationship must be weaker for every possible selection of middle settings.
When Restricting the Range Does Not Weaken \(r\)
A key qualification is that range restriction does not automatically lower correlation. If every point lies exactly on a straight line, keeping only some of the points that still have different \(x\)-values leaves a perfect linear pattern. In that case, the correlation remains \(1\) for an upward-sloping line or \(-1\) for a downward-sloping line.
Worked Example: A Perfect Linear Pattern Stays Perfect
Consider invented measurements of a device setting \(x\) and an output \(y\): \((1,4),(2,7),(3,10),(4,13),(5,16)\). Each point satisfies \(y=3x+1\). Compare all five points with the three points whose \(x\)-values are 2, 3, and 4.
All five points: The means are \(\bar{x}=3\) and \(\bar{y}=10\). The \(x\)-deviations are \(-2,-1,0,1,2\), so \(S_{xx}=10\). Since each \(y\)-deviation is three times its corresponding \(x\)-deviation, \(S_{xy}=3(10)=30\) and \(S_{yy}=9(10)=90\).
Three middle points: The points are \((2,7),(3,10),(4,13)\), with means \(\bar{x}=3\) and \(\bar{y}=10\). The \(x\)-deviations are \(-1,0,1\), so \(S_{xx}=2\). The \(y\)-deviations are \(-3,0,3\), so \(S_{yy}=18\) and \(S_{xy}=6\).
Interpret in context: The device setting and output have a perfect positive linear association in both the full set and the selected middle range. The range restriction reduces the spread of \(x\), but there is no scatter away from the line to make the pattern less linear. If a selected set had only one distinct \(x\)-value, however, its correlation would be undefined because \(x\) would have no variation.
How to Read a Correlation From a Restricted Sample
When a correlation is calculated from observations selected over a narrow \(x\)-range, interpret it as the linear association among the observations that remain. Do not treat it as if it necessarily summarizes the full range of cases that might have been observed. This is especially important when the data were gathered from a group selected for having similar \(x\)-values, or when a report includes only a limited interval of \(x\).
- Check the range of \(x\). Compare the observed \(x\)-values with the range relevant to the question. A restricted range can make an association appear weaker.
- Describe the actual cases. As in “Writing Descriptions in Context,” name the cases and both quantitative variables when interpreting the association.
- Do not infer the full-range correlation from the subset alone. Without the omitted observations, the correlation for the broader range cannot be calculated from the restricted correlation alone.
- Check the form and unusual points. As in “Correlation Measures Only Linear Association” and “How Outliers Change the Correlation,” a curved pattern or an unusual point can also affect \(r\). A narrow range is not the only possible explanation for a low correlation.
- Avoid causal conclusions. A change in correlation after restricting the range is a change in a summary of the observed data; by itself, it does not show that selecting cases caused a change in the underlying relationship.
Common Mistakes and AP Exam Tips
- Claiming range restriction always lowers correlation. It often can weaken the observed linear association when scatter remains, but it is not guaranteed. The exact points in the subset matter, and a perfect line remains perfect if \(x\) still varies.
- Calling the restricted correlation the correlation for all cases. A value of \(r\) calculated from a narrow interval describes only the observations included in that calculation. State which cases or \(x\)-range the result represents.
- Confusing a smaller range of \(x\) with a change in the measurement units. As discussed in “Correlation Has No Units,” converting units does not change \(r\). Range restriction is different: it changes which observations are included.
- Assuming a lower \(r\) proves a different underlying relationship. A smaller observed correlation can result from having less variation in \(x\) relative to the scatter. It does not, on its own, establish that the overall relationship has changed.
- Reporting only the number. A full-credit interpretation identifies the direction and strength of the linear association, names both variables and their units or meanings, and specifies the restricted cases when relevant.
For example, a careful response could say: “Among the five sites with \(x\)-values from 3 to 7, \(x\) and \(y\) have a moderately strong positive linear association (\(r\approx0.6202\)); this is weaker than the very strong positive association among all 11 sites.” This directly compares the data sets without claiming that the restricted correlation applies outside its observed range.
Check Your Understanding
Use the ideas in this tutorial to explain what a restricted-range correlation can and cannot show.
- Why might a narrow range of \(x\)-values make the observed linear association appear weaker when vertical scatter remains?
- A correlation is calculated using only observations with \(x\)-values between 20 and 30. What should an interpretation say about the cases represented by this \(r\)?
- Can selecting a narrow range of \(x\)-values leave \(r\) unchanged? Describe one situation in which it would.
- Why can the restricted correlation not, by itself, determine the correlation for all cases in a broader \(x\)-range?
- A report gives a lower correlation after excluding observations outside a narrow \(x\)-interval. Name one careful conclusion and one conclusion that the comparison alone does not justify.