Tutorials › AP Statistics › Comparing Strength Using Correlations

Correlation · Tutorial 837 of 1000

Comparing Strength Using Correlations

Compare linear association strength by comparing the magnitudes of the correlations, while using their signs to describe direction.

Intermediate 9 min read

What You'll Learn

  • Compare correlation strength by comparing absolute values, not signed values
  • Explain why a correlation of negative 0.9 is stronger than positive 0.7
  • Separate the direction of an association from its strength
  • Calculate and compare correlations from paired data and summary quantities
  • Describe what a comparison does and does not establish in context

Compare Magnitude for Strength, Sign for Direction

In “Strong Correlation Does Not Mean a Good Model,” you saw why a high correlation alone does not establish that a straight line describes every feature of the data. When the question is specifically which of two linear associations is stronger, the key is to compare the magnitudes of the correlations.

The sign of \(r\) describes direction: a positive \(r\) indicates a positive linear association, and a negative \(r\) indicates a negative linear association. The magnitude, or absolute value, of \(r\) describes how closely the points follow a straight-line pattern. Because \(r\) ranges from \(-1\) to \(1\), a value closer to either endpoint represents a stronger linear association than a value closer to zero.

Key rule: To compare the strength of two linear associations, compare \(|r|\), the absolute values of their correlations. The larger absolute value indicates the stronger linear association. Use the signs separately to state the directions.

For example, compare \(r=-0.9\) with \(r=0.7\). Their magnitudes are \(|-0.9|=0.9\) and \(|0.7|=0.7\). Since \(0.9>0.7\), the association with \(r=-0.9\) is stronger in a linear sense. It is negative, while the association with \(r=0.7\) is positive. The negative sign does not make the first association weaker; it tells you that the direction is downward rather than upward.

This comparison is about how closely each set of points follows its own straight-line pattern, not about how steep its line is. As covered in “Correlation Has No Units,” \(r\) is unitless. A slope, by contrast, has units and depends on the variables’ scales. A steeper slope is not necessarily a stronger association.

A Reliable Comparison Method

When you are given two correlations, work through the same short sequence each time. First identify which correlation belongs to which pair of variables. Next compare their absolute values. Finally, state the result in context, including each association’s direction. This avoids treating a negative sign as a measure of weakness or overlooking which variables are being compared.

1
Identify each correlation.
Match each \(r\) to its pair of quantitative variables and the situation it describes.
2
Compare the magnitudes.
Find the absolute values. The larger \(|r|\) indicates the stronger observed linear association.
3
State direction and context.
Use the sign of each \(r\) to describe whether the association is positive or negative, and name the variables.
4
Keep the claim within its limits.
Describe the observed linear associations; do not infer causation or claim that one setting must be more predictable in every way.

As in “Correlation Measures Only Linear Association,” a small \(|r|\) does not rule out every kind of relationship. A curved pattern may be strong even though its linear correlation is small. Before relying on \(r\), consider the scatterplot’s form and any unusual points, as emphasized in “Judging Strength of an Association” and “How Outliers Change the Correlation.”

Worked Example: Why \(r=-0.9\) Is Stronger Than \(r=0.7\)

Two fictional scatterplots summarize different pairs of quantitative variables. In the first, the correlation between outdoor temperature and household heating use is \(r=-0.9\). In the second, the correlation between weekly practice time and a skills score is \(r=0.7\). Which observed linear association is stronger, and what are their directions?

Compare the absolute values:

$$ |-0.9|=0.9 \qquad\text{and}\qquad |0.7|=0.7 $$

Because \(0.9>0.7\), the temperature and heating-use data have the stronger observed linear association. Its negative direction means that higher outdoor temperatures tend to go with lower household heating use. The practice-time and skills-score data have a positive association: greater practice time tends to go with higher scores. The comparison is about linear strength, not whether either relationship is causal.

A concise response in context would be: “The association between outdoor temperature and household heating use is stronger because \(|-0.9|=0.9\) is greater than \(|0.7|=0.7\). It is negative, whereas the association between practice time and skills score is positive.” This explicitly separates the comparison of strength from the description of direction.

Comparing Correlations Calculated from Data

When the correlations are not supplied, use the paired observations to find each \(r\), then compare their absolute values. In “What the Correlation Coefficient Measures” and “Calculating \(r\) From Standardized Values,” you learned that \(r\) summarizes the direction and strength of a linear association between two quantitative variables. The following calculation applies that idea to two invented sets of observations.

Worked Example: Compare Two Tutoring Cohorts

Two fictional tutoring cohorts each include five students. For each student, a program records practice hours \(x\) and a quiz score \(y\), in points. The data are invented for this example. Which cohort has the stronger linear association between practice hours and quiz score?

Practice hours \(x\)Cohort A scoreCohort B score
198.5
278
31211
4910
51312.5

For both cohorts, \(\bar{x}=3\), so the \(x\)-deviations are \(-2,-1,0,1,2\), and \(S_{xx}=(-2)^2+(-1)^2+0^2+1^2+2^2=10\). Both sets of scores have mean \(10\). For Cohort A, the score deviations are \(-1,-3,2,-1,3\). The sum of their squares is \(S_{yy}=1+9+4+1+9=24\). The sum of the products of corresponding deviations is \(S_{xy}=2+3+0-1+6=10\).

Using \(r=S_{xy}/\sqrt{S_{xx}S_{yy}}\), Cohort A’s correlation is:

$$ r_A=\frac{10}{\sqrt{10(24)}} =\frac{10}{\sqrt{240}} \approx 0.6455 $$

For Cohort B, the score deviations are \(-1.5,-2,1,0,2.5\). Their squared values sum to \(S_{yy}=2.25+4+1+0+6.25=13.5\). The corresponding products with the \(x\)-deviations sum to \(S_{xy}=3+2+0+0+5=10\). Therefore:

$$ r_B=\frac{10}{\sqrt{10(13.5)}} =\frac{10}{\sqrt{135}} \approx 0.8607 $$

Both correlations are positive, so both cohorts show positive linear associations between practice hours and quiz score in these observations. To compare strength, use the magnitudes: \(|r_A|=0.6455\) and \(|r_B|=0.8607\). Since \(0.8607>0.6455\), Cohort B has the stronger observed linear association. This comparison does not show that practice time caused scores to increase, nor does it establish that the same pattern applies to other students.

What Comparisons Can and Cannot Tell You

A comparison of two sample correlations describes the relative strength of the two observed linear patterns. It does not, by itself, establish that the difference between the correlations is meaningful in a broader population. The AP-level task here is to compare the given or calculated \(r\) values and interpret what they summarize—not to perform an inference procedure for comparing correlations.

Also keep the two scatterplots in view. The larger \(|r|\) identifies the stronger linear association according to the correlation summaries, but \(r\) does not describe every feature of either plot. One plot might show curvature or an influential outlier. In that case, describe the plot’s form and unusual features as well as the correlation. The previous tutorial, “Strong Correlation Does Not Mean a Good Model,” explains why a high \(|r|\) is not enough to certify a linear model.

When comparing settings with different variables or units, avoid vague statements such as “the first relationship is better.” State exactly what is stronger: the observed linear association between the named variables. Correlation has no units, which makes comparing magnitudes possible, but it does not make the contexts interchangeable. The comparison does not say which outcome matters more, which model has more practical value, or whether either association should be used to make predictions.

Worked Example: Compare Associations in Different Settings

A fictional environmental data set has correlation \(r=-0.82\) between daily outdoor temperature and building heating demand. A separate fictional transportation data set has correlation \(r=0.76\) between the number of bicycles on a path and the number of pedestrians there. Compare the strengths and directions of the observed linear associations.

The magnitudes are \(|-0.82|=0.82\) and \(|0.76|=0.76\). Since \(0.82>0.76\), the temperature and heating-demand data show the stronger observed linear association. The first association is negative: higher temperatures tend to accompany lower heating demand. The second is positive: higher bicycle counts tend to accompany higher pedestrian counts.

A careful conclusion is: “The observed linear association between outdoor temperature and building heating demand is stronger because its correlation has the larger absolute value, \(0.82\) rather than \(0.76\). The first association is negative and the second is positive.” The conclusion compares the two sample patterns; it does not say that temperature changes cause heating demand to change or that bicycles cause more pedestrians.

Common Mistakes and AP Exam Tips

  • Comparing the signed numbers instead of their magnitudes. Since \(-0.9<0.7\), a student might incorrectly say that \(0.7\) is stronger. Compare \(|-0.9|=0.9\) with \(|0.7|=0.7\), so the negative correlation represents the stronger linear association.
  • Calling a negative correlation weak because it is negative. The sign gives direction, not strength. A correlation close to \(-1\) can represent a very strong negative linear association.
  • Confusing strength with slope. Do not compare how steep lines look to decide which association is stronger. Compare \(|r|\); slope depends on the variables and their units.
  • Leaving out the contexts or directions. “The first one is stronger” is incomplete if the variables are not clear. Name both pairs and describe positive or negative direction.
  • Claiming causation or a universal result. A strong observed correlation does not show that one variable caused the other to change, and a sample comparison does not automatically establish the same ordering in a wider population.
  • Treating \(r\) as a complete description of a scatterplot. Correlation summarizes linear association. Check form and unusual points; a curved pattern or influential observation can make a simple comparison incomplete.

For full-credit communication, show or state both absolute values, identify which is larger, and explain the direction of each association in context. For example: “Because \(|-0.9|=0.9>|0.7|=0.7\), the first pair of variables has the stronger observed linear association. That association is negative; the second is positive.”

Key takeaway: Compare the absolute values of correlations to compare linear strength. The larger \(|r|\) indicates the stronger observed linear association; the sign of \(r\) separately gives its direction. In particular, \(r=-0.9\) represents a stronger linear association than \(r=0.7\), even though their directions differ.

Check Your Understanding

For each comparison, focus on magnitude for strength and sign for direction.

  1. Which is stronger: \(r=-0.63\) or \(r=0.58\)? State both directions.
  2. A student says that \(r=-0.91\) is weaker than \(r=0.84\) because \(-0.91\) is the smaller number. What comparison should the student make instead?
  3. Two positive correlations are \(0.72\) and \(0.72\). What can you say about their relative linear strength, and what can you not conclude from that equality alone?
  4. Why should you inspect the scatterplots as well as compare \(|r|\)?
  5. Write a context-based sentence comparing \(r=-0.82\) for temperature and heating demand with \(r=0.76\) for bicycle and pedestrian counts.