Tutorials › AP Statistics › Correlation Requires Quantitative Variables

Correlation · Tutorial 829 of 1000

Correlation Requires Quantitative Variables

Understand why category labels do not make a variable quantitative, and learn to recognize when a reported correlation is misleading.

Intermediate 9 min read

What You'll Learn

  • Distinguish quantitative measurements from categorical labels that happen to be written as numbers.
  • Explain why coding categories can make a calculated correlation depend on arbitrary code choices.
  • Identify misleading claims that treat category codes as meaningful measurements.
  • Choose an appropriate way to summarize categorical and quantitative variables.
  • Describe the limits of correlation when one or both variables are categorical.

Why the Type of Variable Matters

In “Effect of Changing Units on \(r\),” you saw that shifts and rescaling affect the correlation in specific ways. Those properties assume that the variables represent quantities for which arithmetic makes sense. Before calculating \(r\), check what the values mean—not just whether they look like numbers.

A quantitative variable records a numerical amount or measurement, such as time in minutes, distance in kilometers, or number of visits. A categorical variable records which group or label an observation belongs to, such as device type, neighborhood, or favorite fruit. A category can be written using a number without becoming a quantitative measurement. For example, if a survey records “1 = bus, 2 = bicycle, 3 = walking,” those numbers are labels, not amounts of transportation.

Definition: In AP Statistics, \(r\) summarizes the direction and strength of the linear association between two quantitative variables. Its calculation uses numerical values, means, and deviations from means. Category labels do not have meaningful numerical distances or deviations, so an ordinary interpretation of \(r\) is not appropriate for them.

The issue is not that a calculator is physically unable to process category codes. If you enter codes such as 1, 2, and 3, a calculator can produce a numerical result. The issue is that the result may reflect the arbitrary coding scheme rather than a meaningful association between quantities. A computed number is not automatically a meaningful statistical summary.

For a quantitative variable, differences have a meaningful interpretation: 8 minutes is 3 minutes longer than 5 minutes. For nominal categories such as tea, coffee, and water, there is no meaningful statement that coffee is “one unit more” than tea. Their numerical labels can be reassigned without changing the actual observations. If that reassignment changes the calculated \(r\), the apparent correlation is tied to the labels, not to a quantitative relationship.

What Goes Wrong When Categories Receive Numbers

Suppose a fictional service desk records an interface category and the time, in minutes, each customer spends completing a task. The interface categories are labeled A, B, and C. The task times are quantitative, but the interface type is categorical. To see why coding can be misleading, consider six invented observations:

Interface categoryTask time (minutes)
A1
A2
B2
B5
C4
C6

The categories have no natural numerical scale. Still, imagine assigning A = 1, B = 2, and C = 3, then calculating a correlation with task time. That creates a numerical variable from labels. The resulting arithmetic can be checked, but the code values do not measure an amount of interface type.

Worked Example: A Correlation That Depends on Category Codes

Use the six service-desk observations above. First assign A = 1, B = 2, and C = 3. Calculate the resulting \(r\), then change the coding to A = 1, B = 3, and C = 2 and calculate again. The category labels and task times have not changed.

First coding. The coded category values are \(1,1,2,2,3,3\), with mean \(\bar{x}=2\). The task times have mean \(\bar{y}=20/6=10/3\) minutes. The category deviations are \(-1,-1,0,0,1,1\). The task-time deviations are \(-7/3,-4/3,-4/3,5/3,2/3,8/3\). Therefore, \(S_{xx}=4\), and the sum of paired deviation products is

$$ S_{xy}= \frac{7}{3}+\frac{4}{3}+0+0+\frac{2}{3}+\frac{8}{3} =7 $$

The sum of squared task-time deviations is

$$ S_{yy} =\frac{49+16+16+25+4+64}{9} =\frac{174}{9} =\frac{58}{3} $$

So the correlation for these codes is

$$ r=\frac{S_{xy}}{\sqrt{S_{xx}S_{yy}}} =\frac{7}{\sqrt{4(58/3)}} \approx 0.796 $$

Second coding. With A = 1, B = 3, and C = 2, the codes are \(1,1,3,3,2,2\), still with mean 2. The category deviations are now \(-1,-1,1,1,0,0\). The task-time values and their deviations have not changed, so \(S_{yy}=58/3\). The new sums are \(S_{xx}=4\) and \(S_{xy}=7/3+4/3-4/3+5/3=4\). Thus

$$ r=\frac{4}{\sqrt{4(58/3)}}\approx 0.455 $$

The calculated correlations differ even though the observations themselves are identical. Only the arbitrary assignment of numbers to the interface categories changed. Neither result is a suitable ordinary correlation between interface type and task time. A better analysis would compare task-time distributions across the three categories or report each category’s times and appropriate summaries.

This example illustrates an important distinction: \(r\) is unchanged by certain transformations of genuine quantitative variables, as discussed in the earlier tutorial on changing units. But changing the codes assigned to nominal categories is not simply changing measurement units. There is no meaningful numerical scale for those categories to preserve.

Spotting Misuse in a Report or Question

When someone reports a correlation, identify the two variables and ask what each recorded value represents. A useful audit is to check whether the values are measurements or counts with meaningful arithmetic, or instead labels that identify groups. Also ask whether the claimed conclusion depends on treating the spacing between values as meaningful.

AP Exam Tip: Do not decide that a variable is quantitative merely because its data appear as digits. Ask whether differences and averages have a meaningful interpretation. Student ID numbers, postal codes, jersey numbers, and codes for product types are labels, not measurements.

A variable can be categorical even when its categories have an order. For example, a response of “low,” “medium,” or “high” has an ordering, but the gaps between those categories are not automatically equal numerical amounts. Coding them 1, 2, and 3 imposes equal spacing that may not be justified. A numerical rating scale may be treated differently if its design and the analysis support treating scores as quantitative; do not assume that every ordered response is automatically a measurement with equal intervals.

Likewise, counting categories does not necessarily turn them into quantities. The number of customers choosing each flavor is a quantitative count for each flavor, but “flavor” itself is categorical. Keep clear which variable is being analyzed: the category label, the count, or another measurement.

Worked Example: Checking a Claim About Jersey Numbers

A fictional youth league records players’ jersey numbers and their sprint times in seconds. A summary says, “There is a strong positive correlation: players with larger jersey numbers tend to have longer sprint times.” Is that conclusion supported by a meaningful correlation?

Identify the variables. Sprint time in seconds is quantitative: differences in time are meaningful. Jersey number is categorical identification. A player wearing number 30 does not have twice the jersey-number amount of a player wearing number 15, and the numerical gap between jersey numbers does not measure a difference in player characteristics.

Audit the claim. If the league renumbered players while keeping the same players and sprint times, the reported correlation could change. Thus, the direction and strength would describe the chosen labels, not a meaningful linear relationship. The statement that larger jersey numbers tend to go with longer sprint times is not a justified interpretation of \(r\).

A better report would describe sprint times directly or investigate a meaningful quantitative variable, such as players’ practice time, if that information were available. Jersey number should not be used as a quantitative explanatory variable simply because it is numeric-looking.

Choosing a Better Description

If one variable is categorical and the other is quantitative, keep the groups distinct and summarize the quantitative measurements within each group. Depending on the purpose, compare side-by-side plots, medians, means, or spreads. State the groups and the measurement units. Do not replace the group labels with arbitrary codes and then interpret the result as a linear trend.

If both variables are categorical, organize the observations by category combinations in a two-way table or compare the relevant proportions. Correlation is not the appropriate summary because neither variable supplies quantitative values for a linear relationship. The specific display or analysis depends on the question, but the first step is to preserve the categorical structure rather than invent numerical distances.

Worked Example: Choosing a Summary for Study Format and Quiz Score

A fictional teacher records each student’s study format—paper notes, videos, or practice questions—and a quiz score out of 20. A student proposes coding the formats 1, 2, and 3 and calculating \(r\) with quiz score. What should the analysis do instead?

Classify the variables. Study format is categorical and quiz score is quantitative. The numerical order assigned to the formats has no inherent meaning; videos are not “one unit more” than paper notes, and practice questions are not necessarily “one unit more” than videos.

Choose a suitable description. Compare the quiz-score distributions for the three formats. A side-by-side dotplot or boxplot can show differences in center, spread, and unusual scores. A table of group sizes and mean or median scores can also be useful, with the chosen summaries named clearly and scores kept in points out of 20.

For example, a careful descriptive statement might say, “In this class, the median quiz score was highest among students who reported using practice questions.” That describes the observed groups without pretending the study formats lie on a numerical scale. It also avoids claiming that study format caused the score pattern; a correlation or group comparison alone does not establish causation.

When both variables are quantitative, \(r\) can be a useful numerical summary of linear association, as covered in “What the Correlation Coefficient Measures.” Even then, the scatterplot matters: \(r\) does not summarize every feature of a relationship, and an unusual point or nonlinear pattern can affect how it should be understood.

Common Mistakes and Full-Credit Communication

  • Equating digits with quantities. A jersey number or region code is still a category label. Explain what the value represents before classifying the variable.
  • Assuming ordered categories have equal gaps. “Low,” “medium,” and “high” have an order, but that alone does not make the step from low to medium quantitatively equal to the step from medium to high.
  • Accepting a calculator result as proof of validity. A calculator can compute with arbitrary codes. State why those codes do or do not represent a meaningful quantitative scale.
  • Calling a category-number association positive or negative without checking the coding. If a category’s code changes, the apparent direction can change too. A correlation based on arbitrary labels does not support a stable contextual interpretation.
  • Rejecting all uses of numbers attached to categories without examining what is measured. A count, duration, or score can be quantitative even when it is organized by category. Identify the variable itself, not just the table or chart where it appears.

A full-credit response should name the variable type and explain the consequence. For example: “Jersey number is a categorical identifier, not a quantitative measurement. Its numerical spacing is arbitrary, so a correlation with sprint time would not have a meaningful interpretation.” For a categorical explanatory variable and quantitative response, add an appropriate alternative, such as comparing score distributions across groups.

Key takeaway: \(r\) is intended to summarize linear association between two quantitative variables. Numeric category codes do not make categories quantitative; if the code assignments are arbitrary, the calculated correlation can be misleading. Identify the variable types before interpreting \(r\).

Check Your Understanding

For each situation, decide whether an ordinary correlation is an appropriate summary and explain your reasoning.

  1. A researcher records weekly exercise time in minutes and resting heart rate in beats per minute. Are both variables quantitative?
  2. A club records member ID numbers and the number of meetings each member attended. Which variable is categorical and which is quantitative?
  3. A survey records favorite fruit as 1 = apple, 2 = pear, and 3 = orange, along with age in years. Why could the correlation depend on the coding?
  4. Students choose one of three study formats and receive a score out of 25. Name an appropriate way to describe the scores across formats.
  5. A response variable is recorded as “low,” “medium,” or “high” and coded 1, 2, or 3. What additional assumption would be needed to treat the codes as quantitative measurements?