Tutorials › AP Statistics › Is There an Association Between Two Categorical Variables

Categorical tables and summaries · Tutorial 30 of 1000

Is There an Association Between Two Categorical Variables

Learn to identify association by comparing conditional distributions across groups, and distinguish a clear pattern from a weak one without claiming that association proves causation.

Beginner 9 min read

What You'll Learn

  • Define association between two categorical variables using differing conditional distributions.
  • Identify which group-specific distributions to compare in a two-way table.
  • Describe a clear association pattern using conditional percentages and context.
  • Recognize a weak pattern when conditional distributions differ only slightly.
  • Explain what similar conditional distributions suggest about association in the data.
  • Distinguish an observed association from a cause-and-effect claim.

When Do Two Categorical Variables Show an Association?

A two-way table can show more than the number of individuals in each combination of categories. It can also help us decide whether knowing one variable’s category gives us information about the distribution of the other variable. In this tutorial, we use association to describe that relationship.

As in the earlier tutorials on Reading a Two-Way Table and Conditional Distributions by Row, each inside cell counts individuals in two categories, and a conditional distribution describes one variable within a specified group. To look for association, compare the conditional distributions of one variable across the categories of the other variable.

Definition: In a data set, two categorical variables are associated when the conditional distributions of one variable differ across the categories of the other variable. If those conditional distributions are the same, the variables show no association in that data set.

For instance, suppose a table records participants’ workshop format and whether they completed a short course. We could compare the completion distributions for people in each format. If completion percentages differ across formats, the table shows an association between workshop format and course completion in the observed data. The size and clarity of the differences matter: distributions that are far apart show a clearer pattern than distributions that are only slightly different.

A table’s counts alone can be misleading when group totals differ. The relevant comparison is between conditional distributions, not just between raw counts. Use the correct denominator for each group, following the group wording as in Choosing the Correct Denominator for a Percentage.

A Practical Way to Look for a Pattern

Choose one variable whose distribution you want to compare, then describe that distribution separately within each category of the other variable. For a table organized with groups in rows, this usually means finding the row percentages for each row. The percentages in each conditional distribution should add to 100%, apart from small rounding differences.

1
Name the two variables.
Confirm that both variables are categorical and state what each category represents.
2
Choose the groups to compare.
Identify the categories of one variable, such as workshop format or type of residence.
3
Compare conditional distributions.
Within each group, describe the percentages in the categories of the other variable. Use each group’s total as its denominator.
4
Describe the pattern.
Say whether the distributions differ clearly, differ only slightly, or are nearly the same, and name the variables and groups in context.

There is no single percentage-point cutoff that automatically makes a pattern “clear” or “weak.” Consider how far apart the conditional distributions are overall and describe the observed differences accurately. The labels are descriptions of the pattern, not formal tests or proof that a relationship exists in a larger population.

Worked Examples: Clear, Weak, and Little Apparent Association

Worked Example: A Clear Pattern in Course Completion

A fictional training program records the workshop format and whether each participant completes a short course. Compare the completion distributions for the two formats and describe the association shown in the table.

Workshop formatCompletedDid not completeTotal
In person721890
Online365490
1
State.
The variables are workshop format and course completion. We want to see whether the completion distribution differs between participants in the in-person and online formats.
2
Plan.
Compare the conditional distributions of completion within each format. Use each format’s row total as the denominator, not the grand total.
3
Do.
For in-person participants, the percentages are \(72/90\times100\%=80\%\) completed and \(18/90\times100\%=20\%\) did not complete. For online participants, they are \(36/90\times100\%=40\%\) completed and \(54/90\times100\%=60\%\) did not complete.
4
Conclude.
In this program’s data, the completion distributions differ clearly by workshop format: 80% of in-person participants completed the course, compared with 40% of online participants. The table shows an association between workshop format and completion among these participants.

The difference in the completion percentages is \(80\%-40\%=40\) percentage points. The matching differences in the “did not complete” percentages tell the same story: 20% versus 60%. Both conditional distributions provide evidence of a clear pattern in this table.

The comparison is not based on the number 72 being larger than 36 by itself. The two workshop groups have equal totals here, but conditional percentages are still the right way to describe their distributions. In another table, unequal group totals could make raw counts especially hard to compare.

Worked Example: A Weak Pattern in Preferred Study Location

A fictional student survey records whether a student usually studies at home, in a library, or somewhere else, along with whether the student is in an afternoon or evening study group. Compare study-location distributions across the two groups.

Study groupAt homeIn a librarySomewhere elseTotal
Afternoon523018100
Evening493219100

Because each row total is 100, the row counts also give the row percentages directly. Among afternoon-group students, 52% usually study at home, 30% in a library, and 18% somewhere else. Among evening-group students, the corresponding percentages are 49%, 32%, and 19%.

$$ \begin{aligned} \text{Afternoon: } & 52\%,\ 30\%,\ 18\% \\ \text{Evening: } & 49\%,\ 32\%,\ 19\% \end{aligned} $$

The distributions are not identical, but they are quite similar. The home percentages differ by 3 percentage points, the library percentages by 2 percentage points, and the “somewhere else” percentages by 1 percentage point. We would describe this as a weak pattern, or little apparent association, between study group and usual study location in these survey data.

“Weak” does not mean that the variables have been proven unrelated. It describes the small differences visible in the table. A careful conclusion names the groups and variable rather than claiming that the distributions are exactly equal.

Worked Example: Similar Conditional Distributions Despite Different Group Sizes

A fictional community survey records housing type and whether a resident uses a reusable shopping bag most of the time. Compare the bag-use distributions for apartment and house residents. The group sizes are different, so use conditional percentages rather than comparing counts alone.

Housing typeUsually uses a reusable bagDoes not usually use oneTotal
Apartment3070100
House60140200

Among apartment residents, the percentage who usually use a reusable bag is \(30/100\times100\%=30\%\), and the percentage who do not usually use one is \(70/100\times100\%=70\%\). Among house residents, the corresponding percentages are \(60/200\times100\%=30\%\) and \(140/200\times100\%=70\%\).

The conditional distributions are identical: 30% usually use a reusable bag and 70% do not, in each housing group. The counts differ because there are twice as many house residents as apartment residents in this sample. In these data, housing type and usual reusable-bag use show no apparent association.

The conclusion is about the observed data. It does not prove that the variables are independent in every community or that their distributions would always be exactly the same in another sample.

What Association Does—and Does Not—Tell Us

Association is a description of how variables vary together in the data. If conditional distributions differ, the category of one variable is related to the distribution of the other variable in the table. If they are similar, the table shows little apparent association. This way of thinking applies whether the table has two categories per variable or several.

The same table can be examined in either direction. In the workshop example, we compared completion within each format. We could instead compare the format distribution among participants who completed the course and among those who did not. For describing whether there is an association, the key question remains whether conditional distributions differ. The percentages and wording change with the direction of conditioning, so state clearly which groups your percentages describe.

Association does not establish cause and effect. In the workshop example, participants may not have been assigned randomly to formats, and other differences between participants could be related to completion. The table alone cannot show that format caused the difference. Even when an association is clear, describe what was observed rather than claiming a cause that the data do not establish.

Also distinguish a sample description from a claim about a broader population. Here, the examples describe fictional survey or program data. A visible pattern in a table is evidence of an association in those data; deciding whether it generalizes to a population requires attention to how the data were collected and, in later work, appropriate inference.

Common Mistakes and AP Exam Tips

  • Comparing counts instead of distributions. A group with more individuals may have larger counts in several categories simply because its total is larger. Compare conditional percentages within the groups.
  • Using the grand total as the denominator for a group distribution. To describe study locations among afternoon students, divide by the afternoon total. The grand total would describe a joint share of all students, not the conditional distribution within the afternoon group.
  • Calling any difference a clear association. Conditional distributions can differ slightly. Describe the size and pattern of the differences; do not exaggerate a weak pattern.
  • Claiming there is definitely no relationship from similar percentages. Say that the distributions are similar or that there is little apparent association in the data. A table does not prove exact equality in a larger population.
  • Reversing the group and outcome in the interpretation. “80% of in-person participants completed” is not the same statement as “80% of participants who completed were in person.” Name the denominator group explicitly.
  • Claiming causation from association. “The variables are associated in these data” is a descriptive conclusion. “One variable caused the other” requires evidence that a table by itself may not provide.

For a strong AP response, identify the variables, compare conditional distributions using the appropriate group totals, and state whether the pattern is clear, weak, or nearly absent. Include the relevant percentages and interpret them in context. Avoid saying simply “there is a relationship” without showing what differs, and do not turn a descriptive association into a causal claim.

Key takeaway: Two categorical variables are associated in a data set when the conditional distributions of one variable differ across categories of the other. Compare percentages within the groups, describe the strength of the visible pattern, and remember that association alone does not establish causation.

Check Your Understanding

For each question, focus on the conditional distributions and describe what the table shows in context.

  1. In the workshop example, what percentage of online participants did not complete the course? What does that percentage describe?
  2. In the study-location example, which category has the largest percentage-point difference between the afternoon and evening groups? Is the overall pattern clear or weak?
  3. In the reusable-bag example, why are the counts 30 and 60 not evidence by themselves of different bag-use distributions?
  4. Suppose two groups have conditional distributions of 25% “yes” and 75% “no” in each group. What would you say about association in the data?
  5. Write a conclusion about the workshop table that describes its association without claiming that workshop format caused course completion.