When Do Two Categorical Variables Show an Association?
A two-way table can show more than the number of individuals in each combination of categories. It can also help us decide whether knowing one variable’s category gives us information about the distribution of the other variable. In this tutorial, we use association to describe that relationship.
As in the earlier tutorials on Reading a Two-Way Table and Conditional Distributions by Row, each inside cell counts individuals in two categories, and a conditional distribution describes one variable within a specified group. To look for association, compare the conditional distributions of one variable across the categories of the other variable.
For instance, suppose a table records participants’ workshop format and whether they completed a short course. We could compare the completion distributions for people in each format. If completion percentages differ across formats, the table shows an association between workshop format and course completion in the observed data. The size and clarity of the differences matter: distributions that are far apart show a clearer pattern than distributions that are only slightly different.
A table’s counts alone can be misleading when group totals differ. The relevant comparison is between conditional distributions, not just between raw counts. Use the correct denominator for each group, following the group wording as in Choosing the Correct Denominator for a Percentage.
A Practical Way to Look for a Pattern
Choose one variable whose distribution you want to compare, then describe that distribution separately within each category of the other variable. For a table organized with groups in rows, this usually means finding the row percentages for each row. The percentages in each conditional distribution should add to 100%, apart from small rounding differences.
Confirm that both variables are categorical and state what each category represents.
Identify the categories of one variable, such as workshop format or type of residence.
Within each group, describe the percentages in the categories of the other variable. Use each group’s total as its denominator.
Say whether the distributions differ clearly, differ only slightly, or are nearly the same, and name the variables and groups in context.
There is no single percentage-point cutoff that automatically makes a pattern “clear” or “weak.” Consider how far apart the conditional distributions are overall and describe the observed differences accurately. The labels are descriptions of the pattern, not formal tests or proof that a relationship exists in a larger population.
Worked Examples: Clear, Weak, and Little Apparent Association
Worked Example: A Clear Pattern in Course Completion
A fictional training program records the workshop format and whether each participant completes a short course. Compare the completion distributions for the two formats and describe the association shown in the table.
| Workshop format | Completed | Did not complete | Total |
|---|---|---|---|
| In person | 72 | 18 | 90 |
| Online | 36 | 54 | 90 |
The variables are workshop format and course completion. We want to see whether the completion distribution differs between participants in the in-person and online formats.
Compare the conditional distributions of completion within each format. Use each format’s row total as the denominator, not the grand total.
For in-person participants, the percentages are \(72/90\times100\%=80\%\) completed and \(18/90\times100\%=20\%\) did not complete. For online participants, they are \(36/90\times100\%=40\%\) completed and \(54/90\times100\%=60\%\) did not complete.
In this program’s data, the completion distributions differ clearly by workshop format: 80% of in-person participants completed the course, compared with 40% of online participants. The table shows an association between workshop format and completion among these participants.
The difference in the completion percentages is \(80\%-40\%=40\) percentage points. The matching differences in the “did not complete” percentages tell the same story: 20% versus 60%. Both conditional distributions provide evidence of a clear pattern in this table.
The comparison is not based on the number 72 being larger than 36 by itself. The two workshop groups have equal totals here, but conditional percentages are still the right way to describe their distributions. In another table, unequal group totals could make raw counts especially hard to compare.
Worked Example: A Weak Pattern in Preferred Study Location
A fictional student survey records whether a student usually studies at home, in a library, or somewhere else, along with whether the student is in an afternoon or evening study group. Compare study-location distributions across the two groups.
| Study group | At home | In a library | Somewhere else | Total |
|---|---|---|---|---|
| Afternoon | 52 | 30 | 18 | 100 |
| Evening | 49 | 32 | 19 | 100 |
Because each row total is 100, the row counts also give the row percentages directly. Among afternoon-group students, 52% usually study at home, 30% in a library, and 18% somewhere else. Among evening-group students, the corresponding percentages are 49%, 32%, and 19%.
The distributions are not identical, but they are quite similar. The home percentages differ by 3 percentage points, the library percentages by 2 percentage points, and the “somewhere else” percentages by 1 percentage point. We would describe this as a weak pattern, or little apparent association, between study group and usual study location in these survey data.
“Weak” does not mean that the variables have been proven unrelated. It describes the small differences visible in the table. A careful conclusion names the groups and variable rather than claiming that the distributions are exactly equal.
Worked Example: Similar Conditional Distributions Despite Different Group Sizes
A fictional community survey records housing type and whether a resident uses a reusable shopping bag most of the time. Compare the bag-use distributions for apartment and house residents. The group sizes are different, so use conditional percentages rather than comparing counts alone.
| Housing type | Usually uses a reusable bag | Does not usually use one | Total |
|---|---|---|---|
| Apartment | 30 | 70 | 100 |
| House | 60 | 140 | 200 |
Among apartment residents, the percentage who usually use a reusable bag is \(30/100\times100\%=30\%\), and the percentage who do not usually use one is \(70/100\times100\%=70\%\). Among house residents, the corresponding percentages are \(60/200\times100\%=30\%\) and \(140/200\times100\%=70\%\).
The conditional distributions are identical: 30% usually use a reusable bag and 70% do not, in each housing group. The counts differ because there are twice as many house residents as apartment residents in this sample. In these data, housing type and usual reusable-bag use show no apparent association.
The conclusion is about the observed data. It does not prove that the variables are independent in every community or that their distributions would always be exactly the same in another sample.
What Association Does—and Does Not—Tell Us
Association is a description of how variables vary together in the data. If conditional distributions differ, the category of one variable is related to the distribution of the other variable in the table. If they are similar, the table shows little apparent association. This way of thinking applies whether the table has two categories per variable or several.
The same table can be examined in either direction. In the workshop example, we compared completion within each format. We could instead compare the format distribution among participants who completed the course and among those who did not. For describing whether there is an association, the key question remains whether conditional distributions differ. The percentages and wording change with the direction of conditioning, so state clearly which groups your percentages describe.
Association does not establish cause and effect. In the workshop example, participants may not have been assigned randomly to formats, and other differences between participants could be related to completion. The table alone cannot show that format caused the difference. Even when an association is clear, describe what was observed rather than claiming a cause that the data do not establish.
Also distinguish a sample description from a claim about a broader population. Here, the examples describe fictional survey or program data. A visible pattern in a table is evidence of an association in those data; deciding whether it generalizes to a population requires attention to how the data were collected and, in later work, appropriate inference.
Common Mistakes and AP Exam Tips
- Comparing counts instead of distributions. A group with more individuals may have larger counts in several categories simply because its total is larger. Compare conditional percentages within the groups.
- Using the grand total as the denominator for a group distribution. To describe study locations among afternoon students, divide by the afternoon total. The grand total would describe a joint share of all students, not the conditional distribution within the afternoon group.
- Calling any difference a clear association. Conditional distributions can differ slightly. Describe the size and pattern of the differences; do not exaggerate a weak pattern.
- Claiming there is definitely no relationship from similar percentages. Say that the distributions are similar or that there is little apparent association in the data. A table does not prove exact equality in a larger population.
- Reversing the group and outcome in the interpretation. “80% of in-person participants completed” is not the same statement as “80% of participants who completed were in person.” Name the denominator group explicitly.
- Claiming causation from association. “The variables are associated in these data” is a descriptive conclusion. “One variable caused the other” requires evidence that a table by itself may not provide.
For a strong AP response, identify the variables, compare conditional distributions using the appropriate group totals, and state whether the pattern is clear, weak, or nearly absent. Include the relevant percentages and interpret them in context. Avoid saying simply “there is a relationship” without showing what differs, and do not turn a descriptive association into a causal claim.
Check Your Understanding
For each question, focus on the conditional distributions and describe what the table shows in context.
- In the workshop example, what percentage of online participants did not complete the course? What does that percentage describe?
- In the study-location example, which category has the largest percentage-point difference between the afternoon and evening groups? Is the overall pattern clear or weak?
- In the reusable-bag example, why are the counts 30 and 60 not evidence by themselves of different bag-use distributions?
- Suppose two groups have conditional distributions of 25% “yes” and 75% “no” in each group. What would you say about association in the data?
- Write a conclusion about the workshop table that describes its association without claiming that workshop format caused course completion.