Tutorials › AP Statistics › Independence Versus Association in Two-Way Tables

Categorical tables and summaries · Tutorial 32 of 1000

Independence Versus Association in Two-Way Tables

Compare conditional distributions across groups to decide whether a two-way table shows no apparent association, a near-independent pattern, or a clear difference.

Beginner 10 min read

What You'll Learn

  • Explain why equal conditional distributions indicate no association in a two-way table.
  • Compare each category’s conditional percentage across groups using consistent denominators.
  • Distinguish exact equality from a near-independent pattern.
  • Describe a table with clearly different conditional distributions as showing association.
  • Write conclusions in context without confusing association with causation.

When Conditional Distributions Match

In Is There an Association Between Two Categorical Variables and Comparing Conditional Percentages in Context, you learned to compare the distribution of one categorical variable across groups defined by another. This tutorial uses that comparison to identify a special pattern: if the conditional distributions are the same across the groups, the variables show no association in the data.

The key is to compare percentages within groups, not just the counts in the cells. As in Conditional Distributions by Row, a row percentage uses its row total as the denominator. If the rows represent the groups you are comparing, calculate the conditional distribution for each row and look for differences across rows.

Definition: In a two-way table, the variables show no association in the displayed data when the conditional distributions of one variable are the same across all categories of the other variable. If those distributions differ, the variables are associated in the displayed data.

For example, if 60% of people in each of two groups choose option A and 40% choose option B, the conditional distributions match. Knowing which group a person is in does not help distinguish the response distribution in this table. If the percentages differ from group to group, the response distribution varies with group membership, so the table shows an association.

In real data, conditional percentages are often close but not exactly equal. A near match suggests little or no apparent association in the table, while a more pronounced difference suggests a clearer association. There is no universal percentage-point cutoff that separates “near” from “different”: describe the pattern and its size in the context rather than treating a small difference as automatically meaningful or meaningless.

Key idea: Compare like with like: use the same outcome categories and calculate each conditional distribution using the total for its own group. Equal distributions indicate no association in the table; differences indicate association.

A Consistent Way to Compare Distributions

Before comparing, identify which variable defines the groups and which variable’s distribution you want to compare. Keep that choice consistent. For example, if the question is whether preference differs by grade, find the preference distribution within each grade. Do not compare one row percentage with a column percentage or use the grand total as the denominator.

For an outcome with several categories, compare the percentages for all categories, not just the one that looks most interesting. A conditional distribution includes every category and totals 100% (apart from possible rounding). If one category’s percentage is the same across two groups in a two-category outcome, the other category’s percentage must also be the same, because the two percentages add to 100%. With more categories, comparing one percentage alone may miss differences elsewhere.

1
Name the groups and outcome categories.
Identify the variable that defines the groups and the variable whose conditional distribution you are comparing.
2
Calculate conditional percentages.
Within each group, divide each cell count by that group’s total. Use the same type of denominator for every group.
3
Compare the full distributions.
Check whether the percentages for each outcome category are equal, close, or noticeably different across groups.
4
State the pattern in context.
Describe whether the distributions match or differ, and name the variables and groups. Do not turn a descriptive association into a claim of cause and effect.

A table’s counts can differ even when its conditional distributions match. Groups can have different totals, so matching percentages may correspond to different counts. That is why conditional percentages—not raw cell counts—are the right basis for deciding whether the displayed variables are associated.

Worked Examples: Matching and Differing Distributions

Worked Example: Exactly Matching Distributions Across Two Groups

A fictional recreation survey asks people in two neighborhoods whether they prefer an indoor or outdoor activity. Compare the activity-preference distributions by neighborhood.

NeighborhoodIndoorOutdoorTotal
North483280
South7248120

Use each neighborhood’s total as the denominator. In North, the indoor percentage is \(48/80\times100\%=60\%\), and the outdoor percentage is \(32/80\times100\%=40\%\). In South, the indoor percentage is \(72/120\times100\%=60\%\), and the outdoor percentage is \(48/120\times100\%=40\%\). The percentages in each row add to 100%.

NeighborhoodIndoorOutdoor
North60%40%
South60%40%

The conditional distributions match exactly. The number who prefer indoor activities differs—48 in North and 72 in South—but the group totals also differ. The appropriate comparison is that 60% in each neighborhood prefer indoor activities and 40% prefer outdoor activities.

Conclusion: In this survey, activity preference and neighborhood show no association in the displayed data because the conditional distributions of activity preference are identical. This describes the surveyed people; the table does not establish that neighborhood has no relationship with activity preference in every broader population.

Worked Example: A Near-Independent Pattern

A fictional school survey asks students in two lunch periods whether they usually bring lunch from home or get lunch elsewhere. Determine whether the conditional distributions suggest an association between lunch period and lunch source.

Lunch periodBring from homeGet lunch elsewhereTotal
First394180
Second413980

For the first period, the percentage who bring lunch from home is \(39/80\times100\%=48.75\%\); the percentage who get lunch elsewhere is \(41/80\times100\%=51.25\%\). For the second period, the corresponding percentages are \(41/80\times100\%=51.25\%\) and \(39/80\times100\%=48.75\%\). Each distribution totals 100%.

The percentage who bring lunch from home is \(51.25\%-48.75\%=2.5\) percentage points higher in the second period. The percentage who get lunch elsewhere is 2.5 percentage points higher in the first period. These are small differences, and the two conditional distributions are close, though not exactly equal.

Conclusion: In this fictional survey, lunch source has a near-independent pattern across lunch periods: the conditional distributions are similar, with differences of 2.5 percentage points in each category. It would be accurate to call the observed association small or not apparent from these percentages, but not to claim that the distributions are exactly equal. The table alone does not show why the small differences occurred.

Worked Example: Clearly Different Distributions

A fictional technology club surveys members about their preferred way to learn a new device feature. Members are grouped by how often they have used similar devices. Compare the learning-preference distributions across the three experience groups.

Experience with similar devicesWritten guideShort videoHands-on demoTotal
Rarely702010100
Sometimes453520100
Often253540100

Because each experience group has a total of 100, each count is the same number as its row percentage: for example, \(70/100\times100\%=70\%\). The conditional distributions are therefore 70%, 20%, and 10% for the rarely group; 45%, 35%, and 20% for the sometimes group; and 25%, 35%, and 40% for the often group. Each distribution sums to 100%.

The percentages differ across experience groups. The percentage preferring a written guide falls from 70% among members who rarely use similar devices to 25% among those who often use them, a difference of \(70\%-25\%=45\) percentage points. Meanwhile, hands-on-demo preference rises from 10% to 40%, a difference of 30 percentage points.

Conclusion: The table shows an association between experience group and learning preference among the surveyed club members because the conditional distributions differ substantially. This is a descriptive finding: it does not establish that experience caused the different preferences. Other differences among members could also be related to learning preference.

How to Describe Exact and Near Matches

Use language that reflects what the table actually shows. When percentages match, say that the conditional distributions are the same and that the variables show no association in the displayed data. When percentages are close, call the pattern near-independent or describe the distributions as similar, and report the differences that support that description. When percentages differ substantially, describe the association and identify which categories account for the difference.

“Independent” can sound like a claim about every person in a larger population. At this stage, use it carefully to describe the table’s pattern. Exact equality in a sample table is a descriptive result, not proof that the corresponding population distributions are exactly equal. Likewise, a small difference in a table does not by itself prove there is no association in a population. This tutorial is about describing the observed conditional distributions, not drawing an inferential conclusion.

The comparison does not depend on which variable is placed in rows or columns, but the conditional percentages do depend on which variable you condition on. For a question about preference by neighborhood, compare preference percentages within each neighborhood. You could instead calculate the neighborhood distribution within each preference category, but that answers a different conditional question. State your direction of comparison so the reader knows which groups form the denominators.

Common Mistakes and AP Exam Tips

  • Comparing raw counts instead of conditional percentages. Counts such as 48 and 72 look different, but the example’s corresponding percentages are both 60%. Divide by each group’s total before deciding whether the distributions differ.
  • Using the grand total as the denominator. A conditional percentage for one lunch period uses that period’s total, not the total number of students surveyed. Name the group and use its total.
  • Checking only one category in a multi-category outcome. Similar percentages for one category do not guarantee that all other categories match. Compare the full conditional distributions.
  • Calling close percentages exactly equal. A 2.5-percentage-point difference is small in the example, but it is not zero. Say “similar” or “near-independent,” and give the observed difference when useful.
  • Claiming that equal sample percentages prove population independence. The table describes the individuals represented. Do not make a population-level claim unless an appropriate inference procedure supports it.
  • Claiming causation from an association. A difference between experience groups does not show that experience caused the preferences. Report the observed pattern without inventing an explanation.
  • Leaving the context out. “The distributions differ” is incomplete on its own. Name the variables, the groups being compared, and the outcome categories that show the difference.

For a strong AP response, state the conditional percentages or clearly describe how they compare, identify the groups and categories, and use “no association,” “near-independent,” or “associated” with care. Keep the conclusion descriptive and tied to the table.

Key takeaway: Equal conditional distributions indicate no association in a two-way table. Similar but unequal distributions show a near-independent pattern, while noticeable differences show an association. Compare within-group percentages and explain the pattern in context.

Check Your Understanding

Use conditional percentages to decide whether each pattern suggests matching, near-matching, or differing distributions.

  1. A table has two groups of 50 people. In each group, 30 choose option A and 20 choose option B. What are the conditional distributions, and what do they indicate?
  2. In a survey, 24 of 60 morning-shift workers and 28 of 70 evening-shift workers choose a particular response. Find and compare the conditional percentages for that response. What can you say about this category’s percentage across the groups?
  3. Why can two groups have different counts in a category but still have equal conditional percentages?
  4. A three-category outcome has conditional distributions of 50%, 30%, and 20% in one group, and 50%, 25%, and 25% in another. Are the full distributions equal? Explain.
  5. Why would it be too strong to conclude that a small difference in a sample table proves there is no association in the larger population?