Tutorials › AP Statistics › Interpreting Association Versus Independence in a Table

Chi-square tests for categorical data · Tutorial 548 of 1000

Interpreting Association Versus Independence in a Table

Compare proportions within categories—not just counts—to describe whether two categorical variables appear related in a sample.

Intermediate 9 min read

What You'll Learn

  • Explain independence by comparing conditional proportions across categories.
  • Distinguish a sample pattern of association from evidence about a population.
  • Choose the appropriate row or column total as the denominator.
  • Recognize why unequal group sizes make raw counts misleading.
  • Describe patterns in two-way tables without claiming causation.

What a Two-Way Table Can Show

In “Stating Hypotheses for a Test of Independence,” you learned that the null hypothesis says two categorical variables are independent in the population, while the alternative says they are associated. Before carrying out a test, you can inspect the sample table for a pattern: do the proportions in one variable’s categories change across the categories of the other variable?

A two-way table records counts for combinations of categories. As explained in “Categorical Variables and Two-Way Tables,” each interior cell represents observations that fall into one row category and one column category. The counts are useful, but they do not by themselves show whether the variables appear related. A larger group will often have larger counts simply because it contains more observations.

Definition: Two categorical variables are independent when knowing the category of one variable does not change the distribution of the other variable. In a sample table, the variables appear independent when the conditional proportions are approximately the same across the categories being compared. They appear associated when those conditional proportions differ.

A conditional proportion is a proportion calculated within a specified category. For example, to describe the proportion of bus riders who arrive on time, divide the number of bus riders who arrive on time by the total number of bus riders. The bus-rider total is the denominator because the question is about the on-time proportion among bus riders.

For this kind of comparison, calculate the same conditional proportion for each category of the variable you are conditioning on. If the percentages are similar, the sample does not show a clear association in that comparison. If they differ, the sample shows a pattern of association. When a variable has more than two categories, compare the full set of conditional proportions, not just one cell.

$$ \text{Conditional proportion}= \frac{\text{count in the chosen cell}}{\text{total for the category being conditioned on}} $$

This is a descriptive comparison of the observed sample. It does not establish whether an apparent difference is larger than could reasonably occur by chance in a population. A chi-square test of independence addresses that inferential question; this tutorial focuses on reading the pattern before testing.

Compare Proportions Using a Consistent Denominator

For a table with row categories and column categories, one convenient approach is to compare the distribution across columns within each row. Divide each cell by its row total, then compare the resulting row percentages. You could instead compare column percentages by dividing each cell by its column total. Either orientation can be useful, as long as the denominator matches the question and you compare consistently.

The choice of orientation does not change whether the variables are independent: independence is a relationship between the two variables and is not directional. However, the numbers you calculate and the wording of your description depend on which variable’s distribution you are comparing across the categories of the other variable.

Reading rule: State which category the proportions are “among.” Then use that category’s total as the denominator in every comparison. Do not compare counts from groups of different sizes when the question is about relative proportions.

If the conditional proportions match exactly, the sample table displays no association pattern. In real data, they will often be close rather than identical. Conversely, a difference in conditional proportions describes an association in the sample, but not necessarily a meaningful or statistically convincing association in the population. The formal test and its conditions are needed for an inference conclusion.

Worked Example: Coffee and Reported Alertness

Worked Example: Coffee and Reported Alertness

A student group surveys 100 classmates, recording whether each student had coffee that morning and whether the student reports high or low alertness. The results are organized below.

High alertnessLow alertnessTotal
Had coffee362460
No coffee241640
Total6040100

To compare the alertness distributions among students who had coffee and students who did not, divide each cell by its row total. Among students who had coffee, the high-alertness proportion is \(36/60=0.60\), or 60%, and the low-alertness proportion is \(24/60=0.40\), or 40%.

Among students who did not have coffee, the high-alertness proportion is \(24/40=0.60\), or 60%, and the low-alertness proportion is \(16/40=0.40\), or 40%. The conditional distributions match exactly. The observed difference in high-alertness proportions is \(60\%-60\%=0\) percentage points.

Interpretation: In this sample, reported alertness has the same distribution among students who had coffee and students who did not. The table shows no apparent association between coffee status and reported alertness. This descriptive result does not prove that the variables are independent in the population.

Notice that the counts differ: there are 36 high-alertness students who had coffee and 24 who did not. Those counts alone might look like a difference, but the group totals are also different. Comparing conditional proportions accounts for those different group sizes.

Worked Example: Commute Method and Arriving on Time

Worked Example: Commute Method and Arriving on Time

A survey of 140 workers records their usual commute method and whether they usually arrive on time. The sample results are:

On timeLateTotal
Bus421860
Car641680
Total10634140

Compare the on-time proportions within the two commute groups. For bus commuters, the proportion usually on time is \(42/60=0.70\), or 70%. For car commuters, it is \(64/80=0.80\), or 80%.

The difference is \(80\%-70\%=10\) percentage points, with the on-time proportion higher among car commuters in this sample. The late proportions also differ: \(18/60=0.30\), or 30%, for bus commuters and \(16/80=0.20\), or 20%, for car commuters. These comparisons show an association pattern in the sample between commute method and usual arrival status.

Interpretation: The sample shows a higher proportion of workers who usually arrive on time among car commuters than among bus commuters. The table alone does not show whether this difference is convincing evidence of an association in the population. It also does not show that commute method caused the difference; other factors could be involved.

Here, the number of on-time car commuters (64) exceeds the number of on-time bus commuters (42), but that is not the most informative comparison: the car group is larger. The conditional proportions, 80% and 70%, describe the difference while accounting for the group totals.

Worked Example: Time of Park Visit and Main Activity

Worked Example: Time of Park Visit and Main Activity

A random sample of park visitors is classified by time of visit and main activity. Each person is counted once.

ExerciseSocial visitCommute through parkTotal
Morning3012850
Afternoon18221050
Total483418100

Because the row totals are both 50, divide each entry in a row by 50 to find the activity distribution for that visit time. In the morning, the proportions are \(30/50=0.60\) exercising, \(12/50=0.24\) on a social visit, and \(8/50=0.16\) commuting through the park.

In the afternoon, the proportions are \(18/50=0.36\) exercising, \(22/50=0.44\) on a social visit, and \(10/50=0.20\) commuting through the park. The distributions are not alike. For example, the exercise proportions differ by \(60\%-36\%=24\) percentage points, and the social-visit proportions differ by \(24\%-44\%=-20\) percentage points.

Interpretation: In this sample, the distribution of main activity differs between morning and afternoon visitors, so time of visit and main activity appear associated. The differences across several categories make the pattern easier to see than a comparison of one count alone. This describes the sampled visitors; a test would be needed to assess evidence about an association in the population.

The row totals happen to be equal, but that is not required for comparing conditional proportions. If the morning and afternoon totals differed, the same method would apply: divide each cell by its own row total before comparing the activity distributions.

From a Table Pattern to a Careful Conclusion

The table gives a useful first look, but keep three levels of statement separate. First, the conditional proportions describe the sample. Second, a visible difference suggests a sample association. Third, a statistical test can evaluate whether the sample provides convincing evidence of an association in the population, subject to the study design and test conditions.

In “Stating Hypotheses for a Test of Independence,” the population hypotheses were \(H_0\): the two variables are independent, and \(H_a\): the variables are associated. A table’s conditional proportions can help you understand the kind of pattern those hypotheses refer to, but they do not replace the hypotheses or the test.

This distinction also prevents overstatement. A sample can show different conditional proportions even when the population variables are independent, because samples vary. A test later in this module accounts for that variability. Likewise, a statistically convincing association would not by itself establish cause and effect. As in other observational studies, a relationship can reflect other variables or the way the sample was collected.

Key takeaway: To judge whether two categorical variables appear related in a table, compare conditional proportions using the relevant row or column totals. Similar distributions suggest no visible sample association; different distributions show a sample association pattern. Neither result alone proves independence, establishes population association, or demonstrates causation.

Common Mistakes and AP Exam Communication

A clear AP response names the categories being compared, shows the relevant denominators, and describes the pattern in context. Use cautious wording such as “in this sample” or “the variables appear associated.” Do not use a table pattern alone to claim that the population variables are independent or associated.

  • Comparing raw counts: A group with more observations may have more counts in every cell. Compare proportions within groups instead.
  • Using the grand total as the denominator: Dividing a cell by the full sample size gives its share of the entire sample, not the conditional proportion within a group. Use the total for the category you are conditioning on.
  • Changing denominators mid-comparison: To compare the outcome distribution across row groups, divide every cell by its own row total. State what the proportions are “among.”
  • Calling a sample pattern proof of population independence: Similar sample proportions only suggest no visible association. The formal test assesses evidence about the population.
  • Claiming causation from association: A two-way table describes how variables vary together. A relationship in an observational sample does not establish that one variable caused the other.
  • Ignoring categories in a multi-category variable: When a variable has several categories, compare the full conditional distribution. Looking at one cell may miss a different pattern elsewhere in the table.
AP Exam Tip: Write a sentence such as, “Among [category of one variable], [percentage] were in [category of the other variable], compared with [percentage] among [comparison category]. This suggests [little apparent association / an association pattern] in the sample.” Include the actual denominators in your calculations, and do not turn a descriptive comparison into a population or causal conclusion.

Check Your Understanding

For each situation, decide what conditional proportions to compare and describe what a difference or similarity would mean.

  1. A table records whether students participate in a school club and whether they attend a school event. Which denominator would you use to compare event attendance among club participants and nonparticipants?
  2. A table shows 45 positive responses among 75 people in Group A and 30 positive responses among 50 people in Group B. Calculate and compare the conditional proportions. What does the sample show?
  3. Explain why comparing the number of late arrivals in two commute groups is not necessarily as informative as comparing the proportions late within each group.
  4. If conditional distributions are similar across all categories in a sample table, what can you say about the sample pattern? What can you not conclude about the population?
  5. A sample table shows different proportions of preferred activity across age categories. Why does that pattern alone not establish that age causes the preference difference?