Tutorials › AP Statistics › Checking Independence in a Two-Way Table

Independence and unions · Tutorial 285 of 1000

Checking Independence in a Two-Way Table

Compare within-row percentages with overall column percentages to see whether knowing a row category changes the distribution across column categories.

Beginner 8 min read

What You'll Learn

  • Calculate conditional row percentages using each row total as the denominator.
  • Calculate overall column percentages using the grand total.
  • Compare corresponding percentages across every row and column category.
  • Describe whether the table’s conditional distributions match the overall distribution.
  • Avoid confusing a descriptive pattern in a table with a conclusion about a larger population.

Compare Within-Row Percentages with Overall Column Percentages

In Testing Independence with the Product Rule, you checked whether a joint probability equals the product of two marginal probabilities. A two-way table offers another way to look for the same independence pattern: compare the conditional percentages within each row with the overall percentages in the columns.

Suppose the rows represent groups and the columns represent possible outcomes or categories. A row percentage answers a conditional question: among the individuals in this row group, what percentage belongs to each column category? An overall column percentage answers a marginal question: in the full group, what percentage belongs to each column category? If the variables behave independently in the distribution represented by the table, knowing the row category does not change the column distribution. The row percentages should therefore match the corresponding overall column percentages.

Definition: For a two-way table with groups in rows and categories in columns, the row percentage for a cell is the cell count divided by its row total. The overall column percentage is the column total divided by the grand total. Independence is indicated when the conditional row percentages for each column category match the corresponding overall column percentages in every row.

This is a descriptive check of the table, not a significance test. It tells you whether the observed conditional distributions match the overall distribution in the group summarized by the table. As in the earlier tutorial on The Definition \(P(A\mid B)=P(A)\), independence means that the probability of an outcome does not change when the condition is known. Here, a row category supplies the condition, and the column categories supply the outcomes being compared.

Which Denominator Should You Use?

The key is to keep the two kinds of percentages distinct. To calculate a row percentage, divide by the total at the end of that row. To calculate an overall column percentage, divide by the grand total. Those different denominators answer different questions, and that is why comparing the results is useful.

$$ \text{Row percentage} =\frac{\text{cell count}}{\text{row total}}\times 100\% $$
$$ \text{Overall column percentage} =\frac{\text{column total}}{\text{grand total}}\times 100\% $$

For example, if a row is “weekday visitors,” the percentage in its “chooses a digital receipt” cell is the percentage of weekday visitors who chose a digital receipt. In probability notation, it is a conditional probability such as \(P(\text{digital receipt}\mid\text{weekday visitor})\). By contrast, the overall digital-receipt percentage is \(P(\text{digital receipt})\), calculated from everyone in the table.

If each row’s distribution across the columns matches the overall column distribution, the row category does not appear to change the distribution of the column variable. If one or more row distributions differ, the variables do not behave independently in the table. Check all the column categories, not just one. In a two-column table, the two percentages in a row must add to 100%, so a difference in one category necessarily appears as an opposite difference in the other. With more columns, each category contributes information about the overall distribution.

1
Identify what the rows and columns represent.
Name the row groups and the column categories so the conditional comparison is clear.
2
Calculate the row percentages.
For every row, divide each cell count by that row’s total. Each row’s percentages should sum to 100%, allowing for rounding.
3
Calculate the overall column percentages.
Divide each column total by the grand total. These percentages describe the column distribution in the full group.
4
Compare and describe the pattern.
Compare each row percentage with the overall percentage for the same column category. State whether the distributions match and identify the variables and group in context.

The row percentages are also called conditional distributions because each one is calculated within a specified row group. The overall column percentages form the marginal distribution of the column variable. This vocabulary helps you explain exactly what you compared: conditional distributions by row versus the marginal distribution for the whole table.

Worked Example: Visitor Type and Receipt Choice

Worked Example: Visitor Type and Receipt Choice

A fictional shop records whether 200 customers visited on a weekday or a weekend and whether they chose a digital or printed receipt. Let \(A\) represent visitor type and \(B\) represent receipt choice. The table summarizes the customers.

Digital receiptPrinted receiptTotal
Weekday364480
Weekend5466120
Total90110200

State. We will check whether receipt choice behaves independently of visitor type in these 200 customers by comparing the receipt distribution within each visitor-type row with the overall receipt distribution.

Plan. Divide each cell count by its row total to find conditional row percentages. Then divide each column total by the grand total to find overall column percentages. Compare the percentages for digital receipts and printed receipts.

Do. Among weekday visitors, \(36/80=0.45=45\%\) chose a digital receipt, and \(44/80=0.55=55\%\) chose a printed receipt. Among weekend visitors, \(54/120=0.45=45\%\) chose a digital receipt, and \(66/120=0.55=55\%\) chose a printed receipt. Overall, \(90/200=0.45=45\%\) chose a digital receipt, and \(110/200=0.55=55\%\) chose a printed receipt. Thus, both row distributions match the overall column distribution.

Conclude. In this group of 200 customers, the receipt-choice distribution is the same for weekday and weekend visitors. The two variables behave independently in the table.

Worked Example: Reusable Bottle Use and Travel Method

Worked Example: Reusable Bottle Use and Travel Method

A fictional community survey records whether 100 participants usually travel to a recreation center by bicycle or another method and whether they bring a reusable water bottle. Let \(A\) represent travel method and \(B\) represent bottle use.

Brings a reusable bottleDoes not bring oneTotal
Travels by bicycle401050
Travels another way203050
Total6040100

State. We will compare bottle use within each travel-method group with the overall bottle-use percentages to check whether the variables behave independently among these 100 participants.

Plan. Use each row total of 50 to calculate the two conditional distributions. Use the grand total of 100 to calculate the overall percentages for bringing and not bringing a bottle.

Do. Among participants who travel by bicycle, \(40/50=0.80=80\%\) bring a reusable bottle and \(10/50=0.20=20\%\) do not. Among participants who travel another way, \(20/50=0.40=40\%\) bring a bottle and \(30/50=0.60=60\%\) do not. Overall, \(60/100=0.60=60\%\) bring a bottle and \(40/100=0.40=40\%\) do not. The bicycle group’s row percentages, 80% and 20%, do not match the overall 60% and 40%. The other-travel group’s 40% and 60% also do not match.

Conclude. Reusable-bottle use and travel method do not behave independently in these 100 participants: the within-group percentages differ from the overall percentages. In this table, participants who travel by bicycle are more likely to bring a reusable bottle than participants who travel another way. This describes the association in the surveyed group; the table alone does not establish why the pattern occurs or whether it applies to a larger population.

Worked Example: Alert Preferences Across Three Groups

Worked Example: Alert Preferences Across Three Groups

A fictional town survey records residents’ preferred format for community alerts and their usual language for receiving local information. The rows represent language groups, and the columns represent alert preferences. We will compare each group’s preferences with the overall percentages for all 150 residents.

Text messageEmailRecorded callTotal
Group A3618660
Group B20201050
Group C14121440
Total705030150

The overall percentages are \(70/150=0.4667\), or about 46.7%, for text messages; \(50/150=0.3333\), or about 33.3%, for email; and \(30/150=0.20\), or 20%, for recorded calls. The row percentages are:

  • Group A: \(36/60=60\%\) prefer text messages, \(18/60=30\%\) prefer email, and \(6/60=10\%\) prefer recorded calls.
  • Group B: \(20/50=40\%\) prefer text messages, \(20/50=40\%\) prefer email, and \(10/50=20\%\) prefer recorded calls.
  • Group C: \(14/40=35\%\) prefer text messages, \(12/40=30\%\) prefer email, and \(14/40=35\%\) prefer recorded calls.

The three row distributions do not match the overall distribution of about 46.7%, 33.3%, and 20%. For example, 35% of Group C prefers recorded calls, compared with 20% overall, while only 10% of Group A prefers recorded calls. The comparison across all three categories shows that the pattern is not limited to one isolated cell.

Therefore, alert preference and language group do not behave independently in this table of 150 residents. The conditional distribution of alert preferences differs across the language groups. This is a description of the fictional survey table, not a claim that language causes a particular preference.

Common Mistakes and AP Exam Tips

  • Using the grand total for a row percentage. A row percentage is conditional on membership in that row, so its denominator is the row total. Dividing by the grand total instead gives a joint percentage from the full table.
  • Using a row total for an overall column percentage. The overall percentage uses the column total divided by the grand total. It describes everyone in the table, not just one row group.
  • Comparing unrelated categories. Compare the row percentage for “digital receipt” with the overall digital-receipt percentage, not with the overall printed-receipt percentage.
  • Checking only one row or one column. Independence requires the conditional distributions to match the overall distribution across every row and column category. A match in one cell alone is not enough.
  • Calling any small difference proof of dependence in a population. A table of observed data describes its group. If it is a sample, this descriptive comparison alone does not establish what is true for a wider population.
  • Writing only “they are dependent.” Explain what differs. A full-credit response identifies the variables and group, states that the conditional row percentages do not match the overall column percentages, and describes a relevant difference in context.
AP Exam Tip: Make the denominator visible in your work. Write a conditional percentage as a cell count divided by its row total, and an overall percentage as a column total divided by the grand total. Then compare matching categories and give a conclusion in context.

Key Takeaway

To check independence in a two-way table with groups in rows, compare each row’s conditional percentages with the corresponding overall column percentages. Use the row total for each conditional percentage and the grand total for each overall percentage. Matching distributions indicate that the variables behave independently in the table; differing distributions indicate an association in the group summarized.

Key takeaway: Independence means the column distribution does not change from row to row. Compare all conditional row percentages with the overall column percentages, and describe the pattern in context.

Check Your Understanding

For each question, identify the relevant denominator, compare the conditional row percentages with the overall column percentages, and state what the comparison shows.

  1. A table has two rows, each with a total of 40. In the first row, 12 people choose option X; in the second row, 12 people choose option X. What percentage in each row chooses X, and what is the overall percentage if the grand total is 80?
  2. A group of 120 participants is divided into two equal row groups. In the first row, 45 choose category A; in the second row, 30 choose category A. Find both row percentages and the overall category-A percentage. Do the row percentages match the overall percentage?
  3. Explain why a cell count divided by its row total answers a different question from the same cell count divided by the grand total.
  4. In a table with three column categories, why should you compare the full row distribution with the overall column distribution instead of checking just one cell?
  5. A sample table shows different conditional row percentages and overall column percentages. State a suitable conclusion about the sample and one claim this comparison alone cannot establish.