When Two Proportions Are Not Enough
In “Categorical Variables and Two-Way Tables,” you learned how a table organizes counts for combinations of two categorical variables. The earlier tutorials on two-proportion z-tests used that kind of categorical data to compare two groups on a yes-or-no outcome. But what if there are three or more groups, or the outcome has several categories? A single two-proportion z-test cannot answer the overall question.
One tempting approach is to run a separate two-proportion z-test for every pair of groups. That creates many tests, each focused on only one pair. The chance of finding at least one apparently significant difference by chance can increase as the number of tests grows. More importantly, a collection of pairwise tests does not directly answer the broad question: Do the groups’ categorical outcome distributions differ overall?
A chi-square test provides an overall comparison using the counts in a two-way table. For comparing distributions across separate groups, this is called a chi-square test of homogeneity. It can compare a yes-or-no outcome across several groups, or compare an outcome with several categories across several groups.
When the outcome has just two categories, such as “completed” and “did not complete,” comparing outcome distributions is equivalent to asking whether the groups’ success proportions are equal. The chi-square approach handles the whole comparison at once. With more than two outcome categories, it can compare the full distribution rather than reducing the question to one chosen category.
How the Overall Comparison Works
The chi-square test starts with the observed counts in the table. Under the null hypothesis that all groups have the same outcome distribution, it calculates an expected count for each cell. An expected count is the count predicted by that null model, not a count observed in the data.
For a table comparing groups in rows and outcome categories in columns, the expected count in a cell is found from its row total, column total, and grand total. This is the same calculation for every cell.
The test statistic summarizes how far the observed counts are from their expected counts. Each cell contributes a squared difference divided by its expected count. Squaring means that differences in either direction add to the statistic; dividing by the expected count scales the difference relative to the size of that count.
A larger \(\chi^2\) means the table’s observed counts depart more from the pattern expected under the null hypothesis. The p-value is the probability, assuming the null hypothesis is true, of getting a chi-square statistic at least as large as the observed one. The test uses the right tail of a chi-square distribution.
The degrees of freedom depend on the number of groups and outcome categories. If the table has \(r\) rows of groups and \(c\) columns of outcomes, the degrees of freedom are \((r-1)(c-1)\). This accounts for how many cell counts can vary once the margins are fixed.
The expected-count condition is checked using the null model, not by looking only at the observed cell counts. If any expected count is too small, the chi-square approximation may not be appropriate. As in earlier proportion procedures, the design matters: random assignment can support a causal conclusion about assigned treatments, while random sampling can support generalization to the population sampled.
Worked Example: Comparing Completion Rates in Three Programs
Worked Example: Comparing Completion Rates in Three Programs
A fictional organization randomly assigns 360 participants to one of three support programs. Each participant is recorded once as either having completed a training course or not having completed it. The counts are:
| Program | Completed | Did not complete | Row total |
|---|---|---|---|
| A | 54 | 66 | 120 |
| B | 42 | 78 | 120 |
| C | 24 | 96 | 120 |
| Column total | 120 | 240 | 360 |
State. Let the outcome distribution for a program be the proportions of participants who completed and did not complete the course. The null hypothesis is that this distribution is the same for all three programs. The alternative hypothesis is that at least one program’s outcome distribution differs. Because the outcome has two categories, this is also a test of whether the three completion proportions are all equal.
Plan. A chi-square test of homogeneity is appropriate for comparing the categorical outcome across three groups. The experiment used random assignment, and each participant was assigned to one program and counted once, supporting independent observations. The expected counts under the null are calculated below; they must each be at least 5.
Do. For Program A, the expected number who completed is \((120)(120)/360=40\), and the expected number who did not complete is \((120)(240)/360=80\). The row totals for Programs B and C are also 120, so each of their expected counts is 40 completed and 80 not completed. All six expected counts are at least 5.
The chi-square statistic is:
There are \(r=3\) groups and \(c=2\) outcome categories, so the degrees of freedom are \((3-1)(2-1)=2\). For \(\chi^2=17.10\) with 2 degrees of freedom, the right-tail p-value is approximately \(0.0002\), rounded.
Conclude. Because the p-value is less than \(0.05\), reject the null hypothesis. The data provide convincing evidence that completion outcomes are not distributed the same way across all three programs. This overall test does not identify exactly which programs differ; it establishes evidence of a difference somewhere among the three.
More Than Two Outcome Categories
The chi-square test is not limited to yes-or-no outcomes. If participants can fall into several outcome categories, the table includes a column for each category. The test then assesses whether the overall distribution across those categories is the same in each group.
Worked Example: Preferences Across Three Interface Designs
Worked Example: Preferences Across Three Interface Designs
In a fictional experiment, 270 volunteers are randomly assigned to one of three versions of a scheduling app. Each person selects one feature they would most like to see improved: reminders, calendar view, or task list. The observed counts are:
| App version | Reminders | Calendar view | Task list | Row total |
|---|---|---|---|---|
| Version 1 | 40 | 30 | 20 | 90 |
| Version 2 | 30 | 30 | 30 | 90 |
| Version 3 | 20 | 30 | 40 | 90 |
| Column total | 90 | 90 | 90 | 270 |
The null hypothesis is that the distribution of selected features is the same for all three app versions. The alternative is that at least one version has a different distribution. A chi-square test of homogeneity matches this question because there are three groups and three outcome categories.
Random assignment supports comparing the versions, and each volunteer contributes one response to one cell. To find an expected count, use the Version 1 row total and the Reminders column total: \((90)(90)/270=30\). Every row and column has the same total, so all nine expected counts are 30, satisfying the requirement that each be at least 5.
The observed counts in the middle column all equal their expected counts, so those cells contribute zero. In the Reminders column, the deviations from expected are \(10, 0,\) and \(-10\); in the Task list column, they are \(-10, 0,\) and \(10\). Thus:
With 3 groups and 3 outcome categories, the degrees of freedom are \((3-1)(3-1)=4\). The right-tail p-value for \(\chi^2=20.00\) with 4 degrees of freedom is approximately \(0.0005\), rounded.
At the \(0.05\) significance level, reject the null hypothesis. The experiment provides convincing evidence that the distribution of feature preferences differs among the three app versions. The chi-square test does not say which feature or which pair of versions accounts for the overall difference. The table shows the pattern of observed counts, but the test conclusion remains an overall one.
Why Not Run Every Pairwise z-Test?
With three groups, there are three possible pairs to compare; with four groups, there are six. Each two-proportion z-test asks about one pair and, for a binary outcome, uses the earlier course’s methods for testing equality of two population proportions. Running many such tests creates multiple opportunities to reject a true null hypothesis. Even if every pair truly has the same proportion, a long collection of tests makes at least one small p-value more likely than a single test at the same significance level.
The chi-square test instead begins with one overall null hypothesis: all groups share the same categorical distribution. It produces one test statistic and one p-value for that overall question. This does not mean that the chi-square test reveals every detail. A small p-value signals that the observed table is inconsistent with the overall null model, but it does not, by itself, identify the differing groups or categories.
That distinction helps choose a method. If the question is specifically about two groups and a yes-or-no outcome, a two-proportion z-test directly addresses the comparison, as covered in the earlier tutorials. If the question compares three or more groups on a categorical outcome—or compares distributions with several categories—an overall chi-square test is the appropriate starting point. Do not treat a significant overall result as proof that every group differs from every other group.
Worked Example: Four Volunteer Teams and a Yes-or-No Outcome
Worked Example: Four Volunteer Teams and a Yes-or-No Outcome
A fictional community project randomly assigns 400 volunteers to one of four reminder schedules, with 100 volunteers per schedule. The outcome is whether each volunteer submits a report by the deadline.
| Schedule | Submitted | Did not submit | Row total |
|---|---|---|---|
| A | 60 | 40 | 100 |
| B | 50 | 50 | 100 |
| C | 40 | 60 | 100 |
| D | 30 | 70 | 100 |
| Column total | 180 | 220 | 400 |
The overall null hypothesis says that all four schedules have the same submission proportion. The alternative says that at least one schedule has a different submission proportion. Random assignment and one recorded outcome per volunteer support the design conditions. Under the null, Schedule A’s expected submitted count is \((100)(180)/400=45\), and its expected non-submitted count is \((100)(220)/400=55\). The same expected counts apply to each schedule, and all are at least 5.
For Schedule A, the cell contributions are \((60-45)^2/45=5.000\) and \((40-55)^2/55\approx4.091\). For Schedule B, they are \(25/45\approx0.556\) and \(25/55\approx0.455\). Schedule C has those same two contributions. Schedule D has the same contributions as Schedule A. Therefore:
There are 4 groups and 2 outcome categories, giving \((4-1)(2-1)=3\) degrees of freedom. The right-tail p-value is approximately \(0.0002\), rounded. At the \(0.05\) significance level, reject the null hypothesis: there is convincing evidence that submission proportions are not all equal across the four schedules. The test does not establish which particular schedules differ. A set of six pairwise z-tests would ask six separate questions; this chi-square test answers one overall question.
Common Mistakes and AP Exam Tips
- Running pairwise tests instead of addressing the overall question: State whether the research question is about one pair or about all groups together. Use a chi-square test for the overall comparison across several groups.
- Using observed counts as expected counts: Expected counts come from the null model and the table margins. Show the row-total-times-column-total calculation.
- Checking observed counts against the expected-count rule: The condition concerns every expected count, not whether every observed cell count is at least 5.
- Using a two-sided or left-tail area for the p-value: A chi-square statistic is nonnegative, and unusually large values are evidence against the null, so the p-value is a right-tail area.
- Claiming the test identifies the differing groups: A significant result supports an overall difference in distributions. It does not tell you which group or category caused it.
- Ignoring how the data were collected: Random assignment can support a causal conclusion about assigned conditions; random sampling can support generalization. The test statistic alone does not establish either.
Key Takeaway
A two-proportion z-test compares two groups on a binary outcome. When the question concerns three or more groups, or a categorical outcome with several categories, a chi-square test of homogeneity provides one overall comparison. It uses expected counts under a common-distribution null model and measures the departure of observed counts from those expectations.
Check Your Understanding
Answer each question using the ideas in this tutorial.
- A study compares a yes-or-no outcome across four groups. What is the overall question a chi-square test of homogeneity can answer?
- A table has a row total of 50, a column total of 40, and a grand total of 200. Calculate the expected count for that cell.
- A table compares three groups across four outcome categories. Find the degrees of freedom for the chi-square test.
- Why can running a separate two-proportion z-test for every pair of groups be a problem?
- A chi-square test gives a small p-value. What can you conclude about the group distributions, and what can you not conclude about specific pairs?