Tutorials › AP Statistics › Why Chi-Square Is Needed Beyond Two Proportions

Chi-square tests for categorical data · Tutorial 542 of 1000

Why Chi-Square Is Needed Beyond Two Proportions

See how a chi-square test evaluates an overall difference in categorical distributions across several groups—and why repeating two-proportion z-tests is not the same question.

Intermediate 10 min read

What You'll Learn

  • Identify when a comparison involves more than two groups or more than two outcome categories.
  • Explain why a set of pairwise two-proportion z-tests does not replace one overall test.
  • Calculate expected counts for a chi-square test comparing group distributions.
  • Use the chi-square statistic and degrees of freedom to assess an overall pattern.
  • Check the design and expected-count conditions before drawing a conclusion.
  • State what a significant chi-square result does—and does not—show.

When Two Proportions Are Not Enough

In “Categorical Variables and Two-Way Tables,” you learned how a table organizes counts for combinations of two categorical variables. The earlier tutorials on two-proportion z-tests used that kind of categorical data to compare two groups on a yes-or-no outcome. But what if there are three or more groups, or the outcome has several categories? A single two-proportion z-test cannot answer the overall question.

One tempting approach is to run a separate two-proportion z-test for every pair of groups. That creates many tests, each focused on only one pair. The chance of finding at least one apparently significant difference by chance can increase as the number of tests grows. More importantly, a collection of pairwise tests does not directly answer the broad question: Do the groups’ categorical outcome distributions differ overall?

A chi-square test provides an overall comparison using the counts in a two-way table. For comparing distributions across separate groups, this is called a chi-square test of homogeneity. It can compare a yes-or-no outcome across several groups, or compare an outcome with several categories across several groups.

Definition: A chi-square test of homogeneity evaluates whether the distribution of a categorical outcome is the same across several populations or groups. It compares the observed table counts with the counts expected if the groups have the same outcome distribution.

When the outcome has just two categories, such as “completed” and “did not complete,” comparing outcome distributions is equivalent to asking whether the groups’ success proportions are equal. The chi-square approach handles the whole comparison at once. With more than two outcome categories, it can compare the full distribution rather than reducing the question to one chosen category.

How the Overall Comparison Works

The chi-square test starts with the observed counts in the table. Under the null hypothesis that all groups have the same outcome distribution, it calculates an expected count for each cell. An expected count is the count predicted by that null model, not a count observed in the data.

For a table comparing groups in rows and outcome categories in columns, the expected count in a cell is found from its row total, column total, and grand total. This is the same calculation for every cell.

$$ \text{Expected count}=\frac{(\text{row total})(\text{column total})}{\text{grand total}} $$

The test statistic summarizes how far the observed counts are from their expected counts. Each cell contributes a squared difference divided by its expected count. Squaring means that differences in either direction add to the statistic; dividing by the expected count scales the difference relative to the size of that count.

$$ \chi^2=\sum \frac{(\text{observed count}-\text{expected count})^2}{\text{expected count}} $$

A larger \(\chi^2\) means the table’s observed counts depart more from the pattern expected under the null hypothesis. The p-value is the probability, assuming the null hypothesis is true, of getting a chi-square statistic at least as large as the observed one. The test uses the right tail of a chi-square distribution.

The degrees of freedom depend on the number of groups and outcome categories. If the table has \(r\) rows of groups and \(c\) columns of outcomes, the degrees of freedom are \((r-1)(c-1)\). This accounts for how many cell counts can vary once the margins are fixed.

Conditions: Check that the data come from random samples or an appropriate randomized experiment; that the observations are independent and each observation contributes to one cell; and that every expected count is at least 5. For random sampling without replacement, also check the 10% condition for each sample. In a randomized experiment, random assignment supports the comparison of treatment groups.

The expected-count condition is checked using the null model, not by looking only at the observed cell counts. If any expected count is too small, the chi-square approximation may not be appropriate. As in earlier proportion procedures, the design matters: random assignment can support a causal conclusion about assigned treatments, while random sampling can support generalization to the population sampled.

Worked Example: Comparing Completion Rates in Three Programs

Worked Example: Comparing Completion Rates in Three Programs

A fictional organization randomly assigns 360 participants to one of three support programs. Each participant is recorded once as either having completed a training course or not having completed it. The counts are:

ProgramCompletedDid not completeRow total
A5466120
B4278120
C2496120
Column total120240360

State. Let the outcome distribution for a program be the proportions of participants who completed and did not complete the course. The null hypothesis is that this distribution is the same for all three programs. The alternative hypothesis is that at least one program’s outcome distribution differs. Because the outcome has two categories, this is also a test of whether the three completion proportions are all equal.

Plan. A chi-square test of homogeneity is appropriate for comparing the categorical outcome across three groups. The experiment used random assignment, and each participant was assigned to one program and counted once, supporting independent observations. The expected counts under the null are calculated below; they must each be at least 5.

Do. For Program A, the expected number who completed is \((120)(120)/360=40\), and the expected number who did not complete is \((120)(240)/360=80\). The row totals for Programs B and C are also 120, so each of their expected counts is 40 completed and 80 not completed. All six expected counts are at least 5.

The chi-square statistic is:

$$ \chi^2= \frac{(54-40)^2}{40}+\frac{(66-80)^2}{80} +\frac{(42-40)^2}{40}+\frac{(78-80)^2}{80} +\frac{(24-40)^2}{40}+\frac{(96-80)^2}{80} =17.10 $$

There are \(r=3\) groups and \(c=2\) outcome categories, so the degrees of freedom are \((3-1)(2-1)=2\). For \(\chi^2=17.10\) with 2 degrees of freedom, the right-tail p-value is approximately \(0.0002\), rounded.

Conclude. Because the p-value is less than \(0.05\), reject the null hypothesis. The data provide convincing evidence that completion outcomes are not distributed the same way across all three programs. This overall test does not identify exactly which programs differ; it establishes evidence of a difference somewhere among the three.

More Than Two Outcome Categories

The chi-square test is not limited to yes-or-no outcomes. If participants can fall into several outcome categories, the table includes a column for each category. The test then assesses whether the overall distribution across those categories is the same in each group.

Worked Example: Preferences Across Three Interface Designs

Worked Example: Preferences Across Three Interface Designs

In a fictional experiment, 270 volunteers are randomly assigned to one of three versions of a scheduling app. Each person selects one feature they would most like to see improved: reminders, calendar view, or task list. The observed counts are:

App versionRemindersCalendar viewTask listRow total
Version 140302090
Version 230303090
Version 320304090
Column total909090270

The null hypothesis is that the distribution of selected features is the same for all three app versions. The alternative is that at least one version has a different distribution. A chi-square test of homogeneity matches this question because there are three groups and three outcome categories.

Random assignment supports comparing the versions, and each volunteer contributes one response to one cell. To find an expected count, use the Version 1 row total and the Reminders column total: \((90)(90)/270=30\). Every row and column has the same total, so all nine expected counts are 30, satisfying the requirement that each be at least 5.

The observed counts in the middle column all equal their expected counts, so those cells contribute zero. In the Reminders column, the deviations from expected are \(10, 0,\) and \(-10\); in the Task list column, they are \(-10, 0,\) and \(10\). Thus:

$$ \chi^2=6\left(\frac{(10)^2}{30}\right) =6\left(\frac{100}{30}\right) =20.00 $$

With 3 groups and 3 outcome categories, the degrees of freedom are \((3-1)(3-1)=4\). The right-tail p-value for \(\chi^2=20.00\) with 4 degrees of freedom is approximately \(0.0005\), rounded.

At the \(0.05\) significance level, reject the null hypothesis. The experiment provides convincing evidence that the distribution of feature preferences differs among the three app versions. The chi-square test does not say which feature or which pair of versions accounts for the overall difference. The table shows the pattern of observed counts, but the test conclusion remains an overall one.

Why Not Run Every Pairwise z-Test?

With three groups, there are three possible pairs to compare; with four groups, there are six. Each two-proportion z-test asks about one pair and, for a binary outcome, uses the earlier course’s methods for testing equality of two population proportions. Running many such tests creates multiple opportunities to reject a true null hypothesis. Even if every pair truly has the same proportion, a long collection of tests makes at least one small p-value more likely than a single test at the same significance level.

The chi-square test instead begins with one overall null hypothesis: all groups share the same categorical distribution. It produces one test statistic and one p-value for that overall question. This does not mean that the chi-square test reveals every detail. A small p-value signals that the observed table is inconsistent with the overall null model, but it does not, by itself, identify the differing groups or categories.

That distinction helps choose a method. If the question is specifically about two groups and a yes-or-no outcome, a two-proportion z-test directly addresses the comparison, as covered in the earlier tutorials. If the question compares three or more groups on a categorical outcome—or compares distributions with several categories—an overall chi-square test is the appropriate starting point. Do not treat a significant overall result as proof that every group differs from every other group.

Worked Example: Four Volunteer Teams and a Yes-or-No Outcome

Worked Example: Four Volunteer Teams and a Yes-or-No Outcome

A fictional community project randomly assigns 400 volunteers to one of four reminder schedules, with 100 volunteers per schedule. The outcome is whether each volunteer submits a report by the deadline.

ScheduleSubmittedDid not submitRow total
A6040100
B5050100
C4060100
D3070100
Column total180220400

The overall null hypothesis says that all four schedules have the same submission proportion. The alternative says that at least one schedule has a different submission proportion. Random assignment and one recorded outcome per volunteer support the design conditions. Under the null, Schedule A’s expected submitted count is \((100)(180)/400=45\), and its expected non-submitted count is \((100)(220)/400=55\). The same expected counts apply to each schedule, and all are at least 5.

For Schedule A, the cell contributions are \((60-45)^2/45=5.000\) and \((40-55)^2/55\approx4.091\). For Schedule B, they are \(25/45\approx0.556\) and \(25/55\approx0.455\). Schedule C has those same two contributions. Schedule D has the same contributions as Schedule A. Therefore:

$$ \chi^2=2(5.000+4.091)+2(0.556+0.455)\approx20.20 $$

There are 4 groups and 2 outcome categories, giving \((4-1)(2-1)=3\) degrees of freedom. The right-tail p-value is approximately \(0.0002\), rounded. At the \(0.05\) significance level, reject the null hypothesis: there is convincing evidence that submission proportions are not all equal across the four schedules. The test does not establish which particular schedules differ. A set of six pairwise z-tests would ask six separate questions; this chi-square test answers one overall question.

Common Mistakes and AP Exam Tips

  • Running pairwise tests instead of addressing the overall question: State whether the research question is about one pair or about all groups together. Use a chi-square test for the overall comparison across several groups.
  • Using observed counts as expected counts: Expected counts come from the null model and the table margins. Show the row-total-times-column-total calculation.
  • Checking observed counts against the expected-count rule: The condition concerns every expected count, not whether every observed cell count is at least 5.
  • Using a two-sided or left-tail area for the p-value: A chi-square statistic is nonnegative, and unusually large values are evidence against the null, so the p-value is a right-tail area.
  • Claiming the test identifies the differing groups: A significant result supports an overall difference in distributions. It does not tell you which group or category caused it.
  • Ignoring how the data were collected: Random assignment can support a causal conclusion about assigned conditions; random sampling can support generalization. The test statistic alone does not establish either.
AP Exam Tip: For a complete chi-square response, define the overall null and alternative, name the chi-square test of homogeneity, check the design and every expected count, report the statistic, degrees of freedom and right-tail p-value, then conclude in context. Say “at least one group’s distribution differs” rather than claiming every group differs.

Key Takeaway

A two-proportion z-test compares two groups on a binary outcome. When the question concerns three or more groups, or a categorical outcome with several categories, a chi-square test of homogeneity provides one overall comparison. It uses expected counts under a common-distribution null model and measures the departure of observed counts from those expectations.

Key takeaway: Use a chi-square test to compare categorical distributions across several groups, check that each expected count is at least 5, and interpret a small p-value as evidence of an overall difference—not as proof that every group differs or as identification of the specific differences.

Check Your Understanding

Answer each question using the ideas in this tutorial.

  1. A study compares a yes-or-no outcome across four groups. What is the overall question a chi-square test of homogeneity can answer?
  2. A table has a row total of 50, a column total of 40, and a grand total of 200. Calculate the expected count for that cell.
  3. A table compares three groups across four outcome categories. Find the degrees of freedom for the chi-square test.
  4. Why can running a separate two-proportion z-test for every pair of groups be a problem?
  5. A chi-square test gives a small p-value. What can you conclude about the group distributions, and what can you not conclude about specific pairs?