Comparing Response Distributions Across Groups
In “Chi-Square Test of Independence Full Worked Example,” the data came from one sample in which every person was classified by two categorical variables. This tutorial considers a different design: researchers compare separate groups on the distribution of one categorical response. A chi-square test of homogeneity asks whether that response distribution is the same across the groups.
The method works for two or more groups and two or more response categories. Here, we focus on a full four-step test comparing three groups. The groups might be separate random samples from different populations, or groups created by randomly assigning experimental units to treatments. The study design determines how to check conditions and what the conclusion can say.
The Four Steps for a Homogeneity Test
As in “Stating Hypotheses for a Test of Homogeneity,” state the hypotheses about the populations or groups and the response variable. The alternative does not claim that every group differs from every other group. It says only that at least one group’s response distribution differs.
Under the null hypothesis, all groups share the same response distribution. For each group-and-response cell, the expected count is the group total multiplied by the pooled total for that response category, then divided by the grand total. Equivalently, multiply the group total by the pooled proportion for that response category.
The chi-square statistic adds the contributions from every cell. As in “Degrees of Freedom for a Two-Way Table,” for a table with \(r\) group categories and \(c\) response categories, \(df=(r-1)(c-1)\). The p-value is the upper-tail probability of a chi-square statistic at least as large as the observed statistic, assuming the null hypothesis is true.
Identify the populations or groups and the categorical response. State that the distributions are the same under \(H_0\), and that at least one differs under \(H_a\).
Name a chi-square test of homogeneity. Check random sampling or random assignment, independent observations, the 10% condition when sampling without replacement, and the expected-count condition.
Find each expected count, calculate \(X^2\) and the degrees of freedom, then find the upper-tail p-value.
Compare the p-value with the significance level and state what the evidence says about the response distributions in context.
The expected-count condition requires every expected count to be at least 5. For separate random samples, check the 10% condition for each sample if sampling without replacement. In a randomized experiment, check that experimental units were randomly assigned and that each unit contributes to only one cell; a sampling 10% check is not the relevant condition for assignment.
Worked Examples: Three Groups and One Response
Worked Example: Three Neighborhoods and Preferred Cooling Method
Suppose researchers take separate random samples of 90 households from each of three neighborhoods. Each household identifies its preferred cooling method: fan, air conditioner, or open window. Each neighborhood has 5,000 households. These invented counts summarize the samples:
| Neighborhood | Fan | Air conditioner | Open window | Total |
|---|---|---|---|---|
| North | 45 | 30 | 15 | 90 |
| Central | 30 | 30 | 30 | 90 |
| South | 15 | 30 | 45 | 90 |
| Total | 90 | 90 | 90 | 270 |
State. The populations are households in the North, Central, and South neighborhoods. The response variable is preferred cooling method. \(H_0\): the distribution of preferred cooling method is the same in all three neighborhoods. \(H_a\): the distribution differs for at least one neighborhood.
Plan. Use a chi-square test of homogeneity because three separate samples are compared on one categorical response. The researchers randomly sampled households in each neighborhood, satisfying the random condition. Each household is counted once, and the samples are separate. For each neighborhood, \(90\leq0.10(5{,}000)=500\), so the 10% condition is met. We will check the expected counts before calculating the test statistic.
Do. Each group total is 90, each pooled response total is 90, and the grand total is 270. For example, the expected count for North and fan is \(90(90)/270=30\). The same calculation gives 30 for every cell:
All expected counts are at least 5, so the Large Counts condition is met. The four nonzero observed-minus-expected differences are \(15,-15,-15,15\); the other five differences are 0. Thus:
There are \(r=3\) groups and \(c=3\) response categories, so \(df=(3-1)(3-1)=4\). The upper-tail p-value is approximately \(0.0000049\), rounded to four significant digits. It is less than \(0.0001\).
Conclude. At \(\alpha=0.05\), the p-value is less than 0.05, so reject \(H_0\). The samples provide convincing evidence that the distribution of preferred cooling method is not the same across the three neighborhoods. The test identifies an overall difference; it does not, by itself, establish which specific neighborhoods or methods account for that difference.
Worked Example: Three Study Spaces and Preferred Background Sound
Suppose separate random samples of 60 students are taken from each of three colleges. Each student chooses one preferred study-space sound: quiet, instrumental music, or ambient noise. Each college has 1,500 students. The following invented counts are recorded:
| College | Quiet | Instrumental music | Ambient noise | Total |
|---|---|---|---|---|
| Lake | 22 | 18 | 20 | 60 |
| Hill | 18 | 22 | 20 | 60 |
| Valley | 20 | 20 | 20 | 60 |
| Total | 60 | 60 | 60 | 180 |
State. \(H_0\): the distribution of preferred background sound is the same at the three colleges. \(H_a\): at least one college has a different distribution.
Plan. A chi-square test of homogeneity is appropriate because separate samples are compared on one categorical response. Each sample is random, and each student appears once. The 10% check is \(60\leq0.10(1{,}500)=150\) for each college, so it is met. Expected counts are calculated next.
Do. Each group total and each pooled response total is 60, with a grand total of 180. Every expected count is \(60(60)/180=20\), satisfying the expected-count condition. Four cells differ from expectation by 2 in absolute value; the other five have no difference. Therefore:
The degrees of freedom are \((3-1)(3-1)=4\). The upper-tail p-value is approximately \(0.9384\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.9384>0.05\), so fail to reject \(H_0\). These samples do not provide convincing evidence that preferred background-sound distributions differ among the three colleges. This does not prove that the distributions are identical; it means the observed differences are not strong evidence against the null model.
Worked Example: Three Reminder Designs and Task Completion
Suppose 240 volunteers test one of three reminder designs for a daily planning task. Researchers randomly assign 80 volunteers to each design and record whether each volunteer completes the task that day. The invented results are:
| Reminder design | Completed | Not completed | Total |
|---|---|---|---|
| Text | 50 | 30 | 80 |
| Calendar alert | 40 | 40 | 80 |
| Printed card | 30 | 50 | 80 |
| Total | 120 | 120 | 240 |
State. The groups are the volunteers assigned to the text, calendar-alert, and printed-card reminders. The response is task completion. \(H_0\): the task-completion distribution is the same for all three reminder designs. \(H_a\): at least one design has a different task-completion distribution.
Plan. Use a chi-square test of homogeneity because the randomized experiment compares the distribution of one categorical response across treatment groups. Volunteers were randomly assigned, each volunteer contributes one response to one cell, and the groups are independent. Expected counts will be checked below. Since the groups were formed by random assignment rather than selected as samples without replacement from a finite population, a sampling 10% condition is not applicable.
Do. The expected count for each group’s completed cell is \(80(120)/240=40\). The expected count for each not-completed cell is also \(80(120)/240=40\). All expected counts meet the Large Counts condition. Four cells differ from expectation by 10 in absolute value, and the other two have difference 0:
The degrees of freedom are \((3-1)(2-1)=2\). The upper-tail p-value is approximately \(0.0067\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.0067<0.05\), so reject \(H_0\). The experiment provides convincing evidence that task-completion distributions differ among the three reminder designs for the volunteers who participated. Because volunteers were randomly assigned, the experiment supports a causal comparison of the assigned designs for these volunteers. Random assignment alone does not justify generalizing that comparison to other people; doing so would require an appropriate basis, such as random sampling from the target population.
Common Mistakes and AP Exam Tip
- Using the wrong expected-count formula: For homogeneity, use each group total and the pooled total for the response category. Do not calculate expected counts as though the null hypothesis were independence in one sample.
- Writing hypotheses about every pair of groups: The alternative is that at least one group’s distribution differs. It does not assert that all groups differ from one another.
- Checking observed counts instead of expected counts: The Large Counts condition concerns every expected count under the null model. Calculate those counts and verify that each is at least 5.
- Forgetting the design-specific conditions: For separate samples, identify the random samples, independent observations, and 10% check when sampling without replacement. For an experiment, describe random assignment and independent observations.
- Claiming a large p-value proves equal distributions: “Fail to reject” means the data do not provide convincing evidence of a difference. It is not proof that the distributions are identical.
- Overstating the scope of a randomized experiment: Random assignment supports a causal comparison for the experimental units that participated. It does not, by itself, make those units representative of a broader population.
For full credit, connect the test to the study design, show how the expected counts meet the condition, and state the decision in context. A significant result is an overall finding: without further analysis, do not claim that a particular category or pair of groups is responsible.
Check Your Understanding
Use the group structure, response categories, and study design to answer each question.
- Researchers take separate random samples from three towns and record each resident’s preferred public-transport option. Which chi-square procedure fits, and what does its null hypothesis say?
- A study compares three groups and four response categories. Find the degrees of freedom for the chi-square test.
- For one cell, the group total is 75, the pooled response-category total is 120, and the grand total is 300. Find the expected count.
- A homogeneity test has \(p=0.031\) and \(\alpha=0.05\). State the decision and what it means about the group distributions.
- In a randomized experiment with volunteers, what does random assignment allow researchers to conclude, and what does it not establish about a broader population?