Separate Groups, One Categorical Outcome
In “Chi-Square Test for Independence: Purpose and Setting,” the data came from one sample, with each individual classified by two categorical variables. A chi-square test of homogeneity answers a different kind of question: Do separate populations or groups have the same distribution for one categorical outcome?
For example, a public health team could take separate random samples from several towns and record each resident’s preferred way to receive appointment reminders. The outcome—reminder preference—is one categorical variable. The towns are the groups being compared. The question is whether the distribution of preferences is the same in all the towns.
The groups and the outcome categories play different roles. A group identifies the population or condition being compared, such as a town or a treatment. An outcome category is one possible response, such as text, email, or phone. Each individual belongs to one group and contributes a count to one category of that outcome.
The outcome must be categorical. It may have two categories, such as “completed” and “not completed,” or several categories, such as “text,” “email,” and “phone.” The test compares the whole distribution across groups, not just one selected category. A two-way table organizes the counts, with groups in one dimension and outcome categories in the other.
What the Hypotheses Ask
The null hypothesis for homogeneity says that the population distributions of the outcome are the same in all the groups. The alternative says that at least one group has a different distribution. “At least one” matters: rejecting the null does not mean every group differs from every other group.
For a binary outcome, this question can be expressed as whether the proportions in the outcome categories are the same across groups. With more than two categories, the comparison concerns all the category proportions together. For instance, when comparing preferences for three reminder methods across four towns, the question is not just whether one town has a higher text preference. It is whether the overall pattern of text, email, and phone preferences is the same in all four towns.
The expected-count calculation introduced in “Chi-Square Test for Independence: Purpose and Setting” also applies here. For each cell, the expected count is based on the corresponding row total, column total, and grand total. Under the null hypothesis, these counts describe what would be expected if the outcome distribution were the same across the groups. For a chi-square test, check that every expected count is at least 5.
Recognize the Study Design
A homogeneity setting can arise in two common ways. Researchers may take separate samples from different populations or naturally occurring groups. Or, in an experiment, they may randomly assign individuals to treatment groups and compare a categorical response. In both cases, the groups are established first, and researchers compare the distribution of one outcome across them.
This is different from selecting one sample and measuring two variables on each individual. That latter design asks whether the variables are associated, the setting of a chi-square test of independence. The layout of a two-way table alone does not tell you how the data were collected; identify the design and the research question before choosing the test.
In a study using separate random samples, check the 10% condition for each sample in relation to its population if sampling without replacement. In an experiment, check that each unit is assigned to one treatment group and contributes one categorical outcome. The random assignment supports the experiment’s design; it does not make multiple responses from the same individual independent.
Study design also determines what a conclusion can say. Random sampling can support generalizing to the populations represented by the samples. Random assignment can support a cause-and-effect conclusion about the treatments, if the experiment is appropriately conducted. Neither kind of randomization automatically provides both forms of support.
Worked Example: Reminder Preferences Across Towns
Worked Example: Reminder Preferences Across Towns
A fictional community health team takes separate random samples of residents from three towns. Each sampled resident chooses one preferred reminder method: online message, phone call, or mailed letter. The counts are:
| Town | Online message | Phone call | Mailed letter | Sample size |
|---|---|---|---|---|
| Maple | 60 | 40 | 20 | 120 |
| River | 40 | 35 | 25 | 100 |
| Hill | 20 | 25 | 35 | 80 |
| Column total | 120 | 100 | 80 | 300 |
Identify the setting. The team selected separate samples from three towns and recorded one categorical outcome, reminder preference, for each resident. The question is whether the distribution of reminder preferences is the same across the towns. This is a chi-square test of homogeneity.
State the hypotheses. The null hypothesis is that the distribution of reminder preferences is the same in Maple, River, and Hill. The alternative is that at least one town’s distribution differs.
Check the design and expected counts. The samples were selected randomly, and each resident contributes to exactly one cell. To support the 10% condition, each sample must be at most 10% of its town’s resident population. The prompt does not give the town population sizes, so that condition must be verified before proceeding. Assuming the samples are independent, the observations condition is reasonable. The expected counts are:
Each expected count is at least 5. The expected-count condition is satisfied. The design and research question fit the homogeneity setting; these checks do not, on their own, determine whether the data provide convincing evidence of a difference.
Describe the scope. If a test found convincing evidence that the distributions differ, the random samples could support a conclusion about differences among the represented town populations. The result would not establish why preferences differ.
Worked Example: Responses Across Treatment Groups
Worked Example: Responses Across Treatment Groups
In a fictional experiment, 240 volunteers are randomly assigned in equal numbers to one of three reminder formats. After a week, each volunteer is classified as having completed the requested task, partially completed it, or not completed it. The counts are:
| Reminder format | Completed | Partially completed | Not completed | Group total |
|---|---|---|---|---|
| Short message | 36 | 24 | 20 | 80 |
| Detailed message | 30 | 25 | 25 | 80 |
| No reminder | 24 | 26 | 30 | 80 |
| Column total | 90 | 75 | 75 | 240 |
Identify the setting. The experiment has three treatment groups, and the researchers compare one categorical outcome, task-completion status. The question is whether the outcome distribution is the same across the treatment groups. This is a chi-square test of homogeneity.
State the hypotheses. The null hypothesis is that the distribution of task-completion status is the same for all three reminder formats. The alternative is that at least one format has a different distribution.
Check the design and expected counts. Volunteers were randomly assigned to one treatment group, and each volunteer contributes one outcome category. Assuming one volunteer’s response does not determine another’s response, the observations are independent. The expected counts in every treatment group are \(30\) completed, \(25\) partially completed, and \(25\) not completed. For example, the expected count of completed tasks in the short-message group is \((80)(90)/240=30\). All expected counts are at least 5.
Explain what a conclusion could support. If the test provided convincing evidence that the distributions differ, random assignment would support a conclusion that reminder format affects the distribution of task-completion outcomes for these experimental units. Because the volunteers were not described as a random sample from a broader population, the experiment alone does not establish that the result generalizes to all people who might receive reminders.
Worked Example: A Homogeneity Question With a Sampling Limitation
Worked Example: A Homogeneity Question With a Sampling Limitation
A fictional recreation center compares preferred activity—swimming, exercise classes, or walking—among people who voluntarily answer an online survey. The survey collects 50 responses from each of three locations:
| Location | Swimming | Exercise classes | Walking | Responses |
|---|---|---|---|---|
| East | 25 | 15 | 10 | 50 |
| West | 20 | 14 | 16 | 50 |
| Central | 15 | 16 | 19 | 50 |
| Column total | 60 | 45 | 45 | 150 |
Identify the setting. The survey compares separate groups, the three locations, on one categorical outcome, preferred activity. The research question about whether the activity distributions are the same across locations is a homogeneity question.
State the hypotheses. The null hypothesis is that preferred-activity distributions are the same at all three locations. The alternative is that at least one location has a different distribution.
Check the expected counts and design. Each location has 50 responses. The expected counts in each location are \(20\) for swimming, \(15\) for exercise classes, and \(15\) for walking. For example, the expected swimming count at East is \((50)(60)/150=20\). All expected counts meet the minimum of 5. However, voluntary response is not random sampling. The responses may differ systematically from those of people who did not choose to participate, so the survey does not support generalizing an inferential conclusion to all users of the locations. The expected-count condition does not fix this design limitation.
Explain what the test can and cannot resolve. The table has the structure for comparing outcome distributions, but the sampling method limits the strength and reach of an inference. A test calculation cannot turn a voluntary-response sample into a random sample. The center should describe the survey results cautiously and avoid claiming that they represent every user at each location.
Common Mistakes and AP Exam Communication
A strong response identifies the groups, names the single categorical outcome, and ties the test choice to the research question. It also describes the sampling or assignment design and checks the conditions that matter. Avoid choosing a test just because the data fit in a two-way table.
- Calling the outcome categories the groups: In the reminder example, towns are the groups; online message, phone call, and mailed letter are outcome categories.
- Describing only one category: Homogeneity compares the distribution across all categories. A difference in one category may be part of a broader distributional difference, but the question is not limited to that category.
- Saying that every group must differ: The alternative hypothesis is that at least one group’s distribution differs. A significant result does not establish that all groups differ from one another.
- Confusing expected counts with observed counts: The Large Counts condition concerns expected counts under the null hypothesis. Large observed counts do not substitute for checking expected counts.
- Overstating what the design supports: Random samples can support generalization to represented populations. Random assignment can support a causal conclusion about treatment effects. Neither should be claimed without the corresponding design.
Key Takeaway
A chi-square test of homogeneity is appropriate when researchers compare the distribution of one categorical outcome across separate samples or groups. The groups might be different populations or experimental treatments. The study design determines what population or causal conclusions are justified.
Check Your Understanding
For each situation, identify the groups, the categorical outcome, and what a test of homogeneity would ask.
- Separate random samples of residents from four districts are asked whether they prefer a website, phone app, or printed guide. What are the groups and the outcome?
- A researcher randomly assigns participants to two exercise plans and records whether each participant meets, partly meets, or does not meet a weekly goal. What does the null hypothesis say?
- Why does the alternative hypothesis say “at least one group’s distribution differs” rather than “all groups differ”?
- A table has expected counts of 4, 12, and 18 in one row. Does it satisfy the expected-count condition? Explain.
- A voluntary online survey has large expected counts. Does that alone justify generalizing results to the full population? Why or why not?