Tutorials › AP Statistics › Stating Hypotheses for a Test of Homogeneity

Chi-square tests for categorical data · Tutorial 547 of 1000

Stating Hypotheses for a Test of Homogeneity

Learn to describe matching response distributions across populations under the null and a difference in at least one distribution under the alternative.

Intermediate 9 min read

What You'll Learn

  • Identify the populations or groups and the single categorical response being compared.
  • State the null hypothesis as equal response distributions across all populations.
  • State the alternative as a difference in at least one population’s response distribution.
  • Express homogeneity hypotheses category by category or in words.
  • Avoid confusing equal population distributions with equal sample counts or sample sizes.
  • Distinguish a difference in distributions from a directional or causal claim.

From the Study Design to Hypotheses

In “Stating Hypotheses for a Test of Independence,” you learned to write hypotheses about whether two categorical variables are related in a population. This tutorial focuses on a different design: comparing separate populations or groups on one categorical response. For that question, a chi-square test of homogeneity asks whether the response has the same distribution across the populations.

A response distribution describes the proportions of individuals in each category of a categorical variable. For example, a survey might record one preferred option from three choices. The distribution consists of the population proportion choosing each option. Homogeneity means those category proportions are the same across the populations being compared.

Definition: For a chi-square test of homogeneity, \(H_0\) states that the distribution of one categorical response is the same across the populations or groups. \(H_a\) states that the response distribution differs for at least one population or group.

The hypotheses concern population distributions, not the counts in the samples. Sample counts can differ simply because sample sizes differ. Even with equal sample sizes, sample proportions may vary by chance. The test uses the sample data to assess evidence about the population claim.

As explained in “Chi-Square Test for Homogeneity: Purpose and Setting,” this test fits separate random samples or an appropriate randomized experiment, with one categorical response recorded for each individual. The question is whether that response’s distribution is consistent across the groups.

Writing the Null and Alternative

Begin by naming the groups being compared and the response categories. Then say what the response distribution would look like under each hypothesis. The null describes matching distributions across all groups; the alternative says they are not all the same.

Hypotheses in words: \(H_0\): The distribution of the response is the same across the specified populations or groups. \(H_a\): The distribution of the response differs for at least one of those populations or groups.

For a more formal statement, suppose there are \(k\) populations and the response has \(c\) categories. Let \(\pi_{ij}\) be the true proportion of individuals in population \(i\) who fall in response category \(j\). The null says that for each response category, its population proportion is equal across all populations:

$$ H_0:\ \pi_{1j}=\pi_{2j}=\cdots=\pi_{kj}\text{ for every response category }j $$

The alternative says the distributions are not all identical: at least one population has a different proportion in at least one response category.

$$ H_a:\ \text{the response distributions are not all the same across the populations} $$

The category-by-category form makes “same distribution” precise. If a response has categories such as “walk,” “bike,” and “bus,” the null says the proportion choosing “walk” is the same in each population, the proportion choosing “bike” is the same in each population, and likewise for “bus.” Because the category proportions within each population add to 1, a difference in one category’s proportion means the distributions differ.

In an AP response, contextual wording is usually the clearest way to state the hypotheses. You do not need to list a separate equation for every category if the words clearly say the distribution is the same across all populations under \(H_0\), and differs for at least one population under \(H_a\).

What “Differs” Means

The alternative is non-directional. It does not specify which population has a larger proportion in any category, or which categories account for a difference. It says only that the population distributions are not all the same. A chi-square test of homogeneity can detect a difference in many patterns, so the alternative does not identify one particular pattern in advance.

Do not write the alternative as “the groups have different sample counts.” Counts depend on sample sizes and are observed data, not the population claim. Instead, refer to the population response distributions. Similarly, “the proportions are not equal” is incomplete unless it is clear which response proportions and populations are meant.

The test’s name can be a reminder: homogeneity asks whether the response distribution is homogeneous, or the same, across the populations. In “Independence Versus Homogeneity: Choosing the Right Test,” the design distinction was the key: separate groups compared on one response points to homogeneity, while one sample classified by two variables points to independence. The tables may look alike, but the hypotheses reflect the study design.

Key distinction: In a test of homogeneity, \(H_0\) compares the population distribution of one response across groups. It does not say the sample counts must match, that the groups have equal sample sizes, or that observations were sampled independently.

Worked Example: Preferred Library Study Space

Worked Example: Preferred Library Study Space

A school librarian takes separate random samples of students from three high schools. Each student identifies one preferred library study space: a quiet desk, a group table, or a computer area. The question is whether the distribution of preferred study space is the same at the three schools.

Identify the populations and response. The populations are the students at each of the three high schools. The single categorical response is preferred library study space, with categories quiet desk, group table, and computer area.

State the null hypothesis. \(H_0\): The distribution of preferred library study space is the same among students at all three high schools. In particular, the population proportions preferring each of the three spaces are equal across the schools.

State the alternative hypothesis. \(H_a\): The distribution of preferred library study space differs for at least one of the three high schools. At least one school’s population proportions across the three spaces differ from the others.

The alternative does not predict which school will have more students preferring a quiet desk or any other option. It makes a general claim about a difference in distributions, not a directional claim about one response category.

Worked Example: Seasonal Water-Use Category

Worked Example: Seasonal Water-Use Category

A regional planner takes separate random samples of households in two towns. Each household is classified by its main outdoor water-use category: no regular outdoor use, garden watering, or lawn watering. The planner asks whether the distribution of water-use category is the same in the two towns.

Identify the populations and response. The populations are all households in Town A and all households in Town B. The response is each household’s main outdoor water-use category.

State the null hypothesis. \(H_0\): The distribution of main outdoor water-use category is the same among households in Town A and Town B. The population proportions in the no-regular-use, garden-watering, and lawn-watering categories match between the towns.

State the alternative hypothesis. \(H_a\): The distribution of main outdoor water-use category differs between the households in Town A and Town B. At least one category’s population proportion is different between the towns.

The hypotheses do not claim that the numbers of sampled households in the categories must be equal. If the two samples have different sizes, equal population distributions would generally correspond to different expected sample counts. The hypotheses are about proportions in the populations.

Worked Example: Randomized Reminder Methods

Worked Example: Randomized Reminder Methods

A community health program randomly assigns participants to one of three reminder methods for an appointment: a text message, an email, or a phone call. Each participant’s response is recorded as attended, rescheduled, or did not attend. The program asks whether the distribution of appointment outcomes is the same across the three reminder groups.

Identify the groups and response. The groups are participants assigned to the text, email, and phone-call reminders. The single categorical response is appointment outcome, with categories attended, rescheduled, and did not attend.

State the null hypothesis. \(H_0\): The distribution of appointment outcomes is the same for participants assigned to each of the three reminder methods.

State the alternative hypothesis. \(H_a\): The distribution of appointment outcomes differs for at least one reminder method.

These hypotheses compare the distributions across the assigned groups. Because the groups were created by random assignment, the study design may support a cause-and-effect conclusion if evidence of a difference is later found. That possibility comes from the design, not from the wording of \(H_0\) or \(H_a\) alone.

Common Mistakes and AP Exam Communication

A strong response makes the comparison explicit: which populations or groups, which categorical response, and whether its population distribution is the same or differs. Keep the null and alternative in context and make them match the homogeneity question.

  • Writing “the populations are equal”: This does not identify what is being compared. Say that the distribution of the named response is the same across the populations.
  • Writing “the sample distributions differ”: Sample distributions describe observed data. Hypotheses for an inference test make claims about the population distributions.
  • Comparing counts instead of distributions: Equal distributions mean equal category proportions, not necessarily equal counts. Group sample sizes may differ.
  • Leaving out “at least one” in the alternative: The alternative does not require every group to differ from every other group. Say the distributions differ for at least one population, or that they are not all the same.
  • Making the alternative directional: A homogeneity alternative does not say that one group has more of a particular response. It allows any pattern of distributional difference.
  • Confusing homogeneity with independence: Both procedures use categorical data, but the study designs and hypotheses differ. Use the design to identify the question, as in “Independence Versus Homogeneity: Choosing the Right Test.”
  • Claiming cause and effect from the hypotheses: The hypotheses do not establish causation. The study design determines whether a causal conclusion may be supported.
AP Exam Tip: Name the groups and the response first. Then state \(H_0\) as “the distribution of [response] is the same across [groups]” and \(H_a\) as “the distribution differs for at least one [group].” This wording identifies the population claim without confusing it with sample counts or a predicted direction.

Key Takeaway

A chi-square test of homogeneity evaluates whether one categorical response has the same distribution across separate populations or groups. The null says the category proportions match across all groups; the alternative says the distributions are not all identical.

Key takeaway: Define the populations or groups and the response. Write \(H_0\) as the same response distribution across all groups and \(H_a\) as a difference for at least one group. State population distributions—not sample counts, directions, or causes.

Check Your Understanding

For each setting, identify the groups and response, then write contextual hypotheses for a test of homogeneity.

  1. Separate random samples of customers at three grocery stores report which checkout option they prefer: cashier, self-checkout, or mobile checkout. The question asks whether preferences have the same distribution at the stores.
  2. Researchers randomly assign seedlings to one of three light conditions and record whether each seedling has low, medium, or high growth after a fixed period. State the null and alternative hypotheses.
  3. Separate random samples of residents from two neighborhoods report their main way of getting local news: print, online, radio, or television. Write the hypotheses in context.
  4. Explain why “the numbers choosing each option are equal” is not the correct null hypothesis when the group sample sizes differ.
  5. In a homogeneity test, what does “differs for at least one group” mean, and why is it not a directional prediction?