From a Test Decision to a Conclusion About Distributions
In “Making a Decision in a Chi-Square Test” and “Writing a Conclusion for a Test of Independence,” you learned how to use a p-value and \(\alpha\) to decide whether to reject \(H_0\). This tutorial applies that decision to a test of homogeneity, where separate samples or groups are compared on one categorical response.
As explained in “Stating Hypotheses for a Test of Homogeneity,” the null hypothesis says that the response distribution is the same across all populations or groups being compared. The alternative says that the distributions are not all the same. A conclusion should translate the decision into evidence about those distributions, naming the response and the populations in context.
The phrase at least one matters. A test of homogeneity evaluates the overall pattern across all the samples. A significant result supports a difference somewhere among the populations, but the test’s p-value alone does not identify which populations differ or which response categories account for the difference. Those more specific questions require careful additional analysis; they should not be smuggled into the overall conclusion.
Use the population or group labels that match the study design. For separate random samples, state the conclusion about the populations those samples represent. If the data come from treatment groups in an experiment, refer to those groups or treatments. As in earlier tutorials, the study design and conditions limit what can be concluded.
A Four-Part Conclusion
A complete conclusion links the numerical decision to the research question. Compare the p-value with the stated \(\alpha\), say whether you reject or fail to reject \(H_0\), and explain what that decision indicates about the response distributions in context. Do not write only “the result is significant” or “the distributions are the same.”
Use the significance level specified for the test. If the values are close, make the comparison using the unrounded calculator value.
Reject \(H_0\) when \(p\le\alpha\). Fail to reject \(H_0\) when \(p>\alpha\); do not say “accept \(H_0\).”
For rejection, describe convincing evidence that the response distribution differs across the named populations or groups. For failure to reject, say that the data do not provide convincing evidence of a difference.
Do not claim that all populations differ, that the distributions are proven equal, or that the test has identified a particular category as the source of a difference.
Before writing, check that the conclusion refers to the same response, groups, and alternative hypothesis as the test. Also make sure the study’s conditions have been addressed, as in “Random Condition for Chi-Square Procedures,” “Independence and 10 Percent Condition for Chi-Square,” and “Checking the Expected Count Condition for Chi-Square.” A well-phrased conclusion does not repair a problem with the data collection or conditions.
Worked Example: Evidence That Distributions Differ
Worked Example: Evidence That Distributions Differ
Suppose an invented study takes separate random samples of 60 residents from each of three districts. Each district has 1,200 residents. The response is the resident’s preferred yard-waste option: composting, curbside pickup, or drop-off. The observed counts are:
| District | Composting | Curbside pickup | Drop-off | Total |
|---|---|---|---|---|
| North | 26 | 20 | 14 | 60 |
| Central | 14 | 26 | 20 | 60 |
| South | 20 | 14 | 26 | 60 |
| Total | 60 | 60 | 60 | 180 |
State. The populations are residents of the North, Central, and South districts, and the response is preferred yard-waste option. The null hypothesis is that the distribution of preferred option is the same in all three districts. The alternative is that the distributions are not all the same.
Plan. Use a chi-square test of homogeneity because there are separate random samples from three populations, with one categorical response recorded for each resident. The samples are independent, and each resident contributes to exactly one response category. The 10% condition holds for each sample because \(60\le0.10(1200)=120\). Under the null model, all three districts share the pooled response distribution. Each expected count is \(60(60)/180=20\), so the expected-count condition is met.
Do. In each row, the observed counts differ from their expected counts of 20 by \(6, 0,\) and \(-6\), in some order. Thus, the chi-square statistic is:
As a check, each of the three rows contributes \(3.6\), and \(3(3.6)=10.8\). The degrees of freedom are \((3-1)(3-1)=4\). The upper-tail p-value is approximately \(0.0289\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.0289<0.05\), so reject \(H_0\). The samples provide convincing evidence that the distribution of preferred yard-waste option differs among residents of the North, Central, and South districts. This result indicates that at least one district’s distribution differs; it does not show that every pair of districts differs.
Worked Example: No Convincing Evidence of a Difference
Worked Example: No Convincing Evidence of a Difference
Suppose an invented survey takes separate random samples of 60 customers from each of three regions of a subscription service. Each region has 5,000 customers. The response is the customer’s preferred way to receive service updates: email, text, or app notification. The observed counts are:
| Region | Text | App notification | Total | |
|---|---|---|---|---|
| East | 22 | 19 | 19 | 60 |
| Central | 19 | 22 | 19 | 60 |
| West | 19 | 19 | 22 | 60 |
| Total | 60 | 60 | 60 | 180 |
State. The question is whether the distribution of preferred update method is the same among customers in the East, Central, and West regions. The null hypothesis says the three distributions are the same; the alternative says they are not all the same.
Plan. A chi-square test of homogeneity is appropriate because separate random samples from the three customer populations are classified by one categorical response. Each customer is counted once. The 10% condition holds for every sample because \(60\le0.10(5000)=500\). Each expected count is \(60(60)/180=20\), which meets the expected-count condition.
Do. In each row, the deviations from the expected counts are \(2,-1,-1\), in some order. Therefore:
As a check, each row contributes \(0.3\), so the three rows contribute \(3(0.3)=0.9\). The degrees of freedom are \((3-1)(3-1)=4\). The upper-tail p-value is approximately \(0.9246\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.9246>0.05\), so fail to reject \(H_0\). The samples do not provide convincing evidence that the distribution of preferred service-update method differs among customers in the East, Central, and West regions. This does not prove that the three distributions are identical.
Worked Example: Two Populations and Three Response Categories
Worked Example: Two Populations and Three Response Categories
Suppose an invented environmental survey takes separate random samples of 60 households from each of two watersheds, each with 2,000 households. Each household is classified by its primary method for conserving water: low-flow fixtures, outdoor watering changes, or neither. The observed counts are:
| Watershed | Low-flow fixtures | Outdoor watering changes | Neither | Total |
|---|---|---|---|---|
| Upper | 26 | 20 | 14 | 60 |
| Lower | 14 | 20 | 26 | 60 |
| Total | 40 | 40 | 40 | 120 |
State. The response is primary water-conservation method, and the populations are households in the Upper and Lower watersheds. The null hypothesis is that the response distribution is the same in both watersheds. The alternative is that the distributions differ.
Plan. Use a chi-square test of homogeneity because the data come from separate random samples, one from each watershed. Each household appears in one cell. The 10% condition holds in each population because \(60\le0.10(2000)=200\). Every expected count is \(60(40)/120=20\), so all are at least 5.
Do. The first and third response categories have deviations of \(6\) and \(-6\) in the Upper sample, with the opposite deviations in the Lower sample. The middle category matches its expected count in both samples:
A check is that each row contributes \(3.6\), giving \(2(3.6)=7.2\). The degrees of freedom are \((2-1)(3-1)=2\). The upper-tail p-value is approximately \(0.0273\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.0273<0.05\), so reject \(H_0\). The samples provide convincing evidence that the distribution of primary water-conservation method differs between households in the Upper and Lower watersheds. The conclusion concerns the overall distributions; the chi-square test alone does not establish which specific category accounts for the evidence.
Common Mistakes and AP Exam Tip
- Saying “all the distributions are different” after rejecting: Rejection supports that the distributions are not all the same—at least one differs. It does not establish that every population differs from every other population.
- Claiming the distributions are equal after failing to reject: State that the data do not provide convincing evidence of a difference. A large p-value does not prove the null hypothesis.
- Using vague context: Name the response and the populations or groups. “There is a difference” does not say what differs or where.
- Describing a particular cell as the test conclusion: The overall test evaluates whether the distributions differ. Do not use its p-value alone to claim a particular category or pair of groups is responsible.
- Mixing up independence and homogeneity: For homogeneity, describe differences in a response distribution across populations or groups. As covered in “Independence Versus Homogeneity: Choosing the Right Test,” the study design determines which test and conclusion language fit.
- Leaving out the decision comparison: Include the p-value, \(\alpha\), and the comparison that leads to the decision. Then give the contextual evidence statement.
For a full-credit AP-style conclusion, connect the decision directly to the alternative hypothesis and use the response and populations named in the question. For rejection, say there is convincing evidence that the distributions differ across the populations or groups, while avoiding a claim that all of them differ. For failure to reject, say there is not convincing evidence of a difference; do not claim the distributions have been proven identical.
Check Your Understanding
Use each decision and study context to write a conclusion about the response distributions.
- A test of homogeneity comparing three populations has \(p=0.031\) and \(\alpha=0.05\). What decision should be made, and what does the evidence support?
- A test comparing two regions has \(p=0.18\) and \(\alpha=0.05\). Write a correct contextual conclusion and name one claim to avoid.
- Why is “the distributions are all different” too strong a conclusion after rejecting the null hypothesis?
- What information about the context should a homogeneity conclusion name in addition to the p-value decision?
- Does failing to reject \(H_0\) prove that the response distributions are identical? Explain.