What a Chi-Square Result Can—and Cannot—Tell You
In “Writing a Conclusion for a Test of Homogeneity,” you learned to translate a test decision into evidence about whether response distributions differ across groups. This tutorial focuses on the boundaries of that interpretation. A chi-square test evaluates an overall pattern in a table; its p-value does not, by itself, say which categories differ or explain why the pattern occurred.
A small p-value can provide convincing evidence against the null hypothesis. For a test of homogeneity, that means the data support that the response distributions are not all the same. For a test of independence, it means the data support an association between the two categorical variables. Neither statement automatically identifies a particular cell as responsible, establishes a cause, or describes the size and importance of a difference.
This distinction is important because the chi-square statistic combines contributions from every cell. As covered in “Cell Contributions to Chi-Square” and “Standardized Residuals for Interpretation,” cell contributions and standardized residuals can help describe where the observed table departs most from the null model. They are useful diagnostic clues, but they do not turn the overall p-value into a formal test of a particular cell or pair of groups.
Overall Evidence Is Not a Category-by-Category Answer
Suppose a test of homogeneity compares three groups across four response categories. If the test rejects \(H_0\), the result supports that the response distributions are not all the same. It does not show that every group differs from every other group, nor that every response category has a different proportion across groups. The overall test’s alternative is broad: some part of the distributions differs.
You can look at the table’s row percentages or column percentages to describe the sample pattern. Residuals can also help you notice cells with relatively large departures from expected counts. But describing a visible pattern is not the same as establishing that a particular difference is statistically significant in the population. More focused follow-up analysis is needed for that kind of claim. If many follow-up comparisons are made, the analysis also needs to account for the increased chance of finding an apparently unusual result just by chance.
Keep the levels of claim separate: the overall test gives evidence about the full distribution, while percentages and residuals describe features of the observed table. Don’t present a descriptive observation as if the overall test had confirmed it as a separate finding.
A Significant Result Does Not Automatically Establish Causation
A chi-square test measures how compatible the data are with a null model. It does not reveal why the counts have the pattern they do. In an observational study, individuals or groups are observed rather than assigned to treatments. Other differences among the groups could help explain an association or distributional difference, so the test alone does not show that one variable caused another.
The study design matters. In a randomized experiment, researchers assign experimental units to treatments at random. When the experiment is properly conducted, a difference between treatment groups can support a causal conclusion about the effect of the assigned treatments on the categorical response. That causal support comes from the random assignment and the experiment’s design—not from the chi-square statistic by itself. The test assesses evidence of a difference in the response distributions under the null model.
Even in an experiment, describe what the study supports without claiming certainty. Name the treatment and response, and avoid extending a result beyond the experimental units and conditions studied. Random assignment is not a license to claim that a treatment will have the same effect for every person or in every setting.
Worked Example: A Difference Among Districts Is Not a Cause
Worked Example: A Difference Among Districts Is Not a Cause
Suppose an invented survey takes separate random samples of 60 households from each of two districts. Each district has 1,500 households. The response is the household’s main way of traveling to work: walking, public transit, or driving. The observed counts are:
| District | Walking | Public transit | Driving | Total |
|---|---|---|---|---|
| Ridge | 26 | 20 | 14 | 60 |
| Harbor | 14 | 20 | 26 | 60 |
| Total | 40 | 40 | 40 | 120 |
State. The question is whether the distribution of main commuting method is the same for households in Ridge and Harbor districts. The null hypothesis says the distributions are the same; the alternative says they differ.
Plan. Use a chi-square test of homogeneity because the data come from separate random samples, with one categorical response recorded for each household. Each household contributes to exactly one cell, and the samples are independent. The 10% condition holds for each district because \(60\le0.10(1500)=150\). Under the null model, each expected count is \(60(40)/120=20\), so all expected counts are at least 5.
Do. The observed counts differ from their expected counts by \(6, 0,\) and \(-6\) in the Ridge sample, with the opposite deviations in the Harbor sample. Therefore:
The degrees of freedom are \((2-1)(3-1)=2\). The upper-tail p-value is approximately \(0.0273\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.0273<0.05\), so reject \(H_0\). The samples provide convincing evidence that the distribution of main commuting method differs between households in the two districts. In this sample, walking is more common in Ridge and driving is more common in Harbor. Those percentages describe the observed pattern; the overall test does not establish that either category-specific difference is independently significant. Because this is a survey rather than a randomized experiment, it also does not show that living in one district causes a household to choose a particular commuting method.
Worked Example: A Randomized Experiment and a Causal Claim
Worked Example: A Randomized Experiment and a Causal Claim
Suppose an invented experiment randomly assigns 80 volunteers to use either a new evening screen reminder or their usual routine for two weeks. At the end, each volunteer reports whether their sleep schedule became more regular. The counts are:
| Assigned group | More regular | Not more regular | Total |
|---|---|---|---|
| Screen reminder | 50 | 30 | 80 |
| Usual routine | 35 | 45 | 80 |
| Total | 85 | 75 | 160 |
State. The null hypothesis is that the distribution of reported sleep-schedule outcome is the same for volunteers assigned to the reminder and usual-routine groups. The alternative is that the distributions differ.
Plan. Use a chi-square test of homogeneity because the experiment has separate treatment groups and one categorical response for each volunteer. Volunteers were randomly assigned, each contributes to one cell, and the groups are independent. The expected counts under \(H_0\) are \(80(85)/160=42.5\) and \(80(75)/160=37.5\) in each row. All are at least 5.
Do. For the first row, the observed counts are 7.5 above and below their expected counts. The other row has deviations in the opposite directions. Thus:
The degrees of freedom are \((2-1)(2-1)=1\). The upper-tail p-value is approximately \(0.0175\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.0175<0.05\), so reject \(H_0\). The experiment provides convincing evidence that the distribution of reported sleep-schedule outcome differs between the two assigned groups. Because volunteers were randomly assigned, the experiment can support the conclusion that assignment to the screen-reminder routine affected the distribution of reported outcomes for volunteers under these study conditions. The chi-square test alone did not establish causation; the random assignment is what makes a causal interpretation reasonable.
Worked Example: Failing to Reject Does Not Prove Equality
Worked Example: Failing to Reject Does Not Prove Equality
Suppose an invented survey takes separate random samples of 75 customers from each of three service regions. Each region has 4,000 customers. Customers choose one preferred appointment format: phone, video, or in person.
| Region | Phone | Video | In person | Total |
|---|---|---|---|---|
| North | 27 | 24 | 24 | 75 |
| Central | 24 | 27 | 24 | 75 |
| South | 24 | 24 | 27 | 75 |
| Total | 75 | 75 | 75 | 225 |
State. The null hypothesis is that the distribution of preferred appointment format is the same across the three regions. The alternative is that the distributions are not all the same.
Plan. A chi-square test of homogeneity is appropriate because there are separate random samples from three populations, and each customer is recorded in one response category. The 10% condition holds because \(75\le0.10(4000)=400\) for each region. Every expected count is \(75(75)/225=25\), so the expected-count condition is met.
Do. Each row has one count 2 above its expected count and two counts 1 below. Each row contributes \((2^2+1^2+1^2)/25=0.24\), giving:
The degrees of freedom are \((3-1)(3-1)=4\). The upper-tail p-value is approximately \(0.9488\), rounded to four decimal places.
Conclude. At \(\alpha=0.05\), \(0.9488>0.05\), so fail to reject \(H_0\). The samples do not provide convincing evidence that preferred appointment-format distributions differ across the three regions. This does not prove that the population distributions are identical; the data simply do not provide strong evidence of a difference.
Common Mistakes and AP Exam Tip
- Claiming a specific category differs because the overall test is significant: The test supports an overall difference or association. A category-specific claim needs appropriate follow-up analysis; a cell with a large residual is a clue, not automatically a confirmed finding.
- Saying every group or category differs: Rejection supports that the distributions are not all the same. It does not prove that every pair differs.
- Turning an observational association into a cause: A random sample supports inference to a population, not a causal conclusion. Look for random assignment to treatments before making a causal claim.
- Ignoring the experiment’s design when discussing causation: In a well-designed randomized experiment, a causal interpretation may be justified by random assignment. Be precise that the chi-square test evaluates evidence of a distributional difference and the design supports causal reasoning.
- Claiming equality after failing to reject: Say the data do not provide convincing evidence of a difference. Do not say the null hypothesis has been proved.
- Reporting only “significant” or “not significant”: Include the decision and a conclusion naming the response and the populations or groups. Explain the scope of the evidence in context.
For a full-credit AP-style interpretation, state what the test supports, name the groups and response, and stop where the evidence stops. If the question asks which categories differ, explain that the overall test does not answer that by itself. If it asks whether one factor caused another, use the study design to decide whether a causal conclusion is justified.
Check Your Understanding
Use the study design and the scope of the test result to answer each question.
- A test of homogeneity rejects \(H_0\) when comparing four regions on preferred delivery method. What does the result establish about the distributions, and what does it not establish about pairs of regions?
- A significant chi-square test finds an association between neighborhood and preferred transportation in an observational survey. Can the result show that neighborhood caused the preference? Explain.
- In a randomized experiment, a chi-square test finds different outcome distributions across assigned treatments. What role does the random assignment play in interpreting the result?
- A standardized residual is large in one cell after an overall test rejects. Why should you avoid treating that fact alone as proof of a specific category difference?
- A test has a large p-value. Write a careful conclusion and state one claim that would be too strong.