Which Cells Drive the Chi-Square Statistic?
In “Calculating the Chi-Square Statistic by Hand,” you added one term for every cell in the table. This tutorial looks inside that sum. Each term is a cell’s contribution to the chi-square statistic: it tells you how much that particular cell adds to the total measure of discrepancy from the null model.
A large contribution identifies a cell whose observed count is relatively far from its expected count. Comparing contributions helps locate the table’s largest departures from the null model. But a contribution is always nonnegative, so it does not tell you whether the observed count is above or below expectation. To describe the direction of a departure, use the observed-minus-expected residual \(O-E\), as in “Observed Minus Expected: Computing Residuals.”
The contribution combines two pieces of information: the size of the cell’s residual and its expected count. Squaring makes the contribution nonnegative, while dividing by \(E\) scales the squared difference. As a result, equal-sized residuals can produce different contributions when their expected counts differ.
A useful interpretation routine is: find the largest contributions, check the corresponding residual signs, and describe what the observed counts show in context. Contributions tell you where the discrepancy is concentrated; residual signs tell you which way the observed counts depart from expectation.
A Routine for Interpreting Cell Contributions
For each cell, use its own observed count and expected count in \((O-E)^2/E\).
Identify the largest terms. A larger term means that cell adds more to the overall chi-square statistic.
Look at \(O-E\) for the cells that contribute most. A positive residual means more observations than expected; a negative residual means fewer.
Name the relevant row and column categories and say which observed counts are above or below their null-model expectations. Do not claim that a contribution alone proves significance or causation.
You may also describe a cell’s share of the statistic by dividing its contribution by \(X^2\). This is a descriptive comparison of the terms in the sum, not a measure of the percentage of an association that the cell “causes.” The chi-square test’s p-value, not a contribution ranking, is used to assess how convincing the evidence is against the null model.
Worked Example: Ranking Contributions and Using Their Signs
Worked Example: Ranking Contributions and Using Their Signs
Imagine an invented survey in which 120 visitors to two community programs each report one preferred workshop format. The observed counts are below. Consider the null model that program and preferred format are independent. The expected counts, calculated from the row and column totals, are included for comparison.
| Program | In person | Online | Self-paced | Total |
|---|---|---|---|---|
| Harbor observed | 24 | 6 | 10 | 40 |
| Ridge observed | 36 | 24 | 20 | 80 |
| Total | 60 | 30 | 30 | 120 |
| Harbor expected | 20 | 10 | 10 | 40 |
| Ridge expected | 40 | 20 | 20 | 80 |
For instance, the expected Harbor/in-person count is \((40)(60)/120=20\). The Harbor/online expected count is \((40)(30)/120=10\). Use these expected counts to calculate each cell’s contribution:
The largest contribution is 1.6 for Harbor/online. It is followed by Harbor/in-person and Ridge/online, each contributing 0.8. The other two nonzero contributions are smaller, at 0.4 each. Adding all six terms gives the statistic:
To describe what those contributions mean, check the residuals as well. Harbor/online has residual \(6-10=-4\), so there are fewer Harbor visitors who prefer online workshops than expected under independence. Harbor/in-person has residual \(24-20=4\), so there are more than expected. Ridge/online has residual \(24-20=4\), so that cell is above expectation. The contributions alone would not show these directions because all six are nonnegative.
The statistic is a sum: the Harbor/online cell contributes \(1.6/3.6\approx44.4\%\) of the total, while each \(0.8\) contribution accounts for about \(22.2\%\). These shares help compare the terms in this particular calculation. They do not replace a p-value or establish that the population variables are associated.
Worked Example: Equal Residual Sizes, Different Contributions
Worked Example: Equal Residual Sizes, Different Contributions
An invented survey of 300 people records their usual way of receiving local updates and the type of update they prefer. Use independence as the null model. The observed counts and expected counts are shown below.
| Update source | Alerts | Weekly digest | Community board | Total |
|---|---|---|---|---|
| App observed | 38 | 22 | 40 | 100 |
| Email observed | 52 | 68 | 80 | 200 |
| Total | 90 | 90 | 120 | 300 |
| App expected | 30 | 30 | 40 | 100 |
| Email expected | 60 | 60 | 80 | 200 |
For example, the expected App/alerts count is \((100)(90)/300=30\), and the expected Email/alerts count is \((200)(90)/300=60\). Calculate the contributions:
The App/alerts and App/weekly-digest cells each have a residual with absolute value 8 and contribute about 2.1333. The Email/alerts and Email/weekly-digest cells also have residuals with absolute value 8, but each contributes only about 1.0667. The expected counts differ: 30 for the App cells and 60 for the corresponding Email cells. Dividing the same squared difference, 64, by 30 gives a larger term than dividing it by 60.
Adding the contributions gives \(X^2\approx2.1333+2.1333+0+1.0667+1.0667+0=6.4\). For context, the App/alerts and App/weekly-digest residuals are \(+8\) and \(-8\), respectively. Thus the observed App counts are above expectation for alerts and below expectation for weekly digests. In the Email row, the signs are reversed. This describes the pattern in the table without suggesting that the contribution values themselves show direction.
Worked Example: A Pattern Across Three Groups
Worked Example: A Pattern Across Three Groups
Suppose an invented survey classifies visitors to three recreation centers by the activity area they prefer. Consider the null model that center and preference are independent. Each center has 60 respondents, and each preference category has 90 respondents, so the expected count is \((60)(90)/180=30\) in every cell.
| Center | Indoor | Outdoor | Total |
|---|---|---|---|
| North observed | 40 | 20 | 60 |
| Central observed | 30 | 30 | 60 |
| South observed | 20 | 40 | 60 |
| Total | 90 | 90 | 180 |
| Expected count in each cell | 30 | 30 |
The six contributions are:
The North and South rows contain all the nonzero contributions. Each of their four cells contributes \(3.3333\), while the Central row contributes nothing because both observed counts equal their expected counts. Adding the exact terms gives \(X^2=4(100/30)=400/30\approx13.3333\). Each nonzero cell contributes one quarter of the total.
The residual signs reveal the pattern: North has more indoor preferences and fewer outdoor preferences than expected; South has fewer indoor preferences and more outdoor preferences than expected. Central matches the expected counts for both categories. The row percentages tell the same descriptive story: North is \(40/60=66.7\%\) indoor, Central is \(30/60=50.0\%\), and South is \(20/60=33.3\%\). These sample patterns describe how the observed counts differ from the independence model. Whether they provide convincing evidence of a population association is determined using the chi-square test’s p-value, not by reading a contribution in isolation.
Common Mistakes and AP Exam Tip
- Using a contribution to claim direction: Contributions cannot be negative. Pair a large contribution with \(O-E\): a positive residual means the observed count is above expectation, and a negative residual means it is below.
- Assuming the largest residual must have the largest contribution: The denominator matters. Compare \((O-E)^2/E\), not just \(|O-E|\). Equal residual sizes can yield unequal contributions when expected counts differ.
- Interpreting a small or zero contribution as proof of independence: A small contribution means that cell adds little to the statistic. Independence is a claim about the variables in the population and is assessed using the full table and test.
- Calling the contribution a p-value: A contribution is one term in \(X^2\), not a probability. Use the total statistic, its degrees of freedom, and the chi-square distribution to find the p-value.
- Overstating what the table establishes: Say that the observed counts are above or below the counts expected under the null model. Do not claim that a cell causes an association or that the sample pattern proves a population relationship.
For full-credit interpretation, identify the cell or cells with the largest contributions, give their row and column categories, and use the residual signs to state whether each observed count is above or below expectation. If asked to make an inferential conclusion, use the p-value and describe the evidence in context; do not substitute a contribution ranking for the test conclusion.
Check Your Understanding
Use the contribution formula and, when asked about direction, check the residual as well.
- A cell has \(O=18\) and \(E=12\). Calculate its contribution and state whether the observed count is above or below expectation.
- Two cells both have \(|O-E|=5\). One has \(E=10\), and the other has \(E=25\). Calculate both contributions. Which contributes more, and why?
- A cell contributes \(2.4\) to a table’s statistic, and its residual is negative. What does this tell you about that cell?
- The contributions in a table are \(1.2\), \(0.8\), \(0\), and \(2.0\). Find \(X^2\). What fraction of the statistic comes from the cell contributing \(2.0\)?
- Why is it not enough to identify the largest cell contribution when writing a test conclusion about association?