Measuring the Overall Difference Between Observed and Expected Counts
In “Computing Conditional Distributions From a Two-Way Table,” you used percentages to describe how a categorical distribution varies across groups. A chi-square statistic takes a different approach: it compares the table’s observed counts with the counts expected if the two categorical variables were independent. It summarizes the discrepancies across all the cells in one number.
The observed count, \(O\), is the number recorded in a cell. The expected count, \(E\), is the number we would expect in that cell under the model of independence, while keeping the table’s row and column totals fixed. These are model-based counts, not additional observations.
For a test of independence, calculate the expected count for each cell using its row total, column total, and the grand total:
This formula reflects independence: if the row and column variables are independent, the proportion in a column should be the same within each row. The expected count distributes each row total across the columns according to the overall column proportions.
The numerator, \((O-E)^2\), squares the difference. This makes a cell’s contribution nonnegative whether its observed count is above or below its expected count. Dividing by \(E\) scales the squared difference: a discrepancy of a given size matters more when the expected count is small than when it is large. Adding the contributions measures the table’s overall discrepancy from the independence model.
The statistic alone does not tell you which cells are higher or lower than expected, or which categories account most for the result. Look at the cell-by-cell contributions and the observed counts to understand the pattern. As in “Interpreting Association Versus Independence in a Table,” conditional distributions can help describe how that pattern appears across groups.
A large value is not, by itself, a complete test conclusion. To decide whether the value is unusually large under the independence model, it must be evaluated using the appropriate chi-square reference distribution. Later tutorials develop other parts of that inference process. Here, the goal is to understand how the statistic is built and what its size represents.
Worked Example: Two Categories in Each Variable
Worked Example: Two Categories in Each Variable
Suppose a fictional community survey records whether respondents usually commute by bicycle or another method, and whether they live near or far from the community center. Each respondent is counted once.
| Residence | Bicycle | Another method | Total |
|---|---|---|---|
| Near | 30 | 20 | 50 |
| Far | 20 | 30 | 50 |
| Total | 50 | 50 | 100 |
Under independence, the expected count for the near-and-bicycle cell is \(50 \times 50/100=25\). Each of the other three cells also has expected count \(50 \times 50/100=25\). Every expected count is at least 5, satisfying the expected-count condition for the chi-square procedures discussed earlier in the course.
Now calculate a contribution for each cell. In the near-and-bicycle cell, the observed count is 30 and the expected count is 25, so the contribution is \((30-25)^2/25=25/25=1\). For near-and-another-method, it is \((20-25)^2/25=25/25=1\). The far-row cells contribute \((20-25)^2/25=1\) and \((30-25)^2/25=1\).
Interpretation: The statistic is 4. The observed counts differ from the expected counts in all four cells, with each cell contributing 1 to the total. The calculation does not, on its own, establish whether there is convincing evidence of an association; that requires comparing the statistic with the appropriate reference distribution.
Worked Example: Why the Expected Count Scales a Difference
Worked Example: Why the Expected Count Scales a Difference
A fictional school survey records whether students use a printed or digital planner and whether they are in an earlier or later grade group. The counts are:
| Grade group | Printed | Digital | Total |
|---|---|---|---|
| Earlier grades | 18 | 22 | 40 |
| Later grades | 12 | 48 | 60 |
| Total | 30 | 70 | 100 |
For the earlier-grades-and-printed cell, the expected count is \(40 \times 30/100=12\). For earlier grades and digital, it is \(40 \times 70/100=28\). In the later-grades row, the expected counts are \(60 \times 30/100=18\) and \(60 \times 70/100=42\). The four expected counts, 12, 28, 18, and 42, are all at least 5.
The observed counts differ from their expected counts by \(+6,-6,-6,\) and \(+6\), respectively. Although the absolute difference is 6 in each cell, the contributions are not all the same, because the expected counts differ:
Adding the unrounded contributions gives \(3+36/28+2+36/42\approx7.1429\), rounded to four decimal places. The same-sized count difference contributes more in a cell with a smaller expected count. For example, the difference of 6 contributes 3 when the expected count is 12, but about 0.8571 when the expected count is 42.
Interpretation: The statistic summarizes the total discrepancy across the four cells. The table’s observed counts show where those discrepancies occur: printed-planner counts are higher than expected in the earlier-grades group and lower than expected in the later-grades group. The chi-square statistic itself is nonnegative and does not preserve those directions.
Worked Example: Contributions Across a Larger Table
Worked Example: Contributions Across a Larger Table
A fictional environmental survey records residents’ preferred way to receive local air-quality updates in three neighborhoods. Each neighborhood has 60 respondents.
| Neighborhood | Text message | Community board | Total | |
|---|---|---|---|---|
| North | 30 | 20 | 10 | 60 |
| Central | 20 | 20 | 20 | 60 |
| South | 10 | 20 | 30 | 60 |
| Total | 60 | 60 | 60 | 180 |
Each expected count is \(60 \times 60/180=20\), since each row total and column total is 60 and the grand total is 180. All nine expected counts equal 20, so each is at least 5.
In the North row, the contributions are \((30-20)^2/20=100/20=5\), \((20-20)^2/20=0\), and \((10-20)^2/20=100/20=5\). In the Central row, all three observed counts equal 20, so each contribution is 0. In the South row, the contributions are \(100/20=5\), 0, and \(100/20=5\).
Interpretation: The total is 20 because several cells differ from their expected counts, while the three Central cells match exactly. In this sample, text-message preference is more common than expected in the North and less common than expected in the South; community-board preference shows the reverse pattern. Those descriptions come from comparing observed with expected counts. The statistic adds the size of the discrepancies but does not label their direction.
Reading a Chi-Square Statistic Carefully
Each cell adds zero or a positive amount to \(X^2\). If an observed count equals its expected count, the cell contributes zero. If it differs, squaring the difference makes the contribution positive whether the observed count is above or below the expected count. Therefore, \(X^2\) cannot be negative. If every cell matches its expected count, then \(X^2=0\).
The statistic accumulates discrepancies across the table. A few substantial contributions can produce a large total, as can several moderate contributions. Looking only at \(X^2\) hides this detail, so a useful calculation record includes the expected count and contribution for each cell. That lets you see which parts of the table account for most of the total.
The contribution formula also explains why a raw difference is not enough. A difference of 6 is not automatically equally important wherever it occurs: its contribution depends on the expected count in that cell. And comparing statistics across tables without the relevant reference distributions can be misleading. The statistic’s role is to quantify discrepancy under a particular null model, not to give a universal measure of effect size.
A chi-square statistic also does not identify causation. If the study is observational, a finding of association does not show that one variable caused the other. As established in “Chi-Square Test for Independence: Purpose and Setting,” the interpretation concerns a possible relationship between the two categorical variables in the population, subject to the study design and inference conditions.
Common Mistakes and AP Exam Communication
- Leaving out a cell: The sum includes every cell in the table, not only the ones whose observed counts exceed their expected counts.
- Using \(O-E\) without squaring: The formula squares the difference so positive and negative discrepancies do not cancel.
- Dividing by \(O\) instead of \(E\): Each squared difference is divided by the expected count for that same cell.
- Calling expected counts observed data: Observed counts come from the table; expected counts come from the null model and the table’s margins.
- Interpreting a contribution as directional: A contribution is nonnegative. To say whether the observed count is above or below expectation, compare \(O\) with \(E\) directly.
- Claiming a large statistic proves association: A large value signals greater discrepancy, but a formal conclusion requires assessing how unusual it is under the reference distribution and considering the study design.
Check Your Understanding
For each question, focus on what the cell contributions and their sum tell you about a two-way table.
- A cell has observed count 14 and expected count 10. Calculate its contribution to \(X^2\).
- Why do cells with observed counts below their expected counts still make positive contributions?
- If a cell’s observed count equals its expected count, what is its contribution, and why?
- Two cells have the same absolute difference between observed and expected counts. Explain why their contributions might differ.
- A table has a large chi-square statistic. What does that say about the observed counts, and what additional comparison is needed before making a test conclusion?