What a Cell Residual Tells You
In “Building the Full Expected Counts Table” and “Expected Counts for Homogeneity Tests,” you learned to calculate the count predicted in each cell under a null model. Now compare each observed count with its expected count. This difference is called the observed-minus-expected residual, or simply the cell residual in this tutorial.
A residual is a signed difference: it can be positive, negative, or zero. Its sign tells whether the observed count is above or below what the null model predicts. Its size tells the number of observations separating the observed and expected counts. These differences help describe where a table’s pattern departs from the null model; they do not, by themselves, determine whether the overall evidence is convincing.
For example, if a cell has \(O=18\) and \(E=14\), its residual is \(18-14=4\). The observed count is 4 greater than the count predicted under the null model. If \(O=11\) and \(E=14\), the residual is \(11-14=-3\): the cell contains 3 fewer observations than expected.
Expected counts can be fractional. A residual can therefore also be fractional: if \(O=12\) and \(E=12.5\), then \(O-E=-0.5\). The subtraction still has the same interpretation. It compares the actual count with the average count predicted by the null model, not with a requirement that the prediction be a whole number.
Read the Sign Before the Size
A positive residual means \(O>E\): that cell has more observations than the null model predicts. A negative residual means \(O<E\): it has fewer. A residual of zero means the observed and expected counts match exactly for that cell.
Interpret the sign using both categories that identify the cell. For instance, a positive residual in a “weekday / online booking” cell means that the observed number of weekday respondents who prefer online booking exceeds the count predicted under the null model. It does not mean that online booking is preferred by most respondents, or that weekdays caused the preference. Those claims ask different questions.
The magnitude, or absolute value, of a residual gives the size of the count difference without regard to direction. A residual of \(-7\) has a magnitude of 7: there are 7 fewer observations than expected. A residual of \(+7\) also has magnitude 7, but the observed count is above expectation rather than below it.
Worked Example: Interpreting Residuals in a 2-by-2 Table
Worked Example: Interpreting Residuals in a 2-by-2 Table
Suppose an invented random sample of 120 residents is classified by housing type and whether the resident composts food scraps. The expected counts below have been calculated under the null model that housing type and composting are independent. Each row and column total is 60.
| Housing type | Composts | Does not compost | Total |
|---|---|---|---|
| Apartment | 36 | 24 | 60 |
| House | 24 | 36 | 60 |
| Total | 60 | 60 | 120 |
Because both row totals are 60 and both column totals are 60, the expected count in each cell under independence is \((60)(60)/120=30\). Subtract the expected count from the observed count in each cell:
The positive residual of 6 for apartments and composting means there are 6 more apartment residents who compost in the sample than the independence model predicts. The negative residual of \(-6\) for apartments and not composting means there are 6 fewer than predicted. The two house cells show the corresponding pattern in the opposite direction.
A quick check is to add the residuals within each row and column. The apartment row sums to \(6+(-6)=0\), and the house row sums to \(-6+6=0\). The compost column sums to \(6+(-6)=0\), and the other column also sums to zero. The residuals balance because the expected counts preserve the table margins.
Residuals in a Larger Table
The same subtraction works in tables with more than two rows or columns. First identify the correct expected count for each cell under the null model. Then subtract that expected count from the observed count in the same cell. Keep the cells in a consistent order, or use a residual table with the same row and column labels as the observed table.
As you read a residual table, look across a row or down a column for patterns. Several positive residuals in one column mean those cells have more observations than expected, while negative residuals mean fewer. The balancing of residuals means such patterns are connected: a surplus in some cells must be matched by deficits elsewhere when the margins are fixed.
Worked Example: A 3-by-2 Independence Table
Worked Example: A 3-by-2 Independence Table
An invented random sample of 150 residents at three community centers records each person’s preferred way to attend a workshop: in person or by video. The question is whether preferred format is associated with community center. Expected counts are calculated under the null model of independence.
| Community center | In person | By video | Total |
|---|---|---|---|
| North | 20 | 20 | 40 |
| Central | 20 | 40 | 60 |
| South | 20 | 30 | 50 |
| Total | 60 | 90 | 150 |
For example, the expected North/in-person count is its row total times the in-person column total divided by the grand total. The same method gives the expected count in every cell:
Now subtract each expected count from its observed count:
| Community center | In-person residual | Video residual |
|---|---|---|
| North | 20 - 16 = +4 | 20 - 24 = -4 |
| Central | 20 - 24 = -4 | 40 - 36 = +4 |
| South | 20 - 20 = 0 | 30 - 30 = 0 |
The North center has 4 more observed in-person preferences than expected under independence and 4 fewer video preferences. Central has 4 fewer in-person preferences and 4 more video preferences than expected. South matches the expected counts in both cells.
Check the margins to verify the arithmetic. The in-person residuals add to \(4+(-4)+0=0\); the video residuals add to \(-4+4+0=0\). Each row also sums to zero. For instance, Central’s \(-4+4=0\). These checks can help catch a subtraction error, but they do not show that the variables are independent. The residuals describe the sample’s differences from the null model; a test evaluates how the whole pattern compares with chance variation under that model.
Raw Residual Size Needs Context
A raw residual measures a difference in counts. A difference of 5 is five observations whether it comes from a cell expected to contain 10 observations or a cell expected to contain 100. But those two differences are not necessarily equally notable relative to their expected counts. Therefore, raw residual size alone is not a fair way to rank which cells depart most strongly from the null model.
For example, a residual of \(+5\) in a cell with expected count 10 is half the expected count, while \(+5\) in a cell with expected count 100 is one-twentieth of the expected count. The residuals have the same count magnitude but represent different relative sizes. In later calculations, the chi-square procedure accounts for expected counts when combining cell differences. For now, keep the raw residual and its meaning clear: \(O-E\) is a signed count difference.
Worked Example: Residuals in a Homogeneity Table
Worked Example: Residuals in a Homogeneity Table
Three invented recreation programs take separate samples of 60 participants and ask each person to choose one preferred way to travel to activities: reusable cup, refillable bottle, or neither. The observed counts are shown below. The expected counts are calculated under the null model that all three programs have the same response distribution.
| Program | Reusable cup | Refillable bottle | Neither | Total |
|---|---|---|---|---|
| A | 36 | 16 | 8 | 60 |
| B | 27 | 24 | 9 | 60 |
| C | 27 | 20 | 13 | 60 |
| Total | 90 | 60 | 30 | 180 |
Each program has a row total of 60. For Program A, for example, the expected counts are:
Because all program totals are 60, these same expected counts apply to Programs B and C. Subtracting expected from observed gives:
| Program | Reusable cup residual | Refillable bottle residual | Neither residual |
|---|---|---|---|
| A | 36 - 30 = +6 | 16 - 20 = -4 | 8 - 10 = -2 |
| B | 27 - 30 = -3 | 24 - 20 = +4 | 9 - 10 = -1 |
| C | 27 - 30 = -3 | 20 - 20 = 0 | 13 - 10 = +3 |
For Program A, the observed count of reusable-cup responses is 6 higher than expected under the same-distribution model, while its observed “neither” count is 2 lower. In Program C, the “neither” count is 3 higher than expected, and the bottle count matches expectation.
The residuals sum to zero within each program: for A, \(6-4-2=0\); for B, \(-3+4-1=0\); for C, \(-3+0+3=0\). They also sum to zero in each outcome column: \(6-3-3=0\) for reusable cups, \(-4+4+0=0\) for bottles, and \(-2-1+3=0\) for neither. This gives an arithmetic check and illustrates how the observed totals match the expected totals across the margins.
Common Mistakes and AP Exam Tip
- Reversing the subtraction: The residual here is observed minus expected, \(O-E\), not \(E-O\). Reversing it flips every sign.
- Describing a positive residual as a percentage: A residual of \(+4\) means four more observations than expected, not four percent more. Calculate a percentage separately if the question asks for one.
- Leaving the context out: “The residual is 4” does not identify what was counted. Name the row and column categories and explain what was higher or lower than expected.
- Thinking a negative residual is an impossible count: The residual is a difference, not an observed count. An observed count cannot be negative, but \(O-E\) can be.
- Comparing raw residuals as if they show which cell matters most: A larger count difference is not automatically a larger relative departure. Raw residuals do not adjust for the expected count in each cell.
- Treating residual patterns as a test conclusion: Residuals describe cell-by-cell differences from the null model. Do not claim convincing evidence or make a population conclusion from a single residual alone.
For a clear AP response, show the subtraction for the requested cell, keep the sign, and interpret it with the categories and null model named. If asked to describe a whole table, identify the most relevant positive and negative differences without claiming that raw residual size alone establishes statistical significance.
Check Your Understanding
Use observed minus expected for each cell, and interpret each result in the context given.
- A cell has an observed count of 19 and an expected count of 23. Calculate and interpret its residual.
- What does a residual of \(+2.5\) say about the observed and expected counts? Why can the expected count be fractional?
- In a table, a cell for “East district / prefers cycling” has residual \(-7\). Write a sentence interpreting that residual under the null model.
- Why do the residuals in a row sum to zero when the expected counts preserve that row’s total?
- Two cells have residuals of \(+5\), but their expected counts are 10 and 100. Why should you avoid saying their relative departures from expectation are the same?