Tutorials › AP Statistics › Observed Minus Expected: Computing Residuals

Expected counts and chi-square conclusions · Tutorial 567 of 1000

Observed Minus Expected: Computing Residuals

Calculate the difference between each observed count and its null-model expected count, then interpret what its sign and size say about a table.

Intermediate 9 min read

What You'll Learn

  • Calculate a cell residual as observed count minus expected count.
  • Explain what positive and negative residuals mean in context.
  • Interpret the size of a residual as a difference in counts.
  • Check how residuals balance across rows and columns.
  • Explain why raw residuals alone do not compare relative discrepancies fairly.

What a Cell Residual Tells You

In “Building the Full Expected Counts Table” and “Expected Counts for Homogeneity Tests,” you learned to calculate the count predicted in each cell under a null model. Now compare each observed count with its expected count. This difference is called the observed-minus-expected residual, or simply the cell residual in this tutorial.

A residual is a signed difference: it can be positive, negative, or zero. Its sign tells whether the observed count is above or below what the null model predicts. Its size tells the number of observations separating the observed and expected counts. These differences help describe where a table’s pattern departs from the null model; they do not, by themselves, determine whether the overall evidence is convincing.

Definition: For a cell, the observed-minus-expected residual is the observed count \(O\) minus the expected count \(E\) under the null model.
$$ \text{Residual}=O-E $$

For example, if a cell has \(O=18\) and \(E=14\), its residual is \(18-14=4\). The observed count is 4 greater than the count predicted under the null model. If \(O=11\) and \(E=14\), the residual is \(11-14=-3\): the cell contains 3 fewer observations than expected.

Expected counts can be fractional. A residual can therefore also be fractional: if \(O=12\) and \(E=12.5\), then \(O-E=-0.5\). The subtraction still has the same interpretation. It compares the actual count with the average count predicted by the null model, not with a requirement that the prediction be a whole number.

Read the Sign Before the Size

A positive residual means \(O>E\): that cell has more observations than the null model predicts. A negative residual means \(O<E\): it has fewer. A residual of zero means the observed and expected counts match exactly for that cell.

Interpret the sign using both categories that identify the cell. For instance, a positive residual in a “weekday / online booking” cell means that the observed number of weekday respondents who prefer online booking exceeds the count predicted under the null model. It does not mean that online booking is preferred by most respondents, or that weekdays caused the preference. Those claims ask different questions.

The magnitude, or absolute value, of a residual gives the size of the count difference without regard to direction. A residual of \(-7\) has a magnitude of 7: there are 7 fewer observations than expected. A residual of \(+7\) also has magnitude 7, but the observed count is above expectation rather than below it.

Interpretation guide: State the cell’s row and column categories, give the direction and count difference, and name the null model. For example: “There were 4 more observed selections in this cell than expected if the two variables were independent.”

Worked Example: Interpreting Residuals in a 2-by-2 Table

Worked Example: Interpreting Residuals in a 2-by-2 Table

Suppose an invented random sample of 120 residents is classified by housing type and whether the resident composts food scraps. The expected counts below have been calculated under the null model that housing type and composting are independent. Each row and column total is 60.

Housing typeCompostsDoes not compostTotal
Apartment362460
House243660
Total6060120

Because both row totals are 60 and both column totals are 60, the expected count in each cell under independence is \((60)(60)/120=30\). Subtract the expected count from the observed count in each cell:

$$ \begin{aligned} \text{Apartment and composts: }&36-30=+6\\ \text{Apartment and does not compost: }&24-30=-6\\ \text{House and composts: }&24-30=-6\\ \text{House and does not compost: }&36-30=+6 \end{aligned} $$

The positive residual of 6 for apartments and composting means there are 6 more apartment residents who compost in the sample than the independence model predicts. The negative residual of \(-6\) for apartments and not composting means there are 6 fewer than predicted. The two house cells show the corresponding pattern in the opposite direction.

A quick check is to add the residuals within each row and column. The apartment row sums to \(6+(-6)=0\), and the house row sums to \(-6+6=0\). The compost column sums to \(6+(-6)=0\), and the other column also sums to zero. The residuals balance because the expected counts preserve the table margins.

Residuals in a Larger Table

The same subtraction works in tables with more than two rows or columns. First identify the correct expected count for each cell under the null model. Then subtract that expected count from the observed count in the same cell. Keep the cells in a consistent order, or use a residual table with the same row and column labels as the observed table.

As you read a residual table, look across a row or down a column for patterns. Several positive residuals in one column mean those cells have more observations than expected, while negative residuals mean fewer. The balancing of residuals means such patterns are connected: a surplus in some cells must be matched by deficits elsewhere when the margins are fixed.

Worked Example: A 3-by-2 Independence Table

Worked Example: A 3-by-2 Independence Table

An invented random sample of 150 residents at three community centers records each person’s preferred way to attend a workshop: in person or by video. The question is whether preferred format is associated with community center. Expected counts are calculated under the null model of independence.

Community centerIn personBy videoTotal
North202040
Central204060
South203050
Total6090150

For example, the expected North/in-person count is its row total times the in-person column total divided by the grand total. The same method gives the expected count in every cell:

$$ \begin{aligned} E_{\text{North, in person}}&=\frac{(40)(60)}{150}=16, & E_{\text{North, video}}&=\frac{(40)(90)}{150}=24\\ E_{\text{Central, in person}}&=\frac{(60)(60)}{150}=24, & E_{\text{Central, video}}&=\frac{(60)(90)}{150}=36\\ E_{\text{South, in person}}&=\frac{(50)(60)}{150}=20, & E_{\text{South, video}}&=\frac{(50)(90)}{150}=30 \end{aligned} $$

Now subtract each expected count from its observed count:

Community centerIn-person residualVideo residual
North20 - 16 = +420 - 24 = -4
Central20 - 24 = -440 - 36 = +4
South20 - 20 = 030 - 30 = 0

The North center has 4 more observed in-person preferences than expected under independence and 4 fewer video preferences. Central has 4 fewer in-person preferences and 4 more video preferences than expected. South matches the expected counts in both cells.

Check the margins to verify the arithmetic. The in-person residuals add to \(4+(-4)+0=0\); the video residuals add to \(-4+4+0=0\). Each row also sums to zero. For instance, Central’s \(-4+4=0\). These checks can help catch a subtraction error, but they do not show that the variables are independent. The residuals describe the sample’s differences from the null model; a test evaluates how the whole pattern compares with chance variation under that model.

Raw Residual Size Needs Context

A raw residual measures a difference in counts. A difference of 5 is five observations whether it comes from a cell expected to contain 10 observations or a cell expected to contain 100. But those two differences are not necessarily equally notable relative to their expected counts. Therefore, raw residual size alone is not a fair way to rank which cells depart most strongly from the null model.

For example, a residual of \(+5\) in a cell with expected count 10 is half the expected count, while \(+5\) in a cell with expected count 100 is one-twentieth of the expected count. The residuals have the same count magnitude but represent different relative sizes. In later calculations, the chi-square procedure accounts for expected counts when combining cell differences. For now, keep the raw residual and its meaning clear: \(O-E\) is a signed count difference.

Worked Example: Residuals in a Homogeneity Table

Worked Example: Residuals in a Homogeneity Table

Three invented recreation programs take separate samples of 60 participants and ask each person to choose one preferred way to travel to activities: reusable cup, refillable bottle, or neither. The observed counts are shown below. The expected counts are calculated under the null model that all three programs have the same response distribution.

ProgramReusable cupRefillable bottleNeitherTotal
A3616860
B2724960
C27201360
Total906030180

Each program has a row total of 60. For Program A, for example, the expected counts are:

$$ \frac{(60)(90)}{180}=30,\qquad \frac{(60)(60)}{180}=20,\qquad \frac{(60)(30)}{180}=10 $$

Because all program totals are 60, these same expected counts apply to Programs B and C. Subtracting expected from observed gives:

ProgramReusable cup residualRefillable bottle residualNeither residual
A36 - 30 = +616 - 20 = -48 - 10 = -2
B27 - 30 = -324 - 20 = +49 - 10 = -1
C27 - 30 = -320 - 20 = 013 - 10 = +3

For Program A, the observed count of reusable-cup responses is 6 higher than expected under the same-distribution model, while its observed “neither” count is 2 lower. In Program C, the “neither” count is 3 higher than expected, and the bottle count matches expectation.

The residuals sum to zero within each program: for A, \(6-4-2=0\); for B, \(-3+4-1=0\); for C, \(-3+0+3=0\). They also sum to zero in each outcome column: \(6-3-3=0\) for reusable cups, \(-4+4+0=0\) for bottles, and \(-2-1+3=0\) for neither. This gives an arithmetic check and illustrates how the observed totals match the expected totals across the margins.

Common Mistakes and AP Exam Tip

  • Reversing the subtraction: The residual here is observed minus expected, \(O-E\), not \(E-O\). Reversing it flips every sign.
  • Describing a positive residual as a percentage: A residual of \(+4\) means four more observations than expected, not four percent more. Calculate a percentage separately if the question asks for one.
  • Leaving the context out: “The residual is 4” does not identify what was counted. Name the row and column categories and explain what was higher or lower than expected.
  • Thinking a negative residual is an impossible count: The residual is a difference, not an observed count. An observed count cannot be negative, but \(O-E\) can be.
  • Comparing raw residuals as if they show which cell matters most: A larger count difference is not automatically a larger relative departure. Raw residuals do not adjust for the expected count in each cell.
  • Treating residual patterns as a test conclusion: Residuals describe cell-by-cell differences from the null model. Do not claim convincing evidence or make a population conclusion from a single residual alone.

For a clear AP response, show the subtraction for the requested cell, keep the sign, and interpret it with the categories and null model named. If asked to describe a whole table, identify the most relevant positive and negative differences without claiming that raw residual size alone establishes statistical significance.

Key takeaway: For every cell, calculate \(O-E\). A positive residual means more observations than the null model predicts; a negative residual means fewer. The absolute value is the count difference, but raw residuals alone do not account for the size of the expected count.

Check Your Understanding

Use observed minus expected for each cell, and interpret each result in the context given.

  1. A cell has an observed count of 19 and an expected count of 23. Calculate and interpret its residual.
  2. What does a residual of \(+2.5\) say about the observed and expected counts? Why can the expected count be fractional?
  3. In a table, a cell for “East district / prefers cycling” has residual \(-7\). Write a sentence interpreting that residual under the null model.
  4. Why do the residuals in a row sum to zero when the expected counts preserve that row’s total?
  5. Two cells have residuals of \(+5\), but their expected counts are 10 and 100. Why should you avoid saying their relative departures from expectation are the same?