The Expected-Count Formula
In “What Expected Counts Represent Under Independence,” you learned that expected counts describe the cell counts the null model of independence would produce on average, given the table’s margins. This tutorial turns that idea into a direct calculation. For one cell, you need just three totals: its row total, its column total, and the grand total.
Here, \(E\) stands for the expected count in the cell you are calculating. The row total is the number of observations in that cell’s entire row; the column total is the number in its entire column; and the grand total is the number of observations in the whole table. Use the margins that match the cell’s row and column—not totals from neighboring rows or columns.
The formula is another way to apply the independence idea from the previous tutorial. The column’s share of the whole sample is its column total divided by the grand total. Under independence, that same share is used within each row. Multiplying the row total by that share gives the expected count:
This connection is useful for understanding the calculation: first find the overall proportion in the column, then apply it to the row. The compact formula combines those two steps. Both methods must give the same answer.
For example, if a column contains 30 of the 120 observations, its overall proportion is \(30/120\). Under independence, that proportion applies within every row. If the row contains 40 observations, the expected count in the cell where that row and column meet is \(40(30/120)\).
Worked Example: Row Total 40, Column Total 30, Grand Total 120
Worked Example: Row Total 40, Column Total 30, Grand Total 120
A sample contains 120 people, each classified by two categorical variables. Consider a particular cell whose row contains 40 people and whose column contains 30 people. Find the expected count for that cell under independence.
Identify the margins. The row total is 40, the column total is 30, and the grand total is 120. Substitute these values into the expected-count formula:
The expected count is 10. To check it another way, the column’s share of the whole sample is \(30/120=0.25\). Applying that share to the row total gives:
The two calculations agree. In context, if the variables were independent, the model would expect about 10 of the 40 people in this row to fall in this column category. The value 10 is a model-based expected count, not necessarily the observed count in the sample.
The result is also consistent with the margins: 10 is one-third of the column total of 30, matching the fact that this row accounts for \(40/120=1/3\) of all observations. This is a useful reasonableness check, not a different expected-count rule.
Why the Margins Determine the Calculation
Under independence, the expected proportion in a given column is the same across the rows. The row totals can differ, so the expected counts do not have to be equal from row to row. A larger row will generally have a larger expected count in the same column because the column’s overall proportion is applied to more observations.
The same formula works if you think in the opposite direction. A cell’s expected share of its column is its row total divided by the grand total. Multiplying that share by the column total also gives the formula:
This equivalent interpretation can help catch a mismatch. If you start from the column, use the share of the grand total belonging to the cell’s row. If you start from the row, use the share belonging to the cell’s column. Either route must use the same two margins and produce the same expected count.
The formula uses totals rather than the observed count in the cell. That distinction matters: the expected count is calculated from the margins under the independence model. Substituting the cell’s observed count into the formula changes the calculation and does not produce its expected count.
Worked Example: A Fractional Expected Count
Worked Example: A Fractional Expected Count
A sample of 160 residents is classified by neighborhood and preferred way to receive community updates. A selected neighborhood row has a total of 35 residents. The column for email updates has a total of 48 residents. Find the expected count for the cell where that row and column meet.
The grand total is 160, so the formula gives:
Check using the email column’s overall proportion:
The expected count is 10.5. A sample cannot contain half a resident in a cell, but the expected count can be fractional because it is an average specified by the null model. It is not a count observed in one sample.
Do not round 10.5 to 11 before using it as an expected count. In a chi-square calculation, the expected value is used as calculated; rounding early can change later arithmetic. In particular, the chi-square expected-count condition is checked using expected counts, including decimal values.
Worked Example: Checking the Margins and the Answer
Worked Example: Checking the Margins and the Answer
A sample of 144 customers is classified by shopping location and payment preference. One shopping-location row has 54 customers, and one payment-preference column has 36 customers. Find the expected count for their intersecting cell under independence.
Use the row total of 54, the column total of 36, and the grand total of 144:
For a second check, \(36/144=0.25\), so one quarter of the row total is expected in that column:
The expected count is 13.5 customers. As a margin check, this row is \(54/144=0.375\) of the full sample. Applying that share to the column total gives \(36(0.375)=13.5\), the same result. The checks agree because each uses the same margins.
If your result were larger than the row total of 54 or the column total of 36, that would be a strong signal to recheck the inputs or arithmetic. An expected count for a cell cannot exceed either of its margins: it represents a share of the row and a share of the column. Here, 13.5 is less than both totals, as expected.
A Reliable Calculation Routine
For each requested cell, pause before calculating and match the cell to its margins. A compact routine helps prevent using a plausible-looking but incorrect total:
Identify the row category and column category where the cell sits.
Record that row’s total, that column’s total, and the grand total for the entire table.
Multiply the row and column totals, then divide by the grand total. Keep the value unrounded during intermediate calculations.
Recalculate as row total times column proportion, and check that the value does not exceed either margin.
For the first example, this routine means using the row and column that meet at the selected cell, not simply multiplying two totals that happen to be nearby. A clear written solution can show the substitution in one line and then state what the result means under independence.
Common Mistakes and AP Exam Communication
- Using the wrong total: The denominator is the grand total of the whole table. Dividing by the row total or column total alone does not give the expected count.
- Using unrelated margins: The numerator must use the total for the cell’s own row and the total for its own column. Labeling the three values before substituting helps prevent this error.
- Reversing the idea of a proportion: If using the row-proportion method, divide the column total by the grand total, then multiply by the row total. Do not multiply the row total by the grand total divided by the column total.
- Substituting the observed cell count: The formula uses the margins, not the observed count in that cell. The observed count is compared with the expected count later in a chi-square analysis.
- Rounding too soon: Keep decimal expected counts as calculated. Do not turn 10.5 into 11 simply because the table records people as whole numbers.
- Calling the result a guaranteed count: An expected count describes the null model’s benchmark. It does not claim that the sample will contain exactly that many observations in the cell.
For full-credit communication, name the margins, show the substitution, report the expected count, and interpret it in context. For example: “Using a row total of 40, a column total of 30, and a grand total of 120, the expected count is \(40(30)/120=10\). If the variables were independent, about 10 observations in this row would be expected to fall in this column.”
Check Your Understanding
Use the expected-count formula and the independence interpretation to answer each question.
- A cell has row total 40, column total 30, and grand total 120. Calculate its expected count and show the substitution.
- A row total is 28, a column total is 50, and the grand total is 200. What is the expected count? Check your answer using the column proportion.
- Why can an expected count be a decimal even though an observed count must be a whole number?
- Which three totals are needed to calculate the expected count for a particular cell?
- A student uses the observed cell count in place of the column total. Explain why that does not follow the expected-count formula.