Why Check Expected Counts?
In “Expected Counts for Homogeneity Tests,” you used the null model to calculate expected counts for a table comparing groups. Expected counts are also central to deciding whether a chi-square test is appropriate: before relying on a chi-square reference distribution for a test statistic, check that every expected count is at least 5.
This is a condition about the expected table, not the observed table. An observed cell count might be 0 or 2 even when its expected count is at least 5; the condition is still checked using the expected count. Conversely, a large observed count in one cell does not compensate for an expected count below 5 in another cell.
The expected-count formula depends on the test setting. For independence, as covered in “Expected Count Formula: Row Total Times Column Total Over Grand Total,” use the cell’s row total and column total. For homogeneity, as covered in “Expected Counts for Homogeneity Tests,” use the group total and pooled outcome total. In either setting, the result is a count predicted by the null model, and each cell must be checked.
Why the Minimum Matters
A chi-square test compares observed counts with expected counts using contributions of the form \((O-E)^2/E\). Its p-value is found by comparing the resulting statistic with a chi-square distribution. That comparison works well when the chi-square distribution is a reasonable approximation to the statistic’s sampling distribution under the null hypothesis.
Counts are discrete: a cell can contain only whole observations, and when expected counts are small, possible outcomes are relatively limited. A cell may even have a substantial chance of an observed count far from its expected count relative to the expected count’s size. In that situation, the distribution of the test statistic may not be well represented by the smooth chi-square curve used to calculate the p-value.
When expected counts are sufficiently large, the count patterns under the null model are generally better approximated by the probability model that leads to the chi-square reference distribution. Requiring every expected count to be at least 5 is the AP Statistics rule for checking whether that approximation is reasonable. It is a practical condition, not a promise that the approximation is exact.
Passing this condition does not replace the other checks. As discussed in “Random Condition for Chi-Square Procedures” and “Independence and 10 Percent Condition for Chi-Square,” also consider how the data were collected and whether observations are independent. A table with adequate expected counts does not fix a biased sample or dependent observations.
A Cell-by-Cell Checking Routine
After building the full expected-count table, scan every cell and identify the smallest expected count. If that minimum is at least 5, the expected-count condition is met. If even one expected count is below 5, it is not met.
Confirm that the expected counts were calculated under the null hypothesis for the test being considered.
Compare each expected count with 5. Do not check only the row totals, column totals, or average expected count.
State whether all expected counts are at least 5, and identify the smallest count or any counts below 5.
If the condition is met, say the expected-count condition supports using the chi-square approximation. If it is not met, say that approximation may be unreliable.
Keep decimal expected counts as calculated. An expected count is a model-based average, so it does not have to be a whole number. Do not round a value such as 4.8 up to 5 to make the condition appear satisfied.
Worked Example: Checking a Homogeneity Table
Worked Example: Checking a Homogeneity Table
Suppose separate random samples of patients at three invented clinics are asked which appointment reminder they prefer: text, phone call, or email. Each clinic serves more than ten times its sample size, and each patient is counted once. The observed counts and margins are shown below. Check the expected-count condition for a chi-square test of homogeneity.
| Clinic | Text | Phone call | Total | |
|---|---|---|---|---|
| A | 20 | 10 | 30 | 60 |
| B | 22 | 18 | 40 | 80 |
| C | 30 | 20 | 50 | 100 |
| Pooled total | 72 | 48 | 120 | 240 |
State: The question is whether the distribution of reminder preference is the same across the three clinics. The expected-count condition requires every expected count under that null model to be at least 5.
Plan: Use the homogeneity expected-count formula, multiplying each clinic’s row total by the pooled total for a reminder category and dividing by the grand total. Then compare every result with 5. The samples are described as separate random samples, the clinic groups are independent, each patient contributes one response, and each sample is no more than 10% of its clinic’s patient population because each clinic serves more than ten times its sample size.
Do: The grand total is \(60+80+100=240\). The expected counts are:
The expected-count table is:
| Clinic | Text | Phone call | |
|---|---|---|---|
| A | 18 | 12 | 30 |
| B | 24 | 16 | 40 |
| C | 30 | 20 | 50 |
The smallest expected count is 12, which is at least 5. Every expected count therefore meets the condition.
Conclude: The expected-count condition is met for this chi-square test of homogeneity. This supports using the chi-square approximation, provided the other conditions are also satisfied. It does not, by itself, establish that the reminder preferences differ across clinics.
Worked Example: One Small Expected Count Is Enough
Worked Example: One Small Expected Count Is Enough
Imagine a random sample of 60 members of an invented online community. Each person is classified by whether they use a particular accessibility feature and which of three devices they primarily use. The community has more than ten times as many members as the sample. The observed counts are:
| Phone | Tablet | Computer | Total | |
|---|---|---|---|---|
| Uses feature | 1 | 7 | 7 | 15 |
| Does not use feature | 3 | 19 | 23 | 45 |
| Total | 4 | 26 | 30 | 60 |
For a chi-square test of independence, the expected count for each cell is its row total times its column total divided by the grand total. The six expected counts are:
Two expected counts are below 5: 1 and 3. The expected-count condition is not met, even though the other four expected counts are at least 5 and the grand total is 60. The chi-square approximation may be unreliable for this table. A large total sample size cannot make the condition pass when individual expected cells remain too small.
Notice that the observed count of 1 in the first row and first column is not itself the reason the condition fails. The relevant value is that cell’s expected count, which is also 1 here. In another table, an observed count could be small while its expected count is adequate; the decision must still be based on expected counts.
Worked Example: A Count Exactly at the Minimum
Worked Example: A Count Exactly at the Minimum
Suppose separate random samples of 20 and 80 customers at two invented repair centers report whether they would recommend the service. The pooled totals are 25 “yes” responses and 75 “no” responses, for a grand total of 100. Assume each sample is no more than 10% of its center’s customer population, and each customer contributes one response. Check the expected-count condition for a homogeneity test.
For the sample of 20 customers, the expected counts are:
For the sample of 80 customers, they are:
The four expected counts are 5, 15, 20, and 60. The smallest is exactly 5, so the condition is met: “at least 5” includes a count equal to 5. Do not incorrectly require every expected count to be greater than 5.
Common Mistakes and AP Exam Tip
- Checking observed counts instead: The condition applies to expected counts calculated under the null model. State which expected counts you checked.
- Checking only the average: A high average does not rule out a sparse cell. Inspect every cell and report the smallest expected count.
- Using “greater than 5” rather than “at least 5”: An expected count of exactly 5 meets the AP condition.
- Rounding up a value below 5: Keep the calculated value. For example, 4.8 is below 5 and does not meet the condition.
- Assuming a large grand total is enough: The check is cell by cell. A table can have many observations overall and still have a small expected count.
- Claiming that passing proves the test is valid in every way: This check supports the chi-square approximation; it does not establish random selection, independence, or the other design conditions.
For a full-credit AP response, write a direct statement such as: “The smallest expected count is 6.5, so all expected counts are at least 5 and the expected-count condition is met. This supports using the chi-square approximation, assuming the other conditions are satisfied.” If a count is below 5, identify it and explain that the approximation may be unreliable. Do not claim that the null hypothesis is true or false based on a condition check; this step is about whether the reference distribution is appropriate.
Check Your Understanding
For each question, focus on expected counts under the stated null model.
- A table’s expected counts are 8, 12, 5, and 20. Is the expected-count condition met? Explain using the smallest count.
- A table has one expected count of 4.9 and all its other expected counts exceed 10. Is the condition met? Why or why not?
- Why does a chi-square test check expected counts rather than requiring every observed count to be at least 5?
- What can you conclude if all expected counts are at least 5, and what can you not conclude from that check alone?
- A student says a table meets the condition because its average expected count is 12, although one cell has an expected count of 3. Identify the error.