Tutorials › AP Statistics › Checking the Expected Count Condition for Chi-Square

Expected counts and chi-square conclusions · Tutorial 565 of 1000

Checking the Expected Count Condition for Chi-Square

Practice checking every expected count against the AP Statistics minimum of 5 and explaining what the check tells you about a chi-square test.

Intermediate 9 min read

What You'll Learn

  • State the chi-square expected-count condition using the correct threshold.
  • Check every cell in an expected-count table, not just its total or average.
  • Explain why small expected counts can make a chi-square approximation unreliable.
  • Distinguish expected counts under the null model from observed sample counts.
  • Communicate what passing or failing the condition means for a chi-square test.

Why Check Expected Counts?

In “Expected Counts for Homogeneity Tests,” you used the null model to calculate expected counts for a table comparing groups. Expected counts are also central to deciding whether a chi-square test is appropriate: before relying on a chi-square reference distribution for a test statistic, check that every expected count is at least 5.

This is a condition about the expected table, not the observed table. An observed cell count might be 0 or 2 even when its expected count is at least 5; the condition is still checked using the expected count. Conversely, a large observed count in one cell does not compensate for an expected count below 5 in another cell.

Condition: For a chi-square test of independence or homogeneity, calculate the expected count for every cell under the null model. The expected-count condition is met only when every expected count is at least 5.

The expected-count formula depends on the test setting. For independence, as covered in “Expected Count Formula: Row Total Times Column Total Over Grand Total,” use the cell’s row total and column total. For homogeneity, as covered in “Expected Counts for Homogeneity Tests,” use the group total and pooled outcome total. In either setting, the result is a count predicted by the null model, and each cell must be checked.

Why the Minimum Matters

A chi-square test compares observed counts with expected counts using contributions of the form \((O-E)^2/E\). Its p-value is found by comparing the resulting statistic with a chi-square distribution. That comparison works well when the chi-square distribution is a reasonable approximation to the statistic’s sampling distribution under the null hypothesis.

Counts are discrete: a cell can contain only whole observations, and when expected counts are small, possible outcomes are relatively limited. A cell may even have a substantial chance of an observed count far from its expected count relative to the expected count’s size. In that situation, the distribution of the test statistic may not be well represented by the smooth chi-square curve used to calculate the p-value.

When expected counts are sufficiently large, the count patterns under the null model are generally better approximated by the probability model that leads to the chi-square reference distribution. Requiring every expected count to be at least 5 is the AP Statistics rule for checking whether that approximation is reasonable. It is a practical condition, not a promise that the approximation is exact.

Key idea: The minimum is checked cell by cell because a sparse expected cell can make the chi-square approximation unreliable even if the other expected counts are large.

Passing this condition does not replace the other checks. As discussed in “Random Condition for Chi-Square Procedures” and “Independence and 10 Percent Condition for Chi-Square,” also consider how the data were collected and whether observations are independent. A table with adequate expected counts does not fix a biased sample or dependent observations.

A Cell-by-Cell Checking Routine

After building the full expected-count table, scan every cell and identify the smallest expected count. If that minimum is at least 5, the expected-count condition is met. If even one expected count is below 5, it is not met.

1
Use the null model.
Confirm that the expected counts were calculated under the null hypothesis for the test being considered.
2
Inspect every cell.
Compare each expected count with 5. Do not check only the row totals, column totals, or average expected count.
3
Report the result.
State whether all expected counts are at least 5, and identify the smallest count or any counts below 5.
4
Explain the implication.
If the condition is met, say the expected-count condition supports using the chi-square approximation. If it is not met, say that approximation may be unreliable.

Keep decimal expected counts as calculated. An expected count is a model-based average, so it does not have to be a whole number. Do not round a value such as 4.8 up to 5 to make the condition appear satisfied.

Worked Example: Checking a Homogeneity Table

Worked Example: Checking a Homogeneity Table

Suppose separate random samples of patients at three invented clinics are asked which appointment reminder they prefer: text, phone call, or email. Each clinic serves more than ten times its sample size, and each patient is counted once. The observed counts and margins are shown below. Check the expected-count condition for a chi-square test of homogeneity.

ClinicTextPhone callEmailTotal
A20103060
B22184080
C302050100
Pooled total7248120240

State: The question is whether the distribution of reminder preference is the same across the three clinics. The expected-count condition requires every expected count under that null model to be at least 5.

Plan: Use the homogeneity expected-count formula, multiplying each clinic’s row total by the pooled total for a reminder category and dividing by the grand total. Then compare every result with 5. The samples are described as separate random samples, the clinic groups are independent, each patient contributes one response, and each sample is no more than 10% of its clinic’s patient population because each clinic serves more than ten times its sample size.

Do: The grand total is \(60+80+100=240\). The expected counts are:

$$ \begin{aligned} E_{\text{A,text}}&=\frac{(60)(72)}{240}=18, & E_{\text{A,call}}&=\frac{(60)(48)}{240}=12, & E_{\text{A,email}}&=\frac{(60)(120)}{240}=30,\\ E_{\text{B,text}}&=\frac{(80)(72)}{240}=24, & E_{\text{B,call}}&=\frac{(80)(48)}{240}=16, & E_{\text{B,email}}&=\frac{(80)(120)}{240}=40,\\ E_{\text{C,text}}&=\frac{(100)(72)}{240}=30, & E_{\text{C,call}}&=\frac{(100)(48)}{240}=20, & E_{\text{C,email}}&=\frac{(100)(120)}{240}=50. \end{aligned} $$

The expected-count table is:

ClinicTextPhone callEmail
A181230
B241640
C302050

The smallest expected count is 12, which is at least 5. Every expected count therefore meets the condition.

Conclude: The expected-count condition is met for this chi-square test of homogeneity. This supports using the chi-square approximation, provided the other conditions are also satisfied. It does not, by itself, establish that the reminder preferences differ across clinics.

Worked Example: One Small Expected Count Is Enough

Worked Example: One Small Expected Count Is Enough

Imagine a random sample of 60 members of an invented online community. Each person is classified by whether they use a particular accessibility feature and which of three devices they primarily use. The community has more than ten times as many members as the sample. The observed counts are:

PhoneTabletComputerTotal
Uses feature17715
Does not use feature3192345
Total4263060

For a chi-square test of independence, the expected count for each cell is its row total times its column total divided by the grand total. The six expected counts are:

$$ \begin{aligned} E_{\text{uses, phone}}&=\frac{(15)(4)}{60}=1, & E_{\text{uses, tablet}}&=\frac{(15)(26)}{60}=6.5, & E_{\text{uses, computer}}&=\frac{(15)(30)}{60}=7.5,\\ E_{\text{does not use, phone}}&=\frac{(45)(4)}{60}=3, & E_{\text{does not use, tablet}}&=\frac{(45)(26)}{60}=19.5, & E_{\text{does not use, computer}}&=\frac{(45)(30)}{60}=22.5. \end{aligned} $$

Two expected counts are below 5: 1 and 3. The expected-count condition is not met, even though the other four expected counts are at least 5 and the grand total is 60. The chi-square approximation may be unreliable for this table. A large total sample size cannot make the condition pass when individual expected cells remain too small.

Notice that the observed count of 1 in the first row and first column is not itself the reason the condition fails. The relevant value is that cell’s expected count, which is also 1 here. In another table, an observed count could be small while its expected count is adequate; the decision must still be based on expected counts.

Worked Example: A Count Exactly at the Minimum

Worked Example: A Count Exactly at the Minimum

Suppose separate random samples of 20 and 80 customers at two invented repair centers report whether they would recommend the service. The pooled totals are 25 “yes” responses and 75 “no” responses, for a grand total of 100. Assume each sample is no more than 10% of its center’s customer population, and each customer contributes one response. Check the expected-count condition for a homogeneity test.

For the sample of 20 customers, the expected counts are:

$$ E_{\text{small center, yes}}=\frac{(20)(25)}{100}=5, \qquad E_{\text{small center, no}}=\frac{(20)(75)}{100}=15 $$

For the sample of 80 customers, they are:

$$ E_{\text{large center, yes}}=\frac{(80)(25)}{100}=20, \qquad E_{\text{large center, no}}=\frac{(80)(75)}{100}=60 $$

The four expected counts are 5, 15, 20, and 60. The smallest is exactly 5, so the condition is met: “at least 5” includes a count equal to 5. Do not incorrectly require every expected count to be greater than 5.

Common Mistakes and AP Exam Tip

  • Checking observed counts instead: The condition applies to expected counts calculated under the null model. State which expected counts you checked.
  • Checking only the average: A high average does not rule out a sparse cell. Inspect every cell and report the smallest expected count.
  • Using “greater than 5” rather than “at least 5”: An expected count of exactly 5 meets the AP condition.
  • Rounding up a value below 5: Keep the calculated value. For example, 4.8 is below 5 and does not meet the condition.
  • Assuming a large grand total is enough: The check is cell by cell. A table can have many observations overall and still have a small expected count.
  • Claiming that passing proves the test is valid in every way: This check supports the chi-square approximation; it does not establish random selection, independence, or the other design conditions.

For a full-credit AP response, write a direct statement such as: “The smallest expected count is 6.5, so all expected counts are at least 5 and the expected-count condition is met. This supports using the chi-square approximation, assuming the other conditions are satisfied.” If a count is below 5, identify it and explain that the approximation may be unreliable. Do not claim that the null hypothesis is true or false based on a condition check; this step is about whether the reference distribution is appropriate.

Key takeaway: Calculate expected counts under the null model, inspect every cell, and compare the smallest with 5. All expected counts at least 5 support the chi-square approximation; one below 5 means the expected-count condition is not met.

Check Your Understanding

For each question, focus on expected counts under the stated null model.

  1. A table’s expected counts are 8, 12, 5, and 20. Is the expected-count condition met? Explain using the smallest count.
  2. A table has one expected count of 4.9 and all its other expected counts exceed 10. Is the condition met? Why or why not?
  3. Why does a chi-square test check expected counts rather than requiring every observed count to be at least 5?
  4. What can you conclude if all expected counts are at least 5, and what can you not conclude from that check alone?
  5. A student says a table meets the condition because its average expected count is 12, although one cell has an expected count of 3. Identify the error.