When an Expected Count Is Too Small
In “Checking the Expected Count Condition for Chi-Square,” you learned to check every expected count against 5. If even one is below 5, the expected-count condition is not met, and the usual chi-square approximation may be unreliable. That does not mean the data are useless. It means you should pause before interpreting a chi-square p-value and consider whether the study can be improved in a way that preserves its purpose.
Two possible responses are to combine categories or to collect more data. Neither is an automatic fix. Combining categories changes what the analysis can tell you, so the categories must make sense together. Collecting more data can help, but only if the additional observations come from an appropriate design and provide enough information in the cells that are currently sparse.
Combining Categories Carefully
Combining categories means replacing two or more categories of one variable with a broader category. For example, “text message” and “phone call” might be combined as “direct contact” if the research question is about direct contact versus an online portal. This can increase expected counts in the revised table, but it also removes the distinction between text and phone preferences.
A defensible combination has a meaningful, subject-matter reason. The categories should represent related responses, and the broader category should still answer a useful question. Ideally, the decision is made when planning the study or before examining the results. Combining categories only because one expected count is too small—or because a revised analysis might give a preferred conclusion—can make the analysis misleading.
You may combine response categories when that makes sense for the question. You should not casually combine groups that the study was designed to compare. For example, merging two clinics in a homogeneity study changes which clinics are compared and may hide a difference between them. In an independence study, merging a category of one variable also changes the question about the relationship between the variables.
Worked Example: Combining Response Categories in a 3-by-3 Table
Worked Example: Combining Response Categories in a 3-by-3 Table
Suppose separate random samples of 40 patients are selected from each of three invented clinics. Each patient names one preferred appointment reminder: an online portal, a text message, or a phone call. The clinics each serve more than 400 patients, and each patient contributes one response. Researchers want to compare reminder preferences across the clinics.
| Clinic | Online portal | Text message | Phone call | Total |
|---|---|---|---|---|
| A | 28 | 10 | 2 | 40 |
| B | 25 | 13 | 2 | 40 |
| C | 25 | 13 | 2 | 40 |
| Total | 78 | 36 | 6 | 120 |
State: The research question is whether reminder preference has the same distribution across the three clinics. The original table has three clinic categories and three response categories.
Plan: Use the homogeneity expected-count formula discussed in “Expected Counts for Homogeneity Tests.” Check the expected counts under the null model that the clinics have the same reminder-preference distribution. First, verify the study conditions: the clinics supplied separate random samples; the samples are independent; every patient is counted once; and each sample of 40 is no more than 10% of its clinic’s population because each clinic serves more than 400 patients.
Do: The grand total is 120. For each clinic, the expected count for phone calls is its sample size times the pooled phone-call proportion. For clinic A, for example:
The same calculation gives 2 expected phone-call preferences in clinic B and 2 in clinic C. Since these expected counts are below 5, the expected-count condition is not met in the original 3-by-3 table.
The researchers consider combining text messages and phone calls into “direct contact.” This is reasonable for a revised question about online portals versus direct contact, but it no longer distinguishes between the two direct-contact options. The revised observed counts are:
| Clinic | Online portal | Direct contact | Total |
|---|---|---|---|
| A | 28 | 12 | 40 |
| B | 25 | 15 | 40 |
| C | 25 | 15 | 40 |
| Total | 78 | 42 | 120 |
For each clinic, the expected counts are:
Thus, each clinic has expected counts of 26 and 14 in the revised table, so every expected count is at least 5.
Conclude: The expected-count condition is met for the revised comparison of online portal preference with direct-contact preference. A chi-square test could address whether that two-category distribution differs across the clinics, provided the other conditions remain satisfied. It could not determine whether text or phone calls are preferred, because those responses have been combined.
Collecting More Data
If categories should remain distinct, a reasonable alternative may be to collect more observations. More data can raise expected counts, but the goal is not simply to increase the grand total. The additional observations must come from the relevant populations or groups, and the expected counts must be recalculated using the expanded table.
Planning can use the current pooled proportions as a guide. If a category currently represents a small share of the pooled responses, estimate how many observations would be needed for its expected counts to reach 5. This is only a projection: the proportions in future data may differ from the current ones. Once new observations are collected, recalculate every expected count and check the condition again.
The way the additional data are collected matters. For a homogeneity study, adding observations to only one group may leave small expected counts in other groups. Increasing each sample in a planned way can be more useful. For an independence study, the added observations must still come from the population the study is intended to describe. More data collected through a biased or dependent process do not repair those design problems.
Worked Example: Planning a Larger Homogeneity Study
Worked Example: Planning a Larger Homogeneity Study
Three invented schools take separate random samples of students about their preferred after-school activity: sports, arts, or gardening. The schools each have more than ten times as many students as the sample selected. Each student gives one response. The observed counts are:
| School | Sports | Arts | Gardening | Total |
|---|---|---|---|---|
| A | 25 | 23 | 2 | 50 |
| B | 50 | 46 | 4 | 100 |
| C | 75 | 69 | 6 | 150 |
| Total | 150 | 138 | 12 | 300 |
State: The question is whether the distribution of activity preferences is the same across the three schools. The researchers want to keep all three response categories distinct because they represent meaningfully different activities.
Plan: Calculate the expected counts under the null model of the same preference distribution at all three schools. Check the random-sample, independence, and 10% conditions: the samples are separate random samples, students contribute one response, and each sample is no more than 10% of its school’s population. Then consider whether a planned increase in sample size could raise the small expected counts.
Do: The grand total is 300. The expected gardening counts are:
The expected-count condition fails because two expected counts are below 5. Suppose the schools can increase their samples in the same proportions, tripling each sample size to 150, 300, and 450 students. If the preference proportions stay close to the current pooled proportions, the projected pooled totals are 450 for sports, 414 for arts, and 36 for gardening, with a grand total of 900. The projected gardening expected counts would be:
The other projected expected counts are 75 and 69 for school A, 150 and 138 for school B, and 225 and 207 for school C. Under this projection, every expected count is at least 5.
Conclude: Tripling the sample sizes would be expected to address the small-count problem if the observed category proportions remain similar. It does not guarantee that the condition will be met in the actual expanded data. After collecting the additional random samples, the researchers must build the new table and check its expected counts again.
When Categories Should Stay Separate
Sometimes the categories are important precisely because they are different. In that case, combining them can undermine the purpose of the study. The current data can still help plan a larger study. For example, if three groups have equal sample sizes and a rare response appears in 3 of 90 pooled observations, its current pooled proportion is \(3/90=1/30\). A rough planning target for each group is a sample size \(n\) satisfying \(n(1/30)\ge 5\), or \(n\ge150\). That projection depends on the rare response continuing to occur at roughly the same rate.
A target based on a current estimate is not a substitute for checking the expanded data. The actual pooled proportion could change, and expected counts depend on the final table’s margins. If it is not practical to collect enough appropriate data, state that the expected-count condition is not met and that the usual chi-square approximation may be unreliable. Do not report a standard chi-square conclusion as though the condition had passed.
Worked Example: Keep Important Categories Distinct
Three invented community centers take separate random samples of 30 visitors. Each visitor selects one preferred accessibility format: captions, translated text, or audio description. The centers each serve at least 300 visitors, and each visitor is counted once. The observed counts are:
| Center | Captions | Translated text | Audio description | Total |
|---|---|---|---|---|
| A | 17 | 12 | 1 | 30 |
| B | 16 | 13 | 1 | 30 |
| C | 15 | 14 | 1 | 30 |
| Total | 48 | 39 | 3 | 90 |
State: The question is whether the distribution of preferred accessibility formats is the same across the three centers. The researchers want to preserve all three formats because each represents a distinct access need.
Plan: Check the homogeneity conditions and calculate the expected counts. The samples are separate random samples, each visitor contributes one response, and each sample of 30 is no more than 10% of its center’s population because each center serves at least 300 visitors.
Do: The expected audio-description count in each center is:
The expected count of 1 is below 5, so the condition is not met. Combining audio description with another format would blur different access needs and change the question. As a planning estimate, the pooled audio-description proportion is \(3/90=1/30\). To project at least 5 expected audio-description selections in each equally sized center sample, solve:
If each center sampled 150 visitors and the pooled proportions stayed similar, the projected pooled totals would be 240 for captions, 195 for translated text, and 15 for audio description, out of 450 visitors. Each center’s projected expected counts would be:
Conclude: The current table does not meet the expected-count condition, and combining the formats would not preserve the study’s purpose. Increasing each random sample to about 150 visitors is a planning estimate, not a guarantee. The researchers should check the actual expected counts after collecting the additional data.
Common Mistakes and AP Exam Tip
- Combining categories only to make a count reach 5: Give a subject-matter reason for the broader category and state how the research question changes.
- Assuming a projected sample size guarantees success: Projections depend on future proportions resembling the current ones. Recalculate expected counts from the final table.
- Adding observations wherever it is easiest: The added data must come from the appropriate population or group and preserve the study’s sampling design.
- Changing groups instead of response categories without explanation: Merging groups changes the comparison and may hide differences the study was designed to detect.
- Treating a condition failure as a test result: A small expected count does not show that the null hypothesis is true or false. It indicates that the usual chi-square approximation may be unreliable.
For a strong AP response, identify the small expected count, explain why the usual approximation is a concern, and recommend a change that fits the research question. If categories are combined, describe the revised categories and the information lost. If more data are proposed, explain how the additional sample would be collected and make clear that the expected counts must be checked again.
Check Your Understanding
For each question, consider both the expected counts and the purpose of the study.
- A 3-by-3 homogeneity table has one expected count of 3.2. What should be checked before deciding whether to combine categories?
- Why might combining “text message” and “phone call” change the conclusion a study can address, even if the new table meets the expected-count condition?
- A researcher triples every group’s sample size and expects the rare-category counts to exceed 5. What must the researcher do after collecting the new data?
- Why is combining two clinics not automatically an appropriate way to increase expected counts in a homogeneity test?
- If the categories are important and cannot reasonably be combined, name a possible next step and one limitation of that approach.