Expected Counts When Comparing Groups
In “Building the Full Expected Counts Table,” you calculated expected counts for a table under the null model of independence. A chi-square test of homogeneity uses the same arithmetic, but the study design gives the calculation a different meaning: it compares the distribution of one categorical response across separate groups or populations.
For homogeneity, the null hypothesis says that the response distribution is the same in every group. To represent that shared distribution, combine the observed outcome totals across all the groups. Those pooled column totals determine the expected share of each outcome category in every group. Each group’s expected counts then reflect its own sample size and that common outcome distribution.
Keep the table orientation clear: put groups in rows and response categories in columns. The row total is the sample size for a group. A column total combines the observations in that response category from every group. The grand total is the number of observations across all groups.
The Formula and Its Homogeneity Meaning
The expected-count formula is the same one introduced in “Expected Count Formula: Row Total Times Column Total Over Grand Total.” For a cell, multiply its group’s row total by its outcome category’s pooled column total, then divide by the grand total.
You can also calculate the expected count by first finding the pooled proportion for an outcome category. Multiply that proportion by the sample size of the group. These are two forms of the same calculation:
The fraction in parentheses is the pooled proportion for that outcome. Under the homogeneity null model, the same proportion is applied to every group. A larger group therefore has larger expected counts in each category, while the expected category proportions remain the same across groups.
Do not calculate a separate outcome distribution from each group and use it to predict that same group’s expected counts. That would reproduce the group’s own observed pattern rather than model the common distribution assumed by the null hypothesis. Pooling the outcome totals is what makes the expected table represent the homogeneity null model.
Worked Example: Comparing Three Transit Regions
Worked Example: Comparing Three Transit Regions
Suppose separate samples of transit riders from three regions are asked which of three service improvements they would prioritize. The invented survey results are shown below. Find the expected counts under the null hypothesis that the response distribution is the same in all three regions.
| Region | More frequent service | Later evening service | Lower fares | Total |
|---|---|---|---|---|
| North | 30 | 32 | 18 | 80 |
| Central | 28 | 42 | 30 | 100 |
| South | 32 | 46 | 42 | 120 |
| Pooled total | 90 | 120 | 90 | 300 |
The pooled total for More frequent service is \(30+28+32=90\). The other pooled outcome totals are \(32+42+46=120\) for Later evening service and \(18+30+42=90\) for Lower fares. The grand total is \(80+100+120=300\). Thus the pooled proportions are \(90/300=0.30\), \(120/300=0.40\), and \(90/300=0.30\).
For example, the North region’s expected count for More frequent service is its row total times the pooled proportion for that response:
Apply the same pooled proportions to each region. The calculations for all nine cells are:
The resulting expected-count table is:
| Region | More frequent service | Later evening service | Lower fares | Expected row total |
|---|---|---|---|---|
| North | 24 | 32 | 24 | 80 |
| Central | 30 | 40 | 30 | 100 |
| South | 36 | 48 | 36 | 120 |
| Expected column total | 90 | 120 | 90 | 300 |
The expected row sums are \(24+32+24=80\), \(30+40+30=100\), and \(36+48+36=120\). The expected column sums are \(24+30+36=90\), \(32+40+48=120\), and \(24+30+36=90\). These reproduce the group totals and pooled outcome totals.
For example, the expected proportion choosing More frequent service is \(24/80=0.30\) in North, \(30/100=0.30\) in Central, and \(36/120=0.30\) in South. The common proportion comes from pooling the samples; it is not a claim that the observed proportions must match exactly.
Why Pooling Matters
The expected table represents what the group counts would look like if the response distributions were the same across groups. The pooled column totals provide one estimate of that shared distribution. Each group receives expected counts in proportion to its size.
This is closely related to the idea of a pooled proportion in the earlier tutorials on two-proportion tests: information from groups is combined to represent a common outcome rate under a null hypothesis. Here, however, the response can have more than two categories, and the goal is to build an expected-count table for a chi-square test of homogeneity.
The observed table and expected table serve different purposes. The observed table records what the samples actually produced. The expected table applies the pooled outcome distribution to each group’s sample size. A group with a larger sample will generally have larger expected counts, even though the expected proportions are shared.
Worked Example: Fractional Counts in a Park Survey
Worked Example: Fractional Counts in a Park Survey
Imagine separate samples of visitors at three invented parks are asked how they usually arrive: walking, bicycle, or public transit. The observed counts and margins are below. Calculate the expected counts under the assumption that the arrival-mode distribution is the same at all three parks.
| Park | Walking | Bicycle | Public transit | Total |
|---|---|---|---|---|
| Riverside | 20 | 21 | 9 | 50 |
| Hilltop | 25 | 32 | 13 | 70 |
| Meadow | 28 | 34 | 18 | 80 |
| Pooled total | 73 | 87 | 40 | 200 |
The pooled proportions are \(73/200=0.365\) for Walking, \(87/200=0.435\) for Bicycle, and \(40/200=0.20\) for Public transit. For Riverside, whose sample size is 50, the expected count for Walking is:
Calculate the remaining cells using the same pooled proportions:
The expected-count table is:
| Park | Walking | Bicycle | Public transit | Expected row total |
|---|---|---|---|---|
| Riverside | 18.25 | 21.75 | 10 | 50 |
| Hilltop | 25.55 | 30.45 | 14 | 70 |
| Meadow | 29.20 | 34.80 | 16 | 80 |
| Expected column total | 73.00 | 87.00 | 40 | 200 |
Check the row sums: \(18.25+21.75+10=50\), \(25.55+30.45+14=70\), and \(29.20+34.80+16=80\). Check the columns: \(18.25+25.55+29.20=73\), \(21.75+30.45+34.80=87\), and \(10+14+16=40\). The expected counts add to \(200\), the grand total.
The decimal values are appropriate: an expected count is a model-based average, not a prediction that a fractional number of visitors will actually be observed. Keep the decimals when checking the table rather than rounding cells individually.
Worked Example: Spotting the Wrong Pool
Worked Example: Spotting the Wrong Pool
Suppose samples of 64 and 96 customers at two invented service centers are asked which contact method they prefer: phone, email, or chat. Across both centers, the pooled totals are 72 for Phone, 48 for Email, and 40 for Chat, for a grand total of 160. Find the expected counts and identify why using each center’s own observed proportions would be a mistake.
The pooled proportions are \(72/160=0.45\) for Phone, \(48/160=0.30\) for Email, and \(40/160=0.25\) for Chat. Apply all three proportions to each center’s sample size:
For Center 1, the expected counts sum to \(28.8+19.2+16=64\). For Center 2, they sum to \(43.2+28.8+24=96\). Down the columns, the sums are \(28.8+43.2=72\), \(19.2+28.8=48\), and \(16+24=40\), reproducing the pooled totals.
Under the homogeneity null model, both centers use the same expected proportions: 0.45, 0.30, and 0.25. Using each center’s own observed proportions would make its expected counts mirror its observed distribution. It would not calculate the counts predicted by a shared distribution across the groups.
A Reliable Calculation Routine
Use this order to keep the pooled margins and group totals straight:
Record the observed response counts, each group’s row total, and the pooled outcome totals at the bottom.
Add the group totals or the pooled outcome totals. Both calculations should give the same number.
Divide each pooled column total by the grand total. These proportions describe the common distribution assumed by the homogeneity null model.
Multiply each group’s row total by each pooled outcome proportion, or use the row-total-times-column-total formula.
Expected counts should add to each group total across a row and to each pooled outcome total down a column.
Common Mistakes and AP Exam Communication
- Using a group’s own outcome total in place of a pooled total: The column total must combine that response category across all groups.
- Using the wrong denominator: Divide by the grand total, not by an individual group’s row total. The group total is multiplied by the pooled outcome proportion after the proportion is found.
- Pooling group sizes instead of outcomes: Row totals describe how many observations are in each group. Pool the response-category counts down each outcome column.
- Assuming equal group sizes: The same pooled proportions apply across groups, but expected counts differ when the group totals differ.
- Rounding fractional counts too early: Retain decimals while checking row and column totals.
- Confusing observed and expected counts: Observed counts come from the samples. Expected counts are calculated from the pooled outcome distribution under the null model.
For a clear AP response, state that the expected counts are calculated under the null hypothesis that the response distribution is the same across groups. Show the pooled column totals and grand total, give the formula or pooled proportions, and calculate every cell. Then check that expected row sums match the group totals and expected column sums match the pooled outcome totals.
Check Your Understanding
Use the pooled outcome totals to calculate expected counts under a homogeneity null model.
- Three groups have totals 40, 60, and 100. The pooled totals for two response categories are 80 and 120, and the grand total is 200. Find both expected counts for the group of 60.
- Why are the outcome totals pooled across groups when calculating expected counts for a homogeneity test?
- For group totals 50 and 70, pooled outcome totals 36 and 84, and grand total 120, calculate all four expected counts.
- A student calculates expected counts by applying a different response proportion to each group. Explain why this does not represent the homogeneity null model.
- What row and column totals should a completed expected-count table reproduce?