Tutorials › AP Statistics › What Expected Counts Represent Under Independence

Expected counts and chi-square conclusions · Tutorial 561 of 1000

What Expected Counts Represent Under Independence

Understand expected counts as the table pattern the null model of independence would produce, and see how to interpret them in context.

Intermediate 10 min read

What You'll Learn

  • Explain what an expected count represents under independence
  • Use overall category proportions to reason about expected counts in each group
  • Distinguish expected counts from observed counts and guaranteed outcomes
  • Explain why expected counts can be fractional
  • Check how expected counts relate to the margins and chi-square conditions

What Does an Expected Count Mean?

A two-way table records observed counts: how many individuals in a sample actually fall into each combination of categories. A chi-square test compares those observed counts with another set of counts—the expected counts. To understand what that comparison means, picture the table that would result if the two categorical variables were independent in the population.

As in “Stating Hypotheses for a Test of Independence,” the null hypothesis \(H_0\) says that the variables are independent in the population. Under that model, knowing an individual’s category for one variable does not change the distribution of the other variable. Expected counts describe the cell counts that fit this no-association model while keeping the table’s row and column totals in view.

Definition: An expected count is the count a cell would have, on average, under the null model of independence, given the table’s margins. It is not the observed count and does not promise what will occur in a future sample.

One way to understand the idea is to begin with the overall distribution across the columns. If 40% of all sampled individuals are in one column category, then under independence we would expect about 40% of the individuals in each row to be in that category. Apply the overall column proportions to each row total, and the resulting counts form the expected pattern. The expected pattern has the same row totals and column totals as the observed table.

This is the table-level version of independence: the distribution across columns is the same from row to row. Equivalently, the distribution across rows is the same from column to column. In “Computing Conditional Distributions From a Two-Way Table,” you learned to compare row and column percentages. Expected counts give a count-based way to express the pattern those percentages would follow if there were no association.

Expected Does Not Mean Guaranteed

An expected count is not a prediction that a particular cell must contain exactly that many individuals. It is a benchmark from the null model. Random samples can produce counts above or below that benchmark, even when the population variables really are independent. The chi-square statistic, introduced in “The Chi-Square Statistic Formula,” measures how far the observed table is from its expected pattern across all cells.

Expected counts can also be fractional. A model can assign an expected count of 12.58 to a cell even though a real table cannot contain 12.58 people. This is not an error: the count represents an average under the model, not a literal partial individual. In repeated samples under independence, cell counts vary, and their average can be non-whole.

The margins matter because they describe how many observations are in each row and column. Expected counts preserve those totals while distributing the counts according to independence. That makes the comparison fair: the test looks at whether the observed cell pattern departs from independence, rather than treating the overall sample size or category totals as unexpected.

Key distinction: Observed counts describe what happened in the sample. Expected counts describe what the counts would average under the null model of independence, with the same table margins. A difference between them is not, by itself, proof of an association.

Worked Example: Expected Counts Across Two Routes

Worked Example: Expected Counts Across Two Routes

A transportation planner takes a random sample of 300 riders and records each rider’s usual route and preferred way to receive service alerts. The observed counts are shown below. The overall column totals are 120 for text, 72 for email, and 108 for app notifications.

Usual routeTextEmailApp notificationTotal
Route A552520100
Route B654788200
Total12072108300

Suppose the planner is considering the null model that usual route and preferred alert method are independent. In the full sample, 120 of 300 riders prefer text, so the overall text proportion is \(120/300=0.40\). Under independence, Route A’s expected text count is 40% of its 100 riders: \(100(0.40)=40\). Route B’s expected text count is 40% of its 200 riders: \(200(0.40)=80\).

The same reasoning applies to the other alert methods. Email accounts for \(72/300=0.24\) of the sample, and app notifications account for \(108/300=0.36\). Therefore, the expected counts are:

$$ \begin{aligned} \text{Route A: }&100(0.40)=40,\quad 100(0.24)=24,\quad 100(0.36)=36\\ \text{Route B: }&200(0.40)=80,\quad 200(0.24)=48,\quad 200(0.36)=72 \end{aligned} $$

For example, the expected count of Route B riders who prefer email is 48. This means that if alert preference had the same distribution for riders on each route, the model would assign about 24% of Route B’s 200 riders to email preference. It does not mean that exactly 48 riders must prefer email.

The expected counts add to the same row totals as the observed table: \(40+24+36=100\) for Route A and \(80+48+72=200\) for Route B. They also add to the same column totals: \(40+80=120\) for text, \(24+48=72\) for email, and \(36+72=108\) for app notifications. Under independence, both routes have the same expected alert-method proportions as the overall sample.

Notice that the observed counts do not match the expected counts exactly. For Route A, the observed text count is 55 compared with an expected count of 40. That difference describes a sample departure from the null pattern; deciding whether the table provides convincing evidence of an association requires the chi-square test, not just noticing one cell.

Worked Example: Why an Expected Count Can Be a Decimal

Worked Example: Why an Expected Count Can Be a Decimal

A library takes a random sample of 100 visitors and records age group and preferred program format. The observed table has 37 visitors in the younger group and 63 in the older group. Across all visitors, 34 prefer a virtual program, 33 prefer a hands-on program, and 33 prefer a talk.

Age groupVirtualHands-onTalkTotal
Younger15101237
Older19232163
Total343333100

Under independence, each age group would have the same program-format proportions as the full sample: 34% virtual, 33% hands-on, and 33% talk. For the younger group, the expected counts are found by applying those proportions to its row total of 37:

$$ 37(0.34)=12.58,\qquad 37(0.33)=12.21,\qquad 37(0.33)=12.21 $$

For the older group, use its row total of 63:

$$ 63(0.34)=21.42,\qquad 63(0.33)=20.79,\qquad 63(0.33)=20.79 $$

The expected count 12.58 for younger visitors who prefer a virtual program is meaningful even though a count of people cannot be fractional. It says that the independence model allocates 34% of the 37 younger visitors to that cell, which gives an average count of 12.58. In an actual sample, the observed count is a whole number—in this table, 15.

The decimals also preserve the margins: \(12.58+21.42=34\), the virtual column total. For hands-on programs, \(12.21+20.79=33\), and the same is true for talks. Expected counts are calculations from the null model, so rounding them too early can make their totals appear not to add exactly. Keep enough digits during calculations and round only when reporting.

This example illustrates why an expected count is not a forecast for a specific sample. It is a model-based comparison value. The chi-square condition for expected counts is checked using the expected values, including decimal values, rather than by rounding them to whole people first.

Worked Example: A Complete Expected-Count Plan

Worked Example: A Complete Expected-Count Plan

A district selects a random sample of 300 shuttle riders and records the route taken and whether the shuttle arrived on schedule. The district wants to know whether route and on-schedule status are associated. Assume the district has at least 3,000 shuttle riders, and that each sampled rider contributes one independent observation.

RouteOn scheduleLateTotal
North701080
Central7624100
South9426120
Total24060300

State. \(H_0\): Route and on-schedule status are independent among the district’s shuttle riders. \(H_a\): Route and on-schedule status are associated among the district’s shuttle riders.

Plan. The sample contains individuals classified by two categorical variables, and the question asks about association, so a chi-square test of independence is appropriate. The random condition is met because the district selected a random sample. Treating riders’ observations as independent is reasonable under the stated assumption. The 10% condition is met because \(300\) is no more than 10% of a population of at least \(3{,}000\). The expected-count condition must also be checked.

Do. In the full sample, \(60/300=0.20\), or 20%, of riders were late. Under independence, 20% of the riders on each route would be expected to be late. Thus the expected late counts are \(80(0.20)=16\), \(100(0.20)=20\), and \(120(0.20)=24\). The expected on-schedule proportion is \(240/300=0.80\), giving expected counts \(80(0.80)=64\), \(100(0.80)=80\), and \(120(0.80)=96\).

The expected table is therefore:

RouteExpected on scheduleExpected lateTotal
North641680
Central8020100
South9624120
Total24060300

Every expected count is at least 5, so the expected-count condition is met. Each route has the same expected late proportion, 20%, and the same expected on-schedule proportion, 80%, as the full sample.

Conclude about the expected counts. If route and on-schedule status were independent, the model would expect 16 late riders on the North route, 20 on the Central route, and 24 on the South route. These expected counts provide the benchmark for comparing the observed table. They do not, by themselves, show whether the district has convincing evidence of an association; that judgment requires the test statistic and p-value.

Common Mistakes and What to Say Instead

  • Calling an expected count a prediction: Avoid saying, “The district will have 16 late North-route riders.” Say, “Under independence, the expected count is 16,” and explain that it is a model-based average.
  • Confusing observed and expected counts: Observed counts come directly from the sample. Expected counts come from the no-association model. Identify which kind of count you are discussing.
  • Assuming expected counts must be whole numbers: A fractional expected count is valid. Do not round it before checking the chi-square expected-count condition.
  • Thinking independence means equal counts in every row: Row totals can differ. Independence means the expected proportions are the same across rows, not that each row must have identical counts.
  • Declaring association from one difference: Observed and expected counts will often differ because of sampling variation. Use the chi-square test to assess whether the overall departures provide convincing evidence against independence.
  • Forgetting the margins: The expected table retains the observed row and column totals. If your expected values do not add back to those totals, check the calculations.
AP Exam Tip: In context, describe an expected count as the number the null model of independence would expect in that cell, given the margins. If asked whether that count proves association, clarify that the expected table is a benchmark; the test’s statistic and p-value are needed to evaluate evidence.
Key takeaway: Expected counts show the table pattern that would fit independence. They apply the overall category distribution to each row, preserve the table margins, and provide a comparison for observed counts. They are model-based averages, not guaranteed sample results.

Check Your Understanding

Use the idea of an independence model to answer each question.

  1. In a two-way table, what does an expected count represent, and how is it different from an observed count?
  2. If 30% of all sampled customers choose curbside pickup, what proportion would the independence model use for curbside pickup in each customer-type row?
  3. Why can an expected count be 8.4 even though a cell cannot contain 8.4 people?
  4. If the row totals are different, does independence require the expected counts in those rows to be equal? Explain.
  5. A cell’s observed count is 10 and its expected count is 7. Does that difference alone establish an association? Why or why not?