Tutorials › AP Statistics › Exam Practice on Expected Counts and Chi-Square Conclusions

Expected counts and chi-square conclusions · Tutorial 580 of 1000

Exam Practice on Expected Counts and Chi-Square Conclusions

Learn to organize chi-square free responses so your expected counts, calculations, p-value, and conclusion form one clear argument.

Intermediate 9 min read

What You'll Learn

  • Match an expected-count calculation to the null model and table margins.
  • Show how to calculate the chi-square statistic from cell contributions.
  • Find degrees of freedom from the table dimensions.
  • Interpret a chi-square p-value as an upper-tail probability under the null hypothesis.
  • Write a decision and conclusion that answer the research question in context.
  • Check that each part of a free response is explicit and consistent.

Turn a Chi-Square Prompt Into a Complete Response

In “Chi-Square Test of Homogeneity Full Worked Example,” you followed a test from its hypotheses through its conclusion. This tutorial focuses on a common exam task: explaining the expected counts, statistic, degrees of freedom, p-value, and conclusion clearly enough that each part can be evaluated on its own.

A strong response does not just list calculator output. It shows how the expected counts come from the null model, how the statistic measures departures from those counts, why the degrees of freedom fit the table, and what the p-value says in context. The study design still matters: as in “Independence Versus Homogeneity: Choosing the Right Test,” identify whether the data come from one sample classified twice or separate groups compared on one response.

Key response sequence: Identify the null model; give or calculate expected counts; verify the expected-count condition; show the chi-square contributions and sum; calculate \(df\); report the upper-tail p-value; compare it with \(\alpha\); and conclude in context.

The expected count for a cell depends on the null model, but for either a test of independence or homogeneity it is calculated from the relevant row total and column total:

$$ E=\frac{(\text{row total})(\text{column total})}{\text{grand total}} \qquad\qquad X^2=\sum\frac{(O-E)^2}{E} $$

As covered in “Degrees of Freedom for a Two-Way Table,” a table with \(r\) row categories and \(c\) column categories has \(df=(r-1)(c-1)\). The p-value is the probability, assuming \(H_0\) is true, of getting a chi-square statistic at least as large as the observed statistic. It is an upper-tail probability.

A Useful Free-Response Routine

When a prompt asks for several parts, label them. If it requests only the expected counts, statistic, degrees of freedom, p-value, and conclusion, you do not need to rewrite every detail of the test setup. Still, give enough information to explain why your calculations and conclusion fit the study.

1
Expected counts.
Show the formula and at least one substitution. Give the full expected-count table when practical, then check that every expected count is at least 5.
2
Statistic and degrees of freedom.
Show cell contributions or enough arithmetic to verify their sum. Find \(df\) from the table dimensions, not from the number of observed counts.
3
P-value and decision.
Find the upper-tail probability for the observed statistic and the correct \(df\). Compare it with the stated significance level.
4
Conclusion.
Say whether you reject or fail to reject \(H_0\), then state what the evidence indicates about the population association or group distributions.

If using a calculator, enter the observed counts in the correct table order and use the appropriate chi-square test. The calculator’s expected-count matrix can check your hand calculations. For a p-value from a stated statistic and degrees of freedom, use the upper tail, such as \(\chi^2\mathrm{cdf}(X^2,1\mathrm{E}99,df)\). A right statistic paired with the wrong \(df\), or a lower-tail probability, does not answer the test question.

Worked Examples: From Counts to Conclusions

Worked Example: Two Transit Centers and Express Service

A prompt describes separate random samples of 80 riders from Center A and 120 riders from Center B. Each rider is classified by whether they chose an express or regular service. Each center serves 2,000 riders. The observed counts are shown below. At \(\alpha=0.05\), answer the expected-count, statistic, degrees-of-freedom, p-value, and conclusion parts.

CenterExpressRegularTotal
A503080
B5070120
Total100100200

Expected counts. Under the null model, the express and regular service distributions are the same at both centers. For Center A and express service:

$$ E=\frac{(80)(100)}{200}=40 $$

Using the same row-total, column-total, and grand-total formula for each cell gives:

$$ \begin{array}{c|cc} &\text{Express}&\text{Regular}\\ \text{Center A}&40&40\\ \text{Center B}&60&60 \end{array} $$

Every expected count is at least 5, so the expected-count condition is met. The sampling design also supports the test: the samples are random and separate, and each rider contributes one response. The 10% checks are \(80\leq0.10(2{,}000)=200\) and \(120\leq0.10(2{,}000)=200\).

Statistic and degrees of freedom. The four cell contributions are:

$$ X^2 =\frac{(50-40)^2}{40} +\frac{(30-40)^2}{40} +\frac{(50-60)^2}{60} +\frac{(70-60)^2}{60} =2.5+2.5+1.6667+1.6667 \approx 8.3333 $$

There are two rows and two columns, so \(df=(2-1)(2-1)=1\). With \(X^2\approx8.3333\) and \(df=1\), the upper-tail p-value is approximately \(0.0039\), rounded to four decimal places.

Conclusion. Since \(0.0039<0.05\), reject \(H_0\). The samples provide convincing evidence that the distribution of service choice differs between the two transit centers. This is an overall comparison of the distributions, not a claim that every response category differs by the same amount.

Worked Example: Three Garden Plots and Seedling Condition

A student studies a random sample of 180 seedlings from a large greenhouse. Each seedling is classified by garden plot (East, Middle, or West) and condition (strong, average, or weak). The invented counts are:

PlotStrongAverageWeakTotal
East25201560
Middle20202060
West15202560
Total606060180

The prompt asks whether plot and condition are associated. This is a chi-square test of independence because one sample is classified by two categorical variables. Under \(H_0\), plot and seedling condition are independent.

Expected counts. For East and strong, \(E=(60)(60)/180=20\). The row totals and column totals are each 60, so the expected count is 20 in all nine cells. Every expected count is at least 5. The random sample provides the random condition, each seedling is counted once, and if the greenhouse has at least 1,800 seedlings, the 10% condition is met because \(180\leq0.10(1{,}800)\).

Statistic and degrees of freedom. Four cells differ from expectation by 5 in absolute value; the other five have \(O-E=0\). Thus:

$$ X^2 =4\left(\frac{5^2}{20}\right)+5(0) =4(1.25) =5 $$

There are \(r=3\) plot categories and \(c=3\) condition categories, so \(df=(3-1)(3-1)=4\). For \(X^2=5\) and \(df=4\), the upper-tail p-value is approximately \(0.2873\), rounded to four decimal places.

Conclusion. At \(\alpha=0.05\), \(0.2873>0.05\), so fail to reject \(H_0\). The sample does not provide convincing evidence of an association between garden plot and seedling condition in the greenhouse. This result does not prove that the variables are independent; it says the observed table does not give strong evidence against independence.

Worked Example: Grade Level and Getting to School

An exam-style prompt describes a random sample of 180 students from a large high school. Each student reports a grade level, junior or senior, and a primary way of getting to school: bus, bicycle, or walking. The table gives the sample counts. Test whether grade level and transportation method are associated at \(\alpha=0.05\).

GradeBusBicycleWalkingTotal
Junior40302090
Senior20304090
Total606060180

State. \(H_0\): grade level and primary transportation method are independent among students at this high school. \(H_a\): grade level and primary transportation method are associated among students at this high school.

Plan. Use a chi-square test of independence because one sample is classified by two categorical variables. The sample is random, and each student contributes to one cell. Assuming the school has at least 1,800 students, \(180\leq0.10(1{,}800)=180\), so the 10% condition is met. We will check the expected-count condition.

Do. For junior and bus, the expected count is \(90(60)/180=30\). The same row and column totals give an expected count of 30 for all six cells. All expected counts are at least 5. Four cells differ from expectation by 10 and two have no difference, so:

$$ X^2 =4\left(\frac{10^2}{30}\right)+2(0) =\frac{400}{30} \approx13.3333 $$

There are two rows and three columns, so \(df=(2-1)(3-1)=2\). The upper-tail p-value for \(X^2\approx13.3333\) with 2 degrees of freedom is approximately \(0.0013\), rounded to four decimal places.

Conclude. Since \(0.0013<0.05\), reject \(H_0\). The sample provides convincing evidence of an association between grade level and primary transportation method among students at this high school. Because this is an observational sample, the conclusion does not establish that grade level causes students to choose a particular transportation method.

Common Mistakes and AP Exam Tip

  • Reporting the wrong expected counts: Expected counts come from the null model and the table margins, not from copying observed counts or averaging cells. Show a formula substitution so the method is visible.
  • Adding contributions incorrectly: Each contribution is \((O-E)^2/E\). Square the difference, divide by that cell’s expected count, and then add all cells. Keep enough digits during the calculation and round the final statistic consistently.
  • Using the wrong degrees of freedom: For a two-way table, use \((r-1)(c-1)\). A \(2\)-by-\(3\) table has \(2\) degrees of freedom, not the number of cells or the total sample size.
  • Using the wrong tail: A chi-square p-value is an upper-tail probability. It measures how likely a statistic at least as large as the observed one would be if \(H_0\) were true.
  • Writing a conclusion without context: “Reject the null” is incomplete by itself. Name the population variables or response distributions and state what the evidence supports.
  • Claiming proof: Rejecting \(H_0\) does not prove an alternative claim, and failing to reject does not prove the null. Use “convincing evidence” when rejecting and “do not provide convincing evidence” when failing to reject.

For full credit, make the chain of reasoning easy to follow: expected counts under the null, condition check, \(X^2\), \(df\), upper-tail p-value, decision, and an in-context interpretation. A conclusion should match the study design and avoid claims the test does not establish.

Key takeaway: In a chi-square free response, show where expected counts come from, add the cell contributions carefully, calculate \(df\) from the table dimensions, and interpret the upper-tail p-value in context. Explicit, connected steps make the statistical conclusion clear.

Check Your Understanding

Answer each item as you would in a short free response, showing calculations where appropriate.

  1. A cell has row total 72, column total 90, and grand total 360. Find its expected count and show the substitution.
  2. A table has four row categories and two column categories. Find the degrees of freedom.
  3. A chi-square test gives \(X^2=6.4\), \(df=2\), and \(p=0.0408\). At \(\alpha=0.05\), state the decision and write a conclusion in context for a test of independence.
  4. Explain why a chi-square p-value is an upper-tail probability.
  5. A student writes, “We failed to reject the null, so the variables are independent.” Identify the error and give a more careful statement.