Tutorials › AP Statistics › Chi-Square Test With a Random Experiment

Expected counts and chi-square conclusions · Tutorial 577 of 1000

Chi-Square Test With a Random Experiment

Use a chi-square test to compare treatment-group outcome distributions and connect the result to what random assignment lets you conclude.

Intermediate 10 min read

What You'll Learn

  • Identify when a randomized experiment calls for a chi-square test of homogeneity.
  • State hypotheses about whether treatment groups have the same categorical outcome distribution.
  • Check random assignment, independence, and expected counts.
  • Calculate a chi-square statistic, degrees of freedom, and p-value.
  • Explain how random assignment supports a causal conclusion—and limits on that conclusion.

From Treatment Groups to a Causal Conclusion

In “Interpreting Chi-Square Results Without Overclaiming,” you learned that a chi-square test assesses an overall pattern in categorical counts, while the study design determines whether a causal conclusion is reasonable. Here we focus on randomized experiments: researchers assign experimental units to treatments at random, then compare the distributions of a categorical outcome across the assigned groups.

The relevant procedure is a chi-square test of homogeneity. The hypotheses concern whether the outcome distribution is the same across the treatment groups. A small p-value can provide evidence that the distributions differ. If treatment assignment was random and the experiment was properly conducted, that evidence can support a causal conclusion about the effect of assignment to the treatments on the outcome. The test supplies evidence of a difference; random assignment is what supports the causal interpretation.

Definition: In a randomized experiment with a categorical response, a chi-square test of homogeneity evaluates whether the response distribution is the same across the assigned treatment groups. When the experiment is properly conducted, a difference supported by the test can be interpreted as evidence of an effect of the assigned treatments on the response for the experimental units studied.

Keep two questions distinct. First, do the data provide convincing evidence that the outcome distributions differ? The chi-square test addresses this. Second, can a difference be attributed to treatment assignment? Random assignment supports this causal reasoning by distributing other influences among the groups by chance. Neither question alone establishes that the result will generalize to a broader population: generalizing requires an appropriate basis, such as random sampling.

Set Up the Test for an Experiment

The table layout may look like one used for a test of independence, but the study design tells you which question is being asked. In this setting, separate treatment groups are compared on one categorical response, so the test is one of homogeneity. As covered in “Stating Hypotheses for a Test of Homogeneity,” the null hypothesis says the response distributions are the same across all groups; the alternative says they are not all the same.

For each treatment-and-outcome cell, the expected count is the count predicted under the null hypothesis that all groups share the same outcome distribution. Calculate it as the treatment-group total multiplied by the total for that outcome category, divided by the grand total. Then use all observed and expected counts in the chi-square statistic, as in “The Chi-Square Statistic Formula.”

$$ E=\frac{(\text{treatment-group total})(\text{outcome-category total})}{\text{grand total}} \qquad\text{and}\qquad X^2=\sum \frac{(O-E)^2}{E} $$

Check the conditions before interpreting a p-value. Random assignment supports the random condition for this experiment. The observations should be independent: each experimental unit contributes to exactly one cell, and one unit’s response should not determine another’s. The Large Counts condition for a chi-square test is that every expected count is at least 5. The 10% condition is a sampling condition for random samples drawn without replacement from a finite population; it is not a substitute for random assignment in an experiment.

Conditions: For a chi-square test of homogeneity in a randomized experiment, verify that units were randomly assigned to treatment groups, observations are independent with each unit counted once, and every expected count is at least 5. Consider the 10% condition when the data also involve sampling without replacement from a finite population.

The degrees of freedom are \((r-1)(c-1)\), where \(r\) is the number of treatment groups and \(c\) is the number of outcome categories. The p-value is the upper-tail probability for the chi-square statistic with those degrees of freedom, assuming the null hypothesis is true. As in “Making a Decision in a Chi-Square Test,” compare the p-value with the chosen significance level before stating the conclusion.

1
State.
Define the experimental units, the treatments, and the categorical response. State whether the null hypothesis of equal distributions is being tested against a difference.
2
Plan.
Name a chi-square test of homogeneity and check random assignment, independence, and the expected-count condition.
3
Do.
Calculate expected counts, \(X^2\), degrees of freedom, and the upper-tail p-value.
4
Conclude.
Make the reject-or-fail-to-reject decision, describe the evidence about the outcome distributions, and explain what random assignment permits you to say about treatment effects.

Worked Examples: Testing Treatment-Group Outcomes

Worked Example: Three Growing Treatments and Plant Health

Suppose researchers individually pot 90 seedlings and randomly assign 30 to each of three growing treatments: a standard mixture, a compost mixture, or a mineral mixture. After four weeks, each plant is classified as healthy, showing minor damage, or showing severe damage. The invented counts are:

Assigned treatmentHealthyMinor damageSevere damageTotal
Standard mixture20304090
Compost mixture30303090
Mineral mixture40302090
Total909090270

State. Let the response be each plant’s health category after four weeks. \(H_0\) says the distribution of health category is the same for plants assigned to all three mixtures. \(H_a\) says the distributions are not all the same.

Plan. Use a chi-square test of homogeneity because three assigned treatment groups are compared on one categorical response. The seedlings were randomly assigned. Each plant is counted once, and the individually potted plants are treated as independent experimental units. The expected-count condition can be checked by calculating the expected counts. No random sample from a finite population is described, so a 10% condition is not relevant here.

Do. Each treatment row total is 90, each outcome-column total is 90, and the grand total is 270. Thus every expected count is \(90(90)/270=30\), which is at least 5. The observed counts differ from the expected counts by \(-10,0,10\) in the first row, \(0,0,0\) in the second, and \(10,0,-10\) in the third:

$$ X^2 =6\left(\frac{(10)^2}{30}\right) =6\left(\frac{100}{30}\right) =20 $$

There are \(r=3\) treatment groups and \(c=3\) outcome categories, so \(df=(3-1)(3-1)=4\). The upper-tail p-value is approximately \(0.0005\), rounded to four decimal places.

Conclude. At \(\alpha=0.05\), \(0.0005<0.05\), so reject \(H_0\). The experiment provides convincing evidence that the distribution of plant-health categories differs among the three assigned mixtures. Because the plants were randomly assigned to treatments, the result supports a causal conclusion that the assigned mixture affected the distribution of health outcomes for these plants under the greenhouse conditions studied. The table shows that the observed healthy counts were higher for the mineral mixture and lower for the standard mixture than expected under equal distributions; the overall test does not establish that any one category-specific difference is independently significant. The study also does not, by itself, justify generalizing to all plants or growing conditions.

Worked Example: Random Assignment to Reminder Routines

Suppose 120 adult volunteers are randomly assigned to use either a task reminder or no reminder for one week. The response is whether each volunteer completes a planned daily task on at least five days. The invented results are:

Assigned groupCompleted task goalDid not complete goalTotal
Task reminder362460
No reminder243660
Total6060120

State. The response is whether a volunteer meets the task-completion goal. \(H_0\) says the distribution of this response is the same for volunteers assigned to reminders and no reminders. \(H_a\) says the distributions differ.

Plan. Use a chi-square test of homogeneity because separate treatment groups are compared on a categorical response. Volunteers were randomly assigned, each contributes one response to one cell, and the groups do not overlap. The expected count in each cell is \(60(60)/120=30\), so all expected counts are at least 5. Random assignment—not random sampling—supports the causal interpretation.

Do. Each of the four observed counts differs from its expected count of 30 by 6 in absolute value. Therefore:

$$ X^2 =4\left(\frac{(36-30)^2}{30}\right) =4\left(\frac{36}{30}\right) =4.8 $$

The degrees of freedom are \((2-1)(2-1)=1\). The upper-tail p-value is approximately \(0.0285\), rounded to four decimal places.

Conclude. At \(\alpha=0.05\), \(0.0285<0.05\), so reject \(H_0\). The experiment provides convincing evidence that the distribution of task-goal completion differs between volunteers assigned to the reminder and no-reminder groups. Since assignment was random, the result supports the conclusion that assignment to the reminder routine affected the distribution of task-goal outcomes for the volunteers under these study conditions. In the sample, 36 of 60 volunteers assigned to reminders completed the goal, compared with 24 of 60 assigned to no reminder. The volunteers were not described as a random sample of adults, so do not automatically generalize the causal result to all adults.

Worked Example: A Randomized Experiment With Weak Evidence

Suppose 120 seeds are individually potted and randomly assigned to one of two watering schedules. Researchers classify each seedling’s emergence as early, on schedule, or late. The invented results are:

Assigned scheduleEarlyOn scheduleLateTotal
Schedule A22182060
Schedule B18222060
Total404040120

State and plan. \(H_0\) says the distribution of emergence timing is the same for seeds assigned to the two watering schedules; \(H_a\) says the distributions differ. A chi-square test of homogeneity is appropriate. The seeds were randomly assigned, each seedling is counted once, and the individual pots support treating observations as independent. Every expected count is \(60(40)/120=20\), meeting the Large Counts condition.

Do. The observed counts differ from their expected counts of 20 by \(2,-2,0\) in the first row and \(-2,2,0\) in the second row. Thus:

$$ X^2 =4\left(\frac{2^2}{20}\right) =4\left(\frac{4}{20}\right) =0.8 $$

The degrees of freedom are \((2-1)(3-1)=2\). The upper-tail p-value is approximately \(0.6703\), rounded to four decimal places.

Conclude. At \(\alpha=0.05\), \(0.6703>0.05\), so fail to reject \(H_0\). The experiment does not provide convincing evidence that the emergence-timing distributions differ between seeds assigned to the two watering schedules. Random assignment would support a causal interpretation if convincing evidence of a treatment difference were found; it does not turn a large p-value into evidence that the schedules have identical effects. The result does not prove that the schedules have no effect.

Common Mistakes and AP Exam Tip

  • Calling the procedure a test of independence just because the data are in a two-way table: Separate assigned treatment groups compared on one response call for a test of homogeneity. The design and question—not the table’s appearance—guide the choice.
  • Claiming the chi-square statistic itself proves causation: The statistic and p-value assess evidence against equal outcome distributions. Random assignment is the design feature that supports attributing a difference to the assigned treatment.
  • Making a causal claim after failing to reject: Random assignment permits causal reasoning, but the data still need to provide convincing evidence of a difference. A large p-value does not establish that a treatment has no effect.
  • Generalizing from volunteers to everyone: Random assignment is not random sampling. It supports causal conclusions for the experimental units and conditions studied, but does not automatically make the volunteers representative of a wider population.
  • Overstating what an overall result identifies: Rejecting the null says the distributions are not all the same. It does not establish that every treatment differs from every other treatment or that a particular outcome category has a separately significant difference.
  • Skipping condition checks: State that treatment assignment was random, explain why observations can be treated as independent, and verify that every expected count is at least 5. Do not confuse these checks with the 10% condition for sampling without replacement.

A full-credit conclusion names the response and the treatment groups, gives the decision and evidence in context, and then connects the result to the design. For a randomized experiment with convincing evidence, say that the assigned treatments affected the outcome distribution for the experimental units under the study conditions. Avoid claims about unstudied populations, specific category differences, or certainty unless the design and analysis support them.

Key takeaway: Use a chi-square test of homogeneity to compare categorical outcome distributions across randomly assigned treatment groups. The test evaluates evidence of a difference; random assignment supports the causal conclusion, while the study’s participants and conditions limit how far that conclusion can be generalized.

Check Your Understanding

For each question, distinguish what the chi-square test establishes from what the experimental design supports.

  1. Researchers randomly assign volunteers to three study routines and record one of four completion categories. Which chi-square procedure fits, and what does its null hypothesis say?
  2. A randomized experiment has expected counts of 4, 12, 8, and 16. Is the Large Counts condition met? Explain.
  3. A test rejects equal outcome distributions across randomly assigned treatments. What causal conclusion is supported, and what does the overall test not establish about particular categories?
  4. A randomized experiment with volunteers fails to reject its null hypothesis. Why is “the treatments have no effect” too strong?
  5. Why does random assignment support a causal conclusion but not, by itself, justify generalizing the result to all people?