Start With How the Data Were Collected
In “Chi-Square Test for Independence: Purpose and Setting” and “Chi-Square Test for Homogeneity: Purpose and Setting,” you met two procedures that both use counts in a two-way table. Their tables can look exactly alike. The important difference is not the table’s shape; it is how the individuals entered the study and what question the researchers are asking.
For independence, researchers take one sample from a population and record two categorical variables for each individual. They ask whether those variables are associated in that population. For homogeneity, researchers take separate samples from different populations or create groups through an experiment, then record one categorical outcome for each individual. They ask whether the outcome distribution is the same across the groups.
The distinction can feel subtle because each study records two categories per person in some sense. A homogeneity study records the group or treatment and the outcome. An independence study records two variables about each sampled individual, with neither variable defining a separately sampled group. Focus on the data-collection plan: Was there one sample classified in two ways, or were groups selected or assigned first and then compared?
A Practical Decision Sequence
Read the study description before looking at the table. Identify the unit of observation—the person, animal, object, or other individual that contributes data—and trace how that unit was selected or assigned. Then name the variables and the question. This sequence prevents a common shortcut: choosing a test merely because the table has rows and columns.
Was one sample taken from a population and then classified in two ways? Or were separate samples taken from distinct populations, or individuals assigned to different treatment groups?
For one sample, identify both categorical variables. For separate samples or treatment groups, identify the single categorical outcome recorded for each individual.
“Are these two variables associated?” signals independence. “Is the distribution of this outcome the same across these groups?” signals homogeneity.
Confirm that each individual contributes to one cell, observations are independent, the relevant randomization or sampling conditions are addressed, the 10% condition is checked when sampling without replacement, and every expected count is at least 5.
The final step connects test choice to the conditions discussed in the earlier tutorials. A matching design does not by itself guarantee that a procedure is appropriate: the data still need to meet the relevant conditions. Also distinguish the 10% condition from the expected-count condition. The former concerns sampling without replacement from a population; the latter concerns counts predicted under the null hypothesis.
Why the Table Does Not Tell You the Test
A two-way table organizes counts for two categorical classifications. The row and column labels could describe two characteristics measured on one sample, or they could describe the group membership and outcome in a comparison of separate groups. In either case, the table may have the same number of rows and columns, and the expected-count calculation may use the same row-total, column-total, and grand-total structure.
For independence, the null hypothesis says the two categorical variables are independent, or equivalently, that there is no association between them in the population. For homogeneity, the null hypothesis says the outcome distributions are the same across the populations or groups. The alternative statements differ in the same way: an association for independence; at least one group distribution differing for homogeneity. The procedure name follows the design and question, not a visual feature of the table.
The selection plan also affects what a conclusion can reach. An appropriate random sample can support generalization to the population represented by that sample. Random assignment can support a causal conclusion about treatment effects. These are different benefits of study design. As in the earlier tutorials on the two chi-square settings, neither a small p-value nor large expected counts can repair a biased sampling method.
Worked Example: One Sample, Two Characteristics
Worked Example: One Sample, Two Characteristics
A fictional city researcher selects one random sample of 300 residents. For each resident, the researcher records the person’s usual way of commuting—walking, bicycle, or transit—and whether the person usually travels during peak or off-peak hours. The research question is whether commute method and travel time are associated.
Trace the design. The researcher selected one sample of residents from the city. No separate samples were selected for commute method, and no treatment groups were assigned. Each sampled resident is classified by two categorical variables: commute method and travel time.
Match the question. The question asks whether two variables measured on one sample are associated. This is the setting for a chi-square test of independence, not a test of homogeneity.
Check what would be needed before conducting the test. The random sample supports the Random condition for inference about the city’s residents, provided the sampling process was carried out as described. If residents were sampled without replacement, the sample of 300 should be no more than 10% of the city’s resident population. Each resident should contribute to exactly one combination of commute method and travel time, and one resident’s classification should not determine another’s. Before carrying out the test, calculate the expected counts under the independence null model and verify that each is at least 5.
State the scope. If conditions are met and the test gives convincing evidence of an association, the conclusion concerns an association between usual commute method and usual travel time among the city’s residents. This observational study would not show that commute method causes a person to travel at a particular time.
Worked Example: Separate Samples, One Outcome
Worked Example: Separate Samples, One Outcome
A fictional environmental team takes separate random samples of households in three neighborhoods. Each sampled household reports its main way of reducing food waste: meal planning, freezing leftovers, or composting. The team asks whether the distribution of reported strategies is the same in all three neighborhoods.
Trace the design. The team sampled households separately from three neighborhood populations. The neighborhoods are the groups being compared. The categorical outcome is the household’s main food-waste strategy.
Match the question. The question asks whether one outcome’s distribution is the same across separate populations. This is a chi-square test of homogeneity. Although neighborhood and strategy can be displayed as two table variables, the sampling plan—not the table layout—makes this a homogeneity setting.
Check what would be needed before conducting the test. Each sample should be random from its neighborhood population. If sampling without replacement, check the 10% condition separately for each neighborhood. The samples should be independent, and each household should contribute to exactly one strategy category in one neighborhood. Calculate the expected counts under the null hypothesis that the strategy distributions are the same, and check that every expected count is at least 5.
State the scope. If the conditions are met and the test provides convincing evidence of differing distributions, the random samples can support a conclusion that food-waste strategy distributions are not the same in the represented neighborhood populations. The result would not establish why the distributions differ.
Worked Example: Same Table, Different Study Design
Worked Example: Same Table, Different Study Design
Consider this fictional table of counts about whether a participant completed a short online lesson:
| Group label | Completed | Did not complete | Total |
|---|---|---|---|
| Morning | 42 | 18 | 60 |
| Evening | 35 | 25 | 60 |
| Total | 77 | 43 | 120 |
The numbers and table could arise from either of two studies. The correct test changes with the design, even though the table stays identical.
Version A: one sample classified in two ways. Researchers take one random sample of 120 adult learners from a program. They record each learner’s preferred lesson time (morning or evening) and whether that learner completes the lesson. The question asks whether preference and completion status are associated. One sample was classified by two variables, so the setting is a chi-square test of independence. A preference recorded for each sampled person is not the same as assigning that person to a time condition.
Version B: separate groups compared. Researchers take separate random samples of 60 learners who attend morning sessions and 60 who attend evening sessions, then record completion status for each learner. The question asks whether completion distributions are the same across the two groups. This is a chi-square test of homogeneity. The group samples were formed separately, and completion is the one categorical outcome being compared.
Compare the interpretations. In Version A, a result would address whether two characteristics are associated among the population represented by the one sample. In Version B, a result would address whether the completion distribution differs between the two learner populations. The table itself does not reveal which design occurred; the study description does.
Consider an experiment variation. If 120 learners were randomly assigned to morning or evening lesson times and completion was recorded, the design would also fit homogeneity: treatment groups are compared on one categorical outcome. Random assignment could support a causal conclusion about lesson time for the experimental units, assuming the experiment was appropriately conducted. It would not automatically support generalization to all adult learners unless the study also had an appropriate random sample.
In every version, the researchers must still verify the relevant independence and expected-count conditions before performing a chi-square test. For the two random-sample versions, they must also address the 10% condition when sampling without replacement.
Worked Example: When Neither Standard Setting Fits as Described
Worked Example: When Neither Standard Setting Fits as Described
A fictional sports-science class asks 40 players to try two shoe inserts, one after the other, and records each player’s preferred insert: A or B. The instructor proposes putting the number of A preferences and B preferences into rows for the first insert and columns for the second insert, then using a chi-square test of independence.
Trace the design. There is one set of players, and each player contributes a paired response across two conditions. The two entries associated with a player are linked; they are not two independent categorical measurements from unrelated individuals.
Decide whether either standard setting applies. This is not the usual independence setting of one sample with two categorical variables whose observations are independently classified across individuals. It is also not the usual homogeneity setting of separate independent samples or treatment groups, each contributing one outcome. The repeated measurements on each player create paired categorical data.
Explain the consequence. Do not choose a standard chi-square test of independence or homogeneity just because the instructor can arrange the paired counts in a table. The design must match the assumptions behind the procedure. The instructor should use a method intended for paired categorical responses rather than treating the two responses from each player as independent.
Check the general lesson. The unit of observation matters. For independence and homogeneity, each individual should contribute to one cell in the analysis table, and the observations should be independent. Counting the same individual in linked categories can violate that structure.
Common Mistakes and AP Exam Communication
A concise, full-credit test-choice explanation gives the design, the variables or outcome, and the question that follows from them. Do not stop at “there are two categorical variables” or “this is a two-way table.” Those descriptions do not distinguish the procedures.
- Choosing from the table’s appearance: A table with group labels and outcome categories can represent either design. State how participants were sampled or assigned.
- Calling a group variable an outcome in a homogeneity study: In separate samples, the neighborhood or treatment is the group; the response recorded for each individual is the outcome being compared.
- Calling any pair of categories an independence study: A group and an outcome can form two table dimensions, but when separate samples or treatment groups are compared, homogeneity is the matching setting.
- Ignoring paired or repeated observations: If the same individual supplies linked responses, the ordinary independence and homogeneity setups may not apply. Do not count repeated responses as if they came from separate individuals.
- Confusing random sampling with random assignment: Random sampling supports generalization to a population; random assignment supports causal conclusions about treatments. Say which feature the study has, and do not claim both automatically.
- Treating condition checks as a test choice: Expected counts and the 10% condition matter after identifying the design. Meeting them does not change a homogeneity question into an independence question, or repair a biased sample.
Key Takeaway
Both tests use categorical counts, and both can be organized in a two-way table. The distinction comes from the design: one sample classified by two variables calls for independence; separate samples or treatment groups compared on one outcome call for homogeneity. Check how individuals were selected, assigned, and counted before naming the test.
Check Your Understanding
For each situation, identify the study design and the appropriate choice, if either standard chi-square setting fits. Briefly justify your answer.
- A random sample of commuters reports both its usual transportation method and whether it usually travels alone. Which test setting applies, and what question does it address?
- Separate random samples of visitors from two museums report whether they prefer guided tours, audio guides, or self-guided visits. Which test setting applies?
- Participants are randomly assigned to one of three reminder formats, and researchers record whether each participant completes a task. Which test setting applies, and what feature of the design supports that choice?
- A researcher records each student’s preferred study location and preferred study time from one random sample. Why is the table’s two-variable structure not enough by itself to choose the test?
- Each of 30 volunteers tries two different notification tones, and the same volunteer’s preferred tone is recorded under each condition. Why should the responses not be treated as independent observations for the usual tests?