Turn a Study Description into an Exam Answer
On an exam, a chi-square question may ask you to choose a procedure, state hypotheses, and justify the choice from a short study description. The challenge is not just knowing the names of the tests: it is connecting the study’s design to the population question and then to the right hypotheses and conditions. In “Setup and Plan for a Chi-Square Test,” you learned the parts of a State and Plan. Here, you will practice finding those parts efficiently in exam-style prompts.
Start by tracing how the observations were collected. Was there one sample, with every individual classified by two categorical variables? That design points to a chi-square test of independence if the question asks whether the variables are associated. Or were there separate samples or randomized treatment groups, with one categorical response measured in each? That design points to a chi-square test of homogeneity if the question asks whether the response distribution differs across groups.
Then write hypotheses that match the design and the question. For independence, \(H_0\) says the two variables are independent in the population, and \(H_a\) says they are associated. For homogeneity, \(H_0\) says the response distribution is the same across all groups, and \(H_a\) says at least one group has a different distribution. These are population claims, not statements about whether the sample counts happen to match.
Finally, justify the conditions using the study description. Identify the random sample or random assignment, explain why observations can reasonably be treated as independent, and check the 10% condition when random sampling without replacement makes it relevant. For either chi-square procedure, every expected cell count under the null hypothesis must be at least 5. As in “Setup and Plan for a Chi-Square Test,” do not substitute observed counts for expected counts.
A Compact Plan for Exam Questions
A useful way to organize your reasoning is to keep the procedure choice and the condition justification separate. First decide what the study design is; next write the population hypotheses; finally support the procedure’s conditions. This helps prevent a common error: selecting a familiar test based on the table’s appearance while overlooking how the data were collected.
One sample classified by two categorical variables suggests independence. Separate samples or treatment groups measured on one categorical response suggest homogeneity.
Use “independent” versus “associated” for independence, or “same distribution” versus “at least one differs” for homogeneity.
Refer to the stated random sample or random assignment, explain independence, check the 10% condition when applicable, and verify every expected count is at least 5.
A prompt may not supply every fact needed to establish a condition. In that case, distinguish what the description confirms from what must be assumed. For example, if a random sample is stated but the population size is not, you can say that the 10% condition is met provided the population is at least ten times the sample size. If expected counts are not given, calculate them from the margins when possible. If the table margins are not available, say that the expected-count condition must be checked before proceeding rather than claiming it is satisfied.
Worked Example: One Sample Classified by Two Variables
Worked Example: One Sample Classified by Two Variables
A city selects a random sample of 300 households from its service area. Each household is classified by housing type (apartment, detached house, or townhouse) and preferred emergency-alert method (text or phone call). The city asks whether housing type and preferred alert method are associated among households in its service area. The observed counts are:
| Housing type | Text | Phone call | Total |
|---|---|---|---|
| Apartment | 58 | 42 | 100 |
| Detached house | 62 | 38 | 100 |
| Townhouse | 60 | 40 | 100 |
| Total | 180 | 120 | 300 |
Identify the procedure. This is one random sample, and every household is classified by two categorical variables. The question asks whether those variables are associated, so use a chi-square test of independence.
State the hypotheses. \(H_0\): Housing type and preferred emergency-alert method are independent among households in the city’s service area. \(H_a\): Housing type and preferred emergency-alert method are associated among households in the city’s service area.
Justify the conditions. The random condition is met because the city selected a random sample. Each household contributes one classification for each variable, and the sampled households are distinct, so treating observations as independent is reasonable. Because the households were sampled without replacement, check the 10% condition: the service area must contain at least \(10(300)=3{,}000\) households. Assume it does.
The expected count for apartments preferring text alerts under \(H_0\) is:
Each housing-type row total is 100. The expected counts for text and phone call in each row are \(100(180)/300=60\) and \(100(120)/300=40\), respectively. All six expected counts are at least 5. The chi-square test of independence conditions are met, given the stated population-size assumption.
The key exam clue is not simply that the data fit a two-way table. It is that one sample of households was classified by two variables, and the city’s question is about their association.
Worked Example: Separate Samples Compared on One Response
Worked Example: Separate Samples Compared on One Response
A fictional health network takes separate random samples of 120 patients from each of three clinics. Each patient is asked to choose a preferred appointment format: in person, video, or phone. The network asks whether the distribution of appointment-format preference is the same at all three clinics. The combined response totals are 150 in person, 120 video, and 90 phone, for 360 patients altogether.
| Clinic | In person | Video | Phone | Total |
|---|---|---|---|---|
| North | 55 | 40 | 25 | 120 |
| South | 45 | 38 | 37 | 120 |
| West | 50 | 42 | 28 | 120 |
| Total | 150 | 120 | 90 | 360 |
Identify the procedure. The network used separate random samples from three clinics and measured one categorical response. The question compares the response distributions across the clinics, so use a chi-square test of homogeneity.
State the hypotheses. \(H_0\): The distribution of appointment-format preference is the same for patients at the North, South, and West clinics. \(H_a\): At least one clinic has a different distribution of appointment-format preference.
Justify the conditions. The random condition is met because each clinic provided a separate random sample. The samples are independent if patients do not appear in more than one clinic sample; assume that is true. Because the samples were selected without replacement, check the 10% condition separately for each clinic: each clinic must have at least \(10(120)=1{,}200\) patients in its relevant population. Assume each does.
Under \(H_0\), the expected counts in each clinic row are based on that row’s total of 120 and the combined response totals. For example, the North clinic’s expected count for video preference is \(120(120)/360=40\). The expected counts for in person, video, and phone in each clinic are \(120(150)/360=50\), \(120(120)/360=40\), and \(120(90)/360=30\). All nine expected counts are at least 5.
This is not a test of independence just because the data can be placed in a two-way table. The study collected separate samples from clinics and compares one response distribution across those groups; that design calls for homogeneity.
Worked Example: Random Assignment to Treatment Groups
Worked Example: Random Assignment to Treatment Groups
A fictional gardening program recruits 270 volunteers and randomly assigns 90 to each of three reminder methods: email, text, or no reminder. The outcome is whether each volunteer submits a soil-test kit by a deadline. Across all groups, 135 volunteers submit a kit and 135 do not. The program asks whether submission rates differ among the reminder methods.
| Reminder method | Submitted | Did not submit | Total |
|---|---|---|---|
| 51 | 39 | 90 | |
| Text | 48 | 42 | 90 |
| No reminder | 36 | 54 | 90 |
| Total | 135 | 135 | 270 |
Identify the procedure. The volunteers were randomly assigned to three treatment groups, and one categorical response was measured. To compare the response distributions across the groups, use a chi-square test of homogeneity.
State the hypotheses. \(H_0\): The distribution of soil-test submission outcome is the same for the email, text, and no-reminder groups. \(H_a\): At least one reminder group has a different distribution of soil-test submission outcome.
Justify the conditions. The random condition is met by the stated random assignment. Each volunteer contributes one outcome, so observations can be treated as independent if volunteers complete the task without influencing one another; assume that is reasonable. The 10% condition is not the relevant check because this is a randomized experiment, not a random sample selected without replacement from a finite population.
For each reminder group, the expected count for submission under \(H_0\) is \(90(135)/270=45\), and the expected count for non-submission is \(90(135)/270=45\). Thus, all six expected counts are 45, meeting the expected-count requirement.
Random assignment supports a comparison of the reminder methods’ effects for these volunteers, provided the experiment was carried out appropriately. Because the volunteers were recruited rather than randomly sampled from all gardeners, do not claim that the results automatically generalize to every gardener.
Common Mistakes and Full-Credit Communication
- Choosing from the table shape alone: Both chi-square tests can use a two-way table. Say how the data were collected: one sample classified twice, separate samples, or randomly assigned groups.
- Mixing the hypothesis forms: For independence, write about whether two variables are independent or associated. For homogeneity, write about whether one response distribution is the same across groups or differs for at least one group.
- Writing hypotheses about observed counts: The hypotheses concern a population relationship or population distributions, not whether the sample table has equal counts or percentages.
- Using a directional alternative: The chi-square alternative describes association or some difference across the table. It does not say that one named category is “greater” in a particular group.
- Giving a generic condition statement: “The conditions are met” is not a justification. Refer to the stated random method, explain independence, check the 10% condition when relevant, and state the expected counts or their minimum.
- Claiming a condition the prompt does not establish: If population size, nonoverlapping samples, or independence is not stated, identify the needed assumption. Do not turn an assumption into a fact.
- Using the 10% condition for a randomized experiment: This condition addresses sampling without replacement. For an experiment, explain the random assignment and why outcomes can be treated as independent instead.
Check Your Understanding
For each prompt, identify the procedure and describe the hypotheses and condition checks that would be appropriate.
- A random sample of 400 apartment residents is classified by floor level and preferred package-delivery option. The population contains 8,000 residents. Which test fits a question about whether floor level and preference are associated?
- Separate random samples of students from three schools are asked which of four after-school activities they prefer. Write the homogeneity hypotheses in context and name the procedure.
- Participants are randomly assigned to one of two study environments and classified by whether they complete a practice set. Which condition supports using the experiment’s design, and why is the 10% condition not the relevant check?
- In a homogeneity test, one expected count is 4.8 and all others exceed 5. Is the expected-count condition met? What should the plan say?
- A study description says that a random sample was taken without replacement but does not give the population size. What additional information or assumption is needed to justify the 10% condition?