Audit the Setup Before Calculating
The previous tutorial, “Choosing the Chi-Square Test From a Study Description,” focused on matching a study design to a chi-square test of independence or homogeneity. This tutorial takes the next step: checking whether the data and statements used to set up the test actually match that choice.
Three errors can undermine an otherwise careful response. Students may enter percentages instead of observed counts, write hypotheses for a different test than the one the design calls for, or use the wrong degrees of freedom. Each mistake changes what the procedure means or how its result should be evaluated. A reliable audit starts with the study design, then checks the table, hypotheses, and degrees of freedom against it.
Mistake 1: Using Percentages as Observed Counts
The entries in the observed table must be counts of individuals or cases in each category combination. Percentages and proportions can help describe a categorical distribution, as in “Computing Conditional Distributions From a Two-Way Table,” but they are not the observed counts to enter for a chi-square test.
This distinction matters especially when groups have different sample sizes. A group with 100 individuals and a group with 200 individuals could both have a cell percentage of 30%, but those percentages represent 30 and 60 individuals, respectively. Treating both percentages as equal counts would discard information about how many observations were collected.
If only percentages are shown, first ask what their denominators are and whether the original counts can be recovered. Rounded percentages may not allow exact recovery. Do not treat row percentages or column percentages as a replacement observed table. Once the counts are available, expected counts are calculated from the count table, and the statistic adds the contributions \((O-E)^2/E\) across its cells, as explained in “The Chi-Square Statistic Formula.”
Mistake 2: Hypotheses That Do Not Match the Test
The wording of the hypotheses depends on the design. A test of independence asks whether two categorical variables are independent or associated in a population. A test of homogeneity asks whether the distribution of one categorical response is the same across separate populations or groups. These purposes, and the matching standard hypotheses, were developed in “Stating Hypotheses for a Test of Independence” and “Stating Hypotheses for a Test of Homogeneity.”
For independence, name both variables and the population. The null hypothesis says the variables are independent in that population; the alternative says they are associated. For homogeneity, name the groups or populations and the response. The null hypothesis says the response distribution is the same for all groups; the alternative says that at least one group’s distribution differs.
Do not state a one-sided alternative such as “the proportion is greater” for a chi-square test of independence or homogeneity. The chi-square statistic measures departure from the null in any direction, and the alternative is about association or a difference in distributions—not a specified higher or lower direction. Also, “the distributions are different” belongs to homogeneity wording; “the variables are associated” belongs to independence wording. A table’s appearance alone does not tell you which statement fits: return to the study design.
Mistake 3: Miscounting Degrees of Freedom
For a two-way table with \(r\) row categories and \(c\) column categories, the degrees of freedom are \(df=(r-1)(c-1)\). As covered in “Degrees of Freedom for a Two-Way Table,” count the categories, not the sample size and not the number of observations.
A common error is to subtract one from the total number of cells, \(rc-1\). That is not the degrees-of-freedom formula for a two-way table. Another is to use only the number of rows minus one or only the number of columns minus one. The correct formula accounts for the constraints from both sets of marginal totals.
Worked Example: Counts, Not Percentages, in a Homogeneity Test
Worked Example: Counts, Not Percentages, in a Homogeneity Test
A fictional transit office takes separate random samples of 100 riders on Route A and 200 riders on Route B. Each rider names one preferred ticket option: single-ride, day pass, or monthly pass. The observed counts are shown below.
| Route | Single-ride | Day pass | Monthly pass | Total |
|---|---|---|---|---|
| A | 40 | 30 | 30 | 100 |
| B | 60 | 60 | 80 | 200 |
| Total | 100 | 90 | 110 | 300 |
State. Let the populations be all riders on Routes A and B. The response is preferred ticket option. The question is whether its distribution differs between the two routes.
Plan and check conditions. Separate random samples are compared on one categorical response, so use a chi-square test of homogeneity. The random condition is met by the stated random sampling. Treating riders as independent is reasonable if each rider contributes one response and the samples are independent; for sampling without replacement, the 10% condition requires at least 1,000 riders in the Route A population and 2,000 in the Route B population. Assume both populations meet those requirements. The expected counts, calculated below, are all at least 5.
State hypotheses. \(H_0\): The distribution of preferred ticket option is the same for riders on Routes A and B. \(H_a\): The distribution of preferred ticket option differs between the two routes.
Do. The expected count for a cell is \((\text{row total})(\text{column total})/(\text{table total})\). For example, the expected count for Route A and single-ride is \(100(100)/300=33.333\). The complete expected-count table is:
| Route | Single-ride | Day pass | Monthly pass |
|---|---|---|---|
| A | 33.333 | 30 | 36.667 |
| B | 66.667 | 60 | 73.333 |
Each expected count is at least 5. The statistic is calculated from the observed counts, not from the percentages. For example, the Route A, single-ride contribution is \((40-33.333)^2/33.333=1.333\). Adding all six contributions gives:
There are \(r=2\) rows and \(c=3\) response categories, so \(df=(2-1)(3-1)=2\). The upper-tail p-value is \(P(\chi^2\ge3.818)\approx0.1482\), rounded. At \(\alpha=0.05\), the p-value is greater than the significance level.
Conclude. Fail to reject \(H_0\). The samples do not provide convincing evidence that the distribution of preferred ticket option differs between riders on Routes A and B. This conclusion does not prove the distributions are identical.
The sample percentages are useful for description: Route A’s percentages are 40%, 30%, and 30%, while Route B’s are 30%, 30%, and 40%. But entering those percentages as the observed table would ignore that the sample sizes are 100 and 200. The test correctly uses the counts in the table.
Worked Example: Hypotheses for an Independence Test
Worked Example: Hypotheses for an Independence Test
A fictional parks department selects a random sample of 180 hikers. Each hiker is classified by trail experience (new or experienced) and water-bottle choice (refillable, disposable, or none). The observed counts are:
| Experience | Refillable | Disposable | None | Total |
|---|---|---|---|---|
| New | 28 | 20 | 12 | 60 |
| Experienced | 60 | 36 | 24 | 120 |
| Total | 88 | 56 | 36 | 180 |
Identify the test and state the hypotheses. There is one sample, and each hiker is classified by two categorical variables. The question is whether those variables are associated, so this is a chi-square test of independence. \(H_0\): Trail experience and water-bottle choice are independent among hikers in the population. \(H_a\): Trail experience and water-bottle choice are associated among hikers in the population.
Writing “the water-bottle distributions are the same for new and experienced hikers” may describe a related comparison, but it frames the question as homogeneity across separate groups. Here, the design is one sample classified twice, so the independence hypotheses directly match the design.
Check counts and degrees of freedom. The expected count for new hikers choosing refillable bottles is \(60(88)/180=29.333\). The six expected counts are 29.333, 18.667, 12, 58.667, 37.333, and 24; all are at least 5. The random condition is met by the stated random sample. For the 10% condition, assume the population of hikers contains at least 1,800 individuals. Each hiker contributes one pair of classifications, so the observations are independent under the sampling design and stated condition.
The chi-square statistic can be checked from the six cell contributions. The deviations from expected counts are \(-1.333, 1.333, 0, 1.333, -1.333, 0\), giving contributions \(0.0606, 0.0952, 0, 0.0303, 0.0476, 0\), rounded. Thus \(X^2\approx0.2338\). With two rows and three columns, \(df=(2-1)(3-1)=2\), and the upper-tail p-value is approximately \(0.8897\), rounded.
Conclude. Since \(0.8897>0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence of an association between trail experience and water-bottle choice among hikers in the population.
Worked Example: Getting the Degrees of Freedom Right
Worked Example: Getting the Degrees of Freedom Right
A fictional school district selects separate random samples of students from three campuses and records each student’s preferred after-school activity from four categories. The observed table has three campus rows and four activity columns:
| Campus | Sports | Arts | Clubs | Other | Total |
|---|---|---|---|---|---|
| North | 20 | 15 | 15 | 10 | 60 |
| Central | 15 | 20 | 10 | 15 | 60 |
| South | 10 | 15 | 20 | 15 | 60 |
| Total | 45 | 50 | 45 | 40 | 180 |
Match the question to the hypotheses. Separate campus samples are compared on one categorical response, so the test is homogeneity. \(H_0\): The distribution of preferred after-school activity is the same at all three campuses. \(H_a\): At least one campus has a different distribution of preferred activity.
Count categories, then calculate df. There are \(r=3\) campus categories and \(c=4\) activity categories. Therefore:
The table has 12 interior cells, but \(12-1=11\) is not the correct degrees of freedom. Nor should the answer be 2 or 3 by counting only the rows or only the columns and subtracting one. For a count-condition check, the expected counts in each row are 15, 16.667, 15, and 13.333, since every row total is 60 and the column totals are 45, 50, 45, and 40. All expected counts are at least 5.
Answer. The appropriate degrees of freedom are 6. This is the value to use with the chi-square statistic when finding the p-value; the number of table cells or the sample size does not replace the formula.
Common Mistakes and AP Exam Tip
- Entering percentages as observed data: Use counts in the observed table. If percentages have different group denominators, the same percentage can represent different numbers of observations.
- Mixing up observed and expected counts: Observed counts come from the sample. Expected counts are calculated under the null hypothesis using the table margins. Do not use percentages as either table.
- Writing hypotheses for the wrong design: For one sample classified by two categorical variables, state independence versus association. For separate groups compared on one response, state that the distributions are the same versus that at least one differs.
- Using a directional alternative: A chi-square test of independence or homogeneity does not test whether a single group’s proportion is specifically higher or lower. State association or a difference in distributions.
- Using the wrong df: Count row and column categories, then calculate \((r-1)(c-1)\). Do not use the sample size, total number of cells minus one, or only one table dimension.
- Overstating a nonsignificant result: “Fail to reject” does not establish that variables are independent or that distributions are identical. Say there is not convincing evidence for the alternative in context.
Check Your Understanding
For each situation, identify the setup error or state the correct choice. Include a brief reason.
- Two groups have sample sizes 80 and 160. A student enters their row percentages as the observed counts in a chi-square test. What is wrong, and what data should be used?
- One random sample of residents is classified by neighborhood type and preferred way to receive alerts. Should the null hypothesis state that the two variables are independent or that the response distributions are the same across separate samples?
- A homogeneity table has four groups and three response categories. Calculate the degrees of freedom.
- A two-way table has three rows and four columns. A student says \(df=11\), counting the 12 cells and subtracting one. Correct the value and explain the error.
- A test of independence has a p-value greater than \(\alpha\). Write an appropriate conclusion without claiming that the variables are proven independent.