From a Chi-Square Statistic to a P-Value
In “Standardized Residuals for Interpretation,” you used cell-level information to describe how observed counts depart from expected counts. The overall chi-square statistic combines those cell departures into one measure. To judge how unusual that statistic is under the null model, we find its p-value.
A chi-square statistic is nonnegative, and larger values represent greater overall departure from the expected counts. Therefore, the p-value is the area to the right of the observed statistic in the chi-square distribution with the test’s degrees of freedom. It is not the area to the left, and it is not a two-tail area.
Here, \(X^2\) represents a random variable following the chi-square distribution specified by the null model and the degrees of freedom. Since a calculator’s chi-square cdf gives area to the left between its lower and upper bounds, finding the right-tail p-value requires either subtracting the left area from 1 or setting a very large upper bound.
The degrees of freedom must match the table. As covered in “Degrees of Freedom for a Two-Way Table,” for a table with \(r\) row categories and \(c\) column categories, \(df=(r-1)(c-1)\). Use the number of categories, not the sample size, to choose the chi-square distribution. A correct statistic with the wrong degrees of freedom gives the wrong p-value.
A Routine for Finding the Upper-Tail Area
Use the overall chi-square statistic \(x^2\), calculated from the cell contributions or supplied in the question.
For a two-way table, use \(df=(r-1)(c-1)\), as in “Degrees of Freedom for a Two-Way Table.”
Enter the observed statistic as the lower bound and a very large number as the upper bound, or subtract the area to its left from 1.
Describe the probability of a chi-square statistic at least as large as the observed one, assuming the null hypothesis is true. Then compare it with the stated significance level if a test decision is required.
A chi-square test is right-tailed because large values indicate stronger disagreement between the observed counts and the counts expected under the null hypothesis. A small p-value means that a statistic this large would be unusual if the null model were true. The p-value is not the probability that the null hypothesis is true, nor is it a measure of how strong or important an association is.
Worked Example: Finding a P-Value With Six Degrees of Freedom
Worked Example: Finding a P-Value With Six Degrees of Freedom
An invented study takes separate random samples of residents from three neighborhoods and records each person’s preferred way to receive local alerts: text, email, phone call, or printed notice. The question is whether the distribution of alert preferences is the same in all three neighborhood populations. Suppose the test statistic, calculated from the observed and expected counts, is \(13.28\).
This is a chi-square test of homogeneity with three row categories and four response categories. The degrees of freedom are
Using \(\chi^2\mathrm{cdf}(13.28,1\mathrm{E}99,6)\), or equivalently \(1-\chi^2\mathrm{cdf}(0,13.28,6)\), gives
The p-value is approximately \(0.0388\), rounded to four decimal places. In context, if alert preference has the same distribution in all three neighborhood populations, the probability of obtaining a chi-square statistic of at least \(13.28\) from samples like these is about \(0.0388\).
For a complete test conclusion, check the study and expected-count conditions as well as the p-value. The three groups were selected using separate random samples; each sample contains 80 residents, and each neighborhood population has at least 1,200 residents, so the 10% condition is met because \(80\le0.10(1200)=120\) for each sample. The samples are independent, and each resident contributes to one alert-preference category. The smallest expected count was checked and is \(8.4\), so every expected count is at least 5.
At significance level \(\alpha=0.05\), \(0.0388<0.05\), so reject \(H_0\). The samples provide convincing evidence that the distribution of preferred alert method is not the same across the three neighborhood populations. This conclusion does not identify which specific cells differ most; cell contributions and residuals help describe that pattern.
Worked Example: A P-Value for a Test of Independence
Worked Example: A P-Value for a Test of Independence
Imagine an invented random sample of community-center members. Each member is classified by age group and by whether they usually register for activities online or in person. Suppose the resulting table has two age categories and three registration categories, and the calculated chi-square statistic is \(3.20\). Find the p-value for testing whether age group and registration method are independent in the population.
There are two row categories and three column categories, so the degrees of freedom are
The calculator’s upper-tail calculation is
As a check, for \(df=2\), the right-tail area at \(3.20\) is \(e^{-3.20/2}=e^{-1.6}\approx0.2019\). Thus, the calculator and an independent calculation agree. Assuming independence in the population, there is about a \(0.2019\) probability of obtaining a chi-square statistic at least as large as \(3.20\).
At \(\alpha=0.05\), this p-value is greater than the significance level, so fail to reject \(H_0\). The sample does not provide convincing evidence of an association between age group and registration method. This is not proof that the variables are independent; it means the result is not unusual enough under the null model to reject independence at this significance level.
Worked Example: Keeping the Tail and Degrees of Freedom Straight
Worked Example: Keeping the Tail and Degrees of Freedom Straight
Suppose an invented environmental survey classifies households by one of three heating systems and one of three levels of interest in a home-energy workshop. The chi-square statistic for testing whether heating system and interest level are associated is \(9.50\). Find and interpret the p-value.
The table has three row categories and three column categories. Therefore,
Use the statistic as the lower bound to obtain the right-tail probability:
The answer is not \(\chi^2\mathrm{cdf}(0,9.50,4)\) by itself; that command gives the area to the left. Subtracting that left-tail area from 1 gives the same p-value. In context, if heating system and workshop interest are independent in the population, the probability of a chi-square statistic at least \(9.50\) is about \(0.0497\).
At \(\alpha=0.05\), \(0.0497<0.05\), so reject the null hypothesis of independence. The survey provides convincing evidence of an association between household heating system and interest in the workshop. The conclusion concerns an association, not a cause-and-effect relationship.
Common Mistakes and AP Exam Tip
- Reporting the left-tail area: The cdf gives area to the left between its bounds. For a chi-square p-value, use the area to the right of the observed statistic, either directly with a large upper bound or by subtracting the left area from 1.
- Using the wrong degrees of freedom: Count row and column categories and use \(df=(r-1)(c-1)\). Do not use the sample size as the degrees of freedom.
- Doubling a tail area: Chi-square tests use the upper tail. Do not double the p-value as if the test were a two-sided \(z\)-test.
- Changing the observed statistic or tail direction: Use the overall chi-square statistic from the table and ask for statistics at least as large. A larger chi-square value is farther into the relevant tail.
- Rounding too early: Keep the calculator’s full precision through the tail calculation, then round the p-value as requested. If a p-value is very close to \(\alpha\), use unrounded values to make the decision.
- Overstating the meaning: A p-value is calculated assuming the null hypothesis is true. It is not the probability that the null is true, and it does not reveal which cells account for the overall result.
A full-credit response gives the correct degrees of freedom, shows the upper-tail calculation, reports the p-value with appropriate rounding, and interprets it in context under the null hypothesis. If the question asks for a test conclusion, compare the p-value with \(\alpha\), state “reject” or “fail to reject,” and describe the evidence without claiming that a nonsignificant result proves the null.
Check Your Understanding
For each question, show how the degrees of freedom and upper-tail area are determined.
- A two-way table has three row categories and two column categories. What are its degrees of freedom?
- A chi-square test has statistic \(7.20\) and \(df=2\). Write the calculator cdf expression for its p-value. Do not calculate the final decimal.
- For an observed statistic of \(5.10\) with \(df=4\), a student reports \(\chi^2\mathrm{cdf}(0,5.10,4)\) as the p-value. Explain the error and how to correct it.
- A chi-square test gives \(p=0.071\) at \(\alpha=0.05\). State the decision and explain what the p-value means under the null hypothesis.
- Why would using the degrees of freedom for a two-by-two table be incorrect for a table with three rows and four columns?