Testing Whether a Proportion Differs from a Benchmark
Some questions ask whether a proportion is higher or lower than a specified value. Others ask more generally whether it has changed or differs from that value, without specifying a direction. That second kind of question calls for a two-sided one-proportion test.
As in Choosing One-Sided or Two-Sided Alternatives, decide the direction from the research question, before examining the sample result. For a two-sided test, evidence against the null can come from a sample proportion noticeably above the benchmark or noticeably below it. As in Calculating a One-Proportion z-Test by Hand, calculate the test statistic using the null proportion in the standard error.
The test statistic is calculated in the same way as for a one-sided one-proportion test. The difference is how the p-value is found: a two-sided test counts outcomes at least as extreme as the observed result in both tails of the null distribution.
Because the standard Normal curve is symmetric, the two-sided p-value is twice the area beyond the absolute value of the observed \(z\) statistic. Equivalently, find the smaller tail area beyond \(z_{\mathrm{obs}}\) in the direction of the observed difference, then double it. If the observed statistic is positive, that is the upper tail; if negative, it is the lower tail.
For instance, if \(z_{\mathrm{obs}}=2.00\), the two-sided p-value is twice the area to the right of 2.00. If \(z_{\mathrm{obs}}=-2.00\), it is twice the area to the left of \(-2.00\). In either case, the p-value is approximately 0.0455. The observed direction affects which tail you calculate first, but not the fact that both equally extreme tails count.
The Four Steps for a Two-Sided Test
A complete test response connects the research question, the conditions, the calculation, and the conclusion. This four-step structure also helps prevent a common error: finding only the tail in the observed direction when the alternative is two-sided.
Define \(p\) in context, write \(H_0:p=p_0\) and \(H_a:p\ne p_0\), and identify the significance level \(\alpha\), if one is given.
Identify the one-proportion \(z\)-test. Check the Random condition, the 10% condition when sampling without replacement from a finite population, and the test’s Large Counts condition using \(np_0\) and \(n(1-p_0)\).
Calculate \(\hat{p}\), the null standard error, and \(z_{\mathrm{obs}}\). Find the tail area beyond \(|z_{\mathrm{obs}}|\) and double it. A 1-PropZTest can calculate the statistic and p-value when the two-sided alternative is selected.
Compare the p-value with \(\alpha\). Reject \(H_0\) when the p-value is at most \(\alpha\); otherwise, fail to reject \(H_0\). State whether the data provide convincing evidence that the population proportion differs from the benchmark, in context.
On a TI-84, 1-PropZTest takes \(p_0\), the observed number of successes \(x\), the sample size \(n\), and the alternative. Select the not-equal alternative for a two-sided test. Check that the calculator’s \(\hat{p}=x/n\), \(z\), and p-value fit the data and research question. The calculator does not select the hypotheses or verify the study conditions for you.
Worked Examples
Worked Example: Does a Renewal Rate Differ from 60%?
A fictional service randomly selects 600 customers from a list of 12,000 customers whose subscriptions are due for renewal. Of those selected, 390 renew. Test whether the true proportion of these customers who renew differs from 0.60. Use \(\alpha=0.05\).
State: Let \(p\) be the true proportion of customers on this renewal list who renew their subscriptions. The hypotheses are \(H_0:p=0.60\) and \(H_a:p\ne0.60\). The significance level is \(\alpha=0.05\).
Plan and check conditions: Use a one-proportion \(z\)-test. The customers were randomly selected, so the Random condition is met. The sample was drawn without replacement from 12,000 customers, and \(0.10(12{,}000)=1{,}200\). Since \(600\leq1{,}200\), the 10% condition is met. Under the null, the expected number of renewals is \(np_0=600(0.60)=360\), and the expected number of nonrenewals is \(n(1-p_0)=600(0.40)=240\). Both expected counts are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion is:
Calculate the null standard error and test statistic:
The result is above the null value, so first find the upper-tail area. Then double it for the two-sided alternative:
The value 0.0062097 is rounded; the p-value is also rounded. A TI-84 1-PropZTest with \(p_0=0.60\), \(x=390\), \(n=600\), and the not-equal alternative gives approximately the same statistic and p-value.
Conclude: Assuming the true renewal proportion is 0.60, the probability of obtaining a test statistic at least as far from 0 as 2.50, in either direction, is approximately 0.0124. Since \(0.0124<0.05\), reject \(H_0\). The sample provides convincing evidence that the true renewal proportion among customers on this list differs from 0.60.
Worked Example: A Sample Proportion Below 60%
A fictional community program randomly selects 300 people from a list of 6,000 eligible residents. In the sample, 165 say they would use a proposed weekend service. Test whether the true proportion of eligible residents who would use the service differs from 0.60. Use \(\alpha=0.05\).
State: Let \(p\) be the true proportion of eligible residents on this list who would use the proposed weekend service. The hypotheses are \(H_0:p=0.60\) and \(H_a:p\ne0.60\), with \(\alpha=0.05\).
Plan and check conditions: Use a one-proportion \(z\)-test. The residents were randomly selected, meeting the Random condition. The sample was drawn without replacement from 6,000 residents, and \(0.10(6{,}000)=600\). Since \(300\leq600\), the 10% condition is met. Under the null, the expected number who would use the service is \(300(0.60)=180\), and the expected number who would not is \(300(0.40)=120\). Both are at least 10, so the Large Counts condition is met.
Do: Calculate the sample proportion, null standard error, and test statistic:
The observed statistic is negative, so find the lower-tail area. Double that area to include an equally extreme result in the upper tail:
Conclude: Since \(0.0771>0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence that the true proportion of eligible residents who would use the proposed service differs from 0.60. The sample proportion is below 0.60, but a two-sided test considers departures in either direction, and this result is not sufficiently unusual at the 0.05 significance level.
Worked Example: A Small Difference from 60%
A fictional school district randomly selects 250 families from a list of 5,000 families and asks whether they would use a new online scheduling tool. Of those selected, 157 say yes. Test whether the true proportion of families on the list who would use the tool differs from 0.60. Use \(\alpha=0.10\).
State: Let \(p\) be the true proportion of families on this district list who would use the new online scheduling tool. The hypotheses are \(H_0:p=0.60\) and \(H_a:p\ne0.60\), with \(\alpha=0.10\).
Plan and check conditions: A one-proportion \(z\)-test is appropriate if its conditions hold. The families were randomly selected, meeting the Random condition. The sample was drawn without replacement from 5,000 families, and \(0.10(5{,}000)=500\). Since \(250\leq500\), the 10% condition is met. Under the null, the expected number of families who would use the tool is \(250(0.60)=150\), and the expected number who would not is \(250(0.40)=100\). Both expected counts are at least 10, so the Large Counts condition is met.
Do: The sample proportion and null standard error are:
The test statistic is:
The statistic is positive, so calculate the upper-tail area and double it:
Conclude: Since \(0.3662>0.10\), fail to reject \(H_0\). The sample does not provide convincing evidence that the true proportion of families on this district list who would use the online scheduling tool differs from 0.60. A sample proportion above the benchmark is not, by itself, strong evidence of a difference; the size of the standardized difference and the two-sided p-value matter.
Common Mistakes and What Full Credit Says
The alternative hypothesis determines how extreme results are counted. In a two-sided test, both unusually high and unusually low sample proportions count as evidence against \(H_0\). The doubled tail area is part of the p-value, not an optional adjustment.
- Writing a one-sided alternative for a “differs” question. If the question does not specify higher or lower, use \(H_a:p\ne p_0\). Do not choose a direction just because \(\hat{p}\) happens to be above or below \(p_0\).
- Reporting only one tail. For a two-sided test, the p-value is twice the tail area beyond \(|z_{\mathrm{obs}}|\). Reporting just the area on the observed side answers a one-sided question instead.
- Doubling the wrong area. Double the smaller tail beyond the observed statistic, not the larger area between the statistic and the center of the curve. For a negative statistic, find the area to its left; for a positive statistic, find the area to its right.
- Using the wrong standard error. In a one-proportion test, use \(p_0\) in \(SE_0\), because the test models sampling variation under the null hypothesis. Do not substitute \(\hat{p}\).
- Using observed counts for the test’s Large Counts check. Check \(np_0\) and \(n(1-p_0)\), the expected counts under the null, rather than \(x\) and \(n-x\).
- Interpreting a large p-value as proof of equality. When the p-value is greater than \(\alpha\), say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence of a difference. Do not say that the null has been proved or accepted.
Key Takeaway
A two-sided one-proportion test asks whether the population proportion differs from a benchmark in either direction. Calculate the test statistic with the null standard error, find the area beyond its absolute value, and double that area. Then compare the p-value with \(\alpha\) and describe the evidence in context.
Check Your Understanding
For each question, consider the hypotheses, conditions, tail areas, or conclusion for a two-sided one-proportion test.
- A researcher asks whether the proportion of households using a particular heating source differs from 0.35. Write the null and alternative hypotheses using \(p\).
- For a test of \(H_0:p=0.60\) with \(n=250\), calculate both expected counts for the Large Counts condition. Does the condition hold?
- A two-sided test has observed \(z=-1.25\). Which tail area should you find first, and what should you do with that area to obtain the p-value?
- A two-sided test at \(\alpha=0.05\) gives a p-value of 0.032. State the decision and interpret it in context for a claim that the proportion differs from its null value.
- Why does a two-sided test count results in both tails, even when the observed sample proportion is above the null proportion?