One Test, Four Connected Steps
A two-proportion \(z\)-test is more than a calculation. A complete solution explains which population proportions are being compared, why the procedure is appropriate, how the statistic and p-value were found, and what the evidence means in context. The four steps—State, Plan, Do, Conclude—help keep that reasoning visible.
In “Common Mistakes in Two-Proportion Tests,” you audited choices such as the group order, pooled standard error, and p-value tail. Here, you will bring those choices together in a full response. As established in “Hypotheses for Comparing Two Population Proportions” and “Conditions for a Two-Proportion \(z\)-Test,” the hypotheses describe population proportions, and the conditions must be checked before relying on the test.
The Four-Step Structure
Define \(p_1\) and \(p_2\) as the true proportions for the two groups and the same outcome. Write \(H_0:p_1=p_2\), or equivalently \(H_0:p_1-p_2=0\). Choose \(H_a\) to match the question: \(p_1\ne p_2\), \(p_1>p_2\), or \(p_1<p_2\).
Name the two-proportion \(z\)-test and check its conditions using the study design and the null model. For the test of equal proportions, use the pooled proportion to check the four expected counts.
Calculate the sample proportions, pooled proportion, pooled standard error, and \(z\)-statistic. Find the p-value in the direction or directions specified by \(H_a\).
Compare the p-value with the significance level \(\alpha\), state whether you reject or fail to reject \(H_0\), and describe the evidence about the population proportions in context.
The four steps are a chain of reasoning, not four disconnected labels. In particular, the group definitions in State determine the order in the calculation and conclusion; the alternative determines the p-value tail; and the conditions in Plan justify using the test calculation in Do.
The Large Counts check here is for a test of equal proportions: it uses expected counts under the null model. As explained in “Large Counts for Each Group in Two-Proportion Intervals,” an interval instead checks observed successes and failures. Keep the procedure in mind when choosing which counts to report.
Worked Four-Step Solutions
Worked Example: Testing Whether a Conservation Program Is More Common
Question: A fictional city compares two residential districts. Independent random samples of 120 households are taken from each district; each district has at least 1,200 households. In District 1, 72 sampled households report participating in a water-conservation program. In District 2, 54 report participating. Does District 1 have a higher participation proportion? Test at \(\alpha=0.05\).
State: Let \(p_1\) be the true proportion of households in District 1 that participate in the program, and let \(p_2\) be the true proportion in District 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1>p_2\). The research question asks whether District 1’s population proportion is higher, so the alternative is upper-tailed.
Plan: The two groups are independent random samples, and each household is counted in only one district. Each sample is at most 10% of its district’s population because \(120\le0.10(1{,}200)=120\). Thus, the 10% condition is met. Pool the successes:
Under \(H_0\), the expected successes are \(120(0.525)=63\) in each group, and the expected failures are \(120(1-0.525)=120(0.475)=57\) in each group. All four expected counts—63, 57, 63, and 57—are at least 10. The conditions support a two-proportion \(z\)-test.
Do: The sample proportions are \(\hat{p}_1=72/120=0.60\) and \(\hat{p}_2=54/120=0.45\), so the observed difference is \(0.60-0.45=0.15\). Use the pooled standard error:
The test statistic is:
Because \(H_a:p_1>p_2\), the p-value is the upper-tail area \(P(Z\ge2.3267)\approx0.0100\), rounded to four decimal places. This is the chance, assuming equal population proportions, of getting a standardized difference at least as large as the one observed in the direction specified by the alternative.
Conclude: Since \(0.0100<0.05\), reject \(H_0\). The data provide convincing evidence that the true proportion of households participating in the water-conservation program is higher in District 1 than in District 2.
Worked Example: Testing for Any Difference in E-Receipt Use
Question: A fictional retailer takes independent random samples of 90 customers from one customer population and 110 from another. The populations each contain at least 10 times the respective sample size. In the first sample, 27 customers choose an electronic receipt; in the second, 44 do. Is there a difference between the population proportions? Use \(\alpha=0.05\).
State: Let \(p_1\) be the true proportion of customers in the first population who choose an electronic receipt, and \(p_2\) the corresponding proportion in the second population. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). “Is there a difference?” calls for a two-sided alternative.
Plan: The samples are independent random samples, and each customer contributes to only one group. Each sample is at most 10% of its population by the stated population-size information. Pool the successes and check the null-model expected counts:
For Group 1, the expected successes are \(90(0.355)=31.95\) and expected failures are \(90(0.645)=58.05\). For Group 2, they are \(110(0.355)=39.05\) expected successes and \(110(0.645)=70.95\) expected failures. All four expected counts are at least 10. The two-proportion \(z\)-test is appropriate.
Do: The sample proportions are \(\hat{p}_1=27/90=0.30\) and \(\hat{p}_2=44/110=0.40\), so the observed difference is \(-0.10\). The pooled standard error is:
Then:
For a two-sided alternative, results at least as far from zero as \(-1.4703\) count in either tail. Thus, \(p\text{-value}=2P(Z\le-1.4703)\approx0.1415\), rounded to four decimal places.
Conclude: Since \(0.1415>0.05\), fail to reject \(H_0\). The samples do not provide convincing evidence that the true electronic-receipt proportions differ between the two customer populations. This result does not prove that the population proportions are equal.
Worked Example: Keeping the Direction Set by the Question
Question: A fictional parks department takes independent random samples of 50 visitors from each of two large park populations. In Park 1, 18 visitors report seeing litter on their visit; in Park 2, 22 do. The department asks whether the true proportion reporting litter is lower at Park 1. Test at \(\alpha=0.05\).
State: Let \(p_1\) be the true proportion of visitors to Park 1 who report seeing litter, and \(p_2\) the corresponding proportion for Park 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1<p_2\). The lower-tail alternative follows from the question, regardless of the sample results.
Plan: Both samples are independent random samples, and the groups contain different visitors. Each park population is at least 500 visitors, so each sample of 50 is at most 10% of its population. The pooled proportion is:
Under \(H_0\), each group has \(50(0.40)=20\) expected successes and \(50(0.60)=30\) expected failures. All four expected counts are at least 10, so the Large Counts condition is met.
Do: The sample proportions are \(\hat{p}_1=18/50=0.36\) and \(\hat{p}_2=22/50=0.44\). Their difference is \(-0.08\). The pooled standard error and statistic are:
The alternative is lower-tailed, so the p-value is \(P(Z\le-0.8165)\approx0.2071\), rounded to four decimal places. Do not switch to a two-sided test or a different tail because the observed difference points in a particular direction; the stated research question determines the alternative.
Conclude: Since \(0.2071>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the true proportion of visitors reporting litter is lower at Park 1 than at Park 2.
Common Mistakes and AP Exam Tips
A full-credit response should let a reader follow the reasoning without having to guess what a symbol or number represents. Avoid these common problems:
- Leaving the parameters vague: Define each \(p\) as a true proportion, identify its population or group, and name the shared outcome. “Let \(p_1\) and \(p_2\) be the proportions” is not enough by itself.
- Writing a sample claim as a hypothesis: Hypotheses concern population parameters, not observed sample proportions. State \(H_0:p_1=p_2\), not a claim that \(\hat{p}_1=\hat{p}_2\).
- Listing conditions without evidence: Say how random sampling or random assignment occurred, establish that the groups are independent, explain the 10% condition when relevant, and give all four expected counts for Large Counts.
- Using the interval’s standard error: For a test of equal proportions, use the pooled proportion in \(SE_{\text{pooled}}\). The separate sample proportions belong in the standard error for a two-proportion interval, not this test.
- Letting the observed difference choose the alternative: Choose \(H_a\) from the research question. A lower sample proportion in Group 1 does not authorize changing a preexisting question about whether Group 1 is higher.
- Reporting only a calculator result: A TI-84’s 2-PropZTest can calculate the test statistic and p-value, but it does not justify the conditions or explain the conclusion. Show the setup and enough calculation to make the result interpretable.
- Concluding only about the sample: A test conclusion addresses the population proportions. As explained in “Writing a Conclusion for a Two-Proportion Test,” make the decision about \(H_0\), then describe the evidence for the alternative in context. Do not claim that failing to reject proves equality.
A Final Four-Step Check
Use this sequence as a compact review: State the population proportions and hypotheses; Plan the test and justify its conditions; Do the pooled calculation and find the correct p-value; Conclude with a decision and a contextual statement about evidence. Each step answers a different question, and together they form a complete statistical argument.
Check Your Understanding
For each question, focus on how the four-step structure supports a correct test response.
- A researcher compares two proportions and asks whether Group 1’s proportion is higher. Write the alternative hypothesis using \(p_1\) and \(p_2\). Which p-value tail will be used?
- For a test with \(x_1=32\), \(n_1=80\), \(x_2=28\), and \(n_2=80\), calculate the pooled proportion and the four expected success and failure counts. Does the Large Counts condition hold?
- In a Plan step, what information should support the random, independence, and 10% checks?
- Why does the Large Counts condition for a two-proportion test use expected counts from the pooled proportion rather than only the observed counts?
- A two-sided test gives a p-value of \(0.08\) with \(\alpha=0.05\). State the decision and describe, in general terms, what the result says about evidence for a difference.