Small Choices Can Change a Test
A two-proportion \(z\)-test uses familiar pieces: sample proportions, a standard error, a \(z\)-statistic, and a p-value. Yet a test can go wrong even when the arithmetic looks polished. Students sometimes use the interval’s standard error in a test, check Large Counts with the wrong counts, or choose a one-sided alternative after seeing which group had the larger sample proportion.
As in “Conditions for a Two-Proportion \(z\)-Test” and “Finding the P-Value for a Two-Proportion Test,” the hypotheses and the null model determine how the test works. This tutorial focuses on a practical technique: audit each step against the question being asked. In particular, the equality in \(H_0:p_1=p_2\) means the test estimates one common proportion by pooling the successes and sample sizes.
A frequent source of confusion is that a two-proportion test and a two-proportion confidence interval do not use the same standard error or the same Large Counts check. The test evaluates evidence under the null hypothesis that the proportions are equal. The interval estimates a difference without assuming equality.
| Part of the analysis | Two-proportion test of equal proportions | Two-proportion confidence interval |
|---|---|---|
| Proportion used for the standard error | Pooled proportion from both groups | Separate sample proportions |
| Large Counts check | Four expected counts under the null model | Observed successes and failures in each group |
This distinction is not a minor calculator preference. The pooled standard error describes variability under the null model of equal proportions. The interval’s unpooled standard error, described in “Standard Error for a Difference in Proportions,” uses the two observed sample proportions to estimate variability for the difference.
Three Checks That Prevent Common Errors
Define \(p_1\) and \(p_2\) for the same outcome, then write an alternative that matches the research question. Keep this order in the sample difference, test statistic, calculator inputs, and conclusion.
For \(H_0:p_1=p_2\), use \(\hat{p}_c=(x_1+x_2)/(n_1+n_2)\) and the pooled standard error. Do not substitute the separate sample proportions into the test standard error.
Use the pooled proportion to find expected successes and failures in each group. Then use the alternative hypothesis—not whichever direction the observed difference happens to point—to select the p-value tail.
The Large Counts condition for a two-proportion test requires at least 10 expected successes and 10 expected failures in each group under the null model. For group \(i\), those expected counts are \(n_i\hat{p}_c\) and \(n_i(1-\hat{p}_c)\). Do not check this test condition by counting only the observed successes and failures.
For an interval, the check is different: as explained in “Large Counts for Each Group in Two-Proportion Intervals,” count the observed successes and failures in each group. Remembering which procedure is being used prevents a correct count from being applied to the wrong method.
Worked Examples
Worked Example: Using the Pooled Standard Error
Question: A fictional transit agency takes independent random samples of 40 riders from each of two service areas. In Area 1, 12 riders report that they used a mobile fare option during the past week; in Area 2, 8 do. Assume each area has at least 400 riders. At \(\alpha=0.05\), test whether the population proportions differ.
State: Let \(p_1\) be the true proportion of riders in Area 1 who used the mobile fare option during the past week, and let \(p_2\) be the corresponding proportion in Area 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). The question asks about any difference, so the alternative is two-sided.
Plan: Each group was selected as an independent random sample, and each rider is counted in only one area. Each sample of 40 is at most 10% of its area’s population of at least 400, so the 10% condition is met. The pooled proportion is:
Under the null model, each group has \(40(0.25)=10\) expected successes and \(40(0.75)=30\) expected failures. All four expected counts are at least 10, so the Large Counts condition is met. The conditions support a two-proportion \(z\)-test.
Do: The sample proportions are \(\hat{p}_1=12/40=0.30\) and \(\hat{p}_2=8/40=0.20\), so the observed difference is \(0.10\). For the test, use the pooled standard error:
The test statistic is \(z=(0.30-0.20)/0.09682\approx1.0328\). Because the alternative is two-sided, the p-value is the probability of a standard normal result at least this far from zero in either direction. It is approximately \(0.3017\), rounded.
Conclude: Since \(0.3017>0.05\), fail to reject \(H_0\). The samples do not provide convincing evidence that the true proportions of riders using the mobile fare option differ between the two service areas.
A common wrong turn is to use the separate sample proportions in the test standard error. That would borrow the interval method rather than calculate variability under the equal-proportions null model. The observed difference still supplies the numerator, but the pooled standard error belongs in this test statistic.
Worked Example: Checking Expected Counts and Pooling
Question: A fictional online retailer takes independent random samples of 100 customers from each of two customer populations. After seeing a checkout banner, 45 customers in Group 1 make a purchase; in Group 2, 30 do. Assume each population has at least 1,000 customers. The retailer’s question is whether the banner population has a higher purchase proportion. Test at \(\alpha=0.05\).
State: Let \(p_1\) be the true proportion of customers in the banner population who make a purchase, and \(p_2\) the true proportion in the comparison population. Test \(H_0:p_1=p_2\) against \(H_a:p_1>p_2\). The direction comes from the stated research question, not from the observed results.
Plan: The groups are independent random samples, and all customers have the same defined purchase outcome. Each sample of 100 is at most 10% of its population of at least 1,000, so the 10% condition is met. Pool the successes and sample sizes:
Under \(H_0\), each group has \(100(0.375)=37.5\) expected successes and \(100(0.625)=62.5\) expected failures. Each expected count is at least 10, so the Large Counts condition is met.
Do: The sample proportions are \(\hat{p}_1=0.45\) and \(\hat{p}_2=0.30\), with observed difference \(0.15\). The pooled standard error and statistic are:
For the upper-tail alternative, the p-value is \(P(Z\ge2.1909)\approx0.0142\), rounded. Since \(0.0142<0.05\), reject \(H_0\). The data provide convincing evidence that the true purchase proportion is higher in the banner population.
Why the unpooled calculation is a mistake here: Using the interval-style standard error would give:
That number is not the standard error for this test of equality. It uses the two observed sample proportions separately instead of the one common proportion assumed by \(H_0\). The correct test statistic uses \(SE_{\text{pooled}}\approx0.06847\). When checking your work, identify the procedure before entering values into a formula or calculator.
Worked Example: Choosing the Tail Before Looking at Results
Question: In a fictional randomized experiment, 120 commuters are randomly selected from a pool of more than 1,200 and randomly assigned equally to two app displays. The research question, set before the experiment, is whether Display 1 increases the proportion arriving on time. With Display 1, 18 of 60 commuters arrive on time; with Display 2, 24 of 60 do. Test at \(\alpha=0.05\).
State: Let \(p_1\) be the true proportion of commuters who would arrive on time using Display 1, and \(p_2\) the true proportion using Display 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1>p_2\). The alternative is upper-tailed because the question asks whether Display 1 increases the on-time proportion.
Plan: Commuters were randomly selected and randomly assigned, and each commuter used only one display, so the two groups are independent. The sample of 120 is at most 10% of the pool of more than 1,200, meeting the 10% condition. The pooled proportion is \((18+24)/(60+60)=42/120=0.35\). Under the null, each group has \(60(0.35)=21\) expected successes and \(60(0.65)=39\) expected failures. All four expected counts are at least 10, so the Large Counts condition is met.
Do: The observed proportions are \(\hat{p}_1=18/60=0.30\) and \(\hat{p}_2=24/60=0.40\). The pooled standard error and test statistic are:
The alternative is \(p_1>p_2\), so use the area to the right of the observed statistic: \(P(Z\ge-1.1483)\approx0.8746\), rounded. Do not switch to a lower-tail alternative because the observed difference is negative. If the preselected question had instead been whether Display 1 lowers the on-time proportion, the lower-tail area would be about \(0.1254\); that would be a different research question, not a choice to make after seeing the data.
Conclude: Since \(0.8746>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that Display 1 increases the true on-time proportion. This result does not prove the displays have equal proportions.
Common Mistakes and AP Exam Tips
- Using separate sample proportions for a test: This uses the interval’s unpooled standard error. For a test of equal proportions, calculate the pooled proportion first, then use it in \(SE_{\text{pooled}}\).
- Checking Large Counts with observed counts: For the test, the relevant counts are expected under \(H_0\): \(n_1\hat{p}_c\), \(n_1(1-\hat{p}_c)\), \(n_2\hat{p}_c\), and \(n_2(1-\hat{p}_c)\). Observed successes and failures are the check for the two-proportion interval.
- Choosing the tail from the sample result: An alternative must express the research question. Decide between \(p_1>p_2\), \(p_1<p_2\), and \(p_1\ne p_2\) before using the data to calculate the p-value.
- Reversing group order midway: If \(p_1-p_2\) is defined as Display 1 minus Display 2, keep that order in the sample difference, hypotheses, statistic, and conclusion. A positive \(z\)-statistic has meaning only alongside the stated order.
- Reporting a one-sided p-value for a two-sided question: Match the p-value area to the alternative. A two-sided alternative counts results at least as far from zero in either direction.
- Letting the calculator choose the analysis: On a TI-84, 2-PropZTest can calculate the pooled test statistic and p-value, but it cannot decide whether the conditions, group order, or selected alternative fit the study question. Check inputs and output against your setup.
- Overstating the conclusion: As explained in “Writing a Conclusion for a Two-Proportion Test,” state the decision about \(H_0\), then describe evidence about the population proportions in context. Do not say that the null hypothesis has been proved true.
A Quick Test Audit
Before submitting a two-proportion test, ask: Do \(p_1\) and \(p_2\) describe the same outcome for clearly defined groups? Does the alternative match the question and preserve the group order? Did I pool for both the test standard error and the expected-count check? Did I use the correct tail? Does my conclusion describe evidence about the population rather than merely repeat the sample proportions?
Check Your Understanding
Use the test-versus-interval distinction and the research question to identify the correct choice in each case.
- A two-proportion test has \(x_1=22\), \(n_1=80\), \(x_2=18\), and \(n_2=80\). What pooled proportion should be used for the test?
- For the test in Question 1, calculate the four expected success and failure counts under \(H_0:p_1=p_2\). Does the Large Counts condition hold?
- Why is it a mistake to use \(\hat{p}_1\) and \(\hat{p}_2\) separately in the standard error for a test of equal proportions?
- A question asks whether Group 1 has a higher population proportion than Group 2. The observed sample proportion in Group 1 is lower. Which alternative and p-value tail should be used?
- How does the Large Counts check for a two-proportion interval differ from the check for a two-proportion test?