Use the Calculator to Carry Out the Test
In Finding the P-Value for a Two-Proportion Test, you matched the alternative hypothesis to the appropriate tail or tails of the standard normal distribution. The TI-84’s 2-PropZTest performs the calculations for a two-proportion \(z\)-test and reports the p-value for the alternative you select. You still need to define the groups, enter the counts in the right order, check the conditions, and interpret the output.
The calculator asks for \(x_1,n_1,x_2,n_2\), where \(x_i\) is the number of successes and \(n_i\) is the sample size for Group \(i\). It also asks you to choose an alternative: \(p_1\ne p_2\), \(p_1<p_2\), or \(p_1>p_2\). These choices correspond to the two-sided, lower-tail, and upper-tail tests, respectively.
On a TI-84, open STAT, move to TESTS, and choose 2-PropZTest. The calculator’s input order follows the group order in your hypotheses: enter Group 1’s successes and sample size first, followed by Group 2’s successes and sample size. Do not enter the number of failures where it asks for the number of successes.
Decide what Group 1 and Group 2 represent, and define the same success outcome for both.
Use \(H_0:p_1=p_2\), then choose the alternative that matches the research question.
Use the study design to assess randomness and independence, and check the Large Counts condition using expected counts under the null.
Input \(x_1,n_1,x_2,n_2\) in the order you defined, then select \(p_1\ne p_2\), \(p_1<p_2\), or \(p_1>p_2\).
Identify \(z\), the calculator’s \(p\)-value, \(\hat{p}_1\), \(\hat{p}_2\), and the pooled \(\hat{p}\). Interpret them in light of the hypotheses and context.
The display generally labels the results \(z\), \(p\), \(\hat{p}_1\), \(\hat{p}_2\), and \(\hat{p}\). Here, the calculator’s \(p\) means the p-value; the pooled \(\hat{p}\) is a proportion estimate. They are different quantities despite the similar notation. As explained in Pooled Proportion and Why It Is Used, the pooled estimate combines the successes and sample sizes from both groups because the null hypothesis assumes a common population proportion.
Read the Output in the Right Order
First, check that the counts and alternative on the screen match your written setup. Then read each result by its label. The \(z\)-value is the standardized test statistic: it measures how many pooled standard errors the observed difference \(\hat{p}_1-\hat{p}_2\) is from zero. A positive \(z\) means the Group 1 sample proportion is greater than the Group 2 sample proportion; a negative \(z\) means it is smaller.
The calculator’s \(p\) is the p-value for the alternative you selected. Its value depends on that alternative, not merely on the sign of \(z\). As in Finding the P-Value for a Two-Proportion Test, a two-sided test counts results at least as far from zero in either direction, while a one-sided test counts results in the direction named by its alternative.
The calculator also reports \(\hat{p}_1=x_1/n_1\) and \(\hat{p}_2=x_2/n_2\). These show the sample proportions and help you check the direction and size of the observed difference. The pooled \(\hat{p}\) is different: it combines both groups’ counts. It does not replace the separate sample proportions when describing what was observed.
Worked Examples: Enter, Read, and Interpret
Worked Example: An Upper-Tail Test for Two Programs
Question: In a fictional randomized experiment, 60 of 100 participants assigned to Program A complete a short course, compared with 40 of 100 assigned to Program B. Is there evidence that the completion proportion is higher for Program A? Use 2-PropZTest and interpret its output at \(\alpha=0.05\).
State: Let \(p_1\) and \(p_2\) be the true proportions of participants who would complete the course under Programs A and B, respectively. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1>p_2\). The alternative calls for an upper-tail test.
Plan: Participants were randomly assigned to the two programs, supporting use of a two-proportion test. Assume outcomes for different participants are independent and that assignment creates separate groups. This is a randomized experiment rather than sampling without replacement, so the 10% condition is not needed. Under the null hypothesis, the pooled proportion is \((60+40)/(100+100)=0.50\). Each group therefore has 50 expected successes and 50 expected failures. All four expected counts are at least 10, so the Large Counts condition is met.
Do: On the TI-84, choose 2-PropZTest and enter \(x_1=60\), \(n_1=100\), \(x_2=40\), and \(n_2=100\). Select \(p_1>p_2\). The calculator’s output is approximately \(z=2.8284\), \(p=0.00234\), \(\hat{p}_1=0.60\), \(\hat{p}_2=0.40\), and pooled \(\hat{p}=0.50\).
A hand check uses the pooled standard error from Calculating the Pooled Standard Error:
The observed difference is \(0.60-0.40=0.20\), so the test statistic is:
For the upper-tail alternative, the p-value is the standard normal area to the right of this statistic, approximately \(0.00234\), matching the calculator’s result.
Conclude: Assuming equal completion proportions under the two programs, the probability of obtaining a sample difference at least as large in favor of Program A as the observed difference is about \(0.00234\). Because this p-value is less than \(0.05\), we reject \(H_0\). The experiment provides convincing evidence that the completion proportion is higher under Program A than under Program B.
Worked Example: A Two-Sided Test for Two Gardens
Question: In a fictional survey, 56 of 100 randomly selected gardeners in District 1 and 44 of 100 randomly selected gardeners in District 2 say they compost food scraps. Is there evidence that the population proportions differ?
Set up and check: Let \(p_1\) and \(p_2\) be the true proportions of gardeners in Districts 1 and 2 who compost food scraps. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). The samples are random samples from separate districts, supporting the Random condition and independence between groups. Assume each sample is less than 10% of its district’s gardeners, and that individual responses are independent. The pooled proportion is \((56+44)/200=0.50\), giving 50 expected successes and 50 expected failures in each group. The Large Counts condition is met.
Do: Enter \(56,100,44,100\), in that order, and select \(p_1\ne p_2\). The output is approximately \(z=1.6971\), \(p=0.0897\), \(\hat{p}_1=0.56\), \(\hat{p}_2=0.44\), and pooled \(\hat{p}=0.50\). The positive statistic agrees with the sample difference \(0.56-0.44=0.12\).
To check \(z\), the pooled standard error is:
Thus \(z=0.12/0.0707107\approx1.6971\). Because the alternative is two-sided, the p-value includes both tails beyond \(|1.6971|\); it is approximately \(0.0897\).
Conclude: Assuming the true composting proportions are equal in the two districts, the probability of getting a sample difference at least as far from zero as the observed difference, in either direction, is about \(0.0897\). At \(\alpha=0.05\), we fail to reject \(H_0\). The survey does not provide convincing evidence that the composting proportions differ between the two districts.
Worked Example: An Upper-Tail Alternative with a Negative Statistic
Question: In a fictional survey, 35 of 100 randomly selected customers using Checkout A and 45 of 100 using Checkout B say they would recommend the store. A manager asks whether the recommendation proportion is higher for Checkout A. What does 2-PropZTest report?
Let \(p_1\) and \(p_2\) be the true recommendation proportions for customers using Checkouts A and B. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1>p_2\). Assume the two groups are independent random samples, each less than 10% of its customer population, with independent responses. The pooled proportion is \((35+45)/200=0.40\). The expected successes are \(100(0.40)=40\) in each group, and expected failures are \(100(0.60)=60\) in each group. All expected counts meet the Large Counts condition.
Enter \(35,100,45,100\) and select \(p_1>p_2\). The output is approximately \(z=-1.4434\), \(p=0.9255\), \(\hat{p}_1=0.35\), \(\hat{p}_2=0.45\), and pooled \(\hat{p}=0.40\). The negative \(z\) matches the observed difference \(0.35-0.45=-0.10\). Since the alternative asks for evidence that \(p_1\) is greater, the calculator reports the area to the right of a negative statistic, which is large.
Interpretation: Assuming equal population proportions, the probability of obtaining a sample difference at least as favorable to Checkout A as the observed difference is about \(0.9255\). This large p-value does not provide convincing evidence that Checkout A’s recommendation proportion is higher. The alternative—not the sign of the observed difference—determines which tail the calculator uses.
Common Mistakes and AP Exam Tips
- Reversing the input order: If Group 1 is Checkout A, enter its successes and sample size first. Reversing groups changes the sign of \(z\) and swaps the sample-proportion labels. It can also make a one-sided alternative point in the wrong direction.
- Entering failures instead of successes: The \(x_i\) input is the number of successes, not the number of failures. Check that the success definition is the same for both groups.
- Choosing an alternative after seeing the result: Select the alternative that matches the research question and hypotheses. Do not switch from \(p_1>p_2\) to \(p_1\ne p_2\) simply because the observed statistic has an unexpected sign.
- Confusing the calculator’s \(p\) with pooled \(\hat{p}\): The calculator’s \(p\) is the p-value. The pooled \(\hat{p}\) is a combined sample proportion used under the null hypothesis.
- Treating the calculator as a condition checker: 2-PropZTest returns numbers; it does not establish that the randomization or sampling design is appropriate, that groups are independent, or that expected counts are large enough. Check these conditions yourself.
- Reporting only the display: A complete response identifies the test and alternative, checks conditions, reports and interprets the p-value, and states a conclusion about the population proportions. A calculator screen alone does not explain what the evidence means.
Key Takeaway
2-PropZTest is useful because it calculates the two-proportion test statistic and the p-value for a selected alternative. Its output is meaningful only when the entries, hypotheses, and test conditions are appropriate. Read the labels carefully: \(z\) describes the standardized sample difference, calculator \(p\) is the p-value, and pooled \(\hat{p}\) is the combined proportion under the equal-proportions null model.
Check Your Understanding
Assume the two-proportion test conditions are met unless a question asks you to identify a condition.
- Group 1 has 38 successes in a sample of 80, and Group 2 has 29 successes in a sample of 75. What four counts should be entered in 2-PropZTest?
- For \(H_a:p_1<p_2\), which alternative option should you select on the calculator?
- In a test of \(H_0:p_1=p_2\), what does a positive \(z\)-statistic tell you about the two sample proportions?
- For counts \(x_1=45,n_1=90,x_2=35,n_2=90\), calculate the pooled proportion the calculator should report.
- Why can an upper-tail test have a large p-value when the calculator reports a negative \(z\)?