Tutorials › AP Statistics › Finding the P-Value for a Two-Proportion Test

Two-proportion hypothesis tests · Tutorial 528 of 1000

Finding the P-Value for a Two-Proportion Test

Use the alternative hypothesis and the computed z statistic to find and interpret the correct normal-curve area for a two-proportion test.

Intermediate 9 min read

What You'll Learn

  • Match an upper-tail, lower-tail, or two-sided alternative to the appropriate normal-curve area.
  • Use normalcdf with the computed z statistic to find a two-proportion test’s p-value.
  • Calculate and check a pooled test statistic before finding its p-value.
  • Explain why a negative z statistic can still produce a large p-value for an upper-tail test.
  • Interpret the p-value in context without treating it as the probability that the null hypothesis is true.
  • Write a contextual test conclusion by comparing the p-value with the significance level.

From a Two-Proportion z Statistic to a P-Value

In Large Counts Using Expected Successes and Failures in Each Group, you checked whether a normal model is appropriate for a two-proportion \(z\)-test. Once the conditions are supported and the test statistic has been calculated, the next task is to find the p-value: the area in the appropriate tail or tails of the standard normal curve.

For a test of \(H_0:p_1=p_2\), the statistic \(z\) measures how many pooled standard errors the observed difference \(\hat{p}_1-\hat{p}_2\) is from zero. Under the null hypothesis, the standardized statistic is modeled by a standard normal distribution. The p-value is the probability, assuming \(H_0\) is true, of getting a test statistic at least as extreme as the observed one in the direction or directions specified by \(H_a\).

Definition: The p-value for a two-proportion \(z\)-test is the standard normal area at least as extreme as the observed \(z\), in the direction or directions specified by the alternative hypothesis.

The alternative hypothesis tells you which area to calculate. Keep the group order in mind: for example, a positive \(z\) means the observed difference \(\hat{p}_1-\hat{p}_2\) is positive, but it does not determine the direction of the alternative. That direction was set by the research question and hypotheses.

Match the Alternative to the Tail

If \(H_a:p_1>p_2\), evidence against \(H_0\) is in the upper tail. If \(H_a:p_1<p_2\), evidence is in the lower tail. If \(H_a:p_1\ne p_2\), evidence in either direction counts, so the p-value includes both tails.

Formula: For a standard normal random variable \(Z\) and observed statistic \(z_{\text{obs}}\):
  • For \(H_a:p_1>p_2\), the p-value is \(P(Z\ge z_{\text{obs}})\).
  • For \(H_a:p_1<p_2\), the p-value is \(P(Z\le z_{\text{obs}})\).
  • For \(H_a:p_1\ne p_2\), the p-value is \(2P(Z\ge |z_{\text{obs}}|)\).

On a TI-84 or similar calculator, use normalcdf(lower bound, upper bound, 0, 1). The last two inputs specify the standard normal distribution. Since the calculator needs finite bounds, use a very large positive or negative number such as \(1\text{E}99\) to stand in for infinity.

$$ \begin{aligned} \text{Upper tail: }&\operatorname{normalcdf}(z_{\text{obs}},1\text{E}99,0,1)\\ \text{Lower tail: }&\operatorname{normalcdf}(-1\text{E}99,z_{\text{obs}},0,1)\\ \text{Two tails: }&2\operatorname{normalcdf}(|z_{\text{obs}}|,1\text{E}99,0,1) \end{aligned} $$

For a two-sided test, the standard normal curve is symmetric around zero. Thus, the two tail areas beyond \(z_{\text{obs}}\) and \(-z_{\text{obs}}\) are equal. You can calculate one tail beyond \(|z_{\text{obs}}|\) and double it, or calculate each tail and add the areas. Do not double a one-sided p-value unless the test is two-sided.

The sign of \(z_{\text{obs}}\) matters for a one-sided test. If \(H_a:p_1>p_2\) but the observed \(z\) is negative, the sample difference points away from the alternative. The upper-tail area will then be large. A large p-value is not an error; it reflects that the observed direction is not the direction counted as evidence by \(H_a\).

Worked Examples: Find the Correct Normal Area

Worked Example: An Upper-Tail Test

Question: In a fictional randomized experiment, 78 of 120 participants assigned to Program 1 and 60 of 120 assigned to Program 2 complete a training module. Is there evidence that the completion proportion is higher for Program 1? Find and interpret the p-value.

State: Let \(p_1\) and \(p_2\) be the true proportions of participants who would complete the module under Programs 1 and 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1>p_2\). The alternative calls for an upper-tail p-value.

Plan: The experiment randomly assigns participants to separate program groups, supporting use of a two-proportion test. Assume each participant’s outcome is independent of the others. Random assignment, rather than sampling without replacement, was used, so the 10% condition is not needed here. The pooled proportion is \(138/240=0.575\), giving expected successes of \(120(0.575)=69\) and expected failures of \(120(0.425)=51\) in each group. All four expected counts are at least 10, so the Large Counts condition is met.

Do: The sample proportions are \(78/120=0.65\) and \(60/120=0.50\). Using the pooled standard error for the test:

$$ SE_{\text{pooled}} =\sqrt{0.575(0.425)\left(\frac{1}{120}+\frac{1}{120}\right)} =\sqrt{0.0040729167} \approx0.06382 $$

The test statistic, rounded to four decimal places, is:

$$ z=\frac{0.65-0.50}{0.0638194} \approx2.3504 $$

Because the alternative is \(p_1>p_2\), find the area to the right of \(2.3504\):

$$ \text{p-value} =\operatorname{normalcdf}(2.3504,1\text{E}99,0,1) \approx0.0094 $$

Conclude: Assuming the two programs have equal completion proportions, the probability of obtaining a sample difference at least as favorable to Program 1 as the observed difference, by chance, is about \(0.0094\). This is the p-value, not the probability that the null hypothesis is true. At \(\alpha=0.05\), the p-value is smaller than the significance level, so we reject \(H_0\). The experiment provides convincing evidence that the completion proportion is higher under Program 1 than under Program 2.

Worked Example: A Lower-Tail Test

Question: In a fictional survey, 48 of 100 randomly selected residents in Town A and 60 of 100 randomly selected residents in Town B say they use a community bike-share service. Is there evidence that Town A’s use proportion is lower? Find the p-value.

State: Let \(p_1\) and \(p_2\) be the true proportions of residents in Town A and Town B, respectively, who use the service. Test \(H_0:p_1=p_2\) against \(H_a:p_1<p_2\). This alternative requires a lower-tail area.

Plan: The two samples are random samples from separate towns, so the Random condition and independence between groups are supported. Assume each sample is less than 10% of its town’s residents, and that responses from different residents are independent. Under the null hypothesis, the pooled proportion is \(108/200=0.54\). The expected successes in each group are \(100(0.54)=54\), and the expected failures are \(100(0.46)=46\). All four expected counts meet the Large Counts condition.

Do: The observed difference is \(0.48-0.60=-0.12\). The pooled standard error is:

$$ SE_{\text{pooled}} =\sqrt{0.54(0.46)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.004968} \approx0.070484 $$

Thus:

$$ z=\frac{-0.12}{0.070484} \approx-1.7025 $$

For the lower-tail alternative, calculate the area to the left of \(-1.7025\):

$$ \text{p-value} =\operatorname{normalcdf}(-1\text{E}99,-1.7025,0,1) \approx0.0443 $$

Conclude: If the true use proportions in the two towns are equal, the probability of obtaining a sample difference at least as low as the observed difference is about \(0.0443\). At \(\alpha=0.05\), we reject \(H_0\). The survey provides convincing evidence that the proportion of residents who use the service is lower in Town A than in Town B.

Worked Example: A Two-Sided Test

Question: In a fictional randomized trial, 60 of 120 participants assigned to a reminder system submit a form on time, compared with 48 of 120 participants assigned to the usual process. Test whether the on-time submission proportions differ, and find the p-value.

State: Let \(p_1\) and \(p_2\) be the true proportions of participants who submit on time under the reminder system and usual process. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). Either a higher or a lower proportion would count as evidence, so both tails are included.

Plan: Assume participants are randomly assigned to separate groups and have independent outcomes. Because this is a randomized experiment, the 10% condition for sampling without replacement does not apply. The pooled proportion is \(108/240=0.45\), giving expected successes of \(120(0.45)=54\) and expected failures of \(120(0.55)=66\) in each group. Each expected count is at least 10.

Do: The sample difference is \(60/120-48/120=0.10\). The pooled standard error and test statistic are:

$$ SE_{\text{pooled}} =\sqrt{0.45(0.55)\left(\frac{1}{120}+\frac{1}{120}\right)} =\sqrt{0.004125} \approx0.06423 $$
$$ z=\frac{0.10}{0.0642262} \approx1.5570 $$

For a two-sided test, double the area to the right of the positive value \(|1.5570|\):

$$ \text{p-value} =2\operatorname{normalcdf}(1.5570,1\text{E}99,0,1) \approx2(0.059744...) \approx0.1195 $$

Conclude: Assuming equal population proportions, the probability of obtaining a difference at least as far from zero as the observed difference, in either direction, is about \(0.1195\). At \(\alpha=0.05\), we fail to reject \(H_0\). The trial does not provide convincing evidence that the on-time submission proportions differ between the reminder system and the usual process.

Common Mistakes and AP Exam Tips

  • Choosing a tail from the sign of \(z\) alone: The alternative hypothesis, not the observed sign, determines which tail counts as evidence. Identify \(H_a\) before entering normalcdf.
  • Using the wrong bound order: For an upper-tail area, the lower bound is \(z_{\text{obs}}\) and the upper bound is \(1\text{E}99\). For a lower-tail area, the bounds are \(-1\text{E}99\) and \(z_{\text{obs}}\).
  • Forgetting the second tail: A two-sided alternative includes results at least as far from zero in both directions. Calculate one tail beyond \(|z_{\text{obs}}|\) and double it, or calculate both areas and add them.
  • Doubling a one-sided p-value automatically: Double the area only when the alternative is two-sided. The same \(z\) can give different p-values under different alternatives.
  • Rounding the test statistic too early: Keep several digits in the standard error and \(z\) when calculating the tail area. Report the p-value rounded appropriately, usually to three or four decimal places.
  • Calling the p-value the probability that \(H_0\) is true: As explained in Misinterpretations of the P-Value, the p-value is a probability about sample results, calculated assuming \(H_0\) is true.
  • Giving only a calculator number: A full-credit response identifies the tail, reports the p-value, and interprets it in context. When asked for a conclusion, also compare the p-value with \(\alpha\), state “reject” or “fail to reject,” and describe the evidence about the population proportions.
AP Exam Tip: Write the alternative and tail choice before showing the normalcdf command. For a two-sided test, show why you doubled the tail area. Then interpret the p-value under the assumption that the null hypothesis is true; do not describe it as the chance that the observed result is “just random” without specifying the null model and the relevant direction or directions.

Key Takeaway

The test statistic tells how far the observed difference is from zero in pooled standard-error units. The alternative hypothesis tells which part of the standard normal curve counts as evidence against the null hypothesis. Use normalcdf to find exactly that area, then interpret it under the assumption that \(H_0\) is true.

Key takeaway: Upper-tail alternatives use the area to the right of \(z\), lower-tail alternatives use the area to the left, and two-sided alternatives use both tails beyond \(|z|\). Match the calculator bounds to \(H_a\), not just to the sign of the statistic.

Check Your Understanding

For each question, assume the two-proportion \(z\)-test conditions have been checked unless the question asks about the p-value method.

  1. For \(H_a:p_1>p_2\) and \(z=1.40\), write the normalcdf command for the p-value.
  2. For \(H_a:p_1<p_2\) and \(z=-2.10\), which tail is used? Write the calculator command.
  3. For \(H_a:p_1\ne p_2\) and \(z=-1.25\), describe how to use symmetry to find the p-value.
  4. If \(H_a:p_1>p_2\) but the observed \(z\) is negative, should the p-value be small or large? Explain why.
  5. In context, what does a p-value of \(0.03\) mean for a two-sided test of two population proportions?