Tutorials › AP Statistics › P-Values for One-Sided Versus Two-Sided Tests

P-values and conclusions for proportions · Tutorial 491 of 1000

P-Values for One-Sided Versus Two-Sided Tests

Learn how the alternative hypothesis determines which tail area counts as a p-value and why choosing the test direction before seeing the data matters.

Intermediate 9 min read

What You'll Learn

  • Calculate and compare one-sided and two-sided p-values from the same one-proportion z statistic
  • Explain why a matching one-sided p-value is half the two-sided p-value under the Normal model
  • Identify why an opposite-direction one-sided p-value is not half the two-sided value
  • Compare each p-value with a preselected significance level and describe the resulting decision
  • Explain why the alternative hypothesis must be selected before examining the sample result

One Sample, Different Questions

In Misinterpretations of the P-Value, you learned that a p-value measures how unusual the observed result, or a more extreme result, would be if the null hypothesis were true. The alternative hypothesis tells us what counts as “more extreme.” That means the same sample result can produce different p-values when the research questions—and therefore the alternatives—are different.

For a one-proportion \(z\)-test, the test statistic locates the observed sample proportion relative to the null proportion. A one-sided test counts results in one direction. A two-sided test counts results at least as far from the null value in either direction. When the observed result is in the direction of a one-sided alternative, the one-sided p-value is half the two-sided p-value under the symmetric Normal model used for the test.

Key idea: If the observed \(z\) statistic points in the direction specified by a one-sided alternative, then \(p_{\text{two-sided}}=2p_{\text{one-sided}}\). If it points in the opposite direction, the one-sided p-value is not half the two-sided p-value.

This relationship does not mean you may choose whichever alternative gives the smaller p-value after seeing the data. As explained in Choosing One-Sided or Two-Sided Alternatives, choose the alternative from the research question before examining the sample result. Whether the question asks “higher,” “lower,” or “different” determines which tail area is relevant.

Why the P-Values Differ by a Factor of Two

Suppose a test statistic is positive. A positive value means the sample proportion is above the null proportion. If the alternative is \(H_a:p>p_0\), the p-value is the area to the right of the observed \(z\). For \(H_a:p\ne p_0\), results at least as far from the null value in either direction count, so the p-value includes an equal-sized area in the left tail.

$$ \begin{aligned} H_a:p>p_0 &: \quad P(Z\geq z_{\mathrm{obs}})\\ H_a:p\ne p_0 &: \quad 2P(Z\geq |z_{\mathrm{obs}}|) \end{aligned} $$

The two tail areas are equal because the standard Normal curve is symmetric around zero. Thus, when \(z_{\mathrm{obs}}>0\), a right-tailed p-value is half the two-sided p-value. Similarly, when \(z_{\mathrm{obs}}<0\), a left-tailed p-value is half the two-sided p-value.

The matching direction matters. If \(z_{\mathrm{obs}}\) is positive but the alternative is \(H_a:p<p_0\), the observed result is in the direction opposite to the alternative. The left-tail p-value will be large, not half the two-sided p-value. The same principle applies when \(z_{\mathrm{obs}}\) is negative and the alternative is \(H_a:p>p_0\).

Worked Examples: Comparing Tail Areas and Decisions

Worked Example: A Higher Proportion or a Difference?

A fictional housing office takes a random sample of 160 households from a list of 3,200. In the sample, 76 households say they would use a proposed recycling pickup. Compare the p-values for testing whether the true proportion of households on the list who would use the service is higher than 40% and whether it differs from 40%. Use \(\alpha=0.05\) for each test.

State: Let \(p\) be the true proportion of households on the office’s list that would use the proposed pickup. The right-tailed test is \(H_0:p=0.40\) versus \(H_a:p>0.40\). The two-sided test is \(H_0:p=0.40\) versus \(H_a:p\ne0.40\). The research question determines which alternative is appropriate; the calculations below compare both questions using the same sample.

Plan: The sample was randomly selected, so the Random condition is met. Since \(160\leq0.10(3{,}200)=320\), the 10% condition is met. Under \(H_0\), the expected numbers of households who would and would not use the pickup are \(160(0.40)=64\) and \(160(0.60)=96\). Both are at least 10, so the Large Counts condition is met. A one-proportion \(z\)-test is appropriate.

Do: The sample proportion is \(\hat{p}=76/160=0.475\). Using the null proportion, the standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.40(1-0.40)}{160}} =\sqrt{0.0015} \approx0.03873,\qquad z=\frac{0.475-0.40}{0.03873} \approx1.9365 $$

For the right-tailed test, the p-value is \(P(Z\geq1.9365)\approx0.0264\). Assuming 40% of households on the list would use the service, a sample result at least this high would occur about 2.64% of the time under the test’s Normal model.

For the two-sided test, the p-value is \(2P(Z\geq|1.9365|)\approx2(0.0264)=0.0528\), rounded. Assuming the true proportion is 40%, a sample result at least this far from 40% in either direction would occur about 5.28% of the time under the model.

Conclude: At \(\alpha=0.05\), the right-tailed test rejects \(H_0\), because \(0.0264\leq0.05\). It provides convincing evidence that more than 40% of households on the list would use the pickup. The two-sided test fails to reject \(H_0\), because \(0.0528>0.05\). It does not provide convincing evidence that the proportion differs from 40%. The p-values differ by a factor of two because the observed sample proportion is above 40%, matching the right-tailed alternative. The difference in decisions is possible because the p-values fall on opposite sides of the chosen significance level.

Worked Example: When the Sample Points the Other Way

A fictional transit agency takes a random sample of 100 riders from a list of 2,000 monthly pass holders. Of those sampled, 53 say they would renew their pass. Compare a test that asks whether the true renewal proportion is below 60% with a test that asks whether it differs from 60%.

State: Let \(p\) be the true proportion of monthly pass holders on the list who would renew. The left-tailed test is \(H_0:p=0.60\) versus \(H_a:p<0.60\). The two-sided test is \(H_0:p=0.60\) versus \(H_a:p\ne0.60\).

Plan: The sample is random, meeting the Random condition. Since \(100\leq0.10(2{,}000)=200\), the 10% condition is met. Under \(H_0\), the expected counts are \(100(0.60)=60\) riders who would renew and \(100(0.40)=40\) who would not. Both counts meet the Large Counts condition.

Do: The sample proportion is \(\hat{p}=53/100=0.53\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.60(1-0.60)}{100}} =\sqrt{0.0024} \approx0.04899,\qquad z=\frac{0.53-0.60}{0.04899} \approx-1.4289 $$

The observed statistic is negative, matching the left-tailed alternative. Its p-value is \(P(Z\leq-1.4289)\approx0.0765\). The two-sided p-value is \(2P(Z\leq-1.4289)\approx0.1530\), rounded. The left-tailed p-value is half the two-sided value because this sample points in the direction of \(H_a:p<0.60\).

For comparison, if the question had instead been whether the proportion is higher than 60%, the right-tailed p-value would be \(P(Z\geq-1.4289)\approx0.9235\). That is not half of the two-sided p-value. The observed result is below 60%, opposite to the right-tailed alternative.

Conclude: At \(\alpha=0.05\), both the left-tailed and two-sided tests fail to reject \(H_0\). The sample does not provide convincing evidence that the renewal proportion is below 60%, nor does it provide convincing evidence that the proportion differs from 60%. The p-values are not probabilities that the null hypothesis is true; they describe sample results under the assumption that it is true.

Worked Example: A Factor-of-Two Difference Can Change the Decision

A fictional school district randomly samples 250 students from a list of 5,000 students. In the sample, 95 say they walk to school. Compare the p-values for testing whether the true proportion of students on the list who walk to school is greater than 32% and whether it differs from 32%. Suppose the significance level, selected before the sample was collected, is \(\alpha=0.025\).

State: Let \(p\) be the true proportion of students on the district’s list who walk to school. The alternatives are \(H_a:p>0.32\) for the right-tailed test and \(H_a:p\ne0.32\) for the two-sided test; both use \(H_0:p=0.32\).

Plan: The sample was randomly selected. The 10% condition is met because \(250\leq0.10(5{,}000)=500\). Under the null hypothesis, the expected counts are \(250(0.32)=80\) students who walk and \(250(0.68)=170\) who do not. Both expected counts are at least 10, so the Large Counts condition is met.

Do: The sample proportion is \(\hat{p}=95/250=0.38\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.32(1-0.32)}{250}} =\sqrt{0.0008704} \approx0.02950,\qquad z=\frac{0.38-0.32}{0.02950} \approx2.0337 $$

The right-tailed p-value is \(P(Z\geq2.0337)\approx0.0210\). The two-sided p-value is \(2P(Z\geq|2.0337|)\approx2(0.0210)=0.0420\), rounded. The observed proportion is above 32%, so the right-tailed p-value is half the two-sided p-value.

Conclude: At the preselected \(\alpha=0.025\), the right-tailed test rejects \(H_0\), since \(0.0210\leq0.025\). It provides convincing evidence that more than 32% of students on the list walk to school. The two-sided test fails to reject \(H_0\), since \(0.0420>0.025\); it does not provide convincing evidence that the proportion differs from 32%. If the preselected level had instead been 0.05, both tests would reject. The factor-of-two relationship changes the p-values, while the decision depends on comparing each p-value with the chosen \(\alpha\).

Common Mistakes and AP Exam Tips

  • Doubling every one-sided p-value automatically: Double a matching one-sided tail area for the two-sided test. First check whether the observed statistic points in the direction of the one-sided alternative.
  • Halving the two-sided p-value for an opposite-direction alternative: If the sample points against the one-sided alternative, its p-value is the area in the other tail and can be much larger. In the transit example, the right-tailed p-value is about 0.9235, not half of 0.1530.
  • Choosing the direction after seeing the sample: The alternative comes from the research question and should be chosen before looking at the result. Selecting the more favorable tail afterward makes the reported p-value misleading.
  • Assuming a factor of two always changes the decision: The two p-values may both be below \(\alpha\), both above \(\alpha\), or fall on opposite sides. Compare each p-value with the preselected significance level rather than assuming the decision.
  • Changing the conclusion without naming the claim: A conclusion should identify the population proportion and the direction or difference being evaluated. State “convincing evidence that the proportion is greater than…” only for the right-tailed alternative, not merely because the sample proportion is high.
AP Exam Tip: Show the alternative hypothesis, identify the tail area it requires, and connect the observed statistic to that direction. For a matching one-sided and two-sided comparison, state that symmetry of the Normal model makes the two-sided p-value twice the one-sided p-value. Then compare each p-value with \(\alpha\) and give a conclusion in context.

Key Takeaway

The p-value depends on what the alternative hypothesis counts as evidence against \(H_0\). When the observed statistic points in the direction of a one-sided alternative, the two-sided p-value is twice the matching one-sided p-value under the Normal model. When it points the other way, that simple relationship does not apply. The alternative must be chosen from the research question before the data are examined.

Key takeaway: Same data and same null hypothesis can produce different p-values because the alternatives define different sets of results as at least as extreme. Choose the alternative first, calculate the corresponding p-value, and compare it with the preselected \(\alpha\).

Check Your Understanding

Use the relationship between tail areas and the alternative hypothesis to answer each question.

  1. A one-proportion test has \(z_{\mathrm{obs}}=1.80\). Explain how the p-value for \(H_a:p>p_0\) relates to the p-value for \(H_a:p\ne p_0\).
  2. A test has \(z_{\mathrm{obs}}=-2.10\). Which one-sided alternative gives a p-value that is half the two-sided p-value: \(H_a:p<p_0\) or \(H_a:p>p_0\)? Explain.
  3. A test has \(z_{\mathrm{obs}}=1.50\) and \(H_a:p<p_0\). Explain why its one-sided p-value is not half the two-sided p-value.
  4. In the housing example, why can the right-tailed test reject at \(\alpha=0.05\) while the two-sided test fails to reject?
  5. Why is it not appropriate to choose a one-sided alternative only after seeing which direction the sample proportion differs from \(p_0\)?