Tutorials › AP Statistics › One-Sided and Two-Sided P-Values Compared

P-values and mean-inference conclusions · Tutorial 751 of 1000

One-Sided and Two-Sided P-Values Compared

Learn when a two-sided t-test p-value is twice a one-sided tail area, how to calculate it, and why the direction of the alternative must be chosen before looking at the data.

Intermediate 11 min read

What You'll Learn

  • Explain why a two-sided t-test counts extreme results in both tails.
  • Convert a one-sided tail area to a two-sided p-value using the observed t statistic.
  • Distinguish the p-values for left-tailed, right-tailed, and two-sided alternatives.
  • Recognize when doubling a one-sided p-value is appropriate and when it is not.
  • Explain why the alternative hypothesis must be selected before examining the sample result.
  • Compare one-sided and two-sided test decisions at a chosen significance level.

One Result, Different Questions

In “Why a t Test Never Proves the Null Mean,” you learned that a test assesses evidence against a null hypothesis; it does not prove that the null mean is true. This tutorial looks more closely at how the alternative hypothesis affects the p-value. For the same sample and t statistic, a two-sided test often has a p-value twice the one-sided tail area—but only when that one-sided area is in the direction of the observed statistic.

The reason is the question being asked. A one-sided alternative asks whether the mean is greater than the null value or whether it is less. A two-sided alternative asks whether the mean differs in either direction. If the observed t statistic is far above zero, evidence against the null in a two-sided test includes both a result at least that far above zero and a result at least that far below zero.

Definition: A one-sided t-test p-value is the probability, assuming \(H_0\) is true, of getting a t statistic at least as extreme as the observed statistic in the direction specified by \(H_a\). A two-sided t-test p-value includes results at least as far from zero as the observed statistic in either direction.

Under the null hypothesis for a t test, the t distribution is symmetric around zero. So, if the observed statistic is positive, the tail area beyond it equals the tail area below the corresponding negative value. A two-sided test adds those equal tail areas. If the observed statistic is negative, the same logic applies in reverse.

$$ \begin{aligned} t_{\text{obs}}>0:\quad p_{\text{two-sided}}&=2P(T\ge t_{\text{obs}}),\\ t_{\text{obs}}<0:\quad p_{\text{two-sided}}&=2P(T\le t_{\text{obs}}),\\ \text{equivalently:}\quad p_{\text{two-sided}}&=2P(T\ge |t_{\text{obs}}|). \end{aligned} $$

Here, \(T\) is the random variable representing the test statistic under \(H_0\), with the degrees of freedom for the test. These formulas use the symmetry of the t distribution. For an observed statistic of zero, each tail area is \(0.5\), so the two-sided p-value is \(1.0000\).

Which Tail Area Should You Double?

The sign of \(t_{\text{obs}}\) shows which direction the sample result went, based on the order in the test statistic. The alternative hypothesis determines which direction counts as evidence. For a one-sample t test of \(H_0:\mu=\mu_0\), a positive statistic means \(\bar{x}>\mu_0\); a negative statistic means \(\bar{x}<\mu_0\).

For a right-tailed alternative, \(H_a:\mu>\mu_0\), the p-value is the area to the right of the observed statistic. For a left-tailed alternative, \(H_a:\mu<\mu_0\), it is the area to the left. For a two-sided alternative, \(H_a:\mu\ne\mu_0\), it is the combined area in both tails beyond the observed statistic’s distance from zero.

Formula: For a t statistic with the appropriate degrees of freedom:
  • If \(H_a:\mu>\mu_0\), use the area to the right of \(t_{\text{obs}}\).
  • If \(H_a:\mu<\mu_0\), use the area to the left of \(t_{\text{obs}}\).
  • If \(H_a:\mu\ne\mu_0\), double the tail area beyond \(|t_{\text{obs}}|\).

The phrase “double the one-sided p-value” needs care. Double the one-sided area that is in the direction of the observed statistic—not necessarily the p-value for whichever one-sided alternative happens to be written down. If \(t_{\text{obs}}\) is positive, that area is the right-tail area. If \(t_{\text{obs}}\) is negative, it is the left-tail area. The p-value for the opposite one-sided alternative will be large, not half the two-sided p-value.

As in “Choosing One-Sided or Two-Sided for Two Means,” the research question determines the alternative. The choice must be made before looking at the sample result. Choosing whichever direction gives the smaller p-value after seeing the sign of \(t_{\text{obs}}\) makes the test procedure more likely to find evidence by chance than the stated significance level allows.

Worked Example: Testing a Mean Battery Life

Worked Example: Testing a Mean Battery Life

A fictional electronics lab takes a random sample of 16 batteries of one model. Their mean operating time is 52.4 hours, with a sample standard deviation of 6.4 hours. The sample is less than 10% of the relevant battery production, and the sample data show no severe skewness or extreme outliers. Test whether the true mean operating time differs from 50 hours. Use \(\alpha=0.05\), and compare the two-sided p-value with the one-sided area in the direction of the observed result.

1
State.
Let \(\mu\) be the true mean operating time, in hours, for batteries of this model in the production period represented by the sample. The hypotheses are \(H_0:\mu=50\) hours and \(H_a:\mu\ne50\) hours. The research question asks about a difference in either direction.
2
Plan and check conditions.
Use a one-sample t test for a population mean. The batteries were randomly sampled, meeting the random condition. The sample is less than 10% of the relevant production, supporting independence under the 10% condition. The data show no severe skewness or extreme outliers, so a t procedure is reasonable. The two-sided alternative matches the question.
3
Do.
The degrees of freedom are \(16-1=15\). The standard error is \(s/\sqrt{n}=6.4/\sqrt{16}=6.4/4=1.6\) hours. Thus
\(t_{\text{obs}}=(\bar{x}-\mu_0)/(s/\sqrt{n})=(52.4-50)/(6.4/\sqrt{16})=2.4/1.6=1.50.\)
Because the statistic is positive, the tail area in its direction is the right-tail area. Using a t distribution with 15 degrees of freedom, \(P(T\ge1.50)\approx0.0771833\). The two-sided p-value is \(2(0.0771833)=0.1543666\), or \(0.1544\) rounded to four decimal places. Since \(0.1544>0.05\), fail to reject \(H_0\).
4
Conclude in context.
The data do not provide convincing evidence that the true mean operating time for batteries of this model differs from 50 hours. The two-sided p-value counts results at least as far from the null mean as the observed result in either direction, which is why it is twice the right-tail area for this positive t statistic.

If the research question had instead been whether the mean operating time is greater than 50 hours, the right-tail p-value would be \(0.0772\), rounded. If a greater mean had been the prespecified claim, that would be the relevant one-sided p-value. But the stated question asks whether the mean differs, so the two-sided p-value is the appropriate one. The calculation does not allow us to switch alternatives after seeing that the sample mean is above 50.

Worked Example: A Negative t Statistic

Worked Example: A Negative t Statistic

A fictional facilities team takes a random sample of 11 rooms and records the number of minutes needed for a scheduled ventilation cycle. The sample mean is 18 minutes, and the sample standard deviation is \(\sqrt{11}\) minutes. The current target mean is 20 minutes. Assume the sample is less than 10% of the rooms in the relevant setting and the data have no severe skewness or extreme outliers. The team had specified in advance that it wanted to test whether the true mean is less than 20 minutes. Compare that one-sided test with a two-sided test of whether the mean differs from 20 minutes, using \(\alpha=0.05\).

State: Let \(\mu\) be the true mean ventilation-cycle time, in minutes, for rooms in this setting. The one-sided hypotheses are \(H_0:\mu=20\) and \(H_a:\mu<20\). The two-sided comparison uses \(H_0:\mu=20\) and \(H_a:\mu\ne20\).

Plan and check conditions: Use a one-sample t test. The sample is random, and the sample is less than 10% of the relevant rooms, supporting the random and 10% conditions. The data show no severe skewness or extreme outliers, so the t procedure is reasonable. The less-than alternative is appropriate here because it was specified before examining the sample.

Do: The degrees of freedom are \(11-1=10\), and the standard error is

$$ \frac{s}{\sqrt{n}} =\frac{\sqrt{11}}{\sqrt{11}} =1\text{ minute}. $$

The test statistic is

$$ t_{\text{obs}} =\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{18-20}{1} =-2.00. $$

For the prespecified left-tailed test, the p-value is the area to the left of \(-2.00\), which is approximately \(0.036694\), or \(0.0367\) rounded to four decimal places. For the two-sided test, the area in the opposite tail is equal by symmetry, so its p-value is \(2(0.036694)=0.073388\), or \(0.0734\) rounded. At \(\alpha=0.05\), the one-sided test rejects \(H_0\), while the two-sided test fails to reject \(H_0\).

Conclude: For the prespecified claim that the mean cycle time is less than 20 minutes, the data provide convincing evidence that the true mean cycle time in this setting is less than 20 minutes. For the broader question of whether the mean differs from 20 minutes in either direction, the data do not provide convincing evidence at the 0.05 significance level. These decisions differ because the alternatives ask different questions and count different tail areas; the analysis should report the alternative that was chosen in advance.

Worked Example: Reading All Three Alternatives

Worked Example: Reading All Three Alternatives

Suppose a calculator gives \(t_{\text{obs}}=-1.50\) with \(df=15\) for a sample mean compared with a hypothesized mean. The sign indicates that the sample mean is below the null value. Work out the p-values for a left-tailed alternative, a right-tailed alternative, and a two-sided alternative.

For \(H_a:\mu<\mu_0\), the observed result is in the direction of the alternative. The left-tail p-value is \(P(T\le-1.50)\approx0.0771833\), or \(0.0772\) rounded. For \(H_a:\mu>\mu_0\), the p-value is the area to the right of \(-1.50\): \(P(T\ge-1.50)\approx0.9228167\), or \(0.9228\) rounded. This is a large p-value because a statistic of \(-1.50\) is not evidence in the greater-than direction.

For \(H_a:\mu\ne\mu_0\), count both tails at least as far from zero as \(1.50\). By symmetry, each tail has area \(0.0771833\), so the two-sided p-value is \(2(0.0771833)=0.1543666\), or \(0.1544\) rounded. The three values can be summarized as follows:

AlternativeArea usedp-value, rounded
\(H_a:\mu<\mu_0\)Left of \(-1.50\)0.0772
\(H_a:\mu>\mu_0\)Right of \(-1.50\)0.9228
\(H_a:\mu\ne\mu_0\)Both tails beyond \(\pm1.50\)0.1544

Interpretation: The two-sided p-value is twice the smaller one-sided tail area here because the observed statistic is negative and the left tail is its direction. It is not twice the right-tailed p-value of \(0.9228\). Doubling that value would not represent the probability of getting a statistic at least as far from zero in either direction.

Common Mistakes and AP Exam Tips

  • Doubling the wrong tail: First note the sign of \(t_{\text{obs}}\). For a positive statistic, double the right-tail area; for a negative statistic, double the left-tail area. A correct response identifies the area used.
  • Using a one-sided alternative after seeing the data: Do not choose “greater than” or “less than” just because the sample mean went in that direction. The direction must come from the research question and be specified before examining the result.
  • Assuming every two-sided p-value is exactly twice every one-sided p-value: It is twice the one-sided tail area in the direction of the observed statistic, under the symmetric t model. The opposite one-sided p-value may be close to 1.
  • Describing a two-sided test as testing only the observed direction: A two-sided alternative asks about a difference in either direction, even if the sample result points only one way. Its p-value includes both tails.
  • Mixing up the p-value and the decision: Calculate the p-value for the stated alternative, then compare it with the prespecified \(\alpha\), as in “Making a Decision From P-Value and Alpha.” Do not change the alternative to get a different decision.
  • Leaving out the context: A full-credit conclusion names the population mean and the direction or difference described by \(H_a\). It says “convincing evidence” when rejecting \(H_0\), and “do not provide convincing evidence” when failing to reject.

A useful check is to imagine the t distribution centered at zero. Mark the observed statistic, then shade the tail in the direction of the one-sided alternative. For a two-sided test, shade the corresponding tail on the other side as well. Because the null t distribution is symmetric, the two shaded tail areas are equal. That picture explains both the doubling rule and why the rule depends on the sign of the observed statistic.

Key takeaway: For a two-sided t test, the p-value counts extreme results in both directions. With a symmetric t distribution, it is twice the one-sided tail area in the direction of \(t_{\text{obs}}\). Choose the alternative before examining the data, and double only the appropriate tail area.

Check Your Understanding

Use the stated alternatives and the observed statistic’s sign to decide which tail area is relevant.

  1. A t test has \(t_{\text{obs}}=1.8\) and a right-tail area of 0.045. What is the two-sided p-value, rounded to three decimal places?
  2. A t test has \(t_{\text{obs}}=-2.1\). Which tail area gives the p-value for \(H_a:\mu<\mu_0\), and which tail area gives it for \(H_a:\mu>\mu_0\)?
  3. Why is the two-sided p-value approximately twice the left-tail p-value when \(t_{\text{obs}}<0\)? State the property of the t distribution that makes this work.
  4. A student first sees a positive t statistic and then changes the planned alternative from \(H_a:\mu\ne\mu_0\) to \(H_a:\mu>\mu_0\). Explain why this is not an appropriate way to choose a test.
  5. For \(t_{\text{obs}}=-1.50\) with 15 degrees of freedom, the left-tail area is 0.0772. What is the two-sided p-value, and why is it not twice the right-tail p-value?