Tutorials › AP Statistics › Finding the P-Value for a Two-Sample t Test

Two-sample t hypothesis tests · Tutorial 727 of 1000

Finding the P-Value for a Two-Sample t Test

Use the test statistic, the alternative hypothesis, and Welch degrees of freedom to calculate and interpret a two-sample t-test p-value with tcdf.

Intermediate 10 min read

What You'll Learn

  • Match a left-, right-, or two-tailed p-value to the alternative hypothesis.
  • Enter the correct test statistic and Welch degrees of freedom in tcdf.
  • Calculate a two-tailed p-value by adding both tail areas or doubling one tail.
  • Understand how reversing the group order changes the statistic and the tail.
  • Interpret p-values in context without treating them as probabilities that the null hypothesis is true.

From Calculator Output to a P-Value

In “Using 2-SampTTest on the Calculator,” you learned how to enter two independent samples and read the test statistic and Welch degrees of freedom from the output. This tutorial focuses on how to find the p-value from those two numbers using the tcdf function.

The p-value is an area under a t-distribution curve. The alternative hypothesis tells you which area to find: a left-tail area, a right-tail area, or the combined area in both tails. The test statistic tells you where the area begins, and the Welch degrees of freedom tell you which t distribution to use.

Definition: A p-value is the probability, assuming the null hypothesis is true, of getting a test statistic at least as extreme as the observed statistic in the direction specified by the alternative hypothesis.

For a two-sample t test, use the \(t\) statistic and Welch \(df\) shown by the calculator. The test statistic is signed: a positive value means \(\bar{x}_1-\bar{x}_2\) is positive, and a negative value means it is negative. The sign and the alternative hypothesis together identify the relevant tail or tails.

Using tcdf for Each Alternative

On a TI-84, tcdf(lower bound, upper bound, df) returns the area under a t curve between the two bounds. Since a t curve extends indefinitely in both directions, use a very large positive or negative number as a practical substitute for infinity. For example, \(1\text{E}99\) and \(-1\text{E}99\) are convenient bounds.

Formula: For a test statistic \(t_{\text{obs}}\) and Welch degrees of freedom \(df\):
  • For \(H_a:\mu_1-\mu_2<0\), find the left-tail area with \(\operatorname{tcdf}(-1\text{E}99,t_{\text{obs}},df)\).
  • For \(H_a:\mu_1-\mu_2>0\), find the right-tail area with \(\operatorname{tcdf}(t_{\text{obs}},1\text{E}99,df)\).
  • For \(H_a:\mu_1-\mu_2\ne0\), find the areas at least as far from zero as the observed statistic in both directions. By symmetry, this is \(2\operatorname{tcdf}(|t_{\text{obs}}|,1\text{E}99,df)\).

The absolute value in the two-tailed formula matters. If \(t_{\text{obs}}\) is negative, the p-value still includes the equally extreme positive tail. Alternatively, calculate both tail areas and add them. The result should be the same, apart from rounding.

The order of the bounds in tcdf is important: the lower bound comes first. For a left-tail area, the observed statistic is the upper bound. For a right-tail area, it is the lower bound. A p-value is an area, so it must be between 0 and 1.

Key idea: Choose the tail from \(H_a\), not from whichever side of zero happens to contain \(t_{\text{obs}}\). Then use the statistic’s sign and the Welch \(df\) to enter the appropriate bounds in tcdf.

Worked Examples

Worked Example: A Two-Tailed Test for Battery Life

A hypothetical product team compares battery life, in hours, for two device settings. Independent random samples of 10 devices per setting produce \(\bar{x}_1=42\), \(\bar{x}_2=41\), and \(s_1=s_2=\sqrt{5}\) hours. The question is whether the population mean battery lives differ.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean battery lives for devices using Setting 1 and Setting 2, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).

Plan: Battery life is quantitative, and different devices are in the two samples, so the groups are independent rather than paired. Suppose the devices were randomly selected from large populations, and each sample is less than 10% of its population. Suppose plots of the battery-life data show no strong skewness or outliers. These conditions support an unpooled two-sample t test. Because the alternative is two-sided, the p-value must include both tails.

Do: First check the standard error and statistic. Each variance contribution is \(s_i^2/n_i=5/10=0.5\), so:

$$ SE=\sqrt{0.5+0.5}=1\text{ hour}, \qquad t=\frac{42-41}{1}=1. $$

The Welch degrees of freedom are:

$$ df= \frac{(0.5+0.5)^2} {\frac{0.5^2}{10-1}+\frac{0.5^2}{10-1}} = \frac{1}{0.5/9} =18. $$

On the calculator, enter \(\operatorname{tcdf}(1,1\text{E}99,18)\) for the upper-tail area. It is about \(0.1653\). The two-sided p-value is twice this area:

$$ p=2\operatorname{tcdf}(1,1\text{E}99,18) \approx 2(0.1653)=0.3306. $$

Conclude: At the 5% significance level, fail to reject \(H_0\), since \(0.3306>0.05\). These samples do not provide convincing evidence that the population mean battery lives differ between the two settings.

Worked Example: A Left-Tailed Test for Plant Growth

A hypothetical greenhouse comparison measures plant growth, in centimeters, after a fixed period under two lighting plans. Independent random samples of 6 plants per plan give \(\bar{x}_1=18\), \(\bar{x}_2=20\), and \(s_1=s_2=\sqrt{3}\) centimeters. The research question is whether mean growth is lower under Plan 1 than under Plan 2.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean plant growth under Plans 1 and 2, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2<0\).

Plan: Growth is quantitative, and the plants in one sample are different from those in the other, so the samples are independent. Suppose each sample was randomly selected, each is less than 10% of its population, and plots show no strong skewness or outliers. These checks support the two-sample t procedure. The left-sided alternative calls for the area to the left of the observed statistic.

Do: Each variance contribution is \(3/6=0.5\). Thus:

$$ SE=\sqrt{0.5+0.5}=1\text{ centimeter}, \qquad t=\frac{18-20}{1}=-2. $$

The Welch degrees of freedom are:

$$ df= \frac{(0.5+0.5)^2} {\frac{0.5^2}{6-1}+\frac{0.5^2}{6-1}} = \frac{1}{0.5/5} =10. $$

For a left-tail alternative, enter \(\operatorname{tcdf}(-1\text{E}99,-2,10)\). This gives \(p\approx0.0367\), rounded. Do not double this value: the alternative specifies only the left tail.

Conclude: At the 5% significance level, reject \(H_0\), since \(0.0367<0.05\). The samples provide convincing evidence that the true mean plant growth under Plan 1 is less than the true mean under Plan 2.

Worked Example: A Right-Tailed Test for Charging Time

A hypothetical comparison measures the time, in minutes, to charge a small device using two power modes. Independent random samples of 6 devices per mode give a mean of 31 minutes and a sample standard deviation of \(\sqrt{3}\) minutes for Mode 1, and a mean of 29 minutes and a sample standard deviation of \(\sqrt{3}\) minutes for Mode 2. The question is whether Mode 1 has a greater mean charging time.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean charging times for Modes 1 and 2, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2>0\).

Plan: Charging time is quantitative, and the devices in the two samples are different, so the groups are independent. Suppose the samples were randomly selected, each is less than 10% of its population, and plots show no strong skewness or outliers. These conditions support an unpooled two-sample t test. Since the alternative is right-tailed, use the area to the right of the statistic.

Do: The variance contributions are \(3/6=0.5\) for each group, so \(SE=\sqrt{0.5+0.5}=1\) minute. The statistic is \(t=(31-29)/1=2\). Welch’s degrees of freedom are:

$$ df= \frac{(0.5+0.5)^2} {\frac{0.5^2}{5}+\frac{0.5^2}{5}} = \frac{1}{0.1} =10. $$

Enter \(\operatorname{tcdf}(2,1\text{E}99,10)\). The right-tail p-value is approximately \(0.0367\), rounded. If the alternative had instead been \(\mu_1-\mu_2<0\), the p-value would be the opposite tail, about \(1-0.0367=0.9633\). That large value would not support the claim that Mode 1 has a lower mean charging time.

Conclude: At the 5% significance level, reject \(H_0\). These samples provide convincing evidence that Mode 1 has a greater population mean charging time than Mode 2.

Interpreting and Checking a tcdf Result

A p-value is calculated under the assumption that the null hypothesis is true. For example, a p-value of \(0.0367\) means that if the population means were equal, the probability of obtaining a test statistic at least as far in the specified direction as the observed statistic would be about \(0.0367\). It does not mean there is a \(3.67\%\) probability that the null hypothesis is true.

For a two-sided test, “at least as extreme” means at least as far from zero in either direction. For a one-sided test, it means at least as far in the direction specified by the alternative. This distinction is why the same \(t\) statistic and \(df\) can produce different p-values under different alternatives.

Use the Welch degrees of freedom from the unpooled test output, including its decimal part when the calculator reports one. Do not substitute \(n_1+n_2-2\): that is not the Welch degrees of freedom. If you calculate \(t\) or \(df\) yourself, retain unrounded values until you enter them in tcdf; early rounding can slightly change the p-value.

Calculator check: A left-tail p-value uses the interval from a very small number to \(t_{\text{obs}}\). A right-tail p-value uses the interval from \(t_{\text{obs}}\) to a very large number. A two-tail p-value adds both extreme areas; the symmetry shortcut doubles the upper-tail area beyond \(|t_{\text{obs}}|\).

Common Mistakes and AP Exam Tips

  • Using the wrong tail: Choose the tail from the stated alternative, not just from the sign of \(t\). A negative statistic does not automatically mean the test is left-tailed.
  • Forgetting the second tail: For \(H_a:\mu_1-\mu_2\ne0\), include both tails. Doubling the area beyond \(|t_{\text{obs}}|\) is a reliable shortcut.
  • Reversing tcdf bounds: The lower bound goes first. Check that your chosen bounds describe the area required by the alternative.
  • Using the wrong degrees of freedom: Use Welch \(df\) from the unpooled two-sample test, not the total sample size or a pooled degrees-of-freedom formula.
  • Misinterpreting the p-value: A p-value is calculated assuming \(H_0\) is true. A full-credit explanation describes the chance of a result at least as extreme as the observed one, in the direction of \(H_a\), and in the context of the study.
  • Reporting only a calculator number: State the approximate p-value and connect it to the research question. For a test conclusion, also compare it with the significance level and say “reject” or “fail to reject” \(H_0\).
Key takeaway: The alternative hypothesis determines which tail area tcdf must calculate. Use the observed two-sample \(t\) statistic and Welch \(df\), include one tail for a one-sided test or both tails for a two-sided test, and interpret the resulting probability under the assumption that \(H_0\) is true.

Check Your Understanding

For each question, identify the area tcdf should calculate and explain what the result represents.

  1. For \(H_a:\mu_1-\mu_2<0\), \(t_{\text{obs}}=-1.6\), and \(df=12\), what tcdf bounds describe the p-value?
  2. For \(H_a:\mu_1-\mu_2>0\), \(t_{\text{obs}}=1.6\), and \(df=12\), what tcdf bounds describe the p-value?
  3. For a two-sided test with \(t_{\text{obs}}=-1.6\) and \(df=12\), how can you use symmetry to calculate the p-value?
  4. If the observed statistic is positive but the alternative is left-tailed, which side of the t curve is used for the p-value?
  5. In context, what does a p-value of \(0.04\) say if the null hypothesis of equal population means is true?